Method for mining tumorigenesis molecular target based on single cell pedigree tracer technology

Through single-cell lineage tracing technology, using barcode marking and lentiviral transfection, the clonal lineage changes of tumor cells are dynamically tracked, which solves the problem that traditional single-cell sequencing technology cannot dynamically track tumor heterogeneity evolution, and achieves accurate identification of molecular targets in the tumor progression stage, improving the target screening accuracy and clinical relevance.

CN120452526APending Publication Date: 2025-08-08CHONGQING UNIV CANCER HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510506607.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional single-cell sequencing technology cannot dynamically track tumor heterogeneity evolution, cannot accurately identify key molecular targets at different tumor progression stages, and lacks cross-platform data integration verification, resulting in low target screening accuracy.

Method used

Using single-cell lineage tracing technology, tumor cells are marked with lentiviral vectors containing barcodes, and unique identity markers are constructed. Through single-cell transcriptome sequencing and dimensionality reduction analysis, combined with TCGA/GEO database to verify the prognosis and pathway correlation of target genes, and dynamically track the clonal lineage changes of tumor cells.

Benefits of technology

It has achieved dynamic tracking of the clonal lineage evolution of tumor cells, accurately identified specific molecular targets at different tumor progression stages, enhanced the biological significance and clinical relevance of the target, and provided a highly reliable molecular target for the early diagnosis and staged treatment of cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452526A_ABST
    Figure CN120452526A_ABST
Patent Text Reader

Abstract

The invention relates to a method for mining a tumorigenesis molecular target based on a single cell pedigree tracing technology, and belongs to the technical field of biology. The invention aims to solve the problems that the traditional single cell sequencing technology cannot dynamically track the clone evolution process of cancer cells and is difficult to accurately identify key molecular targets for driving different stages of tumors. The invention provides a method for mining a tumorigenesis molecular target based on a single cell pedigree tracing technology. The method comprises the following steps: transfecting cancer cells carrying a unique bar code marker by using lentivirus, separating the cells in stages after in-situ tumorigenesis for single cell sequencing, screening the unique marker cells through quality control and dimensionality reduction, and drawing a clone pedigree evolutionary diagram; screening branch point instantaneous up-regulation genes, and verifying prognosis and pathway relevance of target genes by combining with a TCGA / GEO database. According to the invention, dynamic tracking of cancer cell cloning is realized, and staging cancer promoting genes are accurately locked; the biological function and clinical value of the target are verified through cross-platform data, high-credibility molecular targets are provided for cancer stage diagnosis and treatment, and development of accurate treatment strategies is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology and relates to a method for mining tumor progression molecules based on single cell lineage tracing technology. Background Art

[0002] Tumors are highly heterogeneous, not only in terms of the diverse cell types within tumor tissue, including tumor cells, stromal cells, and immune cells, but also in the heterogeneity of tumor cells themselves. A typical example is the presence of a small population of tumor stem cells within tumor cells that possess the ability to initiate, self-renew, and differentiate. These stem cells can differentiate into rapidly proliferating tumor cells, contributing to the rapid growth of the tumor. Furthermore, under specific conditions, differentiated tumor cells can also dedifferentiate into tumor stem cells. Therefore, the dynamic interaction between tumor cells and their microenvironment promotes tumor progression. This research project focuses on comprehensively analyzing the role and dynamics of tumor cells in tumor progression, as well as the regulatory molecules that play a crucial role.

[0003] Traditional single-cell sequencing technology is very useful for analyzing the heterogeneity between different cell populations at a specific point in time. However, the dynamic changes of cells within a certain time range and the evolutionary relationship of clonal lineages cannot be captured by traditional models, so it is impossible to discover more effective gene targets. Lineage tracing refers to the tracking and observation of the differentiation and developmental activities of a single cell and all its descendant cells. It is very important for understanding the developmental process of the body and the occurrence of cancer. Through "cell labeling", a method of combining cell indexes, it is possible to capture clonal history and cell identity in parallel, and through several consecutive rounds of cell labeling, a multi-level lineage tracing tree can be constructed. It is an improvement on the original single-cell technology and has greatly enriched the application of single-cell technology in the field of precision tumor treatment. At the same time, based on the more accurate molecular targets obtained, it also provides more scientific basis for clinicians' diagnosis.

[0004] Scientists use this technology in in vitro animal models, especially mouse experimental animals, to comprehensively track the developmental process of multicellular systems and analyze the transcriptional regulatory processes behind cell fate selection. For example, the research conducted by Dr. Hans-Reimer Rodewald's team in the 2020 journal Cell Stem Cell showed that they used cell labeling, a barcode labeling technology, to classify hematopoietic stem cells according to their developmental fate, and identified their candidate biomarkers and potential cell fate-determining regulatory genes by analyzing the transcriptome characteristics of hematopoietic stem cells. Compared with the traditional single-cell analysis model based on functional classification, it effectively described their developmental biological characteristics. The work published by Dr. Fernando D. Camargo and Dr. Sahand Hormoz in the 2020 journal Cell used CRISPR-Cas9 gene editing technology to trace the lineage of the CARLIN mouse strain during the system development process, thereby analyzing the heterogeneity of cell populations. In addition, Dr. Samantha A. Morris's team published an article in Nature magazine in 2018, showing that they had successfully used direct lineage reprogramming technology to induce fibroblasts into endoderm precursor cells, and identified two different states of differentiation during the fibroblast induction process: the "reprogramming" state and the "dead end" state.

[0005] The limitations of traditional single-cell sequencing technology in liver cancer research lead to the following technical problems: the inability to dynamically track the evolution of tumor heterogeneity. Traditional single-cell sequencing can only statically analyze cell heterogeneity at a certain point in time and cannot capture the clonal evolution of liver cancer cells in the temporal dimension (such as tumor stem cell differentiation and dedifferentiation) and the rules of clonal expansion.

[0006] Single-cell sequencing technology has difficulty in accurately identifying stage-specific cancer-promoting genes, and static analysis cannot distinguish specific gene targets that drive different stages of liver cancer progression (such as the cloning initiation stage and the proliferation and spread stage), resulting in low target screening accuracy.

[0007] Single-cell sequencing technology lacks cross-platform data integration verification, and the separation of single-cell data from clinical databases (such as TCGA) leads to insufficient verification of the clinical significance of target genes. Summary of the Invention

[0008] In view of this, one of the purposes of the present invention is to provide a method for mining tumor progression molecules based on single-cell lineage tracing technology, and the second purpose is to apply the method for mining tumor progression molecules based on single-cell lineage tracing technology in detecting tumor development markers.

[0009] In order to achieve the above object, the present invention provides the following technical solutions:

[0010] The present invention provides a method for mining molecular targets for liver tumorigenesis based on single-cell lineage tracing technology, comprising the following steps:

[0011] S1: Using barcoded lentiviral vectors to label tumor cells and construct tumor cells with unique identity markers;

[0012] S2: Using the tumor cells labeled in step S1 to prepare an orthotopic tumor model in mice, and collecting tumor cells at different stages of tumor development;

[0013] S3: Perform single-cell transcriptome sequencing on the tumor cells in step S2 and implement cell lineage tracing through barcoding;

[0014] S4: Based on the cell lineage map in step S3, dimensionality reduction and pseudo-time series analysis methods are used to identify key biomarker genes at different stages;

[0015] Preferably, the barcode is a random sequence of 8 bases constructed from the pSMAL-CellTag-V1 plasmid;

[0016] Preferably, the vector further comprises a fluorescent protein tag;

[0017] Preferably, the tumor cells are H22 mouse liver cancer cell line;

[0018] Preferably, in step S3, the sequencing data is subjected to dimensionality reduction and quality control operations using R packages Scater and Seurat;

[0019] Preferably, the cell lineages are subjected to pseudo-time sequence analysis based on Monocle software to identify differentially expressed genes;

[0020] Preferably, pathway enrichment analysis of differentially expressed genes is performed based on GO and KEGG databases;

[0021] Furthermore, the application of methods based on single-cell lineage tracing technology to mine tumorigenesis molecular targets in the detection of tumor development markers

[0022] The beneficial effects of the present invention are:

[0023] (1) Dynamically tracking the evolution of tumor heterogeneity

[0024] Through single-cell lineage tracing technology (based on barcode labeling and lentiviral transfection), the limitations of traditional single-cell sequencing, which can only statically analyze heterogeneity at a specific time point, have been overcome. It is possible to track the dynamic changes in the clonal lineage of liver cancer cells from the early stage of liver cancer formation (T2) to the advanced stage (T3), revealing the continuous process of cell fate differentiation. The evolutionary relationship between clonal subpopulations (such as clonet1t2 and clonet1) was identified, and the cell populations involved in tumor formation (persistent clones) and those not involved (disappearing clones) were clearly distinguished, providing precise targets for screening driver genes.

[0025] (2) High-precision target screening

[0026] At T1, 877 cells with unique barcodes were screened to avoid confusion during subsequent tracking and ensure the accuracy of clonal origin. In the early stages (T1-T2), dimensionality reduction pseudo-time analysis was used to identify transiently upregulated genes near branch points (e.g., the top 10 differentially expressed genes). In the late stages (T2-T3), key molecules (e.g., Nt5c and Dnajb1) were identified through differential expression between proliferating and non-proliferating cells. Specific molecular targets were identified at different stages of tumor progression, enhancing the biological significance and clinical relevance of the targets.

[0027] (3) Multi-level functional verification and mechanism analysis

[0028] Using monocle pseudo-time trajectories to reveal dynamic pathway changes during tumor progression (e.g., early transcriptional activity → mid-stage proliferation → late metabolic activity), we elucidate the functional context of molecular targets. Survival analysis and gene co-expression analysis, combined with TCGA and GEO databases, validate the prognostic value and clinical significance of target genes (e.g., Krt18). Target validation across multiple dimensions, from molecular function and pathway regulation to clinical relevance, enhances the reliability of results.

[0029] (4) Technological complementarity

[0030] Key clones are located through lineage mapping, and temporal changes in gene expression within clones are analyzed through pseudo-time analysis. Single-cell data reveal subpopulation-specific targets, while bulk data (such as survival analysis) validate their pan-cancer applicability. This approach balances cellular heterogeneity with population patterns, avoiding bias from a single technique and enhancing comprehensiveness of conclusions.

[0031] (5) Clinical translation potential

[0032] Discover cancer-promoting genes (such as proliferation-related Nt5c) and potential therapeutic targets (such as dedifferentiation-driving genes), and confirm their association with patient prognosis through survival analysis. This provides candidate molecules for early diagnostic markers and stage-specific treatment strategies for cancer, supporting the application of precision medicine.

[0033] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0035] Figure 1 This is the GFP positive ratio detection graph after H22 infection at MOI=7.48;

[0036] Figure 2 Design a flow chart for liver cancer cell tracing technology experiments;

[0037] Figure 3 The source diagram of single-cell sequencing data at three moments;

[0038] Figure 4 Design a flow chart for the overall experiment;

[0039] Figure 5 This is the tSNE image of the single cell dimension reduction at time T1;

[0040] Figure 6 The expression of cell subpopulation marker genes after dimensionality reduction at time T1;

[0041] Figure 7 This is the tSNE image of single cell dimension reduction at time T2;

[0042] Figure 8 The expression of cell subpopulation marker genes after dimensionality reduction at time T2;

[0043] Figure 9 This is the tSNE image of single cell dimension reduction at time T3;

[0044] Figure 10 The expression of cell subpopulation marker genes after dimensionality reduction at time T3;

[0045] Figure 11 Enter the distribution map of liver cancer cell labels at T1;

[0046] Figure 12 Enter the distribution map for the liver cancer cell labels at T2;

[0047] Figure 13 Enter the distribution map of liver cancer cell labels at T3;

[0048] Figure 14A process for screening cells that can be analyzed at three moments based on barcode tracing;

[0049] Figure 15 It is a lineage diagram of cell clones without unique markers from T1 to T3;

[0050] Figure 16 It is a lineage diagram of cell clones that do not have unique markers from time T1 to T2;

[0051] Figure 17 It is a lineage diagram of cell clones with unique markers from T1 to T3;

[0052] Figure 18 It is a lineage diagram of cell clones with unique markers from T1 to T2;

[0053] Figure 19 This is a dimensionality reduction diagram of cells with unique identity tags related to all clones from time T1 to T2 before removing the time batch effect;

[0054] Figure 20 This is a dimensionality reduction diagram of cells with unique identity tags related to all clones from time T1 to T2 after removing the time batch effect;

[0055] Figure 21 This is the result of pseudo-time analysis of two types of clones composed of cells with unique identity markers from T1 to T2;

[0056] Figure 22 This is a time series diagram after pseudo-time analysis of two types of clones composed of cells with unique identity markers from T1 to T2;

[0057] Figure 23 The top 10 genes with transient up-regulated expression at pseudo-temporal branch point 2;

[0058] Figure 24 The pathway changes during liver cancer cell progression based on pseudo-time trajectory;

[0059] Figure 25 Tracking the division and proliferation of tumor cells from T2 to T3;

[0060] Figure 26 This is the differential expression between non-proliferating cells and proliferating cells at T3 time point relative to T2 time point. DETAILED DESCRIPTION

[0061] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0062] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0063] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0064] Example 1

[0065] The tracing library vector pSMAL-CellTag-V1 in this example was purchased from Addgene, catalog number 115643. pSMAL-CellTag-V1 incorporates an 8-bp random DNA tracing sequence into the 3' UTR of the GFP fluorescent marker protein in the pSMAL lentiviral vector. The vector contains 19,973 random tracing sequences. After lentiviral packaging and successful infection of target cells, the GFP and tracing sequence are integrated into the target cell genome and stably inherited by subsequent cells, achieving cell tracing.

[0066] This example uses the above tracing technology to trace mouse liver cancer cells, study the different evolution processes of tumor cells at different stages of liver cancer progression, and discover key genes that initiate the development of liver cancer. This example uses the mouse liver cancer cell line H22. After exploring infection conditions, it was finally determined that when the MOI (Multiplicity of infection, MOI) of 7-8 is selected, the GFP ratio is about 80% ( Figure 1 ).

[0067] After pSMAL-CellTag-V1 was packaged into a lentivirus, 30,000 H22 cells were infected with MOI=7. After 4 days of in vitro proliferation, the tumor cells that were successfully traced by GFP-positive were sorted, and a total of about 200,000 cells were obtained. Among them, 100,000 cells were sent for single-cell sequencing as the transcriptome sequencing data of H22 before in vivo tumor formation, that is, at the T1 moment. Another 100,000 cells were transplanted into the mouse liver in situ. The condition was that each mouse (Balb / c mouse, male, 8 weeks old) was injected with 10,000 cells in situ into the liver, and a total of 10 mice were injected. At two time points of 2 weeks and 4 weeks after the mouse tumor was formed, 5 mouse liver cancer tissues were taken from each mouse as liver cancer cells at T2 and T3 respectively. The liver cancer tissue was digested into single cells by enzymatic digestion, and 60,000 H22 tumor-derived GFP-positive liver cancer cells were sorted by flow cytometry for single-cell sequencing ( Figure 2 ).

[0068] The above three time points, namely T1, T2, and T3, represent the pre-tumor, early-tumor, and late-tumor stages of liver cancer cells, respectively. Post-sequential analysis was used to obtain the changes in the transcriptome of liver cancer evolution.

[0069] Example 2

[0070] Through the single-cell sequencing work in Example 1, transcriptome expression data and cell label expression data were obtained for 11,996 cells at time T1, 8,257 cells at time T2, and 11,015 cells at time T3. The R package Seurat was used to complete the two-step filtering work in sequence. In the first filtering step, genes expressed in at least 3 cells were first filtered out, and then all cells expressing less than 50 genes were filtered out. In the second filtering step, the proportion of mitochondrial gene expression was <= 25% and the proportion of erythrocyte gene expression was <1% as the criteria for screening, and all low-quality cells were initially discarded. Afterwards, the R package Scater was used for more advanced filtering, the main purpose of which was to use the total number of expressed genes in the cell and the total counts of all gene expressions to discard low-quality cells. Through the above filtering operations, a total of 10,169 cells at time T1, 6,309 cells at time T2, and 9,206 cells at time T3 were retained.

[0071] The R package Seurat was used to perform dimensionality reduction on the single-cell sequencing data after quality control. Using standard downstream analysis pipelines, principal component analysis (PCA) was performed on the normalized expression matrix of the highly variable genes identified by the FindVariableGenes function. Based on the results of the PCA, the top 15 PCs with a resolution parameter of 0.8 were selected for the T1 cluster. For the T2 cluster, the top 17 PCs with a resolution parameter of 1.3 were selected. For the T3 cluster, the top 17 PCs with a resolution parameter of 0.9 were selected.

[0072] By observing the cell label distribution of cells at three time points, it was found that the cell label combinations of many cells at time point T1 showed obvious duplication. Therefore, when using the same cell label combination to track these cells to time points T2 and T3, it is impossible to distinguish the specific cell at time point T1 from which they originated. Therefore, a special unique identity labeling mechanism was introduced: only cells that did not contain any cell label duplication at time point T1 were selected for backward tracking. This circumvents the "confusion" problem and lays the foundation for tracking later clonal lineage maps.

[0073] Example 3

[0074] Figure 5 This figure shows the tSNE effect of dimensionality reduction and clustering of all cells after data quality control at time T1. All cells can be divided into 12 subpopulations. As can be seen from the figure, all cells are clustered into a large subpopulation. Figure 6 It was shown that these cells all expressed the epithelial cell markers Epcam and Krt18 genes, indicating that all cells at the T1 moment were tumor epithelial cells.

[0075] Figure 7 The tSNE effect diagram of all cells after data quality control at time T2 is shown. All cells are divided into 13 subgroups. Figure 8 This shows that at T2, in addition to most tumor epithelial cells, there are also a small number of Lyz2+ myeloid cells, Ptprc+ immune cells and Pecam1+ endothelial cells. The presence of these cells may be due to a small number of cherry-negative cells mixed in during flow cytometry sorting. Therefore, these cells need to be eliminated and only the tumor epithelial cells required for analysis are retained.

[0076] Figure 9 The tSNE effect diagram of all cells after data quality control at time T3 is shown. All cells are divided into 13 subgroups. Figure 10 It can be found that at T3, in addition to the vast majority of tumor epithelial cells, there is also a small number of endothelial cells. Therefore, as with the above T2 moment, this small number of cells also needs to be removed.

[0077] Figure 11-13 The data show the distribution of cell labels in quality-controlled tumor cells after quality control and dimensionality reduction analysis from time points T1 to T3. The data show that barcodes have high coverage in cells at each time point. At the same time, in most cells at each time point, each barcode is not always present alone, but often appears in the form of barcode combinations.

[0078] In order to track the changes of each cell at different times and eliminate the possibility that different cells are barcoded with the same barcode, 877 cells with unique identity markers were screened from the cells at time T1.

[0079] like Figure 14 As shown, at time T1, we found 877 cells corresponding to 877 "unique" barcodes. "Unique" means that the barcode sequence in no two cells is exactly the same. The main purpose of finding "unique" cells at time T1 is to ensure that subsequent cell tracing can be traced back to a specific cell at time T1. At time T2, 101 barcodes with the same sequence as at time T1 were traced, corresponding to 170 cells, indicating that 170 cells at time T2 originated from time T1. Similarly, at time T3, a total of 49 identical barcodes were traced, corresponding to 114 cells at time T2. Therefore, these 49 barcodes represent the total number of barcodes traceable from time T1 to time T3. Tracing back to time T1, 49 barcodes correspond to 49 "unique" T1 cells that can be traced back to time T3. Tracing back to time T2 yields 100 cells that can be traced back to time T3. By adding the 49 cells at T1, the 100 cells at T2, and the 114 cells at T3, a total of 263 cells were obtained for analysis. Each of these cells could be tracked from T1 to T3, ensuring continuity and traceability for subsequent analysis.

[0080] The R packages Seurat and Scater were used to perform quality control on the three sequencing data sets. The goal was to remove the "noise" contained in the experimental process and eliminate the impact of the presence of cells with substandard quality and non-tumor cells. After completing the quality control, the cell labels contained in the cells after quality control were counted, and the cell label sequence of each cell with a cell label was obtained. Cells with unique identity markers were selected from the T1 moment and tracked to subsequent moments to obtain T2 and T3 cells that can be used to define cancer stem cells. Forward tracking also obtained T1 cells that can be used to define cancer stem cells. Together with T2 and T3 moments, they constitute the cell data range for subsequent analysis.

[0081] Example 4

[0082] In this embodiment, a clone is defined as follows: if any two cells contain the same cell label sequence, then the two cells belong to the same clone. Based on this concept, if the cells have a unique identity marker, then it can be determined which cell at time T1 in the clone the cells tracked at time T2 and T3 come from, because the cells at time T1 are unique. However, if the cells do not have a unique identity marker, then there is a question: the cells tracked at time T2 and T3 contain the same cell label sequence as the multiple cells at time T1, so how to determine which cell each of these cells comes from at time T1? In view of the existence and unavoidability of this problem, and the fact that it cannot be determined based solely on the cell label sequence, the total number of cells (regardless of whether they have a unique marker) is used as a verification set to verify the conclusions obtained from the analysis within the range of cells with a unique marker. This allows the reliability of the data conclusions to be retained while minimizing the errors caused by such noise.

[0083] Based on the drawn clonal lineage diagram, the concept of cloning based on the clonal lineage diagram was used to divide all cells between T1 and T2 into two categories: clonet1 and clonet1t2. Among them, clonet1 refers to all cells that appear at T1 but not at T2, representing cells that cannot form tumors; clonet1t2 refers to all cells that appear at T1 and also at T2, representing cells involved in tumor formation. Based on this classification concept, we can find a dividing point on the pseudo-time sequence where tumor cells transition from T1 to T2 by first reducing the dimension and then pseudo-time. The changes in gene expression levels before and after this time point are the focus of our attention and subsequent screening.

[0084] We used pseudo-temporal gene enrichment analysis in monocle2 to investigate pathway changes during early liver cancer cell progression (T1-T2). By setting num_clusters, we clustered the heatmap into three clusters. We then performed GO enrichment analysis on genes with significantly upregulated expression (p<0.05) within each cluster. We identified significantly upregulated biological processes at each stage and annotated them in the heatmap. Genes included in the enrichment analysis were also displayed in the heatmap.

[0085] Compared to T1 and T2, cells at T2 and T3 showed no significant differences, primarily as evidenced by the lack of clear distinction between clone2 and clone2t3 cells in the dimensionality reduction and pseudo-time results. Therefore, we did not employ traditional dimensionality reduction and pseudo-time strategies to identify gene targets that function at T3, the late stage of tumor progression. Instead, we focused on the characteristics exhibited by cells at T2 as they transitioned from T3 to T3: proliferation and division. As previously mentioned in the experimental preparation, we noted a significant increase in tumor volume at T3 compared to T2. Subsequent data analysis also revealed that the number of cells collected at T3 was greater than at T2, indicating that T3 represents a point in tumor progression relative to T2, with further expansion in volume and cell number. Therefore, we explored this characteristic to determine whether T2 tumor cells undergo proliferation after passage to T3, thereby identifying gene targets that promote tumor proliferation and division in the later stages of tumor progression.

[0086] Example 5

[0087] Figure 15 As shown in the figure, based on cell label lineage tracing, 146 common cell labels at T1, T2, and T3 were used to plot the clonal lineage evolution of the total cells from T1 to T3. The total number of clones calculated was 1,230, corresponding to 23,600 cells from T1 to T3. These included 10,135 cells at T1, 5,603 cells at T2, and 7,862 cells at T3.

[0088] Depend on Figure 16 As shown in the figure, based on cell label lineage tracing, 183 common cell labels at T1 and T2 were used to plot the clonal lineage evolution of the total cells from T1 to T2. Statistical calculations revealed a total of 1281 clones, corresponding to 15,490 cells from T1 to T2. These included 9,955 cells at T1 and 5,535 cells at T2.

[0089] like Figure 17As shown, first, based on cell label lineage tracking, after taking 146 common cell labels at T1, T2 and T3, a clonal lineage evolution diagram of cells with unique labels within the three time ranges from T1 to T3 was drawn. The 877 uniquely labeled cells at T1 can lead to the corresponding 877 clones. Among these clones, 101 clones contain cells from T2 and later, of which 52 clones can extend to T2, 49 clones can extend to T3, and the remaining clones only contain cells at T1. The total number of cells that can be tracked at the three time points is 1161. From this, it can be seen that most of the cells with unique labels at T1 cannot be transmitted to T2 and T3. In general, the proportion of cells at T2 and T3 that still preserve the cell lineage at T1 is relatively low.

[0090] like Figure 18 As shown in the figure, based on cell label lineage tracing, 183 common cell labels at T1 and T2 were taken to draw a clonal lineage evolution diagram of cells with unique labels from T1 to T2. This is a schematic diagram of the cell clonal lineage of 922 clones corresponding to 922 uniquely labeled cells at T1. The total number of cells that can be traced at the two moments is 1097, of which 820 clones contain only cells at T1 and 102 clones contain cells at T2. It can be concluded that, under the premise of focusing only on the two time points of T1 and T2, the passage efficiency of cells with unique labels shows roughly the same characteristics as the above three moments.

[0091] like Figure 19-20 As shown in the figure, within the range of time points T1 and T2, a total of 1097 uniquely labeled cells were included in the analysis, of which clone1 contained a total of 922 cells and clone1t2 contained a total of 175 cells. As shown in the figure, after using the R package Seurat to complete the dimensionality reduction operation, the umap graph shows that there is a batch effect between the cells at time points T1 and T2, so it is necessary to call the R package harmony to remove the influence of different time batches. Figure 20 This is the effect diagram after removing the batch effect. It can be seen that while the batch effect does not exist, the cells of clone1t2 have a significant clustering phenomenon, indicating that there is a strong and significant similarity between the cells within it.

[0092] like Figure 21-22As shown, the results of the monocle pseudo-time series analysis of the unique identity-marked cells based on the aforementioned dimensionality reduction are shown. It can be observed that the vast majority of clonet1t2 cells are located after the 2 branch points, while the vast majority of clonet1 cells are located before the 2 branch points. Therefore, this indicates that it is necessary to find genes that are significantly transiently upregulated around the 2 branch points. Such genes are considered to be of great significance to the initiation or early stage progression of tumor cells because their expression levels transiently increase as the clones progress on the pseudo-timeline.

[0093] By using the default differential analysis method in the R package monocle, a total of 50 genes with significantly upregulated expression (p < 0.05) after 2 branch points were obtained. Among them, the top 10 genes with the most significant differences were Figure 23 Presented in.

[0094] The cells at T1 and T2 were analyzed and screened for gene targets in the clonal lineage diagram composed of clonet1 and clonet1t2 using dimensionality reduction and pseudo-time methods. The monocle pseudo-time diagram can be used to analyze the pathway changes on the pseudo-time axis during tumor progression.

[0095] like Figure 24 As shown, the entire pseudo-time trajectory can be roughly divided into three stages along the time axis. In the first stage, RNA splicing, mRNA processing, and mRNA metabolic regulation are mainly upregulated, reflecting the upregulation of gene transcription levels at this stage. In the second stage, mitosis-related pathways are upregulated, mainly including chromatin separation, sister chromosome separation, and nuclear separation, indicating that the cells are in a state of continuous proliferation and division. In the third stage, energy synthesis and metabolism are the main focus, and the main upregulated biological processes include precursor metabolism and energy production, glycolysis, and ATP synthesis, indicating that the cells are currently in a relatively active metabolic state.

[0096] like Figures 25-26As shown, of the 129 uniquely identified cells at T2, 79 showed no proliferation at T3, while 50 cells had divided and proliferated to form 180 cells. Comparison of gene expression profiles between the two cell types at T2 and T3 revealed significant upregulation of Krt18, Selenom, Pgls, Gpx4, Eif3h, Rbm3, Eif3f, Atp5c1, Rtraf, Prdx4, Lars2, Nt5c, and Dnajb1. Nt5c and Dnajb1 were significantly upregulated in proliferating cells but not in non-proliferating cells, suggesting that their increased expression levels likely significantly promote the proliferation of liver cancer cells.

[0097] Based on cell label lineage tracing, clonal lineage evolution maps were plotted for cells with unique identity markers and for all cells at T1, T2, and T1, T2, and T3, respectively. Cells were then divided into clone 1 and clone 1t2 between T1 and T2. Dimensionality reduction pseudo-time analysis was then used to identify the branching point between these two clones. Based on changes in gene expression levels before and after the branching point, several potential gene targets that promote early tumor progression were identified. Furthermore, based on cell label tracing, a group of cells that proliferated from T2 to T3 and a group that did not proliferate from T2 to T3 were identified. Differential expression analysis was performed between these two groups of cells at the two time points, identifying potential gene targets that promote late cell proliferation and division. Finally, based on the dimensionality reduction pseudo-time map, signaling pathway changes during the progression of liver cancer cells from clone 1 to clone 1t2, i.e., the early stages of tumor progression, were investigated. The cells underwent a biological process from active transcription to active proliferation and division to active energy synthesis and metabolism. This shows that liver cancer cells have a very active life state in the early stages of tumor development.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for mining molecular targets of tumorigenesis based on single-cell lineage tracing technology, characterized in that: The following steps are involved: S1: Using barcoded lentiviral vectors to label tumor cells and construct tumor cells with unique identity markers; S2: Using the tumor cells labeled in step S1 to prepare an orthotopic tumor model in mice, and collecting tumor cells at different stages of tumor development; S3: Perform single-cell transcriptome sequencing on the tumor cells in step S2 and implement cell lineage tracing through barcoding; S4: Based on the cell lineage map in step S3, dimensionality reduction and pseudo-time series analysis methods are used to identify key biomarker genes at different stages.

2. The method for mining molecular targets of tumorigenesis based on single-cell lineage tracing technology according to claim 1, characterized in that: The barcode is a random sequence of 8 bases constructed from the pSMAL-CellTag-V1 plasmid.

3. The method for mining tumorigenesis molecular targets based on single-cell lineage tracing technology according to claim 1, characterized in that: The vector also includes a fluorescent protein tag.

4. The method for mining tumorigenesis molecular targets based on single-cell lineage tracing technology according to claim 1, characterized in that: The tumor cells are H22 mouse liver cancer cell line.

5. The method for mining tumorigenesis molecular targets based on single-cell lineage tracing technology according to claim 1, characterized in that: In step S3, the sequencing data were subjected to dimensionality reduction and quality control using the R packages Scater and Seurat.

6. The method for mining tumorigenesis molecular targets based on single-cell lineage tracing technology according to claim 1, characterized in that: Cell lineages were analyzed using pseudo-time series analysis using Monocle software to identify differentially expressed genes.

7. The method for mining tumorigenesis molecular targets based on single-cell lineage tracing technology according to claim 6, characterized in that: Pathway enrichment analysis of differentially expressed genes was performed based on the GO and KEGG databases.

8. Use of the method for mining tumorigenesis molecular targets based on single-cell lineage tracing technology according to any one of claims 1 to 7 in detecting tumor development markers.

Citation Information

Cited By

  • Single cell pedigree tracing method

    CN121610565A