A method of single cell lineage tracing
By combining single-cell transcriptome sequencing with bulk DNA targeted sequencing and tag barcode labeling, the problems of RNA instability and the small number of rare progeny cells were solved, achieving low-cost and high-accuracy single-cell lineage tracing.
Patent Information
- Application Number
- CN202610130458.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-07-31
- Estimated Expiration
- 2046-01-30
AI Technical Summary
Existing single-cell lineage tracing methods suffer from problems such as RNA instability and a small number of rare progeny cells, resulting in high sequencing costs and difficulty in reproducibility.
We employed single-cell transcriptome sequencing combined with Bulk DNA targeted sequencing, using barcode tags to label single cells and obtain polyA RNA and Bulk DNA sequencing data from daughter cells. We then used Cell Ranger and Python to process the data, perform hierarchical clustering and UMAP plotting, and trace cell lineages.
It reduces sequencing costs, improves the accuracy and reproducibility of single-cell labeling, and ensures the accuracy of lineage tracing.
Smart Images

Figure CN121610565B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of bioinformatics, and in particular, to a method for single-cell lineage tracing. Background Technology
[0002] The key to single-cell lineage tracing using mutant barcodes obtained through viral vector integration systems lies in establishing a barcode tracking and retrieval system. In hematopoietic stem cell lineage tracing, previous studies have used single-cell transcriptome sequencing of progeny cells to retrieve barcodes and confirmed single-cell markers by the repetition rate of barcodes retrieved from different samples (see Caleb Weinreb et al., Science, 2020, VOL. 367, NO. 6479). However, this method relies entirely on single-cell transcriptome sequencing, resulting in high costs. Other studies have reported achieving barcode retrieval through single-cell transcriptome sequencing combined with bulk RNA-targeted sequencing of progeny cells (see Merve Aksöz et al., Science Immunology, 2024, VOL. 9, NO. 98). This method reduces sequencing costs but still faces challenges such as RNA instability and a small number of rare progeny cells. Therefore, it remains necessary to provide a low-cost, simple, and easily reproducible single-cell lineage tracing method. Summary of the Invention
[0003] Technical problems to be solved:
[0004] The first aspect of this disclosure is to provide a new method for single-cell lineage tracing, addressing the problems of RNA instability and low number of rare progeny cells in existing single-cell lineage tracing methods.
[0005] Technical solution:
[0006] A method for single-cell lineage tracing includes the following steps:
[0007] Step (1) Mark the single cell using a label barcode;
[0008] Step (2) Introduce the single cell with the labeled barcode into the organism;
[0009] In step (3), after the single cell differentiates or proliferates in the organism, polyA RNA of the daughter cells of the single cell is obtained for transcriptome sequencing and library construction to obtain the first set of sequencing data, and the cell barcode corresponding to the daughter cell is obtained at the same time; Bulk DNA of the daughter cells of the single cell is obtained and sequencing is performed on the tag barcode sequence to obtain the second set of sequencing data; cDNA obtained in the transcriptome sequencing process is obtained and sequencing is performed on the cell barcode and tag barcode sequences to obtain the third set of sequencing data.
[0010] Step (4) analyzes the data obtained from sequencing, and the analysis methods include:
[0011] Step (a) The expression matrix of the first group of sequencing results is processed by Cell Ranger software, and the cell barcode file corresponding to the sample is obtained after data quality control.
[0012] Step (b) Read the third group of sequencing results and the cell barcode file line by line, extract the combination with cell barcode and label barcode and count the UMI number corresponding to the combination, and obtain the combination of sample number, cell barcode and label barcode after filtering;
[0013] Step (c) retains the cell barcode with unique tag barcode from the combination obtained in step (b), combines the retained cell barcode with the tag barcode, extracts the corresponding sequencing data from the second group of sequencing data, performs hierarchical clustering, names the categories, and adds the names as lineage preference tags for the cell barcodes to the first group of sequencing results. A UMAP map is drawn using a dimensionality reduction algorithm, and cells that have been tracked to the tag barcode are marked according to the added tags.
[0014] In some implementations, the aforementioned label barcode may be a label barcode based on natural mutations or a mutation barcode based on artificial markings.
[0015] Furthermore, in some embodiments, the aforementioned label barcode can be a manually labeled mutant barcode, wherein the manually labeled mutant barcode is a mutant barcode obtained using CRISPR barcode technology, viral vector integration, or the Cre-loxP recombinase system. In one embodiment, the aforementioned manually labeled mutant barcode can be barcode labeled using lentiviruses.
[0016] In some embodiments, a step (c0) is included between step (b) and step (c) to filter and normalize the second set of sequencing data to obtain the cell barcode and expression matrix for each mature lineage.
[0017] Furthermore, in some embodiments, the above step (c0) is as follows: using the FastP software to filter out low-quality sequences in the data obtained from sequencing, extracting and counting tag barcode sequences from the filtered data using a custom Python function; using the UMIcluster function in the umi_tools library to cluster and merge the tag barcodes, recounting the merged tag barcodes, and normalizing the counting results.
[0018] In some implementations, step (a) above is: using Cell Ranger software to compare and quantify the first group of sequencing results; using the scanpy library in Python to read and process the obtained data, filtering low-quality cell and gene data, and finally obtaining the cell barcode file corresponding to the sample.
[0019] In some implementations, step (b) above can be: reading the third group of sequencing results and the cell barcode file line by line using a custom Python function, extracting the combination of cell barcode and tag barcode and counting the number of times the UMI and the combination appear, filtering according to the number of times the combination appears, using the UMIcluster function in the umi_tools library to cluster and merge the tag barcodes, filtering again according to the UMI count, and finally obtaining the combination of sample number, cell barcode and tag barcode.
[0020] In some implementations, step (c) above may be: retaining cell barcodes with unique tag barcodes, and the tag barcodes do not appear in other combinations of multi-tag barcodes; extracting the corresponding sequencing data from the second set of sequencing data using the retained cell barcode and tag barcode combination, performing hierarchical clustering, naming each category according to its expression, and adding the naming as a lineage preference tag for the cell barcode to the data processed by the scanpy module of Python, drawing a UMAP map, and marking the cells traced to the tag barcodes according to the added tags.
[0021] In some embodiments, the single cell may be a stem cell, cancer cell, immune cell, tissue cell, cell from an organoid, or plant cell. Further, in some embodiments, the single cell may be a stem cell, which may be a pluripotent stem cell, a multipotent stem cell, or a multipotent stem cell. In one embodiment, the single cell is a hematopoietic stem cell (HSC).
[0022] In some embodiments, the hematopoietic stem cells mentioned above may be hematopoietic stem cells derived from umbilical cord blood, mobilized peripheral blood, or bone marrow.
[0023] A second aspect of this disclosure is to provide a data analysis apparatus for a single-cell lineage tracing method, the data analysis apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor;
[0024] When the computer program is executed by the processor, it implements the analysis steps in the above-described method for single-cell lineage tracing.
[0025] A third aspect of this disclosure is to provide a computer-readable storage medium storing a data analysis program for a single-cell lineage tracing method, wherein when the data analysis program for the single-cell lineage tracing method is executed by a processor, it implements the analysis steps included in the above-described single-cell lineage tracing method.
[0026] Beneficial effects:
[0027] In the technical solution disclosed herein, the cost of pure single-cell transcriptome sequencing is reduced by combining single-cell transcriptome sequencing with Bulk DNA targeted sequencing, and the accuracy of single-cell labeling is ensured by a subsequent filtering method that retains unique barcodes. Attached Figure Description
[0028] Figure 1 This is a flow cytometry sorting strategy and representative diagram of umbilical cord blood HSCs in the embodiments of this disclosure;
[0029] Figure 2 This is a schematic diagram illustrating the lentiviral infection efficiency of umbilical cord blood hematopoietic stem cells in an embodiment of this disclosure;
[0030] Figure 3 This is a diagram showing the sorting strategy and representative images of human erythrocytes in immunodeficient mice in this embodiment of the present disclosure.
[0031] Figure 4 This is a sorting strategy and representative diagram of human mature lineages (B, My, and Mk) in immunodeficient mice in the embodiments of this disclosure;
[0032] Figure 5 This is a sorting strategy and representative diagram of human HSPCs in immunodeficient mice in the embodiments of this disclosure;
[0033] Figure 6 This is a diagram showing the plasmid library capacity in an embodiment of this disclosure;
[0034] Figure 7 This is a unique barcode detection diagram for umbilical cord blood hematopoietic stem cell labels in this embodiment of the present disclosure;
[0035] Figure 8 This is a unique barcode detection image for the 293T cell line label in this embodiment of the present disclosure;
[0036] Figures 9A to 9G This is a schematic diagram illustrating the expression levels of characteristic genes in an embodiment of this disclosure, wherein, Figure 9A This diagram illustrates the expression levels of the CD34 and CD38 genes. Figure 9B This diagram illustrates the expression levels of the AVP and CRHBP genes. Figure 9C This diagram illustrates the expression levels of the HLF and SPINK2 genes. Figure 9D This diagram illustrates the expression levels of the GATA1 and KLF1 genes. Figure 9E This diagram illustrates the expression levels of CLC and HDC genes. Figure 9F This is a schematic diagram showing the expression levels of the MPO and AZU1 genes. Figure 9G A schematic diagram showing the expression levels of the IRF8 and CD79A genes;
[0037] Figure 10 This is a Leiden clustering diagram of human hematopoietic stem / progenitor cells in an embodiment of this disclosure;
[0038] Figure 11 This is a schematic diagram of hematopoietic stem cells for lineage fate classification in an embodiment of this disclosure. Detailed Implementation
[0039] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to specific embodiments.
[0040] This disclosure provides a method for single-cell lineage tracing. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired result. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. Furthermore, those skilled in the art can clearly modify or appropriately alter and combine the content described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.
[0041] In this disclosure, unless otherwise stated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprising of," etc., shall be understood to include the stated elements or components without excluding other elements or other components. The term "a," "an," and "the" includes plural indicators. The term "a plurality of" means two or more. The terms "such as," "for example," etc., are intended to refer to exemplary embodiments and are not intended to limit the scope of this disclosure.
[0042] In this disclosure, when a range of values is provided, it should be understood that, unless the context otherwise explicitly indicates otherwise, the range includes endpoints and each intermediate value between the upper and lower limits of the range, as well as any other specified value or intermediate value within the specified range and any value within a smaller range between specified values.
[0043] In this disclosure, the term "about" generally refers to a variation within a range of 0.5% to 10% above or below a specified value, such as a variation within a range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% above or below a specified value.
[0044] In this disclosure, terms such as "one embodiment," "an example," "some embodiments," "a particular embodiment," "related embodiment," "a certain embodiment," "some embodiments," "additional embodiment," or "further embodiment," "further implementation," or "another embodiment," "some other embodiments," mean that at least one feature or characteristic description is included in relation to the embodiment. Therefore, throughout this disclosure, the above phrases do not necessarily refer to the same embodiment. Furthermore, specific features may be combined in any suitable manner in one or more embodiments.
[0045] In this disclosure, unless otherwise stated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Definitions of common molecular biology terms can be found in *Lewin's GENES, Twelfth Edition*, Jocelyn E. Krebs, Elliott S. Goldstein, Stephen T. Kilpatrick, Jones & Bartlett Learning. Definitions of common biochemistry terms can be found in *Lehninger Principles of Biochemistry, Eighth Edition*, David L. Nelson, Michael M. Cox, WH Freeman. Definitions of common cell biology terms can be found in *Molecular Biology of the Cell, Sixth Edition*, Bruce Alberts, Alexander Johnson, Julian Lewis, David Morgan, Martin Raff, Keith Roberts, Peter Walter, Garland Science. Definitions of common genetics terms can be found in *Genetics: Analysis of Genes and Genomes, Eighth Edition*, Daniel L. Hartl, Maryellen Ruvolo, Jones & Bartlett Learning. Definitions and common methods of bioinformatics terms can be found in "Bioinformatics - Updated Features and Applications", Ibrokhim Y Abdurakhmonov, published by Intechopen.
[0046] Unless otherwise specified, the experimental techniques used in this paper employ standard techniques from immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which can be found in standard books such as *Molecular Cloning: A Laboratory Manual* and *Cell Biology: A Laboratory Handbook*.
[0047] definition:
[0048] The term "label barcode" in this disclosure is used interchangeably with "barcode," "unique molecular identifier (UMI)," "label," and "barcode." It generally refers to a label or identifier that conveys or is capable of conveying information about an analyte. A label barcode can be part of the analyte or independent of it. A barcode can be a tag attached to an analyte (e.g., a nucleic acid molecule) or a combination of a tag and endogenous characteristics of the analyte (e.g., the size or terminal sequence of the analyte). Label barcodes may be unique and can have many different formats. For example, label barcodes can include: polynucleotide barcodes, random nucleic acid and / or amino acid sequences, and synthetic nucleic acid and / or amino acid sequences. Label barcodes can be attached to the analyte in a reversible or irreversible manner. For example, a barcode can be added to a fragment of a deoxyribonucleic acid or ribonucleic acid sample before, during, and / or after sequencing. Label barcodes can identify and / or quantify individual sequencing reads.
[0049] In this disclosure, the term "single cell" refers to a single cell consisting of a nucleus, a membrane, and other cellular components.
[0050] The term "transcriptome sequencing" in this disclosure, commonly referred to as RNA-seq, is a technique that utilizes high-throughput sequencing technology to comprehensively and rapidly obtain information on almost all or specific types of RNA sequences and their abundance (expression level) in a specific sample. The typical method involves obtaining cDNA through reverse transcription, ligating universal adapters required for sequencing to both ends of the cDNA fragments to construct a library, and then sequencing the constructed library on a high-throughput sequencer to generate hundreds of millions of short-read sequences, followed by a series of bioinformatics analyses. Transcriptome sequencing can include various types; in some embodiments of this disclosure, the transcriptome sequencing is single-cell sequencing.
[0051] The term "descendant cell" in this disclosure refers to a cell derived from a parent cell (or mother cell). Descendant cells may have the same genetic characteristics as the mother cell (e.g., descendant cells obtained through mitosis) or different genetic characteristics (e.g., cells obtained through differentiation). In some embodiments, the descendant cell is derived from stem cells; for example, the mother cell is a hematopoietic stem cell, and its descendant cells may be B cells, T cells, megakaryocytes, etc.
[0052] The term "Bulk DNA" in this disclosure refers to genomic DNA extracted from a large group of cells (thousands or even millions) or an entire tissue sample. The final sequencing result of the Bulk DNA represents the average state of all the genomes in that group of cells.
[0053] The term "hematopoietic stem and progenitor cells" (HSPCs) in this disclosure refers to hematopoietic stem cells (HSCs) and / or hematopoietic progenitor cells (HPCs), a type of pluripotent stem cell found in bone marrow, peripheral blood, and umbilical cord blood. These cells possess self-renewal capacity and can differentiate into all types of mature blood cells (such as erythrocytes, leukocytes, and platelets) and immune cells, forming the core cell population that maintains the body's lifelong hematopoietic function. Through staged and multi-level proliferation and differentiation, they form progenitor cells of various lineages (such as myeloid progenitor cells and lymphoid progenitor cells), ultimately generating functionally specific blood cells. Hematopoietic progenitor cells, under the regulation of a specific microenvironment and certain factors, proliferate and differentiate into various types of blood cells. They are also a relatively primitive cell type with proliferative capacity, but have lost their multi-lineage differentiation ability and can only proliferate and differentiate into one or a few blood cell lineages in a directed manner; hence, they are also called committed stem cells. Human hematopoietic stem cells and hematopoietic stem and progenitor cells typically express the CD34 surface marker. In some embodiments, the hematopoietic stem cells may be derived from umbilical cord blood or mobilized peripheral blood. In some embodiments, the hematopoietic stem cells may be artificial hematopoietic stem cells or hematopoietic stem cells from other mammals, examples of which include, for example, mice, rabbits, dogs, monkeys, cattle, horses, sheep, etc.
[0054] The term "mobilized peripheral blood" in this disclosure generally refers to blood collected from the peripheral circulatory system of an individual (donor) after intervention through biological or pharmaceutical means. The most notable characteristic of this blood is that it contains a concentration of hematopoietic stem / progenitor cells far exceeding normal physiological levels.
[0055] Application of the method in this disclosure in single-cell lineage tracing technology for artificial hematopoietic stem cells:
[0056] Hematopoietic stem cell transplantation (HSCT) has been widely used and rapidly developed globally as a treatment option for various hematological and non-hematological diseases. In the field of hematopoietic stem cell (HSC) research, significant heterogeneity in self-renewal capacity and lineage differentiation tendency has been clearly observed within the HSC population. Specifically, these stem cells exhibit diverse response patterns due to different sources, in vitro culture conditions, and alterations in the microenvironment.
[0057] Studies have found that the transcriptome of hematopoietic stem / progenitor cells (HSPCs) exhibits significant proliferative characteristics in the early post-transplantation period. This is consistent with previous reports of increased numbers of MPPs and certain directed progenitor cells (such as CMPs and GMPs) in the HSPC population, particularly in the early post-transplantation phase. Mobilized peripheral blood transplantation, due to its higher proportion of HSPCs compared to bone marrow, leads to an increased number of myeloid clones engrafted early in patients, thus accelerating myeloid remodeling. Integration site analysis in non-human primate models revealed the early involvement of hematopoietic stem cells in the remodeling process after transplantation, and that these HSCs exhibit multi-lineage differentiation characteristics. Studies have found that stem cells that maintain hematopoiesis long-term after transplantation possess multi-lineage remodeling potential and are the dominant clones in vivo. Furthermore, studies have reported clonal conversion of hematopoietic stem cells in hematopoietic adverse states such as post-transplantation or acute platelet depletion. A research team led by Professor Peter J. Campbell at the University of Cambridge, through precise quantitative analysis, determined that the number of HSCs surviving long-term in HSCT recipients ranges from 700 to 25,000, with a higher number of HSCs engrafted from younger donors than from older donors. Another study on hematopoietic stem cell transplantation in mice has shown that the number of hematopoietic stem cells transplanted affects their regenerative kinetics after transplantation. Therefore, establishing a human HSC single-cell lineage tracing system can help explore the regenerative fate and molecular characteristics of HSCs after transplantation, thus providing an important theoretical basis for clinical hematopoietic stem cell transplantation.
[0058] Single-cell tracing systems mainly include lentiviral barcoding, the PolyExpress tracing system, CRISPR-based barcoding, mtDNA lineage tracing, and integration site analysis. Most of these tracing systems use mice as experimental subjects. However, Claus Nerlov's team used a lentiviral barcoding system to achieve lineage tracing of human bone marrow Lin-CD34+CD38-CD45RA-CD90+ HSCs after transplantation, verifying the existence of a platelet-biased HSC population in human bone marrow. However, information on the heterogeneity of HSCs from different sources, especially in umbilical cord blood and mobilized peripheral blood—two commonly used clinical grafts—as well as the lineage differentiation and molecular expression characteristics of HSCs after transplantation, remains lacking, limiting our current understanding of clinical graft HSCs. Developing single-cell lineage tracing technology for human hematopoietic stem cells, accurately tracking the differentiation pathway and post-transplantation regeneration fate of individual HSCs, thereby identifying functional HSCs involved in blood and immune reconstitution processes, and systematically studying their regeneration kinetics in different recipient microenvironments, is a key issue that urgently needs to be addressed in the field of clinical hematopoietic stem cell transplantation and regeneration.
[0059] In some implementations, the following steps may be included:
[0060] Step (1) Sorting human umbilical cord blood and mobilized peripheral blood hematopoietic stem cells for single-cell barcode labeling;
[0061] Step (2) Construct an immunodeficient mouse transplantation model using human umbilical cord blood and mobilized peripheral blood hematopoietic stem cells;
[0062] Step (3) involves the sorting and sequencing of labeled positive cells in the bone marrow of immunodeficient mice after transplantation;
[0063] Step (4) Analyze the sequencing data to determine the lineage fate and transcriptome characteristics of transplanted human umbilical cord blood and mobilized peripheral blood hematopoietic stem cells at the single-cell level.
[0064] The analysis includes:
[0065] Step (a) uses a dimensionality reduction algorithm for visualization;
[0066] Step (b) classifies human umbilical cord blood and mobilized peripheral blood hematopoietic stem cells by unsupervised clustering;
[0067] Step (c) classifies the fate of hematopoietic stem / progenitor cells using sequencing data from mature lineage cells.
[0068] Example 1: Identifying the diversity and size of lentivirus barcode libraries
[0069] (1) Dilute the plasmid library containing the tag barcode to 4 ng / μL, prepare the PCR reaction system according to Table 1, and centrifuge briefly after completion;
[0070] Table 1
[0071]
[0072] (2) Place the eight-chain arrangement in a PCR instrument and set the reaction system as follows: 98℃, 30s; 98℃, 10s, 69℃, 15s, 72℃, 20s, 11 cycles; 72℃, 2 min; hold at 4℃.
[0073] (3) After PCR, centrifuge the eight-row sample immediately;
[0074] (4) Using Monarch ® The PCR & DNA Cleanup Kit (5 μg) was used to purify the PCR products. The specific steps were as follows:
[0075] a) Add 5 times the volume of DNA binding buffer to the PCR amplification product into the EP tube and mix well;
[0076] b) Transfer the above mixture to a spin column and place it in a collection tube. Centrifuge at 16000g for 1 minute and discard the liquid in the collection tube.
[0077] c) Add 200 μL DNA Cleaning buffer to the spin column, centrifuge at 16000g for 1 minute, and discard the liquid in the collection tube;
[0078] d) Repeat the previous step;
[0079] e) Discard the collection tube and place the spin column into a new EP tube. Add 10 μL of enzyme-free water, incubate at 25°C for 5 min, and centrifuge at 16000g for 1 min to obtain the purified product.
[0080] (5) The purified product was subjected to PCR sequencing to detect the tag barcode.
[0081] Example 2: Sorting of human umbilical cord blood and mobilized peripheral blood hematopoietic stem cells
[0082] (1) Collect human umbilical cord blood and mobilized peripheral blood samples, enrich hematopoietic stem / progenitor cells with CD34 magnetic beads, and freeze them in a -80℃ freezer or liquid nitrogen;
[0083] (2) The frozen samples were thawed in a 37°C water bath, and 10 mL of PBE containing 10% FBS was added dropwise. The samples were centrifuged at 1500 rpm for 5 min to remove the cryopreservation solution.
[0084] (3) Discard the supernatant, resuspend the cells in 1 mL of PBE, and wash the original tube to make a total volume of 2 mL. Mix well, and take 10 μL of the cell suspension for counting;
[0085] (4) After centrifugation at 1500 rpm for 5 min, discard the supernatant, resuspend the cells in 100 μL, and repeat every 10 minutes. 7 Cells were labeled with antibodies as shown in Table 2 and stored at 4°C in the dark for 30 min.
[0086] (5) Add 2 mL of PBE buffer, centrifuge at 1500 rpm for 5 min at 4℃, and resuspend the cells in PBS buffer (1 mL / 1000 rpm) after centrifugation. 7 Cells were added to DAPI at a ratio of 1:1000 for sorting.
[0087] Table 2
[0088]
[0089] Hematopoietic stem cells with cell markers Lin-CD34+CD38-CD45RA-CD90+ were sorted using a sorting gate method as follows: Figure 1 As shown.
[0090] Example 3: The uniqueness of the tag barcode in different sample replicates was confirmed by infecting 293T cell lines and primary cells with a lentiviral library.
[0091] (1) Add 50 μg / mL Retronectin to each well of a 96-well plate and 500 μL of 50 μg / mL Retronectin to each well of a 6-well plate, and incubate overnight at 4°C;
[0092] (2) Regenectin was recovered, 400 g of hematopoietic stem cells from umbilical cord blood were sorted, centrifuged for 5 min, and resuspended in stem cell culture medium at a concentration of 25,000 / 100 μL. 100 μL / well was added to a 96-well plate. 293 T cells were resuspended in DMEM culture medium at a concentration of 750,000 / 1 mL. 1 mL / well was added to a 6-well plate and pretreated in a 37℃ 5% CO2 incubator for 24 hours.
[0093] (3) Add 10 mg / mL Polybrene 2 μL to 2 mL stem cell culture medium and DMEM culture medium respectively to prepare infection culture medium, discard half of the original culture medium, add 0.2 μL lentivirus to 50 μL infection culture medium to infect umbilical cord blood hematopoietic stem cells, and add 0.5 μL lentivirus to 500 μL infection culture medium to infect 293 T cells;
[0094] (4) After centrifugation at 33℃, 1800 rpm for 90 minutes, continue culturing in a 37℃, 5% CO2 incubator for 2.5-4.5 hours;
[0095] (5) Discard the culture medium in the wells, add 200 μL of stem cell culture medium to the 96-well plate at 37°C, add 2 mL of DMEM culture medium to the 6-well plate, and continue to incubate in a 5% CO2 incubator for 42 hours;
[0096] (6) After mixing by blowing and blowing several times in the well, all cells were aspirated. 293 T cells were resuspended in 1 mL of PBE buffer and 1 μL of DAPI was added. 1000 and 10000 cells / EP tube were sorted by flow cytometer, and 3 tubes were sorted in each case. Umbilical cord blood stem cells were mixed and directly divided into 2 EP tubes.
[0097] (7) Centrifuge all EP tubes at 4°C for 400 g for 5 minutes, and discard the supernatant;
[0098] (8) Add 20 mg / mL proteinase K to sterile enzyme-free water, dilute 1:40, add 25 μL of diluent to each tube, and shake at 55℃ and 800 rpm for 1 hour;
[0099] (9) Heat to 95℃ and 800 rpm for 10 minutes to inactivate proteinase K;
[0100] (10) After cooling on ice, centrifuge briefly and store at -20℃;
[0101] (11) Prepare the PCR reaction system according to Table 3, and centrifuge briefly after completion;
[0102] Table 3
[0103]
[0104] (12) Place the eight-chain arrangement in a PCR instrument and set the reaction system as follows: 98℃, 30s; 98℃, 10s; 69℃, 15s; 72℃, 20s, 32 cycles; 72℃, 2 min; hold at 4℃.
[0105] (13) After PCR, centrifuge the eight-row sample immediately;
[0106] (14) Using Monarch ® The PCR & DNA Cleanup Kit (5 μg) was used to purify the PCR products. The specific steps were as follows:
[0107] a) Add 5 times the volume of DNA binding buffer to the PCR amplification product into the EP tube and mix well;
[0108] b) Transfer the above mixture to a spin column and place it in a collection tube. Centrifuge at 16000 g for 1 minute and discard the liquid in the collection tube.
[0109] c) Add 200 μL DNA Cleaning buffer to the spin column, centrifuge at 16000g for 1 minute, and discard the liquid in the collection tube;
[0110] d) Repeat the previous step;
[0111] e) Discard the collection tube and place the spin column into a new EP tube. Add 10 μL of enzyme-free water, incubate at 25°C for 5 min, and centrifuge at 16000g for 1 min to obtain the purified product.
[0112] (15) The purified product was subjected to PCR sequencing to detect the tag barcode.
[0113] Example 4: Single-cell barcoding via lentiviral infection
[0114] (1) Add 40 μL of 50 μg / mL Retronectin to each well of a 96-well plate and incubate overnight at 4°C;
[0115] (2) Recover Retronectin, sort 400 g of hematopoietic stem cells, centrifuge for 5 min, add stem cell culture medium at a concentration of 70,000 / 100 μL, resuspend 100 μL / well, and pre-treat in a 37℃ 5%CO2 incubator for 24 hours.
[0116] (3) Add 10 mg / mL Polybrene 1 μL to 1 mL of stem cell culture medium to prepare infection culture medium;
[0117] (4) Discard 50 μL of stem cell culture medium, add 0.1 μL of lentivirus to 50 μL of infection culture medium to infect umbilical cord blood hematopoietic stem cells, and add 1 μL of lentivirus to 50 μL of infection culture medium to infect mobilized peripheral blood hematopoietic stem cells.
[0118] (5) After centrifugation at 33℃, 1800 rpm for 90 minutes, continue culturing in a 37℃, 5% CO2 incubator for 2.5-4.5 hours;
[0119] (6) Discard the culture medium in the well and add 200 μL of stem cell culture medium to continue culturing in a 37°C, 5% CO2 incubator.
[0120] Example 5: In vivo transplantation of HSCs labeled in Example 4 and detection of infection efficiency.
[0121] (1) Add 200 μL of fetal bovine serum to 10 mL of PBS buffer to prepare buffer;
[0122] (2) After changing the medium for 12 hours, blow the liquid in the 96-well plate to mix, aspirate it into a 1.5 mL EP tube, add another 200 μL of buffer to wash the well, and repeat twice;
[0123] (3) Centrifuge the EP tube at 4℃ for 400 g for 5 minutes and discard the supernatant;
[0124] (4) Add 1 mL of buffer and repeat the process in (3), then resuspend the cells in 350 μL of buffer;
[0125] (5) The resuspended cells were injected into NBSGW mice via the tail vein;
[0126] (6) At 4, 8, 12 and 16 weeks after mouse transplantation, cut off 3 mm of the tip of the mouse tail and collect 15-20 μL of tail blood into 100 μL LPBE buffer.
[0127] (7) Add 3 mL of red lysate, mix well, let stand for 12 minutes, centrifuge at 1500 rpm for 5 min and discard the supernatant;
[0128] (8) Resuspend the cells in 3 mL of PBE buffer, centrifuge at 1500 rpm for 5 min, discard the supernatant, and repeat every 10 minutes. 7 Cells were labeled with antibodies as shown in Table 4 and stored at 4°C in the dark for 30 min.
[0129] (9) Add 2 mL of PBE buffer, centrifuge at 1500 rpm for 5 min at 4℃, discard the supernatant after centrifugation, resuspend the cells in 100 μL of PBE buffer, add 1 μL of DAPI, and detect the infection efficiency using a flow cytometer. Figure 2 .
[0130] Table 4
[0131]
[0132] Example 6: Collection of progeny cells from the bone marrow of immunodeficient mice transplanted in Example 5
[0133] (1) Sixteen weeks after transplantation, the mice were euthanized by cervical dislocation;
[0134] (2) Take the humerus, ilium, femur and tibia of the mouse and grind them in a mortar with 2 mL of PBE buffer to obtain mouse bone marrow cells. Wash the mortar with 2 mL of PBE buffer.
[0135] (3) Take 400 μL of bone marrow from 4 mL of bone marrow cells and label it with antibodies as shown in Table 5. Protect from light at 4℃ for 30 min.
[0136] Table 5
[0137]
[0138] (4) The remaining bone marrow cells were centrifuged at 1500 rpm for 5 min, the supernatant was discarded, and 3 mL of lysing solution was added and lysed on ice for 10 min.
[0139] (5) After centrifuging at 1500 rpm for 5 min, discard the supernatant, add 2 mL of buffer to resuspend the cells, and wash the erythrocyte lysis solution;
[0140] (6) After centrifuging at 1500 rpm for 5 min, discard the supernatant, resuspend in 100 μL buffer, label with antibody as shown in Table 6, and incubate at 4℃ in the dark for 30 min.
[0141] Table 6
[0142]
[0143] (7) Add 200 μL of fetal bovine serum to 10 mL of DPBS buffer to prepare DPBS buffer;
[0144] (8) After the labeling time is complete, add 2 mL of DPBS buffer and centrifuge at 1500 rpm for 5 min at 4℃. After centrifugation, resuspend the cells in DPBS buffer (1 mL / 10⁻¹²). 7 Cells were added to DAPI at a ratio of 1:1000 and sorted using a sorting machine.
[0145] (9) 100-50,000 cells were sorted from each lineage. The cell types and their markers are as follows:
[0146] Ery: mCD45-hCD45-CD235a+CD71+;
[0147] My: mCD45-hCD45+CD34-CD33+;
[0148] B: mCD45-hCD45+CD34-CD19+;
[0149] Mk: mCD45-hCD45+CD34-CD33-CD19-CD41+;
[0150] HSPC: mCD45-hCD45+CD34+;
[0151] The sorting method for Ery is as follows: Figure 3 The selection method for B, My, and Mk is as follows: Figure 4 The gate method of HSPC sorting is as follows Figure 5 Except for Mk, which was directly sorted due to its small cell count, HSPC and all other mature lineages were sorted for mCherry+ cells.
[0152] Example 7: Library construction and sequencing of cells sorted in Example 6
[0153] (1) The sorted HSPC cells were subjected to 10X Genomics single-cell transcriptome sequencing;
[0154] (2) The sorted mature lineage cells (Ery, My, B, Mk) were centrifuged at 400g for 5 minutes at 4℃, and the supernatant was discarded;
[0155] (3) Add 20 mg / mL proteinase K to sterile enzyme-free water, dilute 1:40, add 25 μL of diluent to each tube, and shake at 55℃ and 800 rpm for 1 hour;
[0156] (4) Heat to 95℃ and 800 rpm for 10 minutes to inactivate proteinase K;
[0157] (5) After cooling on ice, centrifuge briefly and store at -20℃;
[0158] (6) Prepare the PCR reaction system according to Table 7, and centrifuge briefly after completion;
[0159] Table 7
[0160]
[0161] (7) Place the eight-chain arrangement in a PCR instrument and set the reaction system as follows: 98℃, 30s; 98℃, 10s; 69℃, 15s; 72℃, 20s, 32 cycles; 72℃, 2 min; hold at 4℃.
[0162] (8) Centrifuge the eight-row sample immediately after PCR;
[0163] (9) Using Monarch ® The PCR & DNA Cleanup Kit (5 μg) was used to purify the PCR products. The specific steps were as follows:
[0164] f) Add 5 times the volume of DNA binding buffer to the EP tube containing the PCR amplification product and mix well;
[0165] g) Transfer the above mixture to a spin column and place it in a collection tube. Centrifuge at 16000 g for 1 minute and discard the liquid in the collection tube.
[0166] h) Add 200 μL DNA Cleaning buffer to the spin column, centrifuge at 16000 g for 1 minute, and discard the liquid in the collection tube;
[0167] i) Repeat the previous step;
[0168] j) Discard the collection tube and place the spin column into a new EP tube. Add 10 μL of enzyme-free water, incubate at 25°C for 5 min, and centrifuge at 16000g for 1 min to obtain the purified product.
[0169] (10) The purified product was subjected to PCR sequencing to detect the tag barcode.
[0170] Example 8: Nested PCR amplification and library construction of cDNA from single-cell sequencing using targeted tag barcoding.
[0171] (1) The remaining cDNA from single-cell sequencing was stored in eight-tube arrays at a concentration of 10 ng / 3 μL, 4 μL / tube;
[0172] (2) Prepare the first round of PCR reaction system according to Table 8, and enrich separately for the label barcode;
[0173] Table 8
[0174]
[0175] (3) After the PCR system is prepared, centrifuge eight units at a time and place them in the PCR instrument to run the following program: 98℃, 30 s; 98℃, 10 s, 60℃, 15 s, 72℃, 20 s, 16 cycles; 72℃, 2 min; 4℃ maintenance.
[0176] (4) After the program finishes running, remove the eight-row stacks and centrifuge them instantly;
[0177] (5) Purify the first-round PCR product using Monarch. ® The PCR & DNA Cleanup Kit (5 μg) was used to purify the PCR products. The specific steps were as follows:
[0178] a) Add 5 times the volume of DNA binding buffer to the PCR amplification product into the EP tube and mix well;
[0179] b) Transfer the above mixture to a spin column and place it in a collection tube. Centrifuge at 16000 g for 1 minute and discard the liquid in the collection tube.
[0180] c) Add 200 μL DNA Cleaning buffer to the spin column, centrifuge at 16000 g for 1 minute, and discard the liquid in the collection tube;
[0181] d) Repeat the previous step;
[0182] e) Discard the collection tube and place the spin column into a new EP tube. Add 21 μL of enzyme-free water, incubate at 25°C for 5 min, and centrifuge at 16000 g for 1 min to obtain the purified product.
[0183] (6) Prepare the second round of PCR reaction system according to Table 9 and construct the library;
[0184] Table 9
[0185]
[0186] (7) After the PCR system is prepared, centrifuge eight units at a time and place them in the PCR instrument to run the following program: 98℃, 30 s; 98℃, 10 s, 60℃, 15 s, 72℃, 20 s, 8 cycles; 72℃, 2 min; 4℃ maintenance.
[0187] (8) After the program finishes running, remove the eight-row stacks and centrifuge them instantly;
[0188] (9) Purify the second-round PCR product using Monarch. ®The PCR & DNA Cleanup Kit (5 μg) was used to purify the PCR products. The specific steps were as follows:
[0189] a) Add 5 times the volume of DNA binding buffer to the PCR amplification product into the EP tube and mix well;
[0190] b) Transfer the above mixture to a spin column and place it in a collection tube. Centrifuge at 16000 g for 1 minute and discard the liquid in the collection tube.
[0191] c) Add 200 μL DNA Cleaning buffer to the spin column, centrifuge at 16000 g for 1 minute, and discard the liquid in the collection tube;
[0192] d) Repeat the previous step;
[0193] e) Discard the collection tube and place the spin column into a new EP tube. Add 10 μL of enzyme-free water, incubate at 25°C for 5 min, and centrifuge at 16000g for 1 min to obtain the purified product.
[0194] (10) The purified product was sequenced using a self-built library with base imbalance to detect the tag barcode.
[0195] Example 9: Bioinformatics analysis of the sequencing results of the above PCR products
[0196] (1) Data analysis of targeted tag barcode PCR amplification, library construction, and sequencing of plasmid DNA and cell-extracted DNA: FastP software was used to remove adapter sequences from R1 and R2 files, filter low-quality reads, and merge the corresponding reads in R1 and R2 after filtering to generate a merged single-end reads file. The merged reads file was read, and the reads were cut using the fixed sequences at both ends of the tag barcodes to obtain tag barcode sequences and counted. The UMIcluster function in the umi_tools library was used to cluster and merge the tag barcodes, using an edit distance greater than 4 as the clustering parameter. The tag barcode with the largest count in each category was selected as the true tag barcode, and the total count of the tag barcodes in that category was used as the tag barcode count. The total number of reads for each sample was adjusted to one million to normalize the data from different samples. A Venn diagram was plotted on the number of tag barcodes extracted from the plasmid sequencing results. Figure 6 The plasmid library used in the system has a capacity of approximately 200,000 entries, allowing for the labeling of up to 20,000 hematopoietic stem cells. The detection of barcode diversity for hematopoietic stem cell labels is as follows: Figure 7The data indicates that 1,000 label barcodes were detected for each sample, with approximately 50 label barcodes being duplicated, a duplication rate of 2.56%. The uniqueness detection and duplication rate of label barcodes for the 293T cell line are as follows: Figure 8 Based on the above data, it can be reasonably inferred that the label barcode markings have a single-cell level of accuracy.
[0197] (2) Analysis of single-cell transcriptome sequencing data: Cell Ranger software was used to process the single-cell RNA sequencing data to obtain gene expression matrix files, cell barcode files, and gene annotation files. The “sc.read_10x_mtx” function of the scanpy library was used to convert the output file of Cell Ranger into an AnnData object, and cells with mitochondrial genes and Globin genes with a UMI number greater than 10% were filtered out. Cells with a sequencing gene number between 500 and 8000 and a UMI number between 1000 and 60000 were retained. Double cells were removed using “scr.Scrublet” for data cleaning. The dataset was selected using `sc.pp.highly_variable_genes`. Gene expression data were standardized and logarithmically transformed using `sc.pp.normalize_total` and `sc.pp.log1p`. Cell cycle scoring was performed using `sc.tl.score_genes_cell_cycle` to preserve the characteristics of highly variable genes. Data normalization was performed using `sc.pp.regress_out` and `sc.pp.scale`. Principal component analysis (PCA) was used to reduce the dimensionality of the dataset using `sc.tl.pca`, and the variance explained ratio of the PCA was plotted using `sc.pl.pca_variance_ratio`. The dataset was then corrected and integrated using the function `sc.external.pp.harmony_integrate` to integrate single-cell transcriptome data from different batches. The adjacency graph between different cells was constructed using "sc.pp.neighbors" based on the PCA dimensionality reduction results. UMAP dimensionality reduction analysis based on the adjacency graph was performed using "sc.tl.umap", and the expression of characteristic genes could be characterized on the UMAP graph. Figures 9A to 9G Using "sc.tl.leiden" for adjacency graph-based Leiden clustering analysis, such as... Figure 10 It can be seen that the cells in Cluster 5 highly express AVP, CRHBP and HLF genes, which is a group of cells with strong stemness in HSPC.
[0198] (3) Sequencing data analysis after PCR amplification of residual cDNA from single-cell transcriptome sequencing using individual target tag barcodes: Read the target tag barcode sequencing data and the cell barcode file processed by Cell Ranger from the single-cell transcriptome data line by line. Extract the tag barcode from the fixed sequences before and after the tag barcode in the R1 file. Extract the sequencing sequence after the adapter sequence in the corresponding R2 file, which is the cell barcode and UMI. Combine the cell barcode and tag barcode and count the UMI and the number of times the combination occurs. If multiple tag barcodes correspond to the cell barcode in a cell, connect all tag barcodes with "-". Plot the frequency distribution according to the number of times the combination occurs and use inflection point filtering. Use the UMIcluster function in the umi_tools library to cluster and merge the remaining cell barcode and tag barcode combinations. Filter the barcode combinations with UMI < 2 again according to the UMI count to finally obtain the sample, cell barcode, and tag barcode combination.
[0199] (4) Integration of HSPC single-cell sequencing data with mature lineage PCR product sequencing data: The data was screened, and cell barcodes with unique tag barcodes were retained, and these tag barcodes did not appear in other combinations of multi-tag barcodes. The retained cell barcode and tag barcode combination data frame was merged with the mature lineage cell tag barcode data frame obtained in Example 9 (1) according to the tag barcode column, and only the tag barcodes that existed in both were retained. Based on the normalized reads of the tag barcode in the mature lineage, the "linkage" function and the "dendrogram" function were used to realize hierarchical clustering and generate the leaf node order of the dendrogram. Then, the "cut_tree" function was used to cut the dendrogram into cluster categories, named according to the expression of each category, and the name was added to the AnnData object as a tag for the cell barcode. The "sc.pl.umap" function was used to draw the UMAP diagram and the cells that were tracked to the tag barcode were marked according to the added tags. Figure 11 .
[0200] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method of single cell lineage tracing, characterized in that, Includes the following steps: Step (1) Mark the single cells using lentiviral barcodes; Step (2) introduces the single cell with the lentivirus barcode into the organism; In step (3), after the single cell differentiates or proliferates in the organism, polyA RNA of the daughter cells of the single cell is obtained for transcriptome sequencing and library construction to obtain the first set of sequencing data, and the cell barcode corresponding to the daughter cell is obtained at the same time; Bulk DNA of the daughter cells of the single cell is obtained and sequenced for lentivirus barcode sequence to obtain the second set of sequencing data; cDNA obtained in the transcriptome sequencing process is obtained and sequenced for cell barcode and lentivirus barcode sequence to obtain the third set of sequencing data. Step (4) analyzes the data obtained from sequencing, and the analysis methods include: Step (a) The expression matrix of the first group of sequencing data is processed by Cell Ranger software, and the cell barcode file corresponding to the sample is obtained after data quality control. Step (b) Read the third set of sequencing data and the cell barcode file line by line, extract the combination with cell barcode and lentivirus barcode and count the UMI number corresponding to the combination, and obtain the combination of sample number, cell barcode and lentivirus barcode after filtering; Step (c) retains the cell barcode with a unique lentivirus barcode from the combination obtained in step (b), and this lentivirus barcode does not appear in other combinations of multiple lentivirus barcodes. The retained cell barcode and the lentivirus barcode are combined and the corresponding sequencing data is extracted from the second set of sequencing data for hierarchical clustering and the category is named. The name is added to the first set of sequencing data as a lineage preference label for the cell barcode. A UMAP map is drawn using a dimensionality reduction algorithm and the cells that have been tracked to the lentivirus barcode are marked according to the added label. Between steps (b) and (c), there is also a step (c0) to filter and normalize the second set of sequencing data to obtain the tag barcode and expression matrix of each mature lineage.
2. The method of claim 1, wherein, The step (c0) is as follows: use the FastP software to filter out low-quality sequences in the second set of sequencing data, extract lentivirus barcode sequences from the filtered data and count them using a custom Python function; use the UMIcluster function in the umi_tools library to cluster and merge the lentivirus barcodes, recount the merged lentivirus barcodes, and normalize the count results.
3. The method of claim 1, wherein, Step (a) involves: using Cell Ranger software to align and quantify the first set of sequencing data; using the scanpy library in Python to read and process the obtained data, filtering out low-quality cell and gene data, and finally obtaining the cell barcode file corresponding to the sample.
4. The method of claim 1, wherein, Step (b) is as follows: The third set of sequencing data and the cell barcode file are read line by line using a custom Python function. The combination of cell barcode and lentivirus barcode is extracted and the number of times the UMI and the combination appear is counted. After filtering according to the number of times the combination appears, the lentivirus barcode is clustered and merged using the UMIcluster function in the umi_tools library. After filtering again according to the UMI count, the combination of sample number, cell barcode and lentivirus barcode is finally obtained.
5. The method of claim 1, wherein, Step (c) is as follows: Retain cell barcodes with unique lentiviral barcodes, provided that these barcodes do not appear in other combinations of multiple lentiviral barcodes; extract corresponding sequencing data from the second set of sequencing data using the retained cell barcode and lentiviral barcode combinations, perform hierarchical clustering, name each category according to its expression, and use this name as the lineage preference label for the cell barcode; standardize, detect highly variable genes, and perform PCA dimensionality reduction on the data filtered in step (a); after data integration using the Harmony algorithm to remove batch effects, perform UMAP dimensionality reduction on the data and cluster using the Leiden algorithm; add the lineage preference label to the data processed by the scanpy module in Python, draw a UMAP map, and mark the cells traced to the lentiviral barcodes according to the added labels.
6. The method according to any one of claims 1 to 5, characterized in that, The single cell is a stem cell.
7. The method of claim 6, wherein, The stem cells mentioned are hematopoietic stem cells.
8. The method of claim 7, wherein, The hematopoietic stem cells are derived from umbilical cord blood or from mobilized peripheral blood.
9. A data analysis device for single cell lineage tracing method, characterized in that, The data analysis device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor; When the computer program is executed by the processor, it implements the analysis steps in the method for single-cell lineage tracing as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a data analysis program for a single-cell lineage tracing method. When the data analysis program for the single-cell lineage tracing method is executed by a processor, it implements the analysis steps of the single-cell lineage tracing method as described in any one of claims 1 to 8.