A method for distinguishing the genetic material and cell species origin of xenogeneic hybrid samples in chimeric animal models based on single-cell sequencing

Multiple alignment and separation threshold calculations were performed after single-cell sequencing, which solved the problem of distinguishing between host and chimeric cells in chimeric animal models, and achieved accurate distinction between separation without relying on sequencing, reducing cell damage and cost, and expanding the scope of application.

CN119049567BActive Publication Date: 2025-07-22SHENGYUAN LIFE TECHNOLOGY (ZHEJIANG ANJI) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411081163.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-07-22
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

Prior art It is difficult to accurately distinguish between the genetic material and cellular species sources of host and chimeric cells in chimeric animal models without relying on pre-sequencing cell isolation, and traditional methods may lead to altered cell biological characteristics or loss of transcriptome information.

Method used

After single-cell sequencing, the single-cell transcriptome was multi-compared, and the separation threshold was calculated using the multiple-comparison algorithm. The ratio of genetic material between the host and chimeric cells was tested based on statistical distribution, so as to achieve a distinction method that does not rely on cell separation before sequencing.

Benefits of technology

It realizes accurate distinction between host and chimeric cells without changing the biological characteristics of the cells, reducing cell damage and cost, expanding the scope of application, and retaining the native state of the cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119049567B_ABST
    Figure CN119049567B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for distinguishing the genetic material and cell species origin of xenogeneic hybrid samples in chimeric animal models based on single-cell sequencing. This method is based on single-cell sequencing technology, does not rely on cell separation before sequencing, and uses multiple genome alignments to separate multi-species mixed samples. The computational method for separating multi-species mixtures based on multiple genome alignments includes: performing statistical distribution modeling based on the proportion of genetic material in the results of multiple unsupervised alignments of the sample to the reference genome of the target species; predicting the probability of sample origin based on the statistical distribution modeling of the proportion of genetic material. The method for separating multi-species samples based on multiple genome alignments established by the present invention can accurately determine the origin of the genetic material of different species in xenogeneic hybrid biological samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedicine. Specifically, it is a method for distinguishing the genetic material and the cell species origin of xenogeneic hybrid samples in chimeric animal models based on single-cell sequencing. Background Art

[0002] Chimeric animal models are a biotechnology widely used in biomedical research. Generally, they are constructed by transplanting chimeric species cells into model organisms, such as transplanting human stem cells into mice to construct chimeric animal models. This technology is crucial for biomedical research. Due to reasons such as scientific research ethics and cost control, it is difficult to directly obtain data on the effects of drugs or intervention therapies on humans, while in vitro cell experiments are prone to deviations due to the lack of other biological elements in the human body. Therefore, by constructing chimeric animal models of human cells and model organisms, it can better provide a platform to explore the effects of drugs or intervention therapies on human cells and living organisms.

[0003] Single-cell sequencing technology can obtain high-throughput sequencing data at the cell level resolution and map transcriptome features to individual cells. Traditionally, when studying chimeric animal models and performing single-cell sequencing on them, the obtained single-cell sequencing data cannot distinguish between chimeric cells and host cells. A common feasible method is to separate cells based on their biological characteristics, such as surface markers, cytological characteristics, and cell density, before sending the samples for single-cell sequencing, and then perform single-cell sequencing on the purified and separated cells. However, on the one hand, this method cannot completely and cleanly separate host and chimeric cells; on the other hand, it is easy to change the original biological characteristics of cells during the separation process, or cause cell death and fragmentation, which may lead to misjudgment of cell transcriptional information during subsequent sequencing, or even directly lose a large amount of transcriptome information.

[0004] Therefore, it is of great significance to develop a transcriptome separation method that can separate the transcriptomes of host and chimeric cell sources by analyzing the single-cell transcriptome alignment features after sequencing without relying on cell separation before sequencing, which can effectively promote the biological research on the existing single-cell sequencing data of chimeric animal models and the understanding of clinical intervention experiments. Summary of the Invention

[0005] In view of the defects existing in the prior art, the present invention provides a single-cell transcriptome analysis method that does not rely on cell separation before sequencing, but only distinguishes the genetic material and the cell species origin by analyzing the single-cell transcriptome alignment features in the single-cell sequencing results. At the same time, the implementation details of the program based on the multiple alignment algorithm, and the reasonable modifications in key data processing that still conform to the logic of this algorithm due to migration to different species, all fall within the scope of the technical solution of the present invention to be protected.

[0006] The present invention is realized through the following technical solutions:

[0007] The present invention discloses a method for distinguishing the genetic material and the cell species origin of xenogeneic hybrid samples in a chimeric animal model based on single-cell sequencing, specifically including the following steps:

[0008] 1) Obtain tissue samples from unlabeled chimeric animal samples;

[0009] 2) Obtain a cell suspension of mixed species of the tissue samples, without additional labeling, separation, and precipitation treatment, and directly perform single-cell sequencing;

[0010] 3) Obtain the single-cell transcriptome of the chimeric tissue;

[0011] 4) Perform multiple alignments on the single-cell transcriptome of the chimeric tissue;

[0012] 5) Calculate the separation threshold for multiple alignments;

[0013] 6) Mark and separate the host and chimeric cells according to the threshold.

[0014] As a further improvement, the method for distinguishing the genetic material and the cell species origin of the present invention does not rely on cell separation of the chimeric biological model samples before sequencing submission.

[0015] As a further improvement, the multiple alignment algorithm for sequence reads of the present invention does not rely on specific sequences in a specific species, but is based on the statistical distribution test of the cell origin in the alignment results of the host and chimeric species reference genomes.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] The present invention distinguishes the genetic material and the cell origin based on the analysis of the control characteristics of the single-cell sequencing data at the cell level transcriptome after sequencing. The advantages of the present invention in distinguishing the genetic material and the cell origin are mainly manifested in the following aspects:

[0018] 1) It is not limited to a specific host or chimeric species, has a wide application range, and can be applied to chimeric model samples of various species.

[0019] 2) It does not rely on separating the adopted cells before single-cell sequencing, and thus can minimize the damage to important cell biological and physical-chemical characteristics such as cell activity and cell surface markers.

[0020] 3) It does not rely on the prior labeling of cells in the chimeric sample, such as biomarker (introducing specific marker gene variations, etc.), chemical marker (introducing surface fluorescent groups, etc.), isotope labeling (using specific isotopes for specific species during the modeling process, etc.). Therefore, it can minimize the costs of constructing, sampling, and submitting chimeric models. Moreover, it can restore the native state of cells in the chimeric model as much as possible and reduce various interferences. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the analysis steps of the multiplex alignment and separation method for the single-cell sequencing results of the chimeric animal model of the present invention. Among them Figure 1 a) is a schematic diagram of the multi-species chimeric sample collected in the experiment; Figure 1 b) is a schematic diagram of the single-cell dissociation of the multi-species chimeric sample; Figure 1 c) is to obtain the multi-species mixed transcriptome data by library construction; Figure 1 d) is a schematic diagram of the analysis of multiplex alignment of the mixed transcriptome to the host species and the chimeric species to the corresponding reference genomes; Figure 1 e) is a schematic diagram of the analysis of mapping the multiplex alignment results to the single-cell sequencing reads and storing the data in the.BAM format; Figure 1 f) is a schematic diagram of mapping the alignment results of the single-cell sequencing reads in the.BAM file to each cell; Figure 1 g) is a schematic diagram of calculating the proportion of sequencing reads mapped to different reference genomes in each cell; Figure 1 h) is a schematic diagram of statistically analyzing the distribution of the mapping proportions of the reads after multiplex alignment in all cells and determining the separation threshold proportion; Figure 1 i) is a schematic diagram of separating the host species and chimeric species cells in the mixed transcriptome according to the separation threshold proportion.

[0022] Figure 2 It is a schematic diagram of separating human and mouse cells with human bone mesenchymal stem cells (hBMSCs) chimeric mice as an example in Example 1 of the present invention. Among them Figure 2 a) represents a chimeric liver tissue sample of double-species cells extracted from a chimeric mouse transplanted with hBMSCs; Figure 2 b) represents obtaining a species mixed transcriptome by single-cell sequencing library construction of the double-species cell chimeric liver sample; Figure 2 c) represents performing multiplex alignment on the mixed transcriptome and mapping the alignment results to individual cells; Figure 2 d) is to statistically analyze the proportion distribution of the multiplex alignment results of individual cells to determine the threshold for species division; Figure 2 e, f, g) represent the gradients of the thresholds for dividing species cells, Figure 2h) The screening threshold for the chimeric mouse tissue sample is set as the proportion of human reads greater than 90%.

[0023] Figure 3 This is the data of the human - mouse cell separation results in Example 1 of the present invention. Among them Figure 3 a) is the dimensionality - reduction clustering analysis diagram of human - mouse mixed cells in the separation results; Figure 3 b) is the flow - cytometry identification result of human - mouse mixed cells in the separation results.

[0024] Figure 4 This is a comparison schematic diagram between the present invention and the separation and sequencing method based on flow - cytometry separation technology without using the present invention.

[0025] Among them Figure 4 a) is sampling after modeling by transplanting hBMSC cells into mice; Figure 4 b) is identifying human cells according to human - specific markers; Figure 4 c) is separating human cells using a flow cytometer; Figure 4 d) represents sequencing the separated human cells.

[0026] Figure 5 This is a comparison schematic diagram between the present invention and the separation and sequencing method based on gene - editing and magnetic - bead elution technology without using the present invention. Among them Figure 5 a) represents gene modification of human cells; Figure 5 b) represents transplanting the modified hBMSC cells into mice for modeling and sampling; Figure 5 c) represents eluting human cells using modified target - binding magnetic beads; Figure 5 d) represents sequencing the separated human cells. Detailed implementation manners

[0027] The technical solutions of the present invention are further illustrated by specific examples below. The specific examples do not represent limitations on the protection scope of the present invention. Some non - essential modifications and adjustments made by others according to the concept of the present invention still fall within the protection scope of the present invention.

[0028] Reference Figure 1 , the present invention provides a multiple - alignment separation method for single - cell sequencing results of a chimeric animal model, which specifically includes the following steps:

[0029] 1. Obtain a chimeric tissue cell sample as shown in Appendix Figure 1 a): It should be noted that a tissue sample is obtained from an un - additionally - labeled chimeric animal sample, and a tissue - sample mixed - species cell suspension is made without additional labeling, separation, and precipitation treatment.

[0030] 2. As shown in Appendix Figure 1b) Load the cell suspension of the mixed species in the tissue sample processed in step 1 onto the chimeric tissue single cell sequencer.

[0031] 3. As shown in Figure 1 c), construct a library and sequence to obtain the chimeric tissue single cell transcriptome data.

[0032] 4. As shown in Figure 1 d), align with the reference genomes of multiple species:

[0033] 1) Align the obtained chimeric tissue single cell transcriptome data with the standard sequence library of the host species;

[0034] 2) Align the obtained chimeric tissue single cell transcriptome data with the standard sequence library of the chimeric species;

[0035] It should be noted that the standard sequence library of the host species and the standard sequence library of the chimeric species used in the present invention refer to the reference genome sequence libraries that are internationally recognized and widely used for this species; and this genome sequence library should contain all nucleotide sequence contents, including the sequences of genes corresponding to protein-coding proteins and all other non-coding sequences; if this species is not a model species, then the "de novo alignment and assembly reference genome" that is as comprehensive as possible should be used for alignment.

[0036] 5. As shown in Figure 1 e), map the comparison results to the sequencing reads:

[0037] 1) Extract the alignment information of each cell from the alignment results of the host species standard sequence;

[0038] 2) Extract the alignment information of each cell from the alignment results of the chimeric species standard sequence;

[0039] It should be noted that the alignment information of each cell extracted from the above standard sequence alignment results comes from the SAM format in the alignment algorithm results or other files recording the alignment results of each sequencing sequence read, and the extracted fields should fully describe whether the alignment is successful, the alignment position, the alignment quality, and whether there is multiple alignment.

[0040] 6. As shown in Figure 1 f), map the comparison results to individual cells.

[0041] 7. As shown in Figure 1 g), calculate the read mapping ratio between species in each cell:

[0042] 1) Calculate the number of sequence reads successfully aligned to the standard sequence library of the host species for each cell;

[0043] 2) Calculate the number of sequence reads successfully aligned to the standard sequence library of the chimeric species for each cell;

[0044] 3) Calculate the proportion of sequence reads in each cell that are successfully aligned to two standard sequence libraries.

[0045] 8. As shown in Figure 1 h), statistically analyze the read mapping distribution between species in all cells to determine the separation threshold:

[0046] 1) Use the successful proportion of multiple alignments in each cell as a candidate for the threshold.

[0047] 2) Determine that the separation threshold for the successful proportion of multiple alignments is more than 90%, and it can also be adjusted according to the estimated chimerism rate of the sample.

[0048] 9. As shown in Figure 1 i), use the threshold to label and separate host and chimeric cells:

[0049] 1) According to the set host species threshold, label cells with a successful multiple alignment proportion greater than the threshold as host cells.

[0050] 2) According to the labeled cell Barcodes, separate the single-cell transcriptome sequencing data.

[0051] Example 1: Separate human and mouse cells using human bone marrow mesenchymal stem cells (hBMSC) chimeric mice as an example

[0052] Refer to Figures 2 - 5 , the multi-species sample separation method based on multiple genome comparisons of the present invention was applied to compare the single-cell sequencing results of the liver of mice that received hBMSC cell transplantation with traditional chimeric sample separation methods, and the specific situation is as follows:

[0053] Sample from the liver of a mouse model containing human and mouse dual-species cells that received hBMSC transplantation and prepare for loading according to the standard process of the 10xGenome single-cell sequencing platform ( Figure 2 a); Perform single-cell loading, library construction, and sequencing on the human-mouse dual-species chimeric liver sample to obtain a dual-species mixed single-cell transcriptome, and store the transcriptome data in the.BAM file format ( Figure 2 b); Then perform dual-species reference genome alignment on the single-cell transcriptome data stored in the.BAM format to obtain the proportion of human and mouse single-cell sequencing reads between species within each cell transcriptome ( Figure 2 c);

[0054] Draw a multi-species read proportion distribution map based on the proportion of sequencing reads in all cells aligned to the dual-species reference genome ( Figure 2d). In this figure, the horizontal axis indicates the proportion of transcriptome sequencing reads in each cell that are mapped to the human reference genome; the dark line indicates the number of cells with reads successfully mapped to the human reference genome greater than the threshold when the human read proportion indicated by the current horizontal axis is the threshold; the light line indicates the overall proportion of reads successfully mapped to the human reference genome when the human read proportion indicated by the current horizontal axis is the threshold.

[0055] Furthermore, based on the estimated proportion of chimeric human cells in mouse liver, combined with the proportion of successful read alignment and the proportion of successful cell alignment, three observed proportions were set as threshold candidates according to the equal rate interval, median decline interval, and terminal stable value convergence interval between the two proportion curves (indicated by light-colored dotted lines): (i) 50% human read proportion ( Figure 2 e), (ii) 70% human reads ( Figure 2 f), (iii) 90% human reads ratio ( Figure 2 g).

[0056] Furthermore, based on observations of chimeric mouse tissue samples, the screening threshold was set to a human read ratio greater than 90% ( Figure 2 h); cells with a value greater than the threshold (in the lower right solid line box) are marked as human cells, the cell barcode is obtained, and the transcriptome single-cell data that matches the barcode of this group of human cells in the transcriptome are extracted to complete the separation of human cells and mouse cells.

[0057] Finally, the results of human and mouse cell separation in Example 1 of the present invention are as follows Figure 3 shown.

[0058] Figure 3 a shows the dimensionality reduction cluster analysis data of human-mouse mixed cells in the separation results. Specifically, a total of 10310 cells were detected after 10x single-cell sequencing; after dimensionality reduction cluster analysis, two main cell groups i) and ii) that were clearly separated can be seen. After implementing this method, the cells are marked as human cells and mouse cells according to the comparison ratio of the genetic material in each detected cell. As can be seen in the figure, the gray dots are marked as mouse cells after separation, and the black triangle scattered points are human cells.

[0059] Furthermore, it can be more clearly seen that i) all the cells in the group are mouse cells, indicating that the genetic characteristics of this group of mouse cells are relatively obvious. Generally speaking, it is also possible to separate this group of mouse cells relatively smoothly using traditional methods; while ii) there are both human cells and mouse cells in the group, and it is determined that this group is a human-mouse mixed cell population. This result shows that there are no significant separable human and mouse genetic characteristics within this group of cells, and it is very difficult to further extract human cells from this group of cells using traditional methods. However, after multiple alignment and labeling of human and mouse cells using the method of this patent, successful identification has been achieved, and targeted downstream separation and analysis of human cells can be carried out in subsequent analyses.

[0060] For this example, the result is that a total of 901 cells were separated and labeled as human cells, and it is inferred that the current chimerism rate is 8.74%, which is in line with the estimated human cell chimerism rate of 5%-10% in human-mouse chimeras.

[0061] Subsequently, the human-mouse mixed cells in the separation result were analyzed by flow cytometry, and the results are as Figure 3 shown in b. The horizontal axis is the expression intensity of human-specific surface markers, and the vertical axis is the expression intensity of the standard cell viability of the flow cytometer. Inside the black box is a cell subset that is relatively clearly separated from the main cell population. The human cell-specific surface markers of this group of cells are relatively high, and it can be considered as human cells. The proportion of these cells with high expression of human-specific markers is 7.95%, which is in line with the expected proportion of human cells chimerism in the mouse model.

[0062] Furthermore, by comparing with the results of the multi-species sample separation method based on multiple genome ratios in multiple alignment, the proportion of human cells identified by flow cytometry is close, successfully verifying that this method can successfully separate human cells.

[0063] Furthermore, some human cells do not have specific surface markers, so they cannot be separated by flow cytometry. Therefore, the proportion of human cells is slightly lower than the separation result of this method. This also shows that the separation result of this method contains a more comprehensive variety of human cells and is closer to the true species distribution of human cells in the sample.

[0064] The comparison between the present invention and the existing traditional mixed-species cell separation technology is as follows:

[0065] Comparative Example 1: Pre-sequencing separation method based on human cell surface-specific expression markers

[0066] Please refer to Figure 4 the process. Samples were taken from the liver of a mouse model containing human and mouse dual-species cells that had received hBMSC transplantation ( Figure 4 a); human cells were identified according to human cell-specific markers such as hCD45 ( Figure 4b); isolating and capturing human cells with high expression of the marker using a flow cytometer ( Figure 4 c); submitting the captured human cells for single-cell sequencing ( Figure 4 d).

[0067] The advantages of the method of the present invention are reflected in the following aspects:

[0068] Advantage 1: The method of the present invention does not require the separation of human cells before sequencing, the process steps are simpler, and the upfront cost is lower;

[0069] Advantage 2: The method of the present invention includes all types of human cells, without missing human cells lacking specific surface markers;

[0070] Advantage 3: The method of the present invention directly sequences the cell suspension, avoiding damage to cells caused by the separation step.

[0071] Comparative Example 2: A pre-sequencing separation method based on a genetically modified artificial marker

[0072] Please refer to Figure 5 procedure to genetically modify human cells before transplantation to express a specific surface marker ( Figure 5 a); sampling from the liver of a mouse model containing human and mouse dual-species cells that has received hBMSC transplantation ( Figure 5 b); binding the genetically modified human cells with a probe protein and capturing them using magnetic beads ( Figure 5 c); submitting the captured human cells for single-cell sequencing ( Figure 5 d).

[0073] The advantages of the method of the present invention are reflected in the following aspects:

[0074] Advantage 1: The method of the present invention does not require genetic modification of human cells before transplantation, maximizing the preservation of the native biological characteristics of the cells.

[0075] Advantage 2: The method of the present invention has simpler steps and saves the cost of genetically modifying specific human cells.

[0076] Advantage 3: The method of the present invention directly sequences the cell suspension, avoiding damage to cells and artificial changes in cell biological characteristics caused by the binding of the probe protein to magnetic beads and magnetic capture.

Claims

1. A method for separating multi-species samples based on multiple genome comparisons, characterized in that, Establish a method for discriminating the biological origin of biological samples in xenogeneic hybridization. The steps of the method are as follows: 1) Obtain tissue samples from chimeric animal samples without additional labeling; 2) Obtain a cell suspension of mixed species of the tissue samples, without additional labeling, separation, or precipitation treatment, and directly perform single-cell sequencing; 3) Obtain the single-cell transcriptome of the chimeric tissue; 4) Independently perform multiple alignments on each cell of the single-cell transcriptome of the chimeric tissue; 5) Align the sequencing reads of each cell to the standard reference genomes of the host and chimeric species respectively; The alignment information of each cell extracted from the standard sequence alignment results in steps 4 - 5) comes from the file of the one-by-one alignment results of the sequencing sequence reads in the alignment algorithm results. The extracted fields include the alignment success results, alignment positions, alignment quality, and multiple alignment results; 6) Incorporate the alignment records of the host and chimeric species of each cell into statistical distribution analysis to obtain the statistical distribution pattern of the multiple alignment results of all cells; 7) Calculate the separation threshold for multiple alignments; The specific steps of steps 6 - 7) are as follows: a. Statistically analyze the distribution of the proportion of successful alignments between the host and chimeric cells in the multiple alignment results of the single-cell transcriptome of the chimeric tissue among all cells; b. Draw a distribution curve of the proportion of successful alignments based on the probability of multiple alignment results; c. Calculate the change rate of the proportion of successful alignments in each distribution interval according to the distribution curve of the proportion of successful alignments; d. Calculate the separation threshold according to the change rate and referring to the specific experimental observation results of the chimeric sample; 8) Mark and separate the host and chimeric cells according to the threshold. The specific steps are as follows: a. According to the set threshold of the host species, mark the cells with a multiple alignment success proportion greater than this threshold as host cells; b. Separate the single-cell transcriptome sequencing data according to the marking; 2. The multi-species sample separation method based on multiple genome ratios according to claim 1, wherein The specific step a of steps 6 - 7) is as follows: 1) Calculate the number of sequence reads successfully aligned to the standard sequence library of the host species for each cell; 2) Calculate the number of sequence reads successfully aligned to the standard sequence library of the chimeric species for each cell; 3) Calculate the proportion of sequence reads successfully aligned to the two standard sequence libraries for each cell; 4) Map the number of cells to each sequence read proportion interval to calculate the cell proportion; 5) Statistically analyze the cell distribution in each sequence read proportion interval; The specific step d of steps 6 - 7) is as follows: Take the multiple alignment success proportion of each cell as a candidate for the threshold, and determine the separation threshold for the multiple alignment success proportion according to the estimated chimerism rate of the sample and the statistical distribution pattern of the multi-species alignment results of all cell sequencing reads in the single-cell sequencing results.

3. The multi-species sample separation method based on multiple genome ratios according to claim 2, wherein In step d, a threshold selection with a comparison success rate exceeding 90% of the target species is recognized as the cell threshold of that species, and the threshold is also adjusted according to the actual cell chimerism ratio; at the same time, referring to the statistical distribution pattern of the multi-species alignment results of all cell sequencing reads in the single-cell sequencing results, the finally selected cell threshold of that species should be within the range where the change in the comparison success rate has tended to be stable.

Citation Information

Patent Citations

  • PDX model single cell transcriptome data analysis method, device and medium

    CN117912552A