Splitting method and splitting device for single-cell pooled sample sequencing data
By utilizing SNP locus information and sex-specific gene expression from the 1000 Genomes Project, combined with Vireo software, the accuracy and comprehensiveness issues of splitting single-cell mixed sample sequencing data were resolved, enabling efficient cell origination and data analysis without additional experiments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN NUOHEZHIYUAN BIO-INFORMATION TECH CO LTD
- Filing Date
- 2023-08-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing single-cell mixed sample sequencing data splitting techniques rely on additional experimental steps and lack accuracy and comprehensiveness, failing to effectively distinguish cell origins, especially in the absence of individual SNP information, making accurate splitting difficult.
Using SNP locus information based on the 1000 Genomes Project, combined with the expression ratios of sex-specific genes USP9Y and ZFY, data was split using Vireo software, and the origin of individual cells was inferred through Bayesian methods to construct a complete data analysis workflow.
It achieves high-accuracy cell tracing even without individual SNP information, and provides a complete workflow from upstream experiments to downstream data processing, suitable for donor cell identification of valuable clinical samples and post-transplant samples.
Smart Images

Figure CN117079714B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a method and apparatus for splitting single-cell mixed sample sequencing data. Background Technology
[0002] Single-cell sequencing technologies based on droplets or microplates have successfully enabled the rapid, high-throughput construction of single-cell level transcriptomic maps. To reduce costs or due to limitations in initial sample volume, cell suspensions from different samples can be mixed together for single-cell capture. In some clinical samples, such as tissue samples from transplant patients, donor-recipient cell mixing may also occur (all referred to as pooling).
[0003] Currently, there are three main strategies for cell separation in mixed samples:
[0004] (1) Add cell tags before mixing, such as using cell surface protein antibody technology that couples sample tag sequences, such as the CITE-seq technology developed by the New York Genome Research Center; this technology requires the introduction of additional experimental steps, and its data splitting efficiency is low and its uniformity is poor, so this application has not been widely expanded.
[0005] (2) Data splitting based on natural tags present in cells—SNPs. Several data splitting tools have been developed, including Demuxlet, and a flowchart of the splitting process is shown below. Figure 1 As shown, however, Demuxlet analysis requires providing the individual's SNP loci, which will increase the need for additional genome sequencing and sample size.
[0006] (3) Data splitting based on SNP information in public databases. This technical solution does not rely on additional SNP information, but its accuracy and technical details need further investigation. It can only split the data and cannot determine the individual source.
[0007] Furthermore, there is currently a lack of detailed evaluation of the number of SNP sites required for data splitting by various software programs, as well as their accuracy, so the application of these data splitting tools remains limited.
[0008] In summary, current pooled sample separation techniques rely on additional experiments and lack accurate, comprehensive, and well-established procedures.
[0009] In view of this, the present invention is hereby proposed. Summary of the Invention
[0010] The first objective of this invention is to provide a method for splitting single-cell mixed sample sequencing data to solve at least one of the above-mentioned problems.
[0011] A second objective of this invention is to provide a device for splitting single-cell mixed sample sequencing data.
[0012] A third objective of this invention is to provide a processor.
[0013] In a first aspect, the present invention provides a method for splitting sequencing data from a single-cell mixed sample, comprising the following steps:
[0014] a. Single-cell suspensions were captured and sequenced using a single-cell platform. The sequencing data were then used to perform reference genome alignment, cell identification, and gene expression level quantification using Cellranger.
[0015] b. Based on the SNP site information in the 1000 Genomes Project, the cell data obtained in step a were split into two groups of cell data from different sexes;
[0016] c. Based on the proportion of gene expression of sex-specific genes in the two groups of cell data from different sexes, the two groups of cell data were distinguished as originating from male or female samples, respectively;
[0017] The single-cell suspension includes male and female samples;
[0018] The sex-specific genes include at least one of the USP9Y gene or the ZFY gene.
[0019] As a further technical solution, in step a, a 10×Genomics platform is used for single-cell capture, and an Illumina sequencer is used for sequencing.
[0020] As a further technical solution, in step b, based on the SNP site information in the 1000 Genomes Project, scSplit, Vireo, or Souporcell are used to split the cell data identified in step a into two groups of cell data for different sexes.
[0021] As a further technical solution, in step b, based on the SNP site information in the 1000 Genomes Project, Vireo is used to split the cell data identified in step a into two groups of cell data of different sexes.
[0022] As a further technical solution, the sample may include blood, solid tissue, or bone marrow.
[0023] In a second aspect, the present invention provides a device for splitting sequencing data of single-cell mixed samples, including an alignment module, a data splitting module and an identification module;
[0024] The alignment module is used to perform reference genome alignment, cell identification, and gene expression level quantification on single-cell sequencing data of single-cell suspensions using Cellranger.
[0025] The data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project.
[0026] The identification module is used to distinguish whether the two sets of cell data are from male or female samples, based on the proportion of gene expression of sex-specific genes in the two sets of cell data of different sexes.
[0027] The single-cell suspension includes male and female samples;
[0028] The sex-specific genes include at least one of the USP9Y gene or the ZFY gene.
[0029] As a further technical solution, the data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project, using scSplit, Vireo, or Souporcell.
[0030] As a further technical solution, the data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project using Vireo.
[0031] As a further technical solution, the sample may include blood, solid tissue, or bone marrow.
[0032] Thirdly, the present invention provides a processor for running a program, wherein the program executes the above-described method for splitting single-cell mixed sample sequencing data.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1) This invention utilizes SNP data splitting technology to achieve cell tracing in pooled samples even without individual SNP information.
[0035] 2) This invention comprehensively tests the current mainstream data splitting tools, selects the high-accuracy data splitting tool Vireo, and extends it based on existing technology to realize the entire process from sample processing to downstream data analysis, and proves its feasibility on blood, bone marrow and other sample data.
[0036] 3) Based on the split cell tags, this invention can innovatively construct the correspondence between cells and individuals through sex-specific genes, thereby achieving accurate individual data splitting.
[0037] 4) This invention realizes an integrated data analysis process with high process integrity, applicable to a wide range of application scenarios, without the need for additional samples to obtain SNP information, suitable for precious clinical samples, and provides a complete solution for donor cell identification of transplanted samples. Attached Figure Description
[0038] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the mixed sample splitting technique.
[0040] Figure 2 Accuracy test results for different software;
[0041] Figure 3 This is a flowchart of the splitting method of the present invention;
[0042] Figure 4 UMAP clustering diagram for mixed samples;
[0043] Figure 5 Clustering results for BM_M1 and BM_F1 samples after cell splitting. Detailed Implementation
[0044] The embodiments and examples of the present invention will be described in detail below. However, those skilled in the art will understand that the following embodiments and examples are for illustrative purposes only and should not be considered as limiting the scope of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Unless otherwise specified, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0045] In a first aspect, the present invention provides a method for splitting sequencing data from a single-cell mixed sample, comprising the following steps:
[0046] a. Single-cell suspensions were captured and sequenced using a single-cell platform. The sequencing data were then used to perform reference genome alignment, cell identification, and gene expression level quantification using Cellranger.
[0047] b. Based on the SNP site information in the 1000 Genomes Project, the cell data obtained in step a were split into two groups of cell data from different sexes;
[0048] c. Based on the proportion of gene expression of sex-specific genes in the two groups of cell data from different sexes, the two groups of cell data were distinguished as originating from male or female samples, respectively;
[0049] The single-cell suspension includes male and female samples;
[0050] The sex-specific genes include at least one of the USP9Y gene or the ZFY gene.
[0051] The inventors discovered that the USP9Y and ZFY genes are widely expressed in males but less so in females. Based on this, UFP9Y and ZFY were selected as male-specific expression genes, and the split cell populations were sex-labeled.
[0052] It should be noted that the sex-specific genes in this invention include, but are not limited to, at least one of the USP9Y gene or the ZFY gene. Other sex-specific genes that can be used for sample sex identification can also be used to sex-label the split cell population.
[0053] The single-cell mixed sample sequencing data splitting method provided by this invention eliminates the need for additional experimental operations such as protein labeling and genome sequencing. Even when mixed samples originate from different physiological sexes and individual SNP information is unavailable, it can still provide accurate and reliable data splitting, allowing each split data to be analyzed separately downstream. This method overcomes the shortcomings of current mixed sample splitting methods, which rely on individual SNP locus information and lack the correspondence between split cells and individuals. Furthermore, it comprehensively tests and evaluates the accuracy of various splitting software, forming a complete workflow system from upstream experiments to downstream data processing, demonstrating superior performance.
[0054] In some alternative implementations, in step a, single-cell capture is performed using a 10×Genomics platform, and sequencing is performed using an Illumina sequencer.
[0055] In some alternative implementations, in step b, based on the SNP locus information from the 1000 Genomes Project, scSplit, Vireo, or Souporcell is used to split the cell data identified in step a into two groups of cell data for different sexes.
[0056] In some preferred embodiments, in step b, based on the SNP site information from the 1000 Genomes Project, Vireo is used to split the cell data identified in step a into two groups of cell data for different sexes.
[0057] Vireo software can infer the genotype of each cell and, based on a Bayesian model, assign each cell to a corresponding individual by calculating probabilities. This allows Vireo to provide highly accurate results when splitting pooled data.
[0058] The inventors discovered that, based on SNP locus information from the 1000 Genomes Project, Vireo can perform data splitting quickly, with high detection and accuracy.
[0059] In some alternative implementations, the sample includes, but is not limited to, blood, solid tissue, or bone marrow, or other samples well known to those skilled in the art. Solid tissue may be, for example, but is not limited to, tumor tissue, liver tissue, kidney tissue, etc.
[0060] In some preferred embodiments, the method for splitting single-cell ensemble sequencing data is as follows: Figure 3 As shown, it includes:
[0061] (1) The target tissues of two individuals A (female) and B (male) of different physiological sexes were separated to obtain single-cell suspensions.
[0062] (2) Mix the two suspensions to obtain a mixed suspension. Use a commercial high-throughput single-cell platform for single-cell capture and sequencing. It is recommended to use the 10×Genomics platform for single-cell capture. Use an Illumina sequencer for sequencing. After the sequencer runs, the fastq data is compared with Cellranger to obtain the bam file. Cellranger is used to identify the cell list and then the gene expression matrix is obtained after Cellranger processing.
[0063] (3) For the bam file obtained by Cellranger alignment, based on the SNP site information in the 1000 Genomes Project, the cell data was split into two groups of cell data of different sexes using Vireo.
[0064] (4) The cell list and gene expression matrix obtained by Cellranger identification and processing are processed by Seurat to obtain the cell-gene expression matrix.
[0065] (5) For the two sets of cell data after splitting, the qualitative analysis is performed by the proportion of expression of sex-specific genes (USP9Y and ZFY) in the cell-gene expression matrix.
[0066] In a second aspect, the present invention provides a device for splitting sequencing data of single-cell mixed samples, including an alignment module, a data splitting module and an identification module;
[0067] The alignment module is used to perform reference genome alignment, cell identification, and gene expression level quantification on single-cell sequencing data of single-cell suspensions using Cellranger.
[0068] The data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project.
[0069] The identification module is used to distinguish whether the two sets of cell data are from male or female samples, based on the proportion of gene expression of sex-specific genes in the two sets of cell data of different sexes.
[0070] The single-cell suspension includes male and female samples;
[0071] The sex-specific genes include at least one of the USP9Y gene or the ZFY gene.
[0072] The device for splitting single-cell mixed sample sequencing data provided by this invention can achieve cell origination in mixed samples without individual SNP information, and it is highly accurate and fast.
[0073] In some alternative implementations, single-cell capture is performed using the 10×Genomics platform, and sequencing is performed using an Illumina sequencer to obtain single-cell sequencing data of the single-cell suspension.
[0074] In some alternative implementations, the data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on SNP site information from the 1000 Genomes Project, using scSplit, Vireo, or Souporcell.
[0075] In some preferred embodiments, the data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project using Vireo.
[0076] The inventors discovered that, based on SNP locus information from the 1000 Genomes Project, Vireo can perform data splitting quickly, with high detection and accuracy.
[0077] In some alternative implementations, the sample includes, but is not limited to, blood, solid tissue, or bone marrow, or other samples well known to those skilled in the art. Solid tissue may be, for example, but is not limited to, tumor tissue, liver tissue, kidney tissue, etc.
[0078] Thirdly, the present invention provides a processor for running a program, wherein the program executes the above-described method for splitting single-cell mixed sample sequencing data.
[0079] The processor provided by this invention can achieve cell origination in mixed samples without individual SNP information, with high accuracy and fast separation speed.
[0080] The present invention will be further illustrated below with specific embodiments and comparative examples. However, it should be understood that these embodiments are merely for the purpose of more detailed illustration and should not be construed as limiting the present invention in any way.
[0081] Example 1
[0082] (I) Accuracy Testing of Mixed Sample Splitting Software
[0083] 1. Experimental Design
[0084] Two healthy adult bone marrow samples from different biological sexes were collected and named BM_M1 (male) and BM_F1 (female), respectively. For each sample, 10×Genomics capture was performed to construct a single-cell transcriptome library. Raw FastQ data was obtained via Illunima sequencing, and the FastQ data was aligned using Cellranger to obtain BAM files. Using the respective cell tag barcode libraries of the two samples as controls, pseudo-mixed data was constructed by mixing the two sample data. Based on SNP locus information from the 1000 Genomes Project, data splitting was performed using scSplit, Vireo, and Souporcell software. The splitting results were compared with the control barcode library to evaluate the data splitting accuracy.
[0085] 2. Test Results
[0086] 2.1 Building the Cell Label Barcode Library
[0087] The BM_M1 sample contained 7,718 cells, and the BM_F1 sample contained 4,996 cells, totaling 12,714 cells. 39 barcodes were shared by both samples and were considered doublets. This resulted in a total of 12,675 cells with clear taxonomic information, which served as the reference barcode library. This library included 7,679 BM_M1 cells, 4,957 BM_F1 cells, and 39 doublets.
[0088] 2.2 Data Merging to Construct Mixed Sample Data
[0089] The two original data were merged to construct pseudo-mixed data, and a total of 12,068 mixed cells were finally obtained after comparison and quantification.
[0090] Three data splitting software programs were used for testing: scSplit, Vireo, and Souporcell. Vireo is based on a Bayesian algorithm, while scSplit and Souporcell use deconvolution principles for data splitting. The UMAP clustering plots of the splitting results from scSplit, Vireo, and Souporcell are shown below. Figure 4 ( Figure 4 The image shows the following: (Top left is Reference, top right is scSplite splitting result, bottom left is Vireo splitting result, and bottom right is soupocell splitting result).
[0091] 2.3 Test Results
[0092] 2.3.1 BM_M1 Sample Splitting Results
[0093] Table 1. Detailed results of cell separation from samples tested using the BM_M1 software.
[0094]
[0095] (1) Cells identified by software refers to the number of cells identified by software as originating from BM-M1;
[0096] (2) Overlap with Reference refers to the number of cells detected above that are confirmed to be true, that is, the number of cells that overlap with the Reference.
[0097] (3) Cell counts of Reference: The total number of BM_M1 cells in the Reference;
[0098] (4) Detection Rate: The detection rate, i.e. (2) / (3);
[0099] (5) Accuracy Rate: The accuracy rate is (2) / (1).
[0100] Results from BM_M1 detection (Table 1 and Figure 2 As can be seen, scSplit has high accuracy but low detection rate; Vireo and Soupocell have both detection and accuracy rates above 95%.
[0101] 2.3.2 BM_F1 Sample Splitting Results
[0102] Results from BM_F1 detection (Table 2 and Figure 2 As can be seen, the detection rates of the three software programs are similar (83-84%), and their accuracy rates are similar (over 98%).
[0103] Table 2. Detailed results of cell separation for BM_F1 samples from three testing software programs.
[0104]
[0105] 2.3.3 Operating Speed Evaluation
[0106] Of the three testing software programs, Vireo ran the fastest, as detailed in Table 3.
[0107] Table 3. Detailed Running Time of the Three Test Software Programs
[0108] software Time taken (hours) Souporcell 7 scSplit 22 Vireo 3
[0109] 2.3.4 Conclusion
[0110] Based on comprehensive evaluation, Vireo demonstrated the best detection rate and accuracy, and also ran quickly. Therefore, this technical workflow uses Vireo as the mixed sample separation tool.
[0111] (II) Sex-specific gene screening
[0112] The main problem with existing data splitting tools is that they can only split the data, but cannot determine the correspondence between the split cells and the samples. Based on this, we independently developed a method that uses sex-specific genes as markers to effectively distinguish cells from different physiological sexes, thus solving the problem of cell attribution after splitting.
[0113] Based on literature review, sex-specific genes, i.e. genes specifically expressed on the Y chromosome, were collected. Their expression in individuals of different sexes was statistically analyzed, and genes that are widely expressed in males but less expressed in females were screened. The screening results are shown in Table 4. Finally, UFP9Y and ZFY were selected as Y chromosome-specific expression genes, and individual markers were performed on the split cell populations.
[0114] Table 4. Expression of sex-related genes in different samples and screening of candidate marker genes.
[0115]
[0116]
[0117] (III) Data splitting process that does not rely on SNP
[0118] Based on the above experimental results, this invention provides a method for splitting single-cell mixed sample sequencing data. This method does not rely on additional experiments or the acquisition of individual SNP information. Using the high-accuracy Vireo splitting tool and the expression of sex-specific genes UFP9Y and ZFY, it splits mixed sample data from individuals of different sexes and clarifies the individual cell affiliation. Figure 3 As shown, it includes the following steps:
[0119] (1) The target tissues of two individuals A (female) and B (male) of different physiological sexes were separated to obtain single-cell suspensions.
[0120] (2) Mix the two suspensions to obtain a mixed suspension. Use a commercial high-throughput single-cell platform for single-cell capture and sequencing. It is recommended to use the 10×Genomics platform for single-cell capture. Use an Illumina sequencer for sequencing. After the sequencer runs, the fastq data is compared with Cellranger to obtain the bam file. Cellranger is used to identify the cell list and then the gene expression matrix is obtained after Cellranger processing.
[0121] (3) For the bam file obtained by Cellranger alignment, based on the SNP site information in the 1000 Genomes Project, the cell data was split into two groups of cell data of different sexes using Vireo.
[0122] (4) The cell list and gene expression matrix obtained by Cellranger identification and processing are processed by Seurat to obtain the cell-gene expression matrix.
[0123] (5) For the two sets of cell data after splitting, the qualitative analysis is performed by the proportion of expression of sex-specific genes (USP9Y and ZFY) in the cell-gene expression matrix.
[0124] (iv) Data splitting after cell tracing
[0125] Based on the barcode information of each sample after splitting, the matrix file can be split for downstream data filtering, PCA dimensionality reduction, UMAP / tSNE dimensionality reduction clustering (e.g.) Figure 5 As shown in the figure, the left figure represents the clustering results of the BM_M1 sample, and the right figure represents the clustering results of the BM_F1 sample. The system also includes differential enrichment analysis and related result plotting, achieving complete functionality. This workflow is technologically innovative and practical, providing strong support for single-cell level research and applications.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for splitting sequencing data from a single-cell mixed sample, characterized in that, Includes the following steps: a. Single-cell suspensions were captured and sequenced using a single-cell platform. The sequencing data were then used to perform reference genome alignment, cell identification, and gene expression level quantification using Cellranger. b. Based on the SNP locus information in the 1000 Genomes Project, use scSplit, Vireo, or Souporcell to split the cell data obtained in step a into two groups of cell data for different sexes; c. Based on the proportion of gene expression of sex-specific genes in the two groups of cell data from different sexes, the two groups of cell data were distinguished as originating from male or female samples, respectively; The single-cell suspension includes male and female samples; The sex-specific gene is the USP9Y gene.
2. The splitting method according to claim 1, characterized in that, In step a, single-cell capture was performed using the 10×Genomics platform, and sequencing was performed using an Illumina sequencer.
3. The splitting method according to claim 1, characterized in that, In step b, based on the SNP site information from the 1000 Genomes Project, Vireo is used to split the cell data obtained in step a into two groups of cell data for different sexes.
4. The splitting method according to claim 1, characterized in that, The samples may include blood, solid tissue, or bone marrow.
5. A device for splitting sequencing data from a single-cell mixed sample, characterized in that, It includes a comparison module, a data splitting module, and an identification module; The alignment module is used to perform reference genome alignment, cell identification, and gene expression level quantification on single-cell sequencing data of single-cell suspensions using Cellranger. The data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project, using scSplit, Vireo, or Souporcell. The identification module is used to distinguish whether the two sets of cell data are from male or female samples, based on the proportion of gene expression of sex-specific genes in the two sets of cell data of different sexes. The single-cell suspension includes male and female samples; The sex-specific gene is the USP9Y gene.
6. The splitting device according to claim 5, characterized in that, The data splitting module is used to split the cell data identified by the alignment module into two groups of cell data of different sexes based on the SNP site information in the 1000 Genomes Project using Vireo.
7. The splitting device according to claim 5, characterized in that, The samples may include blood, solid tissue, or bone marrow.
8. A processor, characterized in that, The processor is used to run a program, wherein the program executes the method for splitting single-cell mixed sample sequencing data according to any one of claims 1-4.