Single-cell mix-and-seq and sample splitting method based on hla genes

By using a single-cell pooled sequencing method based on HLA genes, the problems of large cell loss, complex operation, and high computational cost in existing technologies have been solved. This method enables efficient splitting and accurate sequencing of clinical samples with small cell counts, expanding the application potential of single-cell pooled sequencing.

CN120998307BActive Publication Date: 2026-02-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511509987.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-06
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing single-cell pooled sequencing methods suffer from problems such as large cell loss, complex operation, limited throughput, and high computational cost when processing clinical samples with sparse cell counts, making them difficult to effectively apply to single-cell pooled sequencing of clinical samples with small cell counts.

Method used

A single-cell pooled sequencing and sample splitting method based on the human leukocyte antigen (HLA) gene was adopted. By obtaining the HLA genotyping information of each sample to be pooled, single-cell sequencing was performed after pooling. The sample origin of each cell was identified based on the HLA gene expression characteristics, and the sequencing sample splitting was completed, avoiding exogenous labeling operations.

Benefits of technology

It improves the utilization rate of precious samples, reduces computational consumption and sequencing costs, enhances the ability to obtain genetic information from low-coverage cells, improves the accuracy and sensitivity of sequencing, and makes large-scale research more feasible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998307B_ABST
    Figure CN120998307B_ABST
Patent Text Reader

Abstract

The application discloses a single-cell mixed sample sequencing and sample splitting method based on HLA genes. The application first acquires HLA gene typing information of each sample to be mixed; secondly, single-cell mixed sample sequencing is performed on the mixed sample to obtain single-cell mixed sample sequencing results of the mixed sample; finally, based on the HLA gene typing information, the sample source of each cell is identified by calculating the expression characteristics of HLA genes in the single-cell mixed sample sequencing sequence, and sample splitting is completed. The application breaks through the technical obstacle that few-cell clinical samples cannot be subjected to single-cell sequencing, greatly improves the utilization rate of these precious samples, significantly reduces the calculation consumption of sample splitting through an efficient HLA gene sequence pseudo-alignment algorithm, improves the ability to obtain genetic information from low-coverage cells by focusing on the HLA region with the highest polymorphism in the human genome, and effectively reduces the sequencing cost of a single sample through mixed sample sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of single-cell sequencing, and particularly relates to a single-cell mixed sequencing method based on human leukocyte antigen (HLA) genes and a sample splitting method. BACKGROUND

[0002] Single-cell transcriptome sequencing technology has been widely used in various basic and clinical research fields of life sciences, and plays an important role in understanding the clinical detection and pathological mechanism of various major diseases such as malignant tumors, organ transplantation, autoimmune diseases, and blood diseases. However, many clinical samples, such as biopsy puncture samples, tissue fluid or body fluid of a specific organ (such as lung lavage fluid, cerebrospinal fluid, urine, etc.), usually contain a small number of cells, which is difficult to meet the minimum cell load requirement (usually >10,000 cells per sample) required for standard single-cell transcriptome sequencing of a single sample. To solve this problem, one strategy is to mix multiple such samples to achieve the total cell amount required for sequencing.

[0003] Traditional single-cell mixed sequencing methods usually rely on exogenous sample markers for data splitting. These markers (such as sample-specific oligonucleotide sequences) are anchored on the cell membrane or nuclear membrane through antibodies, liposomes or click chemistry reactions, etc. By detecting these exogenous markers, the original sample source of each cell can be determined. However, these methods have limitations such as complex labeling procedures, time-consuming operations, and up to about 50% cell loss during multiple centrifugation and washing processes, making them unsuitable for valuable clinical samples with limited cell numbers.

[0004] Another type of single-cell mixed sequencing method does not require exogenous sample markers, and achieves sample splitting of single-cell data through genotype-based demixing algorithms (such as Demuxlet, Vireo and Souporcell). These methods mainly rely on reading and identifying single nucleotide polymorphisms of different samples from scRNA-seq data, but face two key obstacles.

[0005] First, to obtain a sufficient number of single nucleotide mutations with information value, whole genome sequencing and whole genome or exon alignment are usually required, which brings huge computational burden in terms of computing time and required memory.

[0006] Secondly, the inherent sparsity of single-cell transcriptome sequencing data, especially data from 3' end protocol, will result in low and uneven coverage of single nucleotide mutation sites in the transcriptome. This defect reduces the ability to classify cells with low messenger ribonucleic acid content, resulting in a high proportion of unassigned cells and impairing the utility of valuable samples.

[0007] In summary, although existing single-cell pooled sequencing and splitting methods can achieve accurate sample splitting, they generally suffer from limitations such as large cell loss, complex operation, limited throughput, and high computational cost, making them difficult to effectively apply to single-cell pooled sequencing of clinical samples with small cell counts. Summary of the Invention

[0008] This invention addresses the problems of significant cell loss, operational complexity, limited throughput, and high computational cost associated with existing single-cell pooled sequencing technologies when processing clinical samples with sparse cell counts. It proposes a single-cell pooled sequencing and sample splitting method based on the human leukocyte antigen (HLA) gene, comprising the following steps:

[0009] Obtain the HLA genotyping information for each sample to be mixed;

[0010] The samples to be mixed are then subjected to single-cell sequencing to obtain the single-cell sequencing results of the mixed samples.

[0011] Based on the HLA genotyping information, the expression characteristics of HLA genes in the single-cell sequencing sequence are calculated to identify the sample origin of each cell and complete the sequencing sample splitting.

[0012] The beneficial effects of this invention are as follows:

[0013] First, it overcomes the technical barrier of single-cell sequencing for clinical samples with small cell counts, greatly improving the utilization rate of these precious samples. Second, it eliminates the need for complex whole-genome single nucleotide polymorphism (SNP) detection, significantly reducing the computational cost of sample splitting through an efficient HLA gene sequence pseudo-alignment algorithm. Third, by focusing on the most polymorphic HLA regions in the human genome, it maximizes the ability to obtain genetic information from low-coverage cells, thereby improving accuracy and sensitivity. Finally, pooling effectively reduces the sequencing cost of individual samples, making large-scale studies more feasible. Attached Figure Description

[0014] Figure 1 This is a schematic diagram illustrating the principle of a single-cell pooled sequencing method based on HLA genes.

[0015] Figure 2 A schematic diagram illustrating the core principle of the HLA-Splitter sample splitting algorithm based on the HLA gene;

[0016] Figure 3 This refers to the single-cell t-SNE clustering results based on HLA genes in Example 1. Figure 3 (a) shows the actual single-cell sample source markers. Figure 3 (b) shows the results of single-cell sample splitting based on HLA genes;

[0017] Figure 4 Figure 4: Sample enrichment score heatmap arranged by sample label for Example 1;

[0018] Figure 5 Figure 5: Cell proportion of sample split using HLA-Splitter in mixed sample simulation data with different number of cells per sample for Example 1;

[0019] Figure 6 Figure 6: Computational time of HLA-Splitter and several published genotype-based single cell deconvolution algorithms in processing mixed sample simulation data for Example 1;

[0020] Figure 7 Figure 7: Single cell t-SNE clustering results based on HLA genes for Example 2, Figure 7 (a) sample split results of published algorithm Demuxlet, Figure 7 (b) sample split results of HLA-Splitter;

[0021] Figure 8 Figure 8: Heatmap of consistent cells in each sample between HLA-Splitter and Demuxlet for Example 2. DETAILED DESCRIPTION

[0022] The embodiments of the present application are described in detail below with specific reference being made to certain examples. It is intended that the scope of the present application be defined by the appended claims, and not limited by the following examples. The examples are provided to illustrate the present application and should not be construed as limiting the scope of the present application. The technical solutions of the present application can be modified or changed in various ways based on different perspectives and applications, as long as they do not deviate from the spirit and scope of the present application, and should be covered within the scope of the claims of the present application. It should be noted that the following examples and features in the examples can be combined with each other without conflict.

[0023] The present application proposes a single cell mixed sample sequencing and sample splitting method based on HLA genes. The method can quickly and accurately split single cell transcriptome sequencing data from different samples by recognizing endogenous HLA genotypes, and does not require additional exogenous sample labeling.

[0024] As shown in Figure 1 the present application includes the following steps:

[0025] S1: Obtain sample HLA gene typing information

[0026] In the present application, HLA genotyping information is obtained by one or more genotyping techniques, including but not limited to: polymerase chain reaction sequencing, gene chip, high-throughput sequencing, or by obtaining a sample owner's clinical HLA gene analysis report (such as an organ transplant sample), or by rapid identification through high-efficiency methods such as nanopore sequencing technology. For example, in the examples, by extracting sample peripheral blood nucleic acid and constructing a nanopore sequencing library, HLA genotyping of 30 samples was completed at one time, greatly reducing the sequencing time and cost compared to traditional methods.

[0027] S2: Single-cell transcriptome sequencing after sample mixing

[0028] In the present application, the cells used for sequencing can be mixed samples containing a small number of cells, such as biopsy samples, tissue fluid or body fluid samples of specific organs (such as urine, lung lavage fluid, cerebrospinal fluid, etc.), organoid samples, and detection of rare cell subpopulations, with a minimum cell requirement of only 100 cells per sample. The sample can be stored using liquid nitrogen. The number of single mixing samples is usually 20-40 samples of different individuals, and the specific number is determined according to the total number of cells after mixing, and the number of cells recommended for one-time machine is not more than 20,000.

[0029] S3: Sequencing sample splitting based on sample HLA genotype information

[0030] In the present application, sequencing sample splitting based on sample HLA genotype information is achieved by calculating the expression characteristics of HLA genes in single-cell sequencing sequences to identify the sample source of each cell and complete sequencing sample splitting.

[0031] Further, as shown in Figure 2 , the implementation of sequencing sample splitting includes the following steps:

[0032] (1) HLA gene reference database construction

[0033] In the present application, the construction of HLA gene reference database firstly extracts the corresponding gene coding sequence (CoDing Sequence, CDS) from the immunogenetics database (IMGT / HLA) according to the HLA genotype information (four-bit precision, for example, HLA-A*02:01) of each sample. In order to ensure accuracy, the gene coding sequence is standardized and pretreated to remove any redundant information. Through the index function of the gene expression quantification software kallisto, a sample-specific HLA gene reference database is constructed, which is a graph structure (colored de Bruijn graph) for sequence alignment and assembly. The nodes in the graph represent short sequences (k-mers) of length k, and the edges represent the overlapping relationship between k-mers. This graph structure can efficiently store and represent a large number of HLA allele sequences, providing a fast retrieval skeleton for subsequent pseudo-alignment.

[0034] (2) Pseudo-alignment and count matrix generation

[0035] In the present application, the pseudo-alignment and the generation of the count matrix firstly extract the single-cell sequencing sequences on chromosome 6 from the single-cell sequencing result alignment mapping file (BAM) using Samtools software. Further, the BAM file is generated by the official software Cellranger of the 10X single-cell sequencing platform with default parameters, which already contains the mapping information of the single-cell sequencing data on the human genome (GRCh38). Then, these reads are pseudo-aligned to the HLA gene reference database generated in the previous step through the Kallisto-bustools single-cell data analysis pipeline, thereby generating a count matrix reflecting the expression of each HLA gene in each cell (MEX format).

[0036] In the present application, unlike traditional genome alignment, pseudo-alignment does not perform global alignment on a base-by-base basis, but rather decomposes each sequencing read into a series of short sequences (k-mers) of length k. These k-mers are used to traverse the colored de Bruijn graph of the pre-constructed HLA gene reference database. When a k-mer finds a match in the graph, it inherits the "color" of the corresponding node, which represents the HLA allele it belongs to. By integrating the color information from all k-mers of the same read, Kallisto-bustools can quickly and accurately determine which HLA allele or alleles the read is most likely derived from, greatly improving the alignment speed and accuracy.

[0037] (3) Calculate the enrichment score of each cell for each sample to be mixed based on the HLA gene expression count matrix, and realize sample allocation for each cell based on the enrichment score.

[0038] In the present application, let H be the original HLA gene expression count matrix, where M is the total number of cells, N is the total number of alleles, represents the raw expression value of allele j in cell i. Then, the normalized HLA gene expression matrix is obtained by the following formula :

[0039]

[0040] Subsequently, to balance the expression level difference between different HLA alleles, for each allele in the matrix , only its positive expression value is used to calculate the standard deviation , and scaling is performed to obtain the scaled HLA gene expression count matrix used for calculation.

[0041]

[0042] Specifically, in the probabilistic sample assignment based on HLA-Score, let S be a sample enrichment score matrix, where K is the number of mixed samples. The enrichment score of cell i for each mixed sample k is calculated as follows:

[0043]

[0044] Specifically, for each sample k and each HLA gene locus type (including HLA-A, B, C, DRB1, DQB1 and DPB1), where is the set of HLA alleles of mixed sample k at HLA gene locus type , is the homozygous / heterozygous coefficient: when mixed sample k is homozygous at gene locus , ; when it is heterozygous, .

[0045] Specifically, the preliminary predicted sample label of each cell is assigned to the sample that gives it the maximum enrichment score, but it is required that this maximum score must exceed an absolute threshold default (default 0.1), otherwise it is marked as “Undefined”, the formula is as follows:

[0046]

[0047] wherein is the predicted sample label of cell i, is the enrichment score of cell i for sample k. is the sample label corresponding to the maximum enrichment score in cell i, is the absolute threshold of enrichment score.

[0048] (4) Sample label correction and revision

[0049] In the present application, in order to further improve the confidence of the allocation, the reduced HLA gene expression count matrix is subjected to t-SNE dimensionality reduction to generate a two-dimensional embedding plot. For each cell i (excluding those initially labeled as "Undefined"), according to the local neighborhood majority voting strategy, the most frequent sample label among the nearest 100 neighboring nodes in the two-dimensional t-SNE projection is identified for each cell, and the sample label of the cell is corrected and revised. Let be the most frequent sample label among the neighboring nodes, represent the frequency among the neighboring nodes. The final revised label of cell i is determined according to the following formula.

[0050]

[0051] where τ is the confidence threshold of the neighborhood majority label frequency (default 0.5).

[0052] The present application selects HLA genes as the endogenous gene reference for sample splitting, mainly based on the following reasons:

[0053] (1) HLA genes are highly polymorphic regions in the human genome, with extremely high HLA genotype differences between different individuals. This polymorphism makes HLA genotype an ideal individual identification "fingerprint", which can effectively distinguish cells from different individuals;

[0054] (2) HLA genes are expressed in most nucleated cells, which are usually the target of single-cell sequencing;

[0055] (3) As an endogenous gene of the cell itself, no external labeling operation is required, thus completely avoiding the possible cell loss caused by the external labeling process;

[0056] (4) In some clinical situations (e.g. organ transplantation), the HLA typing information of the subject can already exist, and even if additional sequencing is needed, it can be efficiently completed using fast and low-cost technologies such as nanopore sequencing.

[0057] The features and performances of the present application are further described in detail below in combination with examples.

[0058] Example One: Test method performance using simulated single-cell mixed sequencing data sets

[0059] See Figures 3 to 6 This example aims to test the reliability and robustness of the single-cell sequencing data splitting method based on HLA genotypes, and by comparing the sample splitting results of the statistical simulation data set with the actual sample source, to evaluate the accuracy and computational efficiency of the method, including the following steps:

[0060] 1. Construction of simulated single-cell mixed sequencing data sets

[0061] In the examples of the present application, the simulated data set contains single-cell transcriptome data from 40 peripheral blood mononuclear cell (PBMC) samples in the public database Gene Expression Omnibus (Accession number: GSE226602). The HLA genotypes of these samples were obtained in advance from the single-cell transcriptome sequencing data of each donor using the high-resolution HLA typing software arcasHLA (arcasHLA and related typing procedures are well-known techniques in the art).

[0062] Further, in order to construct the mixed simulation data, from the 40 single-cell sequencing data, 10-1000 cells were extracted from each sample using non-repeated random sampling, thereby simulating a data set with a total cell amount of 400-40000 cells. Specifically, the random sampling method uses the random.sample() function of Python to complete the random sampling of cell barcodes, then uses the subset-BAM tool to filter out the corresponding cell gene expression information from the genome alignment file of 10x Genomics, and finally uses the bamtofastq command of 10x Genomics Cell Ranger to convert the BAM file to Fastq sequence file.

[0063] 2. Single-cell sequencing data splitting based on HLA genotypes

[0064] In this embodiment, the simulated single-cell pooled sequencing data was first processed using 10x Genomics' official single-cell sample processing software, CellRanger v7.0.1, for single-cell barcoding, genome alignment, and counting. The reference genome was the human genome GRCh38-2020. The genome alignment file and single-cell barcode file (barcodes.tsv) generated in this step were used as input files for subsequent HLA-Splitter analysis.

[0065] Subsequently, the HLA-Splitter algorithm was used. This algorithm identifies the sample origin of each cell by calculating the expression characteristics of HLA genes in the single-cell sequencing sequence, thereby achieving sequencing sample splitting. The HLA-Splitter algorithm takes the Fastq sequence file obtained from the single-cell sequencing (belonging to second-generation sequencing data) and the HLA genotype information of the five loci (HLA-A, HLA-B, HLA-C, HLA-DRB1, and HLA-DQB1) corresponding to each of the 40 samples as input. The algorithm outputs the sample label assigned to each single cell and compares it with the actual sample of the cell. Cells with correctly assigned sample labels are marked as "Correct", cells with incorrectly assigned sample labels are marked as "Incorrect", and cells without assigned sample labels are marked as "Undefined".

[0066] In this embodiment, t-SNE unsupervised clustering of single cells is performed by comparing the single-cell HLA gene expression matrix. Figure 3 As shown, Figure 3 (a) shows the actual sample label of the cells. Figure 3 (b) shows the sample splitting results of the HLA-Splitter. The two figures are highly similar in cell clustering and sample distribution, intuitively verifying the accuracy of the method. Furthermore, as... Figure 4 As shown, the sample enrichment score heatmap with sample labels further demonstrates that the score exhibits high specificity in cells from different sample sources.

[0067] In the embodiments of this application, the statistical results are shown in Table 1 and Figure 5 As shown, for pooled simulation data with fewer than 50 cells per sample, the splitting accuracy (correct cell ratio) of HLA-Splitter is relatively low. However, as the cell sample size increases, the splitting accuracy of HLA-Splitter gradually increases and stabilizes at around 93%. This clearly demonstrates that HLA-Splitter can effectively split small-cell samples with an average cell count of at least 50 cells.

[0068] Table 1 Splitting results of 40 PBMC samples by HLA-Splitter algorithm (based on different cell sampling amounts)

[0069]

[0070] Further, to further evaluate the performance of HLA-Splitter, the present application compared the computing time required by it and several published genotype-based single-cell deconvolution algorithms (such as Demuxlet, Vireo and Souporcell) in processing mixed sample simulation data, as shown in Figure 6 For different sizes of simulation data sets, the computing time required by HLA-Splitter is significantly better than that of existing algorithms. Specifically, the computing time required by HLA-Splitter to complete the task is 21 seconds, 29 seconds, 47 seconds, 1 minute 45 seconds, 9 minutes 39 seconds and 19 minutes 18 seconds, which is much less than the computing time required by existing algorithms (for example, Vireo with genotype), which takes 9 minutes 30 seconds, 13 minutes 30 seconds, 21 minutes 30 seconds, 37 minutes, 3 hours and 6 hours 18 minutes on the corresponding data sets. This shows that HLA-Splitter significantly shortens the computing time (more than 19 times faster than the closest existing algorithm Vireo), demonstrating its excellent scalability and efficiency in processing large-scale research data.

[0071] In summary, the present embodiment comprehensively tests the performance of the single-cell sample splitting method based on HLA genotype using a simulation data set containing 40 samples. The results show that the present application has high demultiplexing accuracy (stably around 93%) and excellent computing efficiency (processing time is significantly shorter than that of existing algorithms), highlighting its robustness and scalability in processing low-cell input clinical samples.

[0072] Example Two: Single-cell mixed sample sequencing and sample splitting of 30 peripheral blood mononuclear cells (PBMCs)

[0073] Please refer to Figures 7 to 8 , the present embodiment provides the complete process and splitting results of single-cell mixed sample sequencing of peripheral blood mononuclear cells (PBMCs) of 30 healthy people. Compared with published genotype-based single-cell deconvolution algorithms, the present embodiment verifies that in actual single-cell mixed sample experiments, each sample only needs 200-300 cells to be successfully sequenced and split, which greatly expands the application potential of the present method in clinical samples with scarce cell amounts, specifically including the following steps:

[0074] 1. Nanopore-based HLA genotyping

[0075] In this embodiment, nanopore sequencing technology is used for HLA genotyping. The nanopore sequencing kit (PY-DNA101) and PolyseqOne sequencer from Beijing Piyao Biotechnology Co., Ltd. are used in the experiment. The specific process includes: extracting nucleic acids from donor peripheral blood cells, enriching and amplifying HLA-related sequences, constructing a nanopore sequencing library (including adding adenine bases at the end and connecting sequencing adapters), and adding sample-specific nucleic acid tags, thereby realizing nanopore sequencing of 30 samples at a time (further reducing sequencing costs).

[0076] As shown in Table 2, this embodiment compares the HLA typing results based on nanopore sequencing technology (Nanopore) with the results based on NGS genotyping (SBT, as the gold standard method). The accuracy of HLA typing based on nanopore sequencing is 100% at two-bit precision and 96.7-100% at four-bit precision. This embodiment verifies the high accuracy of nanopore sequencing in HLA typing and makes it a reliable input source for HLA-Splitter.

[0077] Table 2. Comparison of consistency of HLA typing results based on SBT gold standard method and nanopore sequencing

[0078]

[0079] 2. 30 cases of peripheral blood mononuclear cell (PBMC) single cell mixed sample sequencing

[0080] In this embodiment, the frozen PBMC cell sample is first thawed in a 37°C water bath; then, the cells are centrifuged to remove the original preservation liquid, and resuspended with complete culture medium (e.g., RPMI 1640) supplemented with 10% (W / V) fetal bovine serum (FBS). Specifically, 30 samples are pooled into one centrifuge tube according to the same cell ratio (i.e., each sample provides approximately equal number of cells). After staining the cells with AOPI staining solution, the proportion and concentration of viable cells are measured to ensure that the viable cell rate is >85% and the cell clumping rate is <10%. Finally, the cell concentration is adjusted to about 500 per microliter.

[0081] In this embodiment, single-cell sequencing was performed according to the standard single-cell sequencing protocol provided by Chromium Single Cell 5' Reagent Kits (10x Genomics), and the target number of cells for sequencing was 20,000 cells. The 5' gene expression library was prepared using the single-cell library construction kit (10x Genomics Chromium Single Cell 5' Library PrepKit). Specifically, the single-cell gene expression library was quality controlled using the Qubit dsDNA HS assay kit (Life Technologies, no. Q32851) and the Agilent 2100 Bioanalyzer system. Finally, Illumina NovaSeq S4 platform was used to complete the second-generation sequencing.

[0082] 3. Single-cell sequencing data splitting based on HLA genotype

[0083] In this embodiment, the single-cell sequencing data obtained in this embodiment was first processed using the single-cell sample processing software Cell Ranger v7.0.1 officially provided by 10x Genomics for single-cell barcode processing, genome alignment and counting; then the HLA-Splitter algorithm of the present application was used, and the second-generation sequencing data file obtained by the single-cell sequencing and the HLA genotype information corresponding to each of the 30 donors were used as input. The output result of the algorithm was the sample label to which each single cell was assigned, and the cells that failed to match successfully were marked as "Undefined".

[0084] 4. Accuracy evaluation of single-cell splitting results

[0085] In this embodiment, the statistical results are shown in Table 3, the total number of cells split and successfully assigned to samples by HLA-Splitter was 12565, accounting for 95.2% of the number of cells obtained by sequencing (13195), and the average proportion of cells in each sample was 3.3%, which was evenly distributed, corresponding to the approximately same proportion of cells in the 30 samples.

[0086] Table 3 Splitting results of 30 PBMC samples by HLA-Splitter algorithm

[0087]

[0088] In this embodiment, in order to further evaluate the accuracy of the single-cell sample splitting algorithm based on HLA genes, and compare with the existing mature method, this embodiment uses the published single-cell demixing algorithm (Demuxlet, which has been widely verified for its accuracy and reliability, and can be used as a comparison reference) to perform sample splitting on the 30 single-cell mixed sequencing data obtained in this embodiment, in order to evaluate the accuracy of the sample splitting results of HLA-Splitter. As shown in Figure 7 , Figure 7 , in (a) of the figure, the sample splitting results of Demuxlet are shown, Figure 7 , in (b) of the figure, the sample splitting results of HLA-Splitter are shown. By comparing the distribution of cells from different samples on the t-SNE two-dimensional graph, this embodiment intuitively demonstrates the high similarity of the two methods in cell clustering and sample distribution, and intuitively verifies the accuracy of the method. Cells belonging to the same sample are clustered together, which indicates that cells from the same sample have high similarity in HLA gene expression, and HLA gene expression of cells from different samples has significant difference. As shown in Figure 8 , by further calculating the proportion of cells consistent with the sample splitting results of the two methods in each sample, the splitting results of the two methods have high consistency, and the average proportion of the same cells is 94.86%, which indicates that the single-cell sample splitting algorithm based on HLA genes proposed in this application has high consistency with the existing SNP-based method, thereby verifying the reliability of the present application.

[0089] In summary, this embodiment uses 30 PBMC samples for single-cell sequencing with equal proportion mixing, and comprehensively tests the performance of the single-cell sample splitting method based on HLA genotypes. The results fully verify the reliability of the present application, and demonstrate its superiority in handling high-throughput, low-cell clinical samples.

[0090] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present application and not to limit, although the technical solutions of the present application are modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered in the scope of the claims of the present application.

Claims

1. A method for single-cell pooled sequencing and sample splitting based on HLA genes, characterized in that, Includes the following steps: Obtain HLA genotyping information for each sample to be mixed; After the samples to be mixed are combined, single-cell sequencing is performed to obtain the single-cell sequencing results of the mixed samples; Based on the HLA genotyping information, the expression characteristics of HLA genes in the single-cell sequencing sequence are calculated to identify the sample origin of each cell and complete the sequencing sample splitting.

2. The method according to claim 1, characterized in that, The HLA genotyping information is obtained through one or more genotyping technologies, including but not limited to: polymerase chain reaction sequencing, gene chips, high-throughput sequencing, or clinical test reports.

3. The method according to claim 1 or 2, characterized in that, The samples to be mixed include clinical samples with low cell counts, including biopsy samples, bronchoalveolar lavage fluid, cerebrospinal fluid, urine, organoid samples, and rare cell subpopulation samples.

4. The method according to claim 1, characterized in that, The sample origin of each cell is identified by calculating the expression characteristics of HLA genes in the single-cell sequencing sequence, including the following steps: Based on the HLA genotype information of the samples to be mixed, a reference database containing the corresponding HLA genes is constructed. The single-cell sequencing results are compared with the HLA gene reference database to generate a matrix reflecting the expression count of each HLA gene in each cell; The enrichment score of each cell for each sample to be mixed is calculated based on the HLA gene expression counting matrix, and the sample is allocated to each cell based on the sample enrichment score. The sample labels of each cell are corrected and modified based on a local neighborhood majority voting strategy.

5. The method according to claim 4, characterized in that, The construction of the HLA gene reference database includes the following steps: First, based on the HLA genotype information of the samples, the corresponding gene coding sequences are extracted from public databases; Secondly, the gene coding sequence is subjected to standardized preprocessing; Finally, an indexing algorithm was used to construct a sample-specific HLA gene reference database.

6. The method according to claim 5, characterized in that, The HLA gene reference database is a graph structure used for sequence alignment and assembly, capable of storing and representing HLA allele sequences.

7. The method according to claim 4, characterized in that, The process of comparing the single-cell sequencing results with the HLA gene reference database and generating an HLA gene expression count matrix includes the following steps: First, single-cell sequencing sequences on chromosome 6 were extracted from single-cell sequencing results; Secondly, the above single-cell sequencing sequences are compared with the HLA gene reference database using a pseudo-alignment algorithm to generate an HLA gene expression count matrix.

8. The method according to claim 7, characterized in that, The pseudo-alignment algorithm decomposes single-cell sequencing reads into k-mers short sequences instead of performing global base-by-base alignment.

9. The method according to claim 4, characterized in that, Calculating the sample enrichment score to achieve sample allocation includes the following steps: First, the HLA gene expression count matrix is ​​standardized and scaled according to the expression values ​​of each allele; Next, based on the scaled expression levels of each HLA allele in each cell, and combined with the homozygous or heterozygous status of the alleles, the enrichment score of each cell for each sample to be mixed was calculated. Finally, the label with the highest sample enrichment score is selected as the sample label for that cell.

10. The method according to claim 4, characterized in that, The sample label correction and modification process uses a dimensionality reduction clustering algorithm to process the scaled HLA gene expression count matrix and generate a two-dimensional embedding graph. Based on the local neighborhood majority voting strategy, the sample label with the highest frequency among the neighboring nodes of each cell is found, and the sample label of that cell is corrected and modified.

Citation Information

Patent Citations

  • Method and device for splitting sequencing data of single-cell mixed sample

    CN117079714A

  • Guided analysis of single cell sequencing data using bulk sequencing data

    US20230061214A1