Single cell mixed sample sequencing and sample splitting method based on HLA (human leukocyte antigen) gene
By using a single-cell pooled sequencing method based on HLA genes, the problems of large cell loss, complex operation, and high computational cost in existing technologies have been solved. This method achieves efficient and accurate splitting of small-cell clinical samples and expands the application potential of single-cell sequencing.
Patent Information
- Application Number
- CN202511509987.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing single-cell pooled sequencing methods suffer from problems such as large cell loss, complex operation, limited throughput, and high computational cost when processing clinical samples with sparse cell counts, making them difficult to effectively apply to single-cell pooled sequencing of clinical samples with small cell counts.
A single-cell pooled sequencing and sample splitting method based on the human leukocyte antigen (HLA) gene was adopted. By obtaining the HLA genotyping information of each sample to be pooled, single-cell sequencing was performed after pooling, and the sample origin of each cell was identified based on the HLA gene expression characteristics, thus achieving sample splitting.
It overcomes the technical barrier that prevents single-cell sequencing of clinical samples with small cell counts, improves the utilization rate of precious samples, reduces computational consumption and sequencing costs, enhances accuracy and sensitivity, and expands the feasibility of large-scale studies.
Smart Images

Figure CN120998307A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of single-cell sequencing technology, and in particular to a method for single-cell pooled sequencing and sample splitting based on the human leukocyte antigen (HLA) gene. Background Technology
[0002] Single-cell transcriptome sequencing technology has been widely used in various basic and clinical research fields in life sciences, playing a crucial role in understanding the clinical detection and pathological mechanisms of many major diseases, including malignant tumors, organ transplantation, autoimmune diseases, and hematological disorders. However, many clinical samples, such as biopsy samples and tissue fluids or body fluids from specific organs (e.g., bronchoalveolar lavage fluid, cerebrospinal fluid, urine), typically contain a low cell count, making it difficult to meet the minimum cell load requirement (usually >10,000 cells per sample) for standard single-cell transcriptome sequencing. To address this issue, one strategy is to pool multiple such samples to achieve the total cell count required for sequencing.
[0003] Traditional single-cell pooled sequencing methods typically rely on exogenous sample markers for data splitting. These markers (such as sample-specific oligonucleotide sequences) are anchored to the cell or nuclear membrane via antibodies, liposomes, or click chemistry. By detecting these exogenous markers, the original sample origin of each cell can be identified. However, these methods have limitations such as complex labeling procedures, time-consuming operations, and the potential for cell loss of up to approximately 50% during multiple rounds of centrifugation and washing, making them unsuitable for precious clinical samples with limited cell numbers.
[0004] Another type of single-cell pooled sequencing method does not require exogenous sample markers and achieves sample splitting of single-cell data through genotype-based unmixing algorithms (such as Demuxlet, Vireo, and Souporcell). These methods mainly rely on reading and identifying single nucleotide polymorphisms from scRNA-seq data, but face two key obstacles.
[0005] First, to obtain a sufficient number of informative single nucleotide mutations, whole-genome sequencing and alignment with the entire genome or exons are usually required, which imposes a huge computational burden in terms of computation time and memory requirements.
[0006] Secondly, the inherent sparsity of single-cell transcriptome sequencing data, particularly data from the 3' end protocol, leads to low and uneven coverage of single nucleotide mutation sites in the transcriptome. This deficiency reduces the ability to classify cells with low messenger RNA content, resulting in a high proportion of unassigned cells and impairing the usability of valuable samples.
[0007] In summary, although existing single-cell pooled sequencing and splitting methods can achieve accurate sample splitting, they generally suffer from limitations such as large cell loss, complex operation, limited throughput, and high computational cost, making them difficult to effectively apply to single-cell pooled sequencing of clinical samples with small cell counts. Summary of the Invention
[0008] This invention addresses the problems of significant cell loss, operational complexity, limited throughput, and high computational cost associated with existing single-cell pooled sequencing technologies when processing clinical samples with sparse cell counts. It proposes a single-cell pooled sequencing and sample splitting method based on the human leukocyte antigen (HLA) gene, comprising the following steps:
[0009] Obtain the HLA genotyping information for each sample to be mixed;
[0010] The samples to be mixed are then subjected to single-cell sequencing to obtain the single-cell sequencing results of the mixed samples.
[0011] Based on the HLA genotyping information, the expression characteristics of HLA genes in the single-cell sequencing sequence are calculated to identify the sample origin of each cell and complete the sequencing sample splitting.
[0012] The beneficial effects of this invention are as follows:
[0013] First, it overcomes the technical barrier of single-cell sequencing for clinical samples with small cell counts, greatly improving the utilization rate of these precious samples. Second, it eliminates the need for complex whole-genome single nucleotide polymorphism (SNP) detection, significantly reducing the computational cost of sample splitting through an efficient HLA gene sequence pseudo-alignment algorithm. Third, by focusing on the most polymorphic HLA regions in the human genome, it maximizes the ability to obtain genetic information from low-coverage cells, thereby improving accuracy and sensitivity. Finally, pooling effectively reduces the sequencing cost of individual samples, making large-scale studies more feasible. Attached Figure Description
[0014] Figure 1 This is a schematic diagram illustrating the principle of a single-cell pooled sequencing method based on HLA genes.
[0015] Figure 2 A schematic diagram illustrating the core principle of the HLA-Splitter sample splitting algorithm based on the HLA gene;
[0016] Figure 3 This refers to the single-cell t-SNE clustering results based on HLA genes in Example 1. Figure 3 (a) shows the actual single-cell sample source markers. Figure 3 (b) shows the results of single-cell sample splitting based on HLA genes;
[0017] Figure 4 This is a heatmap of sample enrichment scores arranged according to sample labels in Example 1;
[0018] Figure 5 The proportion of cells split using HLA-Splitter in the mixed sample simulation data with different cell sampling numbers in Example 1;
[0019] Figure 6 The computation time required by HLA-Splitter and several published genotype-based single-cell unmixing algorithms in Example 1 for processing mixed simulation data;
[0020] Figure 7 This refers to the single-cell t-SNE clustering results based on HLA genes in Example 2. Figure 7 In the middle (a), the sample splitting results of the published algorithm Demuxlet are shown. Figure 7 (b) shows the sample splitting results of the HLA-Splitter;
[0021] Figure 8 This is a heatmap showing the proportion of identical cells in each sample's total cells obtained from the two single-cell unmixing algorithms, HLA-Splitter and Demuxlet, in Example 2. Detailed Implementation
[0022] The following detailed examples illustrate the implementation methods of this application, aiming to help those skilled in the art to better understand the advantages and effects of this application, rather than limiting the scope of protection of this invention. The technical solutions of this application can be modified or changed in various ways based on different viewpoints and applications, and all modifications or changes should be covered within the scope of the claims of this invention, as long as they do not depart from the spirit and scope of this application. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0023] This application proposes a single-cell pooled sequencing and sample splitting method based on HLA genes. This method can quickly and accurately split single-cell transcriptome sequencing data from different samples by identifying endogenous HLA genotypes in cells, without the need for additional exogenous sample labeling.
[0024] like Figure 1 As shown, this application includes the following steps:
[0025] S1: Obtain HLA genotyping information from the sample
[0026] In this application, HLA genotyping information is obtained through one or more genotyping techniques, including but not limited to: polymerase chain reaction sequencing, gene chips, high-throughput sequencing, or obtained through clinical HLA genotyping reports from sample owners (such as organ transplant samples), or through efficient methods such as nanopore sequencing for rapid identification. For example, in this embodiment, by extracting nucleic acid from peripheral blood cells and constructing a nanopore sequencing library, HLA genotyping of 30 samples was completed in one go, significantly reducing sequencing time and cost compared to traditional methods.
[0027] S2: Single-cell transcriptome sequencing was performed after sample mixing.
[0028] In this application, the cell samples to be pooled for sequencing can be clinical samples containing a small number of cells, such as biopsy samples, tissue fluid or body fluid samples from specific organs (e.g., urine, bronchoalveolar lavage fluid, cerebrospinal fluid, etc.), organoid samples, and samples for detecting rare cell subpopulations, with a minimum cell count requirement of only 100 cells / sample. Samples can be cryopreserved using liquid nitrogen. The number of samples pooled at one time is typically 20-40 from different individuals, with the specific number determined based on the total number of cells after pooling. Generally, it is recommended that the number of cells processed at one time not exceed 20,000.
[0029] S3: Sequencing sample splitting based on HLA genotype information of the sample.
[0030] In this application, sequencing sample splitting is achieved based on the HLA genotype information of the samples. By calculating the expression characteristics of HLA genes in single-cell sequencing sequences, the sample origin of each cell is identified, and sequencing sample splitting is completed.
[0031] Furthermore, such as Figure 2 As shown, splitting sequencing samples involves the following steps:
[0032] (1) Construction of HLA gene reference database
[0033] In this application, the construction of the HLA gene reference database begins by extracting the corresponding gene coding sequence (CDS) from the Immunogenetic Database for Genetics (IMGT / HLA) based on the HLA genotype information (four-digit precision, e.g., HLA-A*02:01) of each sample. To ensure accuracy, the gene coding sequences undergo normalization preprocessing to remove any redundant information. Using the indexing function of the gene expression quantification software kallisto, a sample-specific HLA gene reference database is constructed. This database is a graph structure (colored de Bruijn graph) used for sequence alignment and assembly, where nodes represent short sequences of length k (k-mers), and edges represent the overlap between k-mers. This graph structure can efficiently store and represent a large number of HLA allele sequences, providing a rapid retrieval framework for subsequent pseudo-alignments.
[0034] (2) Pseudo-alignment and counting matrix generation
[0035] In this application, the pseudo-alignment and generation of the counting matrix first use Samtools software to extract single-cell sequencing sequences from the Single-Cell Sequencing Results Alignment Map (BAM) file on chromosome 6. Further, the BAM file is generated by running the Cellranger software (official software of the 10X single-cell sequencing platform) with default parameters, which already contains the mapping information of the single-cell sequencing data to the human genome (GRCh38). Next, these sequences are read and pseudo-aligned against the HLA gene reference database generated in the previous step using the Kallisto-bustools single-cell data analysis pipeline, thereby generating a counting matrix (MEX format) reflecting the expression of each HLA gene in each cell.
[0036] In this application, unlike traditional genome alignment, pseudo-alignment does not perform base-by-base global alignment. Instead, it breaks down each sequencing read into a series of short sequences (k-mers) of length k. These k-mers are used to traverse a pre-constructed colorized de Bruijn graph of an HLA gene reference database. When a k-mer finds a match in the graph, it inherits the "color" of the corresponding node, which represents the HLA allele it belongs to. By integrating the color information of all k-mers from the same read, Kallisto-bustools can quickly and accurately determine which HLA allele(s) the read most likely originates from, greatly improving alignment speed and accuracy.
[0037] (3) Calculate the enrichment score of each cell for each sample to be mixed based on the HLA gene expression counting matrix, and allocate samples to each cell based on the enrichment score.
[0038] In this application, let H be the original HLA gene expression counting matrix, where M is the total number of cells and N is the total number of alleles. This represents the original expression value of allele j in cell i. Then, the normalized HLA gene expression matrix is obtained using the following formula. :
[0039]
[0040] Subsequently, to balance the differences in expression levels among different HLA alleles, the expression levels of different HLA alleles were adjusted. Each allele The standard deviation is calculated using only its positive expression values. The matrix is then scaled to obtain the final scaled HLA gene expression count matrix used for calculation. .
[0041]
[0042] Specifically, in the probabilistic sample allocation based on HLA-Score, let S be a... The enrichment score matrix is given by $\mathbf{k}$, where $K$ is the number of samples to be mixed. The enrichment score of cell $i$ for each sample $k$ to be mixed is $\mathbf{k}$. The calculation formula is as follows:
[0043]
[0044] Specifically, for each sample k and each HLA gene locus type (including HLA-A, B C DRB1, DQB1 and DPB1), where The HLA gene locus type of sample k to be mixed. The set of HLA alleles at that location, It is the homozygous / heterozygous coefficient: when the sample k to be mixed is at the gene locus When the molecule is homozygous, When it is a heterozygote, .
[0045] Specifically, the preliminary predicted sample label for each cell Samples are assigned to achieve the maximum enrichment score, but this maximum score must exceed an absolute threshold by default. (Default is 0.1), otherwise it is marked as "Undefined", as shown in the formula below:
[0046]
[0047] in It is the predicted sample label of cell i. It is the enrichment score of cell i for sample k. It is the sample label corresponding to the maximum enrichment score in cell i. It is the absolute threshold of the enrichment score.
[0048] (4) Sample label error correction and modification
[0049] In this application, to further improve the confidence of the assignment, a dimensionality reduction clustering algorithm is used to scale the HLA gene expression count matrix. Perform t-SNE dimensionality reduction to generate a two-dimensional embedding graph. For each cell i (excluding those initially labeled "Undefined"), based on a local neighborhood majority voting strategy, identify the most frequent sample label among each cell's 100 nearest neighbors in the two-dimensional t-SNE projection, and correct and refine the sample label for that cell. Let... The label of the sample that appears most frequently among these neighboring nodes. represent The frequency of occurrence in these adjacent nodes. The final corrected label for cell i. Determined according to the following formula.
[0050]
[0051] Where τ is the confidence threshold for the majority label frequency in the neighborhood (default 0.5).
[0052] The HLA gene was selected as the endogenous gene reference for sample splitting in this invention, mainly for the following reasons:
[0053] (1) HLA genes are highly polymorphic regions in the human genome, and there is a very high degree of HLA genotypic difference between different individuals. This polymorphism makes HLA genotype an ideal "fingerprint" for individual identification, which can effectively distinguish cells from different individuals;
[0054] (2) HLA genes are expressed in the vast majority of nucleated cells, which are usually the target of single-cell sequencing;
[0055] (3) As an endogenous gene of the cell itself, it does not require any exogenous labeling operation, thus completely avoiding cell loss that may occur during the exogenous labeling process;
[0056] (4) In some clinical situations (such as organ transplantation), the subject’s HLA typing information may already exist. Even if additional sequencing is required, it can be completed efficiently using fast and low-cost technologies such as nanopore sequencing.
[0057] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0058] Example 1: Testing method performance using a simulated single-cell pooled sequencing dataset
[0059] Please see Figures 3 to 6 This embodiment aims to test the reliability and robustness of the HLA genotype-based single-cell sequencing data splitting method, and evaluate the accuracy and computational efficiency of the method by comparing the sample splitting results of a statistically simulated dataset with those of actual sample sources. The method includes the following steps:
[0060] 1. Construction of a simulated single-cell pooled sequencing dataset
[0061] In this embodiment of the application, the simulated dataset contains single-cell transcriptome data from 40 peripheral blood mononuclear cell (PBMC) samples from the public database Gene Expression Omnibus (accession number: GSE226602). The HLA genotypes of these samples were pre-obtained from the single-cell transcriptome sequencing data of each donor using the high-resolution HLA typing software arcasHLA (arcasHLA and related typing procedures are well-known in the art).
[0062] Furthermore, to construct a mixed-sample simulation dataset, 10-1000 cells were extracted from each of the 40 single-cell sequencing data points using non-repeating random sampling, thus simulating a dataset with a total cell count of 400-40000 cells. Specifically, the random sampling method used Python's `random.sample()` function to randomly sample cell barcodes. Then, the `subset-BAM` tool was used to filter corresponding cell gene expression information from the 10x Genomics genome alignment files. Finally, the `bamtofastq` command in the 10x Genomics Cell Ranger was used to convert the BAM file into a Fastq sequence file.
[0063] 2. HLA genotype-based splitting of single-cell sequencing data
[0064] In this embodiment, the simulated single-cell pooled sequencing data was first processed using 10x Genomics' official single-cell sample processing software, CellRanger v7.0.1, for single-cell barcoding, genome alignment, and counting. The reference genome was the human genome GRCh38-2020. The genome alignment file and single-cell barcode file (barcodes.tsv) generated in this step were used as input files for subsequent HLA-Splitter analysis.
[0065] Subsequently, the HLA-Splitter algorithm was used. This algorithm identifies the sample origin of each cell by calculating the expression characteristics of HLA genes in the single-cell sequencing sequence, thereby achieving sequencing sample splitting. The HLA-Splitter algorithm takes the Fastq sequence file obtained from the single-cell sequencing (belonging to second-generation sequencing data) and the HLA genotype information of the five loci (HLA-A, HLA-B, HLA-C, HLA-DRB1, and HLA-DQB1) corresponding to each of the 40 samples as input. The algorithm outputs the sample label assigned to each single cell and compares it with the actual sample of the cell. Cells with correctly assigned sample labels are marked as "Correct", cells with incorrectly assigned sample labels are marked as "Incorrect", and cells without assigned sample labels are marked as "Undefined".
[0066] In this embodiment, t-SNE unsupervised clustering of single cells is performed by comparing the single-cell HLA gene expression matrix. Figure 3 As shown, Figure 3 (a) shows the actual sample label of the cells. Figure 3 (b) shows the sample splitting results of the HLA-Splitter. The two figures are highly similar in cell clustering and sample distribution, intuitively verifying the accuracy of the method. Furthermore, as... Figure 4 As shown, the sample enrichment score heatmap with sample labels further demonstrates that the score exhibits high specificity in cells from different sample sources.
[0067] In the embodiments of this application, the statistical results are shown in Table 1 and Figure 5 As shown, for pooled simulation data with fewer than 50 cells per sample, the splitting accuracy (correct cell ratio) of HLA-Splitter is relatively low. However, as the cell sample size increases, the splitting accuracy of HLA-Splitter gradually increases and stabilizes at around 93%. This clearly demonstrates that HLA-Splitter can effectively split small-cell samples with an average cell count of at least 50 cells.
[0068] Table 1. Splitting results of 40 PBMC samples using the HLA-Splitter algorithm (based on different cell sampling volumes)
[0069]
[0070] Furthermore, to further evaluate the performance of HLA-Splitter, this application compares its computation time with several published genotype-based single-cell unmixing algorithms (such as Demuxlet, Vireo, and Souporcell) in processing mixed simulation data, such as... Figure 6 As shown, for simulated datasets of different sizes, HLA-Splitter significantly outperforms existing algorithms in terms of computation time. Specifically, the computation times required for HLA-Splitter to complete the task are 21 seconds, 29 seconds, 47 seconds, 1 minute 45 seconds, 9 minutes 39 seconds, and 19 minutes 18 seconds, respectively, which are much shorter than the computation times required by existing algorithms (e.g., Vireo with genotype), which took 9 minutes 30 seconds, 13 minutes 30 seconds, 21 minutes 30 seconds, 37 minutes, 3 hours, and 6 hours 18 minutes on the corresponding datasets. This indicates that HLA-Splitter significantly reduces computation time (more than 19 times faster than the closest existing algorithm, Vireo), demonstrating its superior scalability and efficiency in processing large-scale research data.
[0071] In summary, this embodiment comprehensively tested the performance of the HLA genotype-based single-cell sample splitting method using a simulated dataset containing 40 samples. The results show that the present invention has high de-reuse accuracy (stable at around 93%) and excellent computational efficiency (processing time significantly shorter than existing algorithms), highlighting its robustness and scalability in handling clinical samples with low cell input.
[0072] Example 2: Single-cell pooled sequencing and sample splitting of 30 peripheral blood mononuclear cells (PBMCs)
[0073] Please see Figures 7 to 8 This embodiment provides a complete workflow and splitting results for single-cell pooled sequencing of peripheral blood mononuclear cells (PBMCs) from 30 healthy individuals. Compared with published genotype-based single-cell unmixing algorithms, this embodiment verifies that in actual single-cell pooled experiments, only 200-300 cells per sample are needed for successful sequencing and splitting. This greatly expands the application potential of this method in clinical samples with scarce cell numbers. The specific steps include:
[0074] 1. HLA genotyping based on nanopores
[0075] This embodiment utilizes nanopore sequencing technology for HLA genotyping. The experiment employed the PY-DNA101 nanopore sequencing kit and PolyseqOne sequencer from Beijing Puyi Biotechnology Co., Ltd. The specific procedure included: extracting nucleic acids from donor peripheral blood cells, enriching and amplifying HLA-related sequences, constructing a nanopore sequencing library (including adding adenine bases to the ends and ligating sequencing adapters), and adding sample-specific nucleic acid tags, thereby enabling nanopore sequencing of 30 samples in a single run (further reducing sequencing costs).
[0076] As shown in Table 2, this example compares the HLA genotyping results based on nanopore sequencing technology with those based on NGS genotyping (SBT, the gold standard method). Nanopore sequencing-based HLA genotyping achieved 100% accuracy at two-digit precision and 96.7-100% accuracy at four-digit precision. This example validates the high accuracy of nanopore sequencing in HLA genotyping and makes it a reliable input source for HLA-Splitter.
[0077] Table 2. Comparison of HLA genotyping results between the SBT gold standard method and nanopore sequencing.
[0078]
[0079] 2. Single-cell pooled sequencing of 30 peripheral blood mononuclear cells (PBMCs)
[0080] In this embodiment, the frozen PBMC cell samples were first thawed in a 37°C water bath. Subsequently, the cells were centrifuged to remove the original storage liquid and resuspended in complete culture medium (e.g., RPMI 1640) supplemented with 10% (w / v) fetal bovine serum (FBS). Specifically, 30 samples were pooled into a single centrifuge tube with the same cell ratio (i.e., each sample provided approximately the same number of cells). After staining the cells with AOPI staining solution, the viable cell percentage and viable cell concentration were measured to ensure a viable cell rate >85% and a cell clumping rate <10%. Finally, the cell concentration was adjusted to approximately 500 cells per microliter.
[0081] In this embodiment, single-cell sequencing was performed according to the standard single-cell sequencing protocol provided by the Chromium Single Cell 5' Reagent Kits (10xGenomics), with a target of 20,000 cells per sequencing run. 5' gene expression libraries were prepared using the single-cell library construction kit (10x Genomics Chromium Single Cell 5' Library PrepKit). Specifically, the single-cell gene expression libraries were quality controlled using the Qubit dsDNA HS analysis kit (Life Technologies, no. Q32851) and an Agilent 2100 bioanalyzer system. Finally, next-generation sequencing was performed using the Illumina NovaSeq S4 platform.
[0082] 3. HLA genotype-based splitting of single-cell sequencing data
[0083] In this embodiment, the single-cell sequencing data obtained in this embodiment is first processed using the 10x Genomics official single-cell sample processing software Cell Ranger v7.0.1 for single-cell barcoding, genome alignment, and counting. Then, the HLA-Splitter algorithm of this invention is used, with the second-generation sequencing data file obtained from the single-cell sequencing and the HLA genotype information corresponding to each of the 30 donors as input. The algorithm outputs a sample label assigned to each single cell; cells that fail to match are marked as "Undefined".
[0084] 4. Accuracy assessment of single-cell dissection results
[0085] In this embodiment, the statistical results are shown in Table 3. The total number of cells successfully split and assigned to the samples by HLA-Splitter was 12,565, accounting for 95.2% of the number of cells obtained from sequencing (13,195). The average cell proportion of each sample was 3.3%, which is a balanced distribution and corresponds to roughly the same cell proportion in the 30 samples.
[0086] Table 3. Splitting results of 30 PBMC samples using the HLA-Splitter algorithm.
[0087]
[0088] In this embodiment, to further evaluate the accuracy of the HLA gene-based single-cell sample splitting algorithm and compare it with existing mature methods, this embodiment uses the published single-cell unmixing algorithm (Demuxlet, whose accuracy and reliability have been widely verified and can be used as a comparative reference) to split the 30 single-cell mixed sequencing data obtained in this embodiment, in order to evaluate the accuracy of the HLA-Splitter sample splitting results. Figure 7 As shown, Figure 7 (a) shows the sample splitting results of Demuxlet. Figure 7 Figure (b) shows the sample splitting results of the HLA-Splitter. This embodiment visually demonstrates the high similarity between the two methods in cell clustering and sample distribution by comparing the distribution of cells from different sample sources on the t-SNE two-dimensional plot, thus intuitively verifying the accuracy of the methods. Cells belonging to the same sample are clustered together, indicating that cells from the same sample source have high similarity in HLA gene expression, and that HLA gene expression of cells from different sample sources has significant differences. Figure 8 As shown in the figure, this embodiment further calculated the proportion of cells in each sample whose splitting results were consistent with the two methods. The splitting results of the two methods were highly consistent, with an average proportion of identical cells of 94.86%. This indicates that the single-cell sample splitting algorithm based on HLA gene proposed in this application is highly consistent with the existing SNP-based method, thus strongly verifying the reliability of the present invention.
[0089] In summary, this embodiment used 30 PBMC samples for proportionally pooled single-cell sequencing to comprehensively test the performance of the HLA genotype-based single-cell sample splitting method. The results fully validate the reliability of the invention and demonstrate its superiority in processing high-throughput, low-cell-volume clinical samples.
[0090] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although modifications or equivalent substitutions may be made to the technical solutions of the present invention with reference to preferred embodiments, they shall not depart from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions shall be covered within the scope of the claims of the present invention.
Claims
1. A method for single-cell pooled sequencing and sample splitting based on HLA genes, characterized in that, Includes the following steps: Obtain HLA genotyping information for each sample to be mixed; After the samples to be mixed are combined, single-cell sequencing is performed to obtain the single-cell sequencing results of the mixed samples; Based on the HLA genotyping information, the expression characteristics of HLA genes in the single-cell sequencing sequence are calculated to identify the sample origin of each cell and complete the sequencing sample splitting.
2. The method according to claim 1, characterized in that, The HLA genotyping information is obtained through one or more genotyping technologies, including but not limited to: polymerase chain reaction sequencing, gene chips, high-throughput sequencing, or clinical test reports.
3. The method according to claim 1 or 2, characterized in that, The samples to be mixed include clinical samples with low cell counts, including biopsy samples, bronchoalveolar lavage fluid, cerebrospinal fluid, urine, organoid samples, and rare cell subpopulation samples.
4. The method according to claim 1, characterized in that, The sample origin of each cell is identified by calculating the expression characteristics of HLA genes in the single-cell sequencing sequence, including the following steps: Based on the HLA genotype information of the samples to be mixed, a reference database containing the corresponding HLA genes is constructed. The single-cell sequencing results are compared with the HLA gene reference database to generate a matrix reflecting the expression count of each HLA gene in each cell; The enrichment score of each cell for each sample to be mixed is calculated based on the HLA gene expression counting matrix, and the sample is allocated to each cell based on the sample enrichment score. The sample labels of each cell are corrected and modified based on a local neighborhood majority voting strategy.
5. The method according to claim 4, characterized in that, The construction of the HLA gene reference database includes the following steps: First, based on the HLA genotype information of the samples, the corresponding gene coding sequences are extracted from public databases; Secondly, the gene coding sequence is subjected to standardized preprocessing; Finally, an indexing algorithm was used to construct a sample-specific HLA gene reference database.
6. The method according to claim 5, characterized in that, The HLA gene reference database is a graph structure used for sequence alignment and assembly, capable of storing and representing HLA allele sequences.
7. The method according to claim 4, characterized in that, The process of comparing the single-cell sequencing results with the HLA gene reference database and generating an HLA gene expression count matrix includes the following steps: First, single-cell sequencing sequences on chromosome 6 were extracted from single-cell sequencing results; Secondly, the above single-cell sequencing sequences are compared with the HLA gene reference database using a pseudo-alignment algorithm to generate an HLA gene expression count matrix.
8. The method according to claim 7, characterized in that, The pseudo-alignment algorithm decomposes single-cell sequencing reads into k-mers short sequences instead of performing global base-by-base alignment.
9. The method according to claim 4, characterized in that, Calculating the sample enrichment score to achieve sample allocation includes the following steps: First, the HLA gene expression count matrix is standardized and scaled according to the expression values of each allele; Next, based on the scaled expression levels of each HLA allele in each cell, and combined with the homozygous or heterozygous status of the alleles, the enrichment score of each cell for each sample to be mixed was calculated. Finally, the label with the highest sample enrichment score is selected as the sample label for that cell.
10. The method according to claim 4, characterized in that, The sample label correction and modification process uses a dimensionality reduction clustering algorithm to process the scaled HLA gene expression count matrix and generate a two-dimensional embedding graph. Based on the local neighborhood majority voting strategy, the sample label with the highest frequency among the neighboring nodes of each cell is found, and the sample label of that cell is corrected and modified.
Citation Information
Patent Citations
Kit for detecting expression typing and expression quantity of HLA-I / II type genes at single cell level and use method of kit
CN114525328A
Method and device for splitting sequencing data of single-cell mixed sample
CN117079714A
Method for detecting genomic abnormalities in cell
CN119343464A
Methods For The Identification, Targeting And Isolation Of Human Dendritic Cell (DC) Precursors "Pre-DC" And Their Uses Thereof
US20190324038A1
Guided analysis of single cell sequencing data using bulk sequencing data
US20230061214A1