A method for analyzing characteristics of b lymphocyte bcr based on single cell multi-omics sequencing
By using single-cell multi-omics sequencing technology, combined with host reference genomes and annotation files, the characteristics of B lymphocytes (BCRs) were analyzed. This solved the problem that existing technologies could not deeply analyze the characteristics of BCR clones, and enabled the assessment of stable BCR expression rates and clonal status of cell subsets during viral infection, providing important evidence for vaccine and drug development.
Patent Information
- Application Number
- CN202111374293.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-11-19
AI Technical Summary
Existing technologies cannot comprehensively and systematically analyze the clonal characteristics of B lymphocytes before and after viral infection, disease severity groups, and before and after treatment. They also cannot provide in-depth understanding of the stable BCR expression rate and clonal status of cell subsets under different disease states, as well as the BCR assembly and VJ gene usage preferences.
We employed a single-cell multi-omics sequencing approach, using single-cell transcriptomics and immunomic sequencing data to construct a gene expression matrix, identify BCR sequences, assess stable BCR expression rates, analyze cell subset clonal states and BCR assembly and VJ gene usage preferences, and combine the host reference genome and annotation files with Cell Ranger and Seurat software for data processing and analysis.
This study enabled in-depth analysis of the characteristics of B lymphocyte BCRs under different disease states, assessed the stable BCR expression rate and clonal status of cell subsets, elucidated the relationship between BCR assembly and VJ gene usage preference, and provided important data support for research on viral infection mechanisms and vaccine development.
Smart Images

Figure CN116153408B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of single-cell transcriptome sequencing, single-cell immunomic sequencing, viral infection and host immune response, and bioinformatics analysis, specifically to a method for analyzing the characteristics of B lymphocyte BCR based on single-cell multi-omics sequencing. Background Technology
[0002] The advent of high-throughput sequencing technology has ushered in a high-throughput era for research in various fields of biology. Traditional RNA-seq, based on large tissue samples, yields the average of all cell types within a tissue. However, biological samples and tissues are often highly complex and heterogeneous systems with diverse cell types. Against this backdrop, single-cell RNA-sequencing (scRNA-seq) emerged, first reported in 2009, marking the beginning of the single-cell era in biological research. With the convenience, maturity, and reduced cost of scRNA-seq technology, its application in various fields of biological research has flourished. The number of papers on single-cell sequencing research has increased dramatically, and the number of cells used in these studies has exploded, from a few hundred or a few thousand cells initially to hundreds of thousands or even millions of cells today.
[0003] With the rapid development of scRNA-seq technology, single-cell research has entered the era of multi-omics, leading to the emergence of numerous single-cell multi-omics sequencing technologies that are effectively utilized to address various problems in the life sciences. Examples include single-cell transcriptome sequencing, single-cell immunomic sequencing, single-cell genome sequencing, and single-cell chromatin accessibility sequencing. In particular, with the advent of in situ hybridization and specific RNA probes, single-cell spatial omics technology has developed rapidly and has been recognized as a technology of the year by several journals, thus holding great promise. Currently, single-cell transcriptome sequencing technology is relatively mature and widely used. A representative example is 10x Genomics' Chromium technology, which sequences transcripts from the 3' or 5' ends. Because it only sequences one end of the transcript, it is relatively inexpensive and has high cell throughput.
[0004] Viral infectious diseases are among the major killers threatening human health. For example, human immunodeficiency virus (HIV) infection can directly destroy the human immune system, causing AIDS (Acquired immune deficiency syndrome); rabies virus (RABV) is a neuroinfectious virus with an incidence rate close to 100%, and there are still no effective drugs or treatments after the onset of the disease.
[0005] When a virus infects the human body, it triggers an antiviral immune response, involving both innate and adaptive immune responses. Innate immunity responds earlier and has a wider range of action than adaptive immunity. Adaptive immunity is further divided into cellular immunity and humoral immunity, with B lymphocytes primarily mediating the humoral immune response. Humoral immunity is a crucial component of the human immune system. After receiving antigen stimulation, B lymphocyte receptors rapidly proliferate and differentiate. Some differentiate into effector B cells, or plasma cells, which protect the body by producing immunoglobulin antibodies to neutralize antigens. Others differentiate into memory B cells, which can rapidly differentiate into plasma cells to produce antibodies to neutralize the same antigen when it reappears. B cells are also professional antigen-presenting cells, presenting antigens to CD4+ T cells via MHC class II molecules on their surface, thereby regulating cellular immunity. Simultaneously, activated B lymphocytes secrete large amounts of cytokines, participating in inflammatory responses and immune regulation. B cell receptors (BCRs) are immunoglobulin molecules on the surface of B lymphocytes. B cell immunoglobulins include five types: IgA, IgG, IgM, IgD, and IgE, each containing two heavy chains (H) and two light chains (κ and λ). Before maturation, B lymphocytes undergo VDJ gene rearrangement in the heavy chain and VJ gene rearrangement in the light chain. Normal IgH, Igκ, and Igλ gene recombination is a relatively random process. After viral infection, B lymphocytes undergo clonal expansion in response to antigen stimulation. The proportion, state, diversity, and preference of BCR clonal expansion play a crucial role in viral clearance and understanding the mechanisms by which the body clears viruses after infection.
[0006] Today, scRNA-seq has been successfully applied to many cutting-edge biomedical fields, including the study of viral infection and host immune responses. scRNA-seq has successfully investigated the mechanisms of various viral infections and revealed the response characteristics of the host's innate and adaptive immune systems, providing valuable insights for vaccine research and clinical treatment. However, current analyses combining single-cell transcriptomics and immunomics have not yet achieved a comprehensive and systematic analysis of the detailed characteristics of BCR clonal patterns in patients with different disease states (before and after infection, disease severity grouping, and disease recovery). Therefore, it is impossible to comprehensively and deeply analyze the stable BCR expression rates and clonal states of various cell subpopulations in different disease states during the disease progression of viral infection, as well as disease severity grouping and BCR assembly and usage preferences before and after treatment. Summary of the Invention
[0007] The purpose of this invention is to address the shortcomings of existing analytical techniques by providing a method for analyzing the characteristics of B lymphocyte BCRs based on single-cell multi-omics sequencing. This method can assess and identify the expression rate of stable BCRs in subpopulations, evaluate the clonal status of each cell subpopulation in different disease states, and analyze the usage preference of BCR assembly single genes and VJ pairing genes in various disease states. This has significant implications for the study of viral infection mechanisms and the development of specific vaccines and drug therapies targeting host immune responses.
[0008] The technical solution of the present invention is as follows:
[0009] A method for analyzing B lymphocyte (BCR) characteristics based on single-cell multi-omics sequencing involves collecting samples from patients in different disease states for single-cell sequencing, followed by analysis through the following steps:
[0010] 1) Constructing single-cell gene expression matrices and cell type annotations based on single-cell transcriptome sequencing data: Download the host reference genome sequence from the corresponding reference genome sequence website, and simultaneously download the corresponding version of the reference genome annotation file (GTF file). Obtain the single-cell gene expression matrix through data alignment and gene expression quantification. Perform data dimensionality reduction and cell type clustering and annotation using single-cell analysis software.
[0011] 2) Assembling BCR sequences and identifying BCR chain constituent genes based on single-cell immunomic sequencing data: From the downloaded host reference genome sequence and gene / functional subunit annotation files, B lymphocyte BCR-related gene sequences and annotation information were specifically extracted to construct the BCR reference genome information. BCR sequences were then quantitatively assembled and BCR chain constituent genes were identified through data alignment and gene expression analysis.
[0012] 3) Identifying the stable BCR expression rate of each cell subpopulation based on expression data: A reasonable threshold is set to filter and denoise the single-cell expression data obtained in step 1), retaining only high-quality single-cell expression data. After BCR sequence quantification and identification, only BCRs with both heavy and light chain data are retained. Finally, only cells with both high-quality single-cell expression data and BCR data are retained, thus obtaining the stable BCR expression rate of each cell subpopulation, i.e., the percentage of cells with detected stable BCRs in each cell subpopulation out of the total number of cells.
[0013] 4) Assess the clonal status of cell subpopulations in different disease states: Divide the data into different disease states, and then group the data in the same disease state according to different cell types, define clonal and non-clonal states, and assess the overall clonal characteristics of different disease states and cell types.
[0014] 5) Analyze the BCR assembly and single gene usage preferences of samples under different disease states: Group different samples according to different disease states, and analyze the characteristics of heavy chain and light chain single gene usage among all samples under specific disease states to analyze the BCR assembly and single gene usage preferences under different disease states.
[0015] 6) Analyze the BCR assembly and VJ gene pairing preferences of samples under different disease states: Group different samples according to different disease states, and analyze the characteristics of heavy chain and light chain VJ pairing among all samples under specific disease states to analyze the BCR assembly and VJ gene pairing preferences under different disease states.
[0016] The requirements for the host reference genome sequence and annotation, as well as the cell type annotation, in step 1) above are as follows:
[0017] The host reference genome sequence and corresponding gene annotation information can be derived from, but are not limited to, human reference genomes hg18 (GRCh36), hg19 (GRCh37), hg20 (GRCh38), or subsets of chromosome sets compiled and maintained by databases such as UCSC, NCBI, Ensemble, Genecode, and 10x Genomics. The host reference genome sequence file should be in standard FASTA format, and the gene annotation file should be in standard GTF format. Cell type grouping can be performed using, but is not limited to, software such as Seurat, Scanpy, and Cellranger Loupe. Cell type annotation is based on classic cell surface or specifically expressed marker genes reported in the literature.
[0018] Step 2) above involves assembling the BCR sequence and identifying the genes that make up the BCR chain, wherein:
[0019] The host BCR sequence and corresponding gene annotations are extracted from the host reference genome sequence and annotation information. Extracted tags primarily include, but are not limited to: `--attribute=gene_biotype:IG_LV_gene \IG_V_gene \IG_V_pseudogene\IG_D_gene \IG_J_gene \IG_J_pseudogene \IG_C_gene \IG_C_pseudogene`. The BCR reference genome sequence file should be in standard FASTA format, and the gene annotation file should be in standard GTF format. Assembling the BCR sequence and identifying the constituent genes of the BCR chain can be done using, but is not limited to, the Cellrangervdj software.
[0020] Step 3) above combines expression data to filter unstable BCRs to identify stable expression rates, where:
[0021] Single-cell expression data underwent filtering and noise reduction to retain only high-quality single-cell expression data. Filtering criteria included, but were not limited to: cells with excessively low gene counts, cells with excessively low UMI counts, cells with excessively high mitochondrial ratios, dual-cell structures with two or more classic cell marker genes, and cell fragments without any obvious marker gene expression. After BCR sequence quantification and identification, only cells with at least one expressible heavy chain (IGH) and one expressible light chain (IGK / IGL) were retained for further analysis. Finally, the intersection of the two sets of cells was taken, retaining only cells with both high-quality expression data and BCR data for downstream analysis.
[0022] Step 4) above assesses the clonal status of each cell subset in different disease states, where:
[0023] After obtaining the BCR sequences, gene composition, and expression levels of different cells, each unique IGH(s)-IGK(s) pair or IGH(s)-IGL(s) pair is defined as a clonal type. A cell carrying a clonal type is considered cloned if it is present in at least two cells. This allows for simultaneous assessment of cases where BCR is not detected, BCR is not clonal, and BCR is clonal in a particular cell type under a specific disease state.
[0024] Step 5 above analyzes the BCR assembly and single-gene usage preferences of each sample under different disease states, where:
[0025] Normal heavy and light chain gene recombination is a relatively random process. After antigen stimulation, clonal amplification occurs, and the heavy and light chains randomly select and rearrange the V(D)J gene to form a complete functional gene. We statistically analyzed the single-gene usage frequency of the V, D, and J genes in different disease states for the heavy chain, and the single-gene usage frequency of the V and J genes in different disease states for the light chain.
[0026] Step 6 above analyzes the BCR assembly and VJ gene pairing preference of each sample under different disease states, where:
[0027] During BCR recombination, the heavy chain undergoes D and J linkage sequentially, followed by V and DJ linkage, before finally linking to the C region gene to form a complete heavy chain functional gene. The light chain, lacking the D segment, directly undergoes V and J linkage, then links to the C region to form a complete light chain functional gene. Therefore, when analyzing both the light and heavy chains simultaneously, we specifically analyze the preference for V and J gene pairings in different disease states.
[0028] This invention addresses the limitations of current scRNA-seq technology in analyzing the role of BCR in overall cell types and disease states during viral infection. It provides a method for analyzing B lymphocyte BCR characteristics based on single-cell multi-omics sequencing. This method can identify stable BCR expression rates in cell subpopulations by combining expression data, assess the clonal status of each cell subpopulation in different disease states, analyze BCR assembly and single-gene usage preferences in samples under different disease states, and analyze BCR assembly and VJ gene pair usage preferences in samples under different disease states. This has significant implications for elucidating the mechanisms of specific viral infection in hosts and for the development of specific vaccines and drugs. Attached Figure Description
[0029] Figure 1 The overall analytical process for analyzing the characteristics of B lymphocytes (BCR) after viral infection is described in this embodiment of the invention.
[0030] Figure 2 In this embodiment of the invention, expression data is combined to identify the overall stable BCR expression rate of each cell subpopulation, wherein: the left figure is a cell classification diagram, and the numbers represent different cell types; the two middle figures show whether stable BCR is detected in the cells; and the right figure shows the percentage of stable BCR signals detected in cell subpopulations of different cell types.
[0031] Figure 3 The clonal status assessment results of various cell subpopulations in different disease states in the embodiments of the present invention.
[0032] Figure 4 The embodiments of this invention present the results of a preference analysis of the BCR assembly single gene in each sample under different disease states.
[0033] Figure 5 The embodiments of the present invention use the results of preference analysis of BCR assembly VJ pairing genes for each sample under different disease states. Detailed Implementation
[0034] The following is a more detailed description of the present invention. The parameters and specific implementation details are used to explain the feasibility and implementation effects of the present invention, and do not constitute a limitation of the present invention.
[0035] This embodiment uses peripheral blood mononuclear cell (PBMC) samples from 13 COVID-19 infected patients (including 4 severe cases, 7 mild cases, and 6 patients in the recovery phase, of whom 4 had paired pre- and post-antiviral treatment samples) and 5 healthy individuals, based on 10x Genomics' Single Cell 5' Library & Gel Beadkit single-cell sequencing. Paired BCR sequencing (Single Cell V(D)J Enrichment kit) was also performed on each sample during transcriptome sequencing.
[0036] 1. Sample library construction and high-throughput single-cell sequencing based on Single Cell 5' Library & Gel Bead Kit (10x Genomics) and Chromium Single Cell A Chip Kit (10x Genomics).
[0037] Cell suspensions (300–600 viable cells / μL, counted by Countstar) were loaded into a Chromium Single-cell controller (10x Genomics) to generate water-in-oil droplets (GEMs). Briefly, single cells were suspended in phosphate-buffered saline (PBS) containing 0.04% bovine serum albumin (BSA), and cells were added to each channel of the chip, with a final recovery rate of approximately 50%. Captured cells were lysed, and the released RNA was labeled in a single GEM via reverse transcription. Reverse transcription was performed on an S1000™ thermal cycler (Bio-Rad Laboratories, Hercules, CA) at 53°C for 45 min, followed by 85°C for 5 min and then held at 4°C. Complementary DNA (cDNA) was generated and amplified, and its quality was assessed using an Agilent 4200 system. scRNA-seq libraries were constructed using the Single Cell 5' Library & Gel Bead Kit, Single Cell V(D)JEnrichment Kit, and Human B Cell (1000016). The libraries were sequenced using an Illumina NovaSeq 6000 sequencer with a paired-end 150 bp (PE150) read strategy.
[0038] 2. Construct gene expression matrices and cell type annotations for single-cell transcriptome data.
[0039] The Cell Ranger (v.3.0.2) software developed by 10x Genomics was used for sequencing data alignment, counting, and construction of single-cell expression matrices. The process mainly consisted of two steps: (1) cellranger mkfastq; (2) cellranger count. In this way, a single-cell expression matrix could be obtained for each sample, with rows representing different genes and columns representing different cells. The filtering, standardization, dimensionality reduction, and clustering of single-cell data were all completed using the R package Seurat (version 3.5.3). The steps were as follows: (1) The “CreateSeuratObject” function was used to construct a Seurat object, retaining only genes expressed in at least 0.1% of cells and cells with at least 200 genes expressed; (2) The “subset” function was used to perform secondary filtering of cells. Cells with fewer than 500 genes, fewer than 800 UMIs, and greater than 10% mitochondrial content were filtered out; (3) The “NormalizeData” function was used to standardize the data; (4) “FindVariableFeatures” was used to construct a gene set with high inter-cell differences; (5) The “ScaleData” function was used to normalize the data and remove the influence of unintended variables; (6) “RunPCA” was used to perform principal component analysis; (7) “FindNeighbors” and “FindClusters” functions were used to calculate the distance between cells and to cluster cells; (8) The “RunUMAP” function was used to perform nonlinear spatial dimensionality reduction of high-dimensional data and to map cells to a two-dimensional space for visualization of the results; (9) Finally, “FindAllMarkers” was used to identify the marker genes of different groups. Cell type identification is mainly based on prior knowledge. The detailed operation of the above steps is referred to the Seurat official tutorial (https: / / satijalab.org / seurat / v3.0 / pbmc3k_tutorial.html).
[0040] 3. BCR sequence assembly, clonal filtering and identification
[0041] BCR assembly and clonoid identification were performed using the Cell Ranger (v.3.0.2) V(D)J workflow and GRCh38 as a reference. After obtaining a list containing clonoid composition and frequencies, only cells with at least one expressible BCR heavy chain (IGH) and one expressible BCR light chain (IGK / IGL) were retained for further analysis. Each unique IGH(s)-IGK(s) or IGH(s)-IGL(s) pair was defined as a clonoid. Cells carrying a clonoid were considered clones if a clonoid was present in at least two cells. The number of cells with a particular clonoid is the clonoid size, indicating the degree of clonalization of the clonoid.
[0042] 4. Analyze the capture efficiency of BCR in different cell subsets and the clonal status under different disease states.
[0043] Using cell barcode information, data with BCR clonal types were projected onto the UMAP plot of expression data. This was then further filtered; if a cell group expressed two or more classic marker genes from different cell types, it suggested that the cell group was likely a twin-cell subpopulation generated during library construction. If a subpopulation's marker genes were mostly mitochondrial-related genes, ribosome genes, or lacked significant genes, these cells often had very low BCR capture efficiency, indicating that the subpopulation was highly likely to be low-quality cells. All twin-cell subpopulations and low-quality cell subpopulations were removed from the overall cell matrix and not subjected to downstream analysis. After obtaining high-quality expression and BCR data, the stable proportion of BCR detection in different cell types was calculated, representing the percentage of cells in that cell type with stable BCR detection out of that cell type. The results are as follows: Figure 2 As shown. Furthermore, different cell types are separated according to different disease states, and the clonal status of a specific cell type in a specific disease state is calculated, including the proportion of cells without detected BCR, the proportion of cells with detected BCR in a non-clonal state, and the proportion of cells with detected BCR in a clonal state. The results are shown in the figure. Figure 3 As shown.
[0044] 5. Analyze the BCR assembly and single-gene usage preferences of samples under different disease states.
[0045] For each sample, we obtained the clonogenic composition of the BCR in each cell of that sample, and extracted gene recombination information from IGH and IGK / IGL in that cell. Then, we statistically analyzed the frequency of V, D, and J genes used in IGH and IGK / IGL across all BCRs in that sample. After statistically analyzing the V(D)J usage in the BCRs of all samples, we grouped the samples according to different disease states and selected the top 5% of frequently used genes in each sample to create frequency maps. This effectively showed the specific and shared V, D, and J genes among different patients, as well as the specific and shared V, D, and J genes across different disease states. This allowed us to analyze the preference for high-frequency V(D)J gene usage that differs between viral infection and healthy states, and between patients who have undergone antiviral treatment. The analysis results are as follows: Figure 4 As shown.
[0046] 6. Analyze the BCR assembly and VJ gene usage preference in different samples under different disease states.
[0047] For each sample, we obtained the clonogenic composition of the BCR in each cell of that sample, and extracted gene recombination pairing information from IGH and IGK / IGL in that cell, and then statistically analyzed the frequency of VJ gene pairings used in IGH and IGK / IGL across all BCRs of that sample. After statistically analyzing the V(D)J usage of BCRs in all samples, the samples were separated according to different disease groups, and frequency maps were plotted for the most frequently used VJ pairing genes (Top 5%) in each sample. This effectively showed the specific and shared VJ gene pairs among different patients, as well as the specific and shared VJ gene pairs among different disease states. This allowed us to analyze the preference for high-frequency VJ gene pair usage that differs between viral infection and healthy states, and between patients and those treated with antiviral therapy. The analysis results are as follows: Figure 5 As shown.
[0048] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made in accordance with the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for analyzing B lymphocyte (BCR) characteristics based on single-cell multi-omics sequencing, which involves collecting samples from patients in different disease states for single-cell sequencing, followed by analysis through the following steps: 1) Construct a single-cell gene expression matrix based on single-cell transcriptome sequencing data and annotate it by cell type; 2) Assemble BCR sequences based on single-cell immunomic sequencing data and identify genes that make up the BCR chain; 3) Identify the stable BCR expression rate of each cell subpopulation by combining expression data: The single-cell expression data obtained in step 1) is filtered and denoised to retain only high-quality single-cell expression data; after assembling and identifying BCR sequences in step 2), only BCRs with both heavy and light chain data are retained; cells with high-quality single-cell expression data and BCRs with both heavy and light chains are identified, thereby obtaining the stable BCR expression rate of each cell subpopulation, that is, the percentage of cells with stable BCRs detected in each cell subpopulation out of the total number of cells; 4) Assess the clonal status of cell subpopulations in different disease states: Define each unique IGH(s)-IGK(s) pair or IGH(s)-IGL(s) pair as a clonal type. If a clonal type exists in at least two cells, the cell carrying that clonal type is considered cloned. At the same time, assess the cases of no BCR, non-clonal BCR, and cloned BCR in each cell subpopulation in different disease states. 5) Analysis of BCR assembly and single gene usage preferences in different disease states: Different samples were grouped according to different disease states. By analyzing the characteristics of heavy chain and light chain single gene usage among all samples in each disease state, the BCR assembly and single gene usage preferences in different disease states were analyzed. Specifically, for each sample, the clonogenic composition of the BCR in each cell of that sample was obtained first, and the gene recombination information in IGH and IGK / IGL of that cell was extracted. Then, the usage frequencies of V, D, and J genes of IGH and V and J genes of IGK / IGL in all BCRs of that sample were counted. After counting the single gene usage frequencies of BCRs of all samples, the samples were grouped according to different disease states, and the high-frequency used genes in each sample were selected to draw frequency maps, showing the V, D, and J genes specific to different disease states and shared between different disease states, thereby analyzing the single gene usage preferences in different disease states. 6) Analysis of BCR assembly and VJ gene pair usage preferences in different disease states: Different samples were grouped according to different disease states. By analyzing the characteristics of heavy and light chain VJ pair usage among all samples in each disease state, the preferences of BCR assembly and VJ gene pair usage in different disease states were analyzed. Specifically, for each sample, the clonogenic composition of the BCR of each cell in that sample was obtained first, and the gene recombination pairing information of IGH and IGK / IGL in that cell was extracted respectively. Then, the frequency of VJ gene pair usage of IGH and IGK / IGL in all BCRs of that sample was counted. Then, the samples were grouped according to different disease states, and the frequently used VJ pairing gene pairs in each sample were selected and plotted to show the specific and shared VJ gene pairs between different disease states, thereby analyzing the preferences of VJ gene pair usage in different disease states.
2. The method as described in claim 1, characterized in that, Step 1) After obtaining single-cell transcriptome sequencing data, download the host reference genome sequence and the corresponding gene annotation file. Obtain the single-cell gene expression matrix through data alignment and gene expression quantification. Perform data dimensionality reduction and cell type clustering and annotation using single-cell transcriptome sequencing analysis software.
3. The method as described in claim 2, characterized in that, In step 1), the host reference genome sequence and the corresponding gene annotation file are derived from the human reference genome hg18, hg19, hg20 or a subset of chromosome sets. The host reference genome sequence file is in standard FASTA format, and the gene annotation file is in standard GTF format. Cell type clustering is performed using Seurat, Scanpy, or Cellranger Loupe software, and cell type annotation is based on classic cell surface or specifically expressed marker genes reported in the literature.
4. The method as described in claim 1, characterized in that, Step 2) In the downloaded host reference genome sequence and gene annotation file, specifically extract the gene sequences and annotation information related to B lymphocyte BCR to construct BCR reference genome information; quantify the BCR sequence through data alignment and gene expression and identify the genes that make up the BCR chain.
5. The method as described in claim 4, characterized in that, Step 2) Extract the host BCR sequence and corresponding gene annotations from the host reference genome sequence and gene annotation file. The extraction tags are: --attribute=gene_biotype:IG_LV_gene \IG_V_gene \IG_V_pseudogene \IG_D_gene \IG_J_gene \IG_J_pseudogene \IG_C_gene \IG_C_pseudogene; Use Cellranger vdj software to assemble the BCR sequence and identify the genes that make up the BCR chain.
6. The method as described in claim 1, characterized in that, In step 3), when filtering single-cell expression data, the criteria for filtering cells include: cells with too low gene count, cells with too low UMI count, cells with too high mitochondrial ratio, dual cells with two or more classic cell marker genes, and cell fragments without any obvious marker gene expression.
Citation Information
Patent Citations
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A