A method for analyzing TCR characteristics of T lymphocytes based on single-cell multi-omics sequencing

The TCR characteristics of T lymphocytes were analyzed through single-cell multiomic sequencing technology, and the problem of TCR changes in viral infection cannot be comprehensively analyzed in the existing technology, and an in-depth evaluation of TCR expression rate and clonal status is achieved, supporting the study of viral infection mechanism and host immune response.

CN115995267BActive Publication Date: 2025-08-15PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111213369.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-19
Publication Date
2025-08-15
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

The prior art cannot comprehensively and in-depth analysis of the stable TCR expression rate and clonal status of each cell subpopulation under different disease states during viral infection, the TCR assembly and use preferences of each sample before and after viral infection, as well as the TCR change status and specificity of the same patient before and after treatment.

Method used

By constructing a single-cell gene expression matrix based on single-cell multiomic sequencing data, annotating and cohortizing cell types, assembling TCR sequences and identifying TCR chain genes, evaluating the stable TCR expression rate of each cell subpopulation, analyzing the TCR assembly and use preferences under different disease states, and the TCR changes in the same patient before and after treatment.

Benefits of technology

A comprehensive analysis of the TCR characteristics of T lymphocytes after viral infection was achieved, the stable TCR expression rate and clonal status of each cell subpopulation were evaluated, and the TCR assembly and use preferences in different disease states were revealed, and important analysis of viral infection mechanisms and research basis for host immune responses was provided, and the development of specific vaccines and drug therapies was supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115995267B_ABST
    Figure CN115995267B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for analyzing TCR characteristics of T lymphocytes based on single-cell multi-omics sequencing. The method performs single-cell sequencing on different patient samples before and after specific viral infection and before and after treatment, constructs a single-cell gene expression matrix based on single-cell transcriptome sequencing data, and performs annotation and grouping of cell types. Based on single-cell immune group sequencing data, TCR sequences are assembled and TCR chain component genes are identified. Then, the expression rate of stable TCR of each cell subset is evaluated and identified, the clonal status of each cell subset in different disease states and the TCR assembly and use preference of each sample are evaluated, and the change characteristics of TCR before and after treatment of the same patient are analyzed. The present invention is of great significance for the analysis of specific viral infection mechanisms and the development of specific vaccines and drug therapies for host immune responses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of single-cell transcriptome sequencing, single-cell immunome sequencing, viral infection and host immune response, and bioinformatics analysis, and specifically to a method for comprehensively analyzing the TCR characteristics of T lymphocytes after viral infection based on single-cell multi-omics data. Background Art

[0002] The advent of high-throughput sequencing technology has ushered in a high-throughput era for research in all areas of biology. Traditional transcriptome sequencing (RNA-seq) based on bulk tissues only obtains average values for all cell types within a tissue. However, biological samples and tissues are often complex and heterogeneous systems with diverse cell types. Against this backdrop, single-cell RNA-sequencing (scRNA-seq) emerged, first reported in 2009, ushering in the single-cell era in biological research. With the increasing convenience and maturity of scRNA-seq technology, as well as its decreasing cost, scRNA-seq has been rapidly applied across various fields of biological research. The number of papers published on single-cell sequencing studies has increased significantly, and the number of cells studied has also exploded, from a few hundred or a few thousand cells to hundreds of thousands or even millions.

[0003] With the rapid development of scRNA-seq technology, single-cell research has ushered in the era of multi-omics. Numerous single-cell multi-omics sequencing technologies have emerged and are being effectively applied to various problems in the life sciences, including single-cell transcriptome sequencing, single-cell immunogenomic sequencing, single-cell genome sequencing, and single-cell chromatin accessibility sequencing. In particular, with the advent of in situ hybridization and specific RNA probes, single-cell spatial omics technologies have experienced rapid development this year, earning high expectations for their potential, having been recognized as Technology of the Year by multiple journals. Currently, single-cell transcriptome sequencing is the most mature and widely used technology, exemplified by 10x Genomics' Chromium technology, which sequences transcripts at either the 3' or 5' end. Because this technology only sequences one end of a transcript, it is relatively inexpensive and offers high cell throughput.

[0004] Viral infections are currently one of the leading threats to human health. For example, infection with the human immunodeficiency virus (HIV) can directly destroy the human immune system, causing acquired immune deficiency syndrome (AIDS). Rabies virus (RABV) is a neurological infection with a near-100% incidence rate, and there are currently no effective medications or treatments for the disease.

[0005] When a virus infects the human body, it triggers an antiviral immune response, encompassing both innate and adaptive immune responses. Innate immune responses are more rapid and encompass a wider range of responses than adaptive responses. Adaptive immune responses are further divided into two categories: cellular and humoral. T lymphocytes primarily mediate cellular immune responses. They originate from lymphoid stem cells in the bone marrow, differentiate and mature in the thymus, and are distributed throughout the body via the blood and lymphatic circulation, mediating cellular immune responses. Initially, T cells are activated by antigen-presenting cells and differentiate into effector T cells and memory T cells. The former kill target cells, while the latter can rapidly differentiate into effector T cells upon the reappearance of the same antigen. T cell activation relies on dual stimulation by the T cell receptor (TCR) and co-stimulatory signals. Activated effector T cells bind to target cells and simultaneously release substances such as interferon, perforin, and granzymes to kill virus-infected or tumor cells. T lymphocytes primarily consist of TCRs composed of α and β chains. During differentiation and maturation, specific recombinases act to connect the α and β chain genes, previously separated and transcriptionally inactive, into a complete, transcriptionally active gene. The general process is as follows: the β chain undergoes D-J ligation, followed by V-DJ ligation, before finally joining with the C region gene to form a complete, functional β chain gene. Rearrangement of the β gene induces rearrangement of the α gene. The α chain lacks a D segment and undergoes direct V-J ligation before joining with the C region to form a complete, functional α chain gene. Normal recombination of the α and β chain TCR genes is a relatively random process. Following viral infection, T lymphocytes undergo clonal expansion. The rate, pattern, diversity, and bias of TCR clonal expansion play a crucial role in understanding viral clearance and the mechanisms of viral clearance following infection.

[0006] scRNA-seq has been successfully applied in many cutting-edge biomedical fields, including the study of viral infection and host immune responses. scRNA-seq has successfully investigated the mechanisms of infection in various viral infections and revealed the characteristics of both innate and adaptive host immune responses, providing valuable insights for vaccine development and clinical treatment. However, current studies of viral infection using either transcriptome sequencing or immunogenomic sequencing techniques rely primarily on quantitative analysis of the composition and transcriptome of specific host cells and immune systems, or analysis of the degree of clonality of specific sequences. Furthermore, combined single-cell transcriptome and immunogenomic analyses have not yet enabled comprehensive and systematic analysis of the TCR clonal characteristics of specific subpopulations across different disease states (before and after infection, before and after treatment), across different patients, or within the same patient across different disease states. Consequently, comprehensive and in-depth analysis of the stable TCR expression rate and clonality of individual cell subpopulations across different disease states, the TCR assembly and usage preferences of individual samples before and after viral infection and treatment, and the changes in the state, specificity, and diversity of TCRs within the same patient before and after treatment are limited. Summary of the Invention

[0007] The purpose of the present invention is to address the shortcomings of the above-mentioned existing analysis technologies and provide a method for comprehensively analyzing the TCR characteristics of T lymphocytes after viral infection based on single-cell multi-omics sequencing data. This method can evaluate and identify the expression rate of stable TCRs in each cell subset, assess the clonal status of each cell subset in different disease states and the TCR assembly and usage preferences of each sample, and analyze the changes in TCR characteristics before and after treatment in the same patient. It has important significance for the study of viral infection mechanisms and the development of specific vaccines and drug therapies targeting host immune responses.

[0008] The technical solutions of the present invention are as follows:

[0009] A method based on single-cell multi-omics sequencing data to comprehensively analyze the TCR characteristics of T lymphocytes after viral infection. Samples from different patients before and after specific viral infection and before and after treatment are collected for single-cell sequencing, and then analyzed through the following steps:

[0010] 1) Construct a single-cell gene expression matrix based on single-cell transcriptome sequencing data and perform cell type annotation and clustering: Download the host reference genome sequence from the corresponding reference genome sequence website, along with the corresponding reference genome gene / functional subunit annotation file (GTF file). Obtain a single-cell gene expression matrix through data alignment and gene expression quantification. Perform data dimensionality reduction, cell type clustering, and annotation using single-cell transcriptome sequencing analysis software.

[0011] 2) Assemble TCR sequences and identify TCR chain component genes based on single-cell immune genome sequencing data: Specifically extract T lymphocyte TCR-related gene sequences and annotation information from the downloaded host reference genome sequence and gene / functional subunit annotation files to construct a TCR reference genome. Assemble TCR sequences and identify TCR chain component genes through data alignment and gene expression quantification.

[0012] 3) Identify the stable TCR expression rate for each cell subset by combining expression data: The single-cell expression data obtained in step 1) are filtered and denoised using appropriate thresholds to retain only high-quality single-cell expression data. After TCR sequence quantification and identification, only TCRs with both α and β chain data are retained. Ultimately, only cells with both high-quality single-cell expression data and TCR data are retained to determine the stable TCR expression rate for each cell subset, i.e., the percentage of cells with detected stable TCRs in each cell subset relative to the total number of cells.

[0013] 4) Assess the clonal status of each cell subpopulation in different disease states: Categorize data by disease state, then group data within the same disease state by cell type, define clonal and non-clonal states, and assess the overall clonal characteristics of different disease states and cell types.

[0014] 5) Analysis of TCR assembly and usage preferences in samples under different disease states: By analyzing the characteristics of TCR α and β chain VJ gene pairing in samples from different patients and different disease states, the TCR assembly and usage preferences of samples under different disease states can be analyzed.

[0015] 6) Analyze the status, specificity, and diversity of TCR changes in the whole patient and in specific cell subsets before and after treatment: For all cell types or specific cell types of interest with significant differences in expression levels and cell proportions, conduct in-depth analysis of the overall TCR clonal changes, TCR clonal changes in all cell subsets, and TCR clonal changes in specific cell subsets in the same patient before and after treatment.

[0016] The requirements for host reference genome sequence and annotation and cell type annotation in step 1) above are as follows:

[0017] The host reference genome sequence and corresponding gene annotation information can be sourced from, but are not limited to, the human reference genome hg18 (GRCh36), hg19 (GRCh37), hg20 (GRCh38), or a subset of the genome, compiled and maintained by databases such as UCSC, NCBI, Ensemble, Genecode, and 10x Genomics. The host reference genome sequence file should be in standard FASTA format, and the gene annotation file should be in standard GTF format. Cell type clustering can be performed using, but is not limited to, software such as Seurat, Scanpy, and Cellranger Loupe. Cell type annotation should be based on classical cell surface marker genes or specifically expressed marker genes reported in the literature.

[0018] Step 2) above) assembles the TCR sequence and identifies the TCR chain component genes, where:

[0019] Host TCR sequences and corresponding gene annotations are extracted from the host reference genome sequence and annotation information. Extracted tags typically include, but are not limited to, the following: --attribute=gene_biotype:TR_V_gene / TR_V_pseudogene / TR_D_gene / TR_J_gene / TR_J_pseudogene / TR_C_gene. TCR reference genome sequence files should be in standard FASTA format, and gene annotation files should be in standard GTF format. TCR sequences can be assembled and TCR chain component genes identified using, but are not limited to, Cellranger vdj software.

[0020] Step 3) above combines expression data to filter unstable TCRs and identify stable expression rates, where:

[0021] Single-cell expression data were filtered and denoised to retain only high-quality single-cell expression data. Filtering criteria included, but were not limited to, cells with low gene counts, cells with low UMI counts, cells with an excessively high mitochondrial ratio, double cells expressing two or more classic cell marker genes, and cell fragments lacking any significant marker gene expression. After TCR sequence assembly and quantitative identification, only cells with at least one expressible TCR α chain (TRA) and one expressible TCR β chain (TRB) were retained for further analysis. Finally, the intersection of the two sets of cells was taken, and only cells with both high-quality expression and TCR data were retained for downstream analysis.

[0022] Step 4) above assesses the clonal status of each cell subpopulation in different disease states, where:

[0023] After obtaining the TCR sequence, gene composition, and expression levels of different cells, each unique TRA(s)-TRB(s) pair is defined as a clonotype. If a clonotype is present in at least two cells, the cells carrying that clonotype are considered clonal. This allows for simultaneous assessment of TCR absence, nonclonal TCRs, and clonal TCRs in a cell type within a given disease state.

[0024] In step 5) above, analyze the TCR assembly and usage preferences of each sample under different disease states, where:

[0025] Normal α and β chain TCR gene recombination is a relatively random process. After antigen stimulation, clonal expansion occurs. The β chain sequentially undergoes D- and J-linking, followed by V- and DJ-linking, before finally connecting to the C region gene to form a complete, functional β chain gene. The α chain lacks a D segment and directly undergoes V- and J-linking before finally connecting to the C region to form a complete, functional α chain gene. Therefore, when analyzing TRA and TRB simultaneously, it is important to specifically analyze the preferences for V and J gene pairing in different samples.

[0026] Step 6) above analyzes the status, specificity, and diversity of TCR changes in the same patient before and after treatment, both overall and in specific subpopulations, where:

[0027] For each clonotype, the number of cells with that clonotype is called its clone size. If a clonotype is present both before and after treatment, it is considered a shared clone and a maintenance clone. If a clonotype is present only before or after treatment, it is considered a specific clone. Sequence analysis identifies different clonotypes, and analysis of clone size and specificity can be used to characterize TCR clonal changes in both overall and specific cell subsets within the same patient before and after treatment.

[0028] The present invention addresses the current deficiencies in the use of scRNA-seq technology to study TCRs in overall cell types, specific cell subsets, disease states, single patients, and the same patient before and after treatment during viral infection. This method provides a method for comprehensively analyzing TCR characteristics of T lymphocytes after viral infection based on single-cell multi-omics data. This method can combine expression data to identify the stable TCR expression rate of each cell subset, evaluate the clonal status of each cell subset in different disease states, analyze the TCR assembly and usage preferences of each sample under different disease states, and analyze the state, specificity, and diversity of TCR changes in the same patient before and after treatment. This method is of great significance for the analysis of the mechanism of specific viral infection of the host and the development of specific vaccines and drugs. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 . The overall analysis process for analyzing TCR characteristics of T lymphocytes after viral infection in the embodiments of the present invention.

[0030] Figure 2 . In the embodiment of the present invention, the expression data is combined to identify the overall stable TCR expression rate of each cell subpopulation, where: the left figure is a cell classification diagram, and the numbers represent different cell types; the middle figure shows whether the cell detects a stable TCR; the right figure shows the percentage of stable TCR signals detected in cell subpopulations of different cell types.

[0031] Figure 3 . Evaluation of the clonal status of each cell subpopulation in different disease states in the embodiments of the present invention.

[0032] Figure 4 . In the embodiments of the present invention, TCR assembly and usage preferences of each sample under different disease states are analyzed.

[0033] Figure 5 . Analysis of overall TCR clone changes in the same patient before and after treatment in the examples of the present invention.

[0034] Figure 6 . Analysis of TCR clone changes in all cell subsets before and after treatment of the same patient in the examples of the present invention.

[0035] Figure 7 . Analysis of TCR clone changes in specific subpopulations of the same patient before and after treatment in the examples of the present invention. DETAILED DESCRIPTION

[0036] The following is a more detailed description of the implementation of the present invention, wherein the parameters and specific implementation details are used to explain the feasibility and implementation effect of the present invention, and do not constitute a limitation of the present invention.

[0037] In this example, CD4 T cells isolated by magnetic beads from peripheral blood mononuclear cells (PBMCs) of 14 HIV-infected patients (including 9 patients before treatment and 8 patients after antiviral treatment, of which 3 patients had paired samples before and after antiviral treatment) and 4 healthy subjects were used. + T cells were used as samples for single-cell sequencing using the 10x Genomics Single Cell 5'Library & GelBead kit. Simultaneously with transcriptome sequencing, paired TCR sequencing was also performed for each sample using the SingleCell V(D)J Enrichment kit.

[0038] 1. Sample library construction and high-throughput single-cell sequencing based on the Single Cell 5' Library & Gel Bead Kit (10x Genomics) and Chromium Single Cell A Chip Kit (10x Genomics)

[0039] CD4 was immediately purified from fresh PBMC using the CD4 (130-045-101, Miltenyi Biotech) isolation kit. + T cells. A cell suspension (300-600 viable cells / μL, counted by Countstar) was loaded onto a Chromium Single-Cell Controller (10x Genomics) to generate water-in-oil droplets (GEMs). Briefly, single cells were suspended in phosphate-buffered saline (PBS) containing 0.04% bovine serum albumin (BSA) and added to each channel of the chip, resulting in a final recovery of approximately 50%. Captured cells were lysed, and the released RNA was labeled in individual GEMs by reverse transcription. Reverse transcription was performed on an S1000™ Thermal Cycler (Bio-Rad Laboratories, Hercules, CA) at 53°C for 45 minutes, followed by 85°C for 5 minutes and a hold at 4°C. Complementary DNA (cDNA) was generated and amplified, and its quality was assessed using an Agilent 4200 system. scRNA-seq libraries were constructed using the Single Cell 5' Library & Gel Bead Kit, the Single Cell V(D)J Enrichment Kit, and Human T Cell (1000005). Libraries were sequenced using an Illumina NovaSeq 6000 sequencer with a paired-end 150 bp (PE150) read strategy.

[0040] 2. Constructing gene expression matrices and cell type annotations for single-cell transcriptome sequencing data

[0041] Sequencing data were aligned, counted, and constructed using the Cell Ranger (v.3.0.2) software developed by 10x Genomics. This was done in two steps: (1) cellranger mkfastq; (2) cellranger count. This yielded a single-cell gene expression matrix for each sample, with rows representing genes and columns representing cells. The R package Seurat (version 3.5.3) was used to filter, normalize, reduce, and cluster the single-cell transcriptome sequencing data. The steps were as follows: (1) the "CreateSeuratObject" function was used to construct the Seurat object, retaining only genes expressed in at least 0.1% of cells and cells with at least 200 gene expressions; (2) the "subset" function was used to perform secondary cell filtering. Cells with fewer than 500 genes, fewer than 1000 UMIs, and greater than 10% mitochondrial content were filtered out; (3) the "NormalizeData" function was used to normalize the data; (4) the "FindVariableFeatures" function was used to construct a gene set with high inter-cell differences; (5) the "ScaleData" function was used to normalize the data and remove the influence of unintended variables; (6) the "RunPCA" function was used to perform principal component analysis; (7) the "FindNeighbors" and "FindClusters" functions were used to calculate inter-cell distances and cluster cells; (8) the "RunTSNE" function was used to perform nonlinear spatial dimensionality reduction of high-dimensional data, mapping cells to a two-dimensional space for visualization of the results; (9) finally, the "FindAllMarkers" function was used to identify marker genes for different groups. The identification of cell types was mainly based on prior knowledge. For detailed instructions on the above steps, please refer to the official Seurat tutorial (https: / / satijalab.org / seurat / v3.0 / pbmc3k_tutorial.html).

[0042] 3. TCR Sequence Assembly, Clonal Type Filtering and Identification

[0043] TCR assembly and clonotype identification were performed using the Cell Ranger (v.3.0.2) V(D)J pipeline program and GRCh38 as a reference. After obtaining a list containing clonotype composition and frequency, only cells with at least one expressible TCR α chain (TRA) and one expressible TCR β chain (TRB) were retained for further analysis. Each unique TRA(s)-TRB(s) pair was defined as a clonotype. If a clonotype was present in at least two cells, the cells carrying that clonotype were considered clonal. The number of cells harboring a clonotype was defined as the clone size of that clonotype, indicating the degree of clonality.

[0044] 4. Analyze TCR capture efficiency of different cell subsets and clonal status in different disease states

[0045] Using cell barcode information (Barcode), the data with TCR clonal types are projected onto the t-SNE graph of the expression data. Re-filtering is performed accordingly. If a group of cells expresses two or more classic marker genes of different cell types, it suggests that the group of cells is most likely a dual cell generated during the library construction process. If the marker genes of a group are mostly mitochondrial-related genes, ribosomal genes, or there are no significant genes, the TCR capture efficiency of these cells is often very low, which suggests that the subpopulation is likely to be a low-quality cell. All dual-cell subpopulations and low-quality cell subpopulations are removed from the total cell matrix and no downstream analysis is performed. After obtaining high-quality expression data and TCR data, the stable ratio of TCR detection in different cell types is calculated, which is the percentage of cells in which stable TCRs are detected in the cell type to the total number of cells of the cell type. The results are as follows. Figure 2 Furthermore, different cell types were separated according to different disease states, and the clonal status of a specific cell type in a specific disease state was calculated, including the proportion of cells with no TCR detected, the proportion of cells with non-clonal TCR detected, and the proportion of cells with clonal TCR detected. The results are shown in Figure 3 shown.

[0046] 5. Analysis of TCR Assembly and Usage Preference in Different Samples under Different Disease States

[0047] For each sample, we obtained the clonal composition of each cell TCR in the sample, extracted the pairing information of the VJ genes in the TRA and TRB of the cell respectively, and then counted the frequency of VJ gene pairing of TRA and TRB in all TCRs of the sample. After counting the TCR VJ pairing usage of all samples, the samples were grouped according to different disease states and different patients, and the high-frequency VJ gene pairs (Top 5) in each sample were selected to draw a graph, which effectively displayed the specific and shared V genes, J genes and VJ gene pairs between different patients, as well as the specific and shared V genes, J genes and VJ gene pairs between different disease states. In this way, the preference for the use of high-frequency VJ gene pairs after viral infection compared with the healthy state, and after antiviral treatment was analyzed. The analysis results are as follows. Figure 4 shown.

[0048] 6. Analyze TCR changes in the whole body and specific cell subsets before and after treatment in the same patient

[0049] For the paired data of the same patient before and after antiviral treatment, we comprehensively analyzed the changes in clones before and after treatment. First, we analyzed the changes in clone types and divided the clones into pre-treatment specific clones, post-treatment specific clones, and clones shared before and after treatment. Scatter plots were used to compare whether the patterns of clone changes in the three paired patients after HIV infection were consistent, and whether the main occurrence was new clone expansion or clone maintenance. The results are as follows Figure 5 We then counted the relative proportions of specific clones and maintenance clones in different cell types under different disease states to reveal the diversity of clonal changes in specific disease states and the specificity of clonal changes in specific cell subsets before and after antiviral treatment. For example, we found that the effector CD4 expressing the GNLY gene + T cell subsets play a major role in clonal maintenance before and after antiviral treatment. Figure 6 Finally, we further analyzed the distribution of different clone sizes in different clone states in the cell subpopulation based on the cell type of interest found. For example, we found that the effector CD4 expressing the GNLY gene + The clones that play the main role in clonal maintenance of T cell subsets before and after antiviral treatment are all large clones. Figure 7 shown.

[0050] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for comprehensively analyzing TCR characteristics of T lymphocytes after viral infection based on single-cell multi-omics sequencing data. Samples from different patients before and after specific viral infection and before and after treatment are collected for single-cell sequencing, and then analyzed through the following steps: 1) Construct a single-cell gene expression matrix based on single-cell transcriptome sequencing data, and perform cell type clustering and annotation; 2) Assemble TCR sequences and identify TCR chain component genes based on single-cell immune genome sequencing data; 3) Identify the stable TCR expression rate of each cell subset by combining expression data: Filter and denoise the single-cell expression data obtained in step 1) to retain only high-quality single-cell expression data; after assembling and identifying TCR sequences in step 2), only retain TCRs with both α and β chain data; Identify cells with high-quality single-cell expression data that have both α and β chains of TCRs, thereby obtaining the stable TCR expression rate of each cell subset, that is, the percentage of cells with stable TCRs detected in each cell subset to the total number of cells; 4) Assessment of the clonality of cell subsets across disease states: Each unique TRA(s)-TRB(s) pair is defined as a clonotype. If a clonotype is present in at least two cells, the cells carrying that clonotype are considered clonal. Furthermore, the presence of no TCR, nonclonal TCR, and clonal TCR in each cell subset across disease states were assessed. 5) Analysis of TCR assembly and usage preferences in samples under different disease states: By analyzing the characteristics of TCR α and β chain VJ gene pairing in samples from different patients and different disease states, the TCR assembly and usage preferences of samples under different disease states can be analyzed; Specifically, for each sample, the clonal composition of each cell TCR in the sample is first obtained, and the pairing information of the VJ genes in the TRA and TRB of the cell is extracted respectively. Then, the frequency of VJ gene pairing of TRA and TRB in all TCRs of the sample is counted. After counting the TCR VJ pairing usage of all samples, the samples are grouped according to different disease states and different patients, and the VJ gene pairs used frequently in each sample are selected and plotted into a graph, showing the specific and shared V gene, J gene and VJ gene pairs between different patients, as well as the specific and shared V gene, J gene and VJ gene pairs between different disease states. In this way, the preference for the use of high-frequency VJ gene pairs after viral infection compared with the healthy state, and after antiviral treatment is analyzed. 6) Analyze the status, specificity, and diversity of TCR changes in the whole body and specific cell subsets in the same patient before and after treatment.

2. The method according to claim 1, wherein Step 1) After obtaining the single-cell transcriptome sequencing data, download the host reference genome sequence and the corresponding gene annotation file. Obtain the single-cell gene expression matrix through data alignment and gene expression quantification. Use single-cell transcriptome sequencing analysis software to perform data dimensionality reduction and cell type clustering and annotation.

3. The method according to claim 2, wherein In step 1), the host reference genome sequence and corresponding gene annotation files are from the human reference genome hg18, hg19, hg20, or a subset of the chromosome set. The host reference genome sequence file is in the standard FASTA format, and the gene annotation file is in the standard GTF format. Cell type clustering is performed using Seurat, Scanpy, or Cellranger Loupe software. Cell type annotation is based on classic cell surface marker genes or specifically expressed landmark genes reported in the literature.

4. The method according to claim 1, wherein Step 2) From the downloaded host reference genome sequence and gene annotation files, specifically extract T lymphocyte TCR-related gene sequences and annotation information to construct TCR reference genome information; quantitatively assemble TCR sequences and identify TCR chain component genes through data alignment and gene expression.

5. The method according to claim 4, wherein Step 2) Extract host TCR sequences and corresponding gene annotations from the host reference genome sequence and gene annotation file using the following extraction tag: --attribute=gene_biotype: TR_V_gene / TR_V_pseudogene / TR_D_gene / TR_J_gene / TR_J_pseudogene / TR_C_gene; Use Cellranger vdj software to assemble TCR sequences and identify TCR chain component genes.

6. The method according to claim 1, wherein In step 3), when filtering single-cell expression data, the criteria for filtering cells include: cells with too low a gene count, cells with too low a UMI count, cells with too high a mitochondrial ratio, double cells with two or more classic cell marker genes, and cell fragments without any obvious marker gene expression.

7. The method according to claim 1, wherein For each clonotype, the number of all cells with that clonotype is called the clone size of that clonotype. If a clonotype exists both before and after treatment, it is a shared clone before and after treatment and is called a maintenance clone. If a clonotype exists specifically before or after treatment, it is called a specific clone. In step 6), different clonotypes are identified through sequence analysis, and the clone size and specificity are analyzed to analyze the characteristics of TCR clonal changes in the whole patient and specific cell subsets before and after treatment.

Citation Information

Patent Citations

  • Kit for amplifying TCR full-length sequence and use thereof

    CN111344418A

  • Systems and methods for multiplexed measurements in single and ensemble cells

    CN112004920A