A single-cell transcriptome dynamic map construction method and immunotherapy prediction device
Patent Information
- Application Number
- CN202310319746.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-03-28
AI Technical Summary
[0005]本发明提供了一种单细胞转录组图谱的构建方法及免疫治疗预测装置,以阐明免疫治疗完全反应的机制,并解决现有技术中cCR评估预测准确性低的技术问题
[0041] The technical solution of this invention involves collecting single-cell data to construct a single-cell transcriptome library, sequencing and analyzing the single-cell transcriptome library to obtain a single-cell transcriptome atlas, thereby constructing a dynamic single-cell transcriptome atlas. This allows for the analysis of the types of immune cells and/or immune cell subtypes in a sample using single-cell sequencing technology, and enables the analysis of changes in the proportion of immune cell subclones during treatment, thus providing a theoretical basis for subsequent immunotherapy prediction.
Smart Images

Figure CN116364231B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a method for constructing a dynamic transcriptome map of a single cell and an immunotherapy prediction device. Background Technology
[0002] Immune checkpoint blockade therapy, represented by PD-1 antibody therapy, has shown considerable efficacy as an emerging anticancer treatment in neoadjuvant therapy for cancer patients. Neoadjuvant therapy can induce regression of the primary tumor and mesangial lymph nodes, and some patients can achieve a pathological complete response (pCR). The local recurrence rate in pCR patients is extremely low, and improving the pCR rate in preoperative treatment is an important means to improve cancer survival. However, the mechanisms of immunotherapy efficacy and resistance are currently unclear. To improve the pCR rate of preoperative immunotherapy, it is necessary to explore specific cell types and their response patterns to treatment.
[0003] Furthermore, pCR is the result of postoperative pathological diagnosis. Currently, the concept of clinical complete remission (cCR) has also been proposed, defined as the absence of tumor residue on both clinical and imaging examinations. As a means of preoperatively assessing pCR, the "watch and wait" series of studies has shown that patients achieving cCR can avoid surgery, have the same survival rate as the surgical group, and even have a lower local recurrence rate. However, cCR based on imaging assessment has certain limitations, leading to low accuracy in predicting cCR, making it difficult to formulate a "watch and wait" strategy and hindering effective prediction of complete response before treatment.
[0004] Therefore, improving the pCR rate and the accuracy of cCR assessment are urgent problems to be solved. To improve the pCR rate of preoperative immunotherapy, it is necessary to provide a dynamic single-cell transcriptome map during the pCR process of immunotherapy to gain a deeper understanding of the mechanism of immunotherapy efficacy, and to provide a theoretical basis for further improving the pCR rate; to improve the accuracy of cCR assessment, a method that can predict the pCR of patients before treatment is needed. Summary of the Invention
[0005] This invention provides a method for constructing a single-cell transcriptome atlas and an immunotherapy prediction device to elucidate the mechanism of complete response to immunotherapy and to solve the technical problem of low accuracy in cCR assessment and prediction in the prior art.
[0006] To address the aforementioned technical problems, embodiments of the present invention provide a method for constructing a dynamic transcriptome map of a single cell, comprising:
[0007] Collect and construct a single-cell transcriptome library based on single-cell data;
[0008] The single-cell transcriptome library was sequenced;
[0009] The data obtained after sequencing are analyzed to construct a single-cell transcriptome atlas;
[0010] Based on the single-cell transcriptome atlas, a dynamic single-cell transcriptome atlas was constructed.
[0011] As a preferred embodiment, the single-cell transcriptome library includes: a 10×Genomics 3' single-cell transcriptome library.
[0012] As a preferred embodiment, the analysis of the data obtained after sequencing specifically includes:
[0013] The sequencing data is visualized using dimensionality reduction algorithms;
[0014] Global cells are classified and annotated using unsupervised clustering;
[0015] Perform characteristic gene expression level analysis;
[0016] Phenotypic proportion analysis;
[0017] Mitochondrial characteristic gene analysis.
[0018] As a preferred embodiment, the visualization of the sequenced data using a dimensionality reduction algorithm specifically includes:
[0019] Each sequencing sample underwent initial filtering, followed by batch processing effect correction and normalization.
[0020] After calculating highly variable genes, dimensionality reduction is performed, and individual cells are projected into a two-dimensional space.
[0021] Individual cells are clustered and grouped, and cell populations are visualized.
[0022] As a preferred embodiment, the characteristic gene expression level analysis specifically includes:
[0023] Detect characteristic genes of cell clusters and screen for characteristic genes and differentially expressed genes;
[0024] Specifically, the expression level of characteristic genes is analyzed by classifying all cell markers using global cell marker genes. The global cell marker genes include: global T / NK cell marker genes, global B cell marker genes, global myeloid cell marker genes, global epithelial cell marker genes, global fibroblast marker genes, and global endothelial cell marker genes.
[0025] As a preferred embodiment, the mitochondrial characteristic gene analysis specifically includes:
[0026] After global cell type identification, a personalized analysis of the mitochondrial high expression threshold is performed, and cells with corresponding high mitochondrial gene expression levels are removed based on the threshold.
[0027] As a preferred embodiment, the construction of a dynamic map of the single-cell transcriptome based on the single-cell transcriptome map specifically includes:
[0028] Based on the single-cell transcriptome map, identify the types of cell subclones;
[0029] Calculate the proportion of each type of cell subclone in the overall annotation of cell types;
[0030] Constructing cell subcloning marker gene tags;
[0031] We selected and compared the differences in the proportion of cell subclones before and after treatment to construct a dynamic map of the single-cell transcriptome.
[0032] As a preferred embodiment, the construction of the cell subcloning marker gene tag specifically includes:
[0033] The marker genes of each cell subclone are identified by a preset procedure; wherein the preset procedure uses the logarithmic change in average expression to identify the markers and uses a preset rank-sum test.
[0034] As a preferred embodiment, the step of selecting and comparing the differences in the proportion of cell subclones before and after treatment to construct a dynamic atlas of the single-cell transcriptome specifically includes:
[0035] Based on the preset rank-sum test, the difference in the proportion of cell subclones before and after treatment is compared. If the significance value of the difference in the proportion of cell subclones is less than the preset value, a dynamic map of the single-cell transcriptome is constructed.
[0036] Accordingly, the present invention also provides an immunotherapy prediction device, comprising: a sequencing module, a quantification module, and a clustering module;
[0037] The sequencing module is used to perform transcriptome sequencing on the collected pre-treatment biopsy tissue to obtain a transcriptome sequencing expression matrix.
[0038] The quantitative module is used to obtain immune cell subclonal marker gene tags based on the above-mentioned single-cell transcriptome dynamic map, thereby obtaining a gene-normalized expression matrix, and to evaluate the abundance of immune cell subclones based on the transcriptome sequencing expression matrix to obtain the relative abundance of each immune cell infiltration; wherein, the single cell is an immune cell;
[0039] The clustering module is used to identify and predict the response type of immunotherapy based on the relative abundance of each immune cell infiltration through unsupervised clustering analysis; wherein the response type of immunotherapy includes complete response and incomplete response.
[0040] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0041] The technical solution of this invention involves collecting single-cell data to construct a single-cell transcriptome library, sequencing and analyzing the single-cell transcriptome library to obtain a single-cell transcriptome atlas, thereby constructing a dynamic single-cell transcriptome atlas. This allows for the analysis of the types of immune cells and / or immune cell subtypes in a sample using single-cell sequencing technology, and enables the analysis of changes in the proportion of immune cell subclones during treatment, thus providing a theoretical basis for subsequent immunotherapy prediction.
[0042] Furthermore, the dynamic single-cell transcriptome atlas enables accurate prediction of complete response to immunotherapy, which helps to compensate for the shortcomings of imaging-based assessment of treatment response, improves the accuracy of prediction, and provides a reference for subsequent clinical decision-making. Compared with existing imaging-based judgments, it improves the accuracy of cCR assessment and prediction, thereby providing a reference for subsequent clinical decision-making and facilitating early adjustment of treatment plans based on prediction results in clinical practice. Attached Figure Description
[0043] Figure 1 : This is a flowchart illustrating the steps of a method for constructing a dynamic transcriptome map of a single cell provided in an embodiment of the present invention;
[0044] Figure 2 : A schematic diagram of the structure of an immunotherapy prediction device provided in an embodiment of the present invention;
[0045] Figure 3 : This is a flowchart of single-cell sequencing data analysis and cell type identification provided in an embodiment of the present invention;
[0046] Figure 4 : This refers to the high-quality global cell type clustering results provided in the embodiments of the present invention;
[0047] Figure 5 This is a dynamic map of the microenvironment single-cell immunotherapy according to an embodiment of the present invention and a schematic diagram of its application in immunotherapy;
[0048] Figure 6 : The results of T / NK cell, B cell, and myeloid cell subclone identification provided in the embodiments of the present invention;
[0049] Figure 7This is a schematic diagram illustrating the dynamic changes of cell subclones in immunotherapy and the main marker genes of each cell, as provided in the embodiments of the present invention.
[0050] Figure 8 : This is a schematic diagram of the unsupervised clustering results based on single-cell sequencing subclone tags provided in an embodiment of the present invention;
[0051] Figure 9 : This is a schematic diagram showing the statistical results of the efficacy of different subtype immunotherapy provided in the embodiments of the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Example 1
[0054] Please refer to Figure 1 The present invention provides a method for constructing a dynamic transcriptome map of a single cell, comprising the following steps S101-S104:
[0055] Step S101: Collect and construct a single-cell transcriptome library based on single-cell data.
[0056] As a preferred embodiment, the single-cell transcriptome library includes: a 10×Genomics 3' single-cell transcriptome library.
[0057] Step S102: Sequencing the single-cell transcriptome library.
[0058] It should be noted that, preferably, the single-cell RNA library is sequenced using an Illumina NovaSeq 6000 or BGISEQ DNBSEQ-T7 sequencer. The library is denatured, diluted, and mixed before sequencing, using the PE150 sequencing mode.
[0059] Step S103: Analyze the data obtained after sequencing to construct a single-cell transcriptome atlas.
[0060] As a preferred embodiment, the analysis of the data obtained after sequencing specifically includes:
[0061] The sequencing data is visualized using dimensionality reduction algorithms; global cells are classified and annotated using unsupervised clustering; characteristic gene expression levels are analyzed; phenotypic proportions are analyzed; and mitochondrial characteristic genes are analyzed.
[0062] As a preferred embodiment, the visualization of the sequenced data using a dimensionality reduction algorithm specifically includes:
[0063] Each sequencing sample was initially filtered, and batch processing effect correction and normalization were performed. After calculating highly variable genes, dimensionality reduction was performed, and individual cells were projected into a two-dimensional space. Individual cells were clustered and grouped, and the cell populations were visualized.
[0064] In this embodiment, the Seurat package (version 4.1.0) is used in R (version 4.1.0) to handle batch effects, dimensionality reduction, and UMI-based clustering analysis. First, initial filtering is performed on each sequencing sample: genes expressed in <3 cells are deleted; preferably, cells expressing too many (>5000) or too few (<500) genes, or containing less than 400 or more than 25000 unique molecular identifiers (UMIs), are also discarded. Then, batch effect correction and normalization are performed using the Seurat v4 function "IntegrateData" with its RPCA (Reintegrated Processing Analysis) function. The data is then processed using "SCTransform," and after calculating highly variable genes based on "Find Variable Features," initial dimensionality reduction is performed using "RunPCA." Individual cells are projected into a two-dimensional space using "RunUMAP," and individual cells are clustered and grouped using "FindNeighbors" and "Find Clusters." Cell population visualization is achieved through UMAP.
[0065] As a preferred embodiment, the step of performing characteristic gene expression level analysis specifically includes:
[0066] The study detects characteristic genes of cell clusters and screens characteristic genes and differentially expressed genes. Among them, the expression level of characteristic genes is analyzed by classifying all cell markers using global cell marker genes. The global cell marker genes include: global T / NK cell marker genes, global B cell marker genes, global myeloid cell marker genes, global epithelial cell marker genes, global fibroblast marker genes, and global endothelial cell marker genes.
[0067] In this embodiment, "FindAll Markers" is used to detect cell cluster characteristic genes. The criteria for screening characteristic genes and differentially expressed genes are a fold change > 1.5 and an adjusted p-value < 0.05, where the p-value is the significance level calculated based on the actual statistical measures. Furthermore, global cell marker genes are used to effectively and accurately classify all cell markers. The marker genes used are as follows:
[0068] Global T / NK cell marker genes: CD3D, CD3E, TRAC, TRBC1; Global B cell marker genes: CD79A, CD79B, MS4A1, TNFRSF17, MZB1; Global myeloid cell marker genes: CD14, CD68; Global epithelial cell marker genes: EPCAM, CD24; Global fibroblast marker genes: COL1A2, COL3A1, MYH11, ACTA2; Global endothelial cell marker genes: VWF, PECAM1.
[0069] As a preferred embodiment, the mitochondrial characteristic gene analysis specifically includes:
[0070] After global cell type identification, a personalized analysis of the mitochondrial high expression threshold is performed, and cells with corresponding high mitochondrial gene expression levels are removed based on the threshold.
[0071] In this embodiment, quality control for removing cells with high mitochondrial expression and cell classification are integrated. After global cell type identification, a personalized analysis of the mitochondrial high expression threshold is performed, and cells with corresponding high mitochondrial gene expression levels are removed based on the threshold. If the mitochondrial gene expression ratio of epithelial cells is greater than 75%, the epithelial cells are removed. For the other five global cell types, the median absolute deviation (MAD) and variance normal distribution are used to fit the mitochondrial gene expression levels, and then cells with expression levels significantly higher than expected are removed (determined by Bonferroni-corrected p<0.05).
[0072] Understandably, the algorithm utilizes Unified Manifold Approximation and Projection (UMAP) dimensionality reduction for visualization, classifies and annotates single cells through unsupervised clustering, analyzes feature gene expression levels, performs phenotypic proportion analysis, and combines global cell types to individually remove cells with high mitochondrial gene expression levels. It is understandable that the construction of the above-mentioned single-cell transcriptome atlas can be applied to single-cell analysis of the tumor microenvironment, including but not limited to: tumor microenvironment cell analysis of various samples such as breast cancer, esophageal cancer, lung cancer, melanoma, nasopharyngeal carcinoma, ovarian cancer, pancreatic cancer, renal cancer, and endometrial cancer. It can also be used for single-cell sequencing technology analysis of immune cells. For example, a single-cell transcriptome library can be constructed for single cells of microsatellite unstable colorectal tumors, the single-cell transcriptome library can be deeply sequenced, and the sequencing data can be visualized using dimensionality reduction algorithms, cells can be classified and annotated through unsupervised clustering, and characteristic gene expression level analysis, phenotypic ratio analysis, and granulocyte characteristic gene analysis can be performed to construct a single-cell transcriptome atlas of microsatellite unstable colorectal tumors containing rich information, which can provide a theoretical basis for exploring tumor microenvironment cells.
[0073] Step S104: Construct a dynamic single-cell transcriptome map based on the single-cell transcriptome atlas.
[0074] It is understandable that single-cell sequencing technology is used to analyze the types of immune cells and / or immune cell subtypes in a sample. This involves single-cell sequencing and sequencing data analysis techniques used in constructing a single-cell transcriptome atlas, enabling the construction of a dynamic single-cell transcriptome atlas based on the sequenced data. In this embodiment, by constructing a dynamic atlas of immune cells and analyzing the changes in the proportion of immune cell subclones during treatment, a single-cell dynamic atlas of immunotherapy is provided, playing a crucial role in elucidating the mechanism of action of immunotherapy.
[0075] As a preferred embodiment, the step of constructing a dynamic map of the single-cell transcriptome based on the single-cell transcriptome map specifically includes:
[0076] Based on the single-cell transcriptome atlas, the types of cell subclones were identified; the proportion of each cell subclone type in the overall cell type annotation was calculated; cell subclone marker gene tags were constructed; and the differences in the proportion of cell subclones before and after treatment were selected and compared to construct a dynamic atlas of the single-cell transcriptome.
[0077] For example, based on a single-cell transcriptome map, immune cell subclones can be obtained from T / NK cells, B cells, and myeloid cells in the global cell types, and cluster analysis can be performed on the three global cell types respectively.
[0078] Furthermore, to identify T / NK subclones, first distinguish CD3 gene-positive T cell lineages (CD4 T cells, CD8 T cells, γδ T cells) from other innate immune cell subsets (CD3 gene-negative); then identify T / NK subclones by analyzing characteristic genes to identify NK cells (expressing KLRD1, KLRC1, FCGR3A, TYROBP), native lymphocytes (expressing KIT and KRT86), and proliferating T cells (expressing proliferation genes MKI67, TOP2A).
[0079] Furthermore, the proliferating T cells showed a mixed population of CD4 and CD8 T cells, which were classified according to the normalized expression of CD4 and CD8A: CD4-Expr > CD8A-Expr, CD4+ prorif.T cells; CD4-Expr ≤ CD8A-Expr, CD8+ Trm-mitotic. After unsupervised clustering analysis, the T / NK subclones were divided into 16 distinct subsets, including 5 CD8 subclones, 7 CD4 subclones, 2 NK subclones, native lymphocyte subclones, and γδT cell subclones. CD8 subclones include intraepithelial lymphocytes (IEL), mucosa-associated inert T cells (MAIT), effector memory cells (Tem), tissue residency memory cells (Trm), and proliferating tissue residency memory cells (Trm-mitotic); CD4 subclones include naive (Tn), follicular helper (Tfh), exhausted (Texh), T helper 1-like (Th1-like), T helper (Th), proliferating (Prolif.), and regulatory T cell (Treg) subclones.
[0080] Further identification of B cell clones involved first distinguishing between plasma cells (expressing MZB1, XBP1, TNFRSF17A, IGHA1, IGHD1) and CD20+ B cells (expressing MS4A1, CD79A, CD79B), followed by further annotation based on marker genes. After unsupervised cluster analysis, B cell subclones were divided into seven distinct subpopulations, including three plasma cell subclones and four CD20+ B cell subclones. Plasma cell subclones included IgA plasma cells (pB-IgA), IgG plasma cells (pB-IgG), and proliferating plasma cells (pB-prolif); CD20+ B cell subclones included primitive (Bn), memory (Bmem), follicular helper (Bfoc), and germinal center (Bgc) subclones.
[0081] Furthermore, to identify myeloid cell subclones, dendritic cell, macrophage, and monocyte subpopulations were first distinguished. The dendritic cell subpopulation expressed dendritic cell marker genes (CD1C, CLEC10A, etc.), highly expressed HLA-DR gene, and low expressed CD14. The macrophage subpopulation highly expressed CD68 or CD163 and low expressed monocyte markers (FCN1, S100A8, S100A12). The monocyte subpopulation expressed monocyte marker genes (FCN1, S100A8, S100A12) but did not express dendritic or macrophage markers. After unsupervised clustering analysis, the myeloid cell subclones could be divided into 10 different subpopulations, including 3 dendritic cell subclones, 5 macrophage subclones, and 2 monocyte-like cell subclones. Dendritic cell subclones include mature dendrites (mDC), plasmacytoid dendrites (pDC), and classical dendrites (cDC); macrophage subclones and monocyte-like cell subclones are named according to the corresponding highly expressed genes.
[0082] As a preferred embodiment, the construction of the cell subcloning marker gene tag specifically includes:
[0083] The marker genes of each cell subclone are identified by a preset procedure; wherein the preset procedure uses the logarithmic change in average expression to identify the markers and uses a preset rank-sum test.
[0084] In this embodiment, the subclonal marker gene tag acquisition method uses the "FindAllMarkers" program in Seurat.V4 to identify the marker gene for each subclone. This program uses the log-fold change (FC) of average expression to identify the marker and uses the Wilcoxon rank-sum test by default (min.pct = 0.25, logfc.threshold = 0.25). The criteria for subclonal marker gene tags are a fold change > 1.5 and an adjusted p-value < 0.05.
[0085] As a preferred embodiment, the step of selecting and comparing the differences in the proportion of cell subclones before and after treatment to construct a dynamic map of the single-cell transcriptome specifically includes:
[0086] Based on the preset rank-sum test, the difference in the proportion of cell subclones before and after treatment is compared. If the significance value of the difference in the proportion of cell subclones is less than the preset value, a dynamic map of the single-cell transcriptome is constructed.
[0087] In this embodiment, samples with complete tumor remission before and after treatment were selected for comparison. The Wilcoxon rank-sum test was used to compare the differences in the proportion of immune cell subclones before and after treatment. A p < 0.05 was considered a significant change, and a dynamic atlas was constructed.
[0088] In this embodiment, in samples with complete tumor remission, the proportions of CD4 regulatory T cells (Tregs), proliferating CD8 tissue-resident memory cells (CD8+Trm-mitotic) subsets, pro-inflammatory monocytes (Mono-IL1 B), IgA plasma cells, and proliferating CD4 (CD4 Prolif) were significantly reduced after treatment, while the proportions of CD8 effector memory cells (CD8 Tem), CD4 helper T cells (CD4 Th), germinal center B cells (Bgc), and mature dendritic cells (mDC) were significantly increased. The proportions of other cell types did not change significantly before and after treatment. Because the experimental group consisted of samples with complete tumor remission, the dynamic atlas of immune cell subclones provided by this invention illustrates the cellular and molecular basis of the mechanism of action of immunotherapy. The dynamic atlas provided by this invention is of significant value in elucidating the mechanism of action of immunotherapy.
[0089] Implementing the above embodiments has the following effects:
[0090] The technical solution of this invention involves collecting single-cell data to construct a single-cell transcriptome library, sequencing and analyzing the single-cell transcriptome library to obtain a single-cell transcriptome atlas, thereby constructing a dynamic single-cell transcriptome atlas. This allows for the analysis of the types of immune cells and / or immune cell subtypes in a sample using single-cell sequencing technology, and enables the analysis of changes in the proportion of immune cell subclones during treatment, thus providing a theoretical basis for subsequent immunotherapy prediction.
[0091] Furthermore, the dynamic single-cell transcriptome atlas enables accurate prediction of complete response to immunotherapy, which helps to compensate for the shortcomings of imaging-based assessment of treatment response, improves the accuracy of prediction, and provides a reference for subsequent clinical decision-making. Compared with existing imaging-based judgments, it improves the accuracy of cCR assessment and prediction, thereby providing a reference for subsequent clinical decision-making and facilitating early adjustment of treatment plans based on prediction results in clinical practice.
[0092] Example 2
[0093] The present invention also provides an immunotherapy prediction device, comprising: a sequencing module 201, a quantification module 202, and a clustering module 203.
[0094] In this embodiment, the prediction of complete response to immunotherapy is achieved by assessing the subclonal abundance pattern of immune cells in the tumor microenvironment, thereby determining whether a sample is likely to produce a complete response to immunotherapy, providing a theoretical basis for related immunotherapy.
[0095] The sequencing module 201 is used to perform transcriptome sequencing on the collected pre-treatment biopsy tissue to obtain a transcriptome sequencing expression matrix.
[0096] In this embodiment, untreated samples were subjected to routine transcriptome sequencing, and quantitative analysis was performed using stringTie software to obtain the TPM expression matrix, i.e., the transcriptome sequencing expression matrix.
[0097] The quantitative module 202 is used to obtain the subclonal marker gene tags of immune cells according to the single-cell transcriptome dynamic map of Example 1, thereby obtaining the gene-normalized expression matrix, and to evaluate the subclonal abundance of immune cells according to the transcriptome sequencing expression matrix to obtain the relative abundance of each immune cell infiltration; wherein, the single cell is an immune cell.
[0098] In this embodiment, based on the method for constructing a dynamic single-cell transcriptome map described in Example 1, 33 immune cell subclonal marker genes were identified. All genes with a fold change > 1.5 and an adjusted p-value < 0.05 were included as tag genes. The relative abundance of the 33 immune cell subclonal infiltrations in the TME was quantified using the ssGSEA method with the "GSVA" R package.
[0099] The clustering module 203 is used to identify and predict the response type of immunotherapy based on the relative abundance of each immune cell infiltration through unsupervised clustering analysis; wherein the response type of immunotherapy includes complete response and incomplete response.
[0100] In this embodiment, NMF unsupervised clustering was used to cluster based on the subclonal abundance of 33 immune cells in the tumor microenvironment; the complete response rate of tumors of different molecular subtypes was compared; 100% complete remission was classified as complete response type (C type), and non-100% complete remission was classified as incomplete response type (P type). For the NMF method, the standard "brunet" option was selected, and 100 iterations were performed. The initial number of clusters k was set to 2 to 5, and the average silhouette width of the common membership matrix was determined using the R package "NMF". Common, dispersed, and silhouette indices were used to determine the optimal number of clusters, and the optimal number of clusters selected was 3.
[0101] It is understood that the embodiments of the present invention provide a method for predicting complete response to immunotherapy, determining the likelihood of a sample responding completely to immunotherapy. For example, pre-treatment biopsy tissue of the target sample is subjected to conventional transcriptome sequencing and assembled using stringTie software to obtain a TPM expression matrix. Based on the 33 immune cell subclonal marker genes described in Embodiment 1, a gene-normalized expression matrix is obtained for assessing the abundance of immune cell subclones. The samples are then included in a sample cohort for clustering, and it is observed whether the target sample clusters into type C (complete response type) to determine the likelihood of a complete response to immunotherapy.
[0102] Furthermore, pre-treatment prediction helps compensate for the limitations of imaging-based assessment of treatment response, improves the accuracy of prediction, and, importantly, provides a reference for subsequent clinical decisions. For patients with a complete response, a "wait-and-see" strategy can be adopted to avoid surgery, preserve organ function to the greatest extent possible, reduce patient suffering, and improve long-term quality of life.
[0103] Implementing the above embodiments has the following effects:
[0104] The immunotherapy complete response prediction method provided in this invention can predict the complete response of patients before treatment. On the one hand, it provides a theoretical basis for immunotherapy and further screens effective populations; on the other hand, pre-treatment prediction helps to compensate for the shortcomings of imaging-based treatment response assessment, improves the accuracy of prediction, provides a reference for subsequent clinical decision-making, and facilitates early adjustment of treatment plans based on test results. Patients with a complete response can adopt a "wait-and-see" strategy, avoiding surgery, preserving organ function to the greatest extent, reducing patient suffering, and improving long-term quality of life. It is convenient to use, has significant clinical implications, and broad application prospects.
[0105] Example 3
[0106] This embodiment collected 40 tumor samples and adjacent normal specimens from 19 patients with microsatellite instability colorectal cancer who had not received immunotherapy, had a complete response to immunotherapy, or had an incomplete response to immunotherapy. Single-cell transcriptome libraries were prepared and sequenced using the Chromium Next GEM single-cell 3' kit from 10x Genomics, and the sequencing data was analyzed. This corresponds to the analysis of the sequencing data in Embodiment 1, including the following steps S1-S3:
[0107] S1: Sample Collection and Separation: Tissue samples were digested with Thermo Scientific enzymes at 37°C for 20–30 minutes. The separated cells were then filtered through a 40 mm cell filter in RPMI-1640 medium (Invitrogen) containing 10% FBS until a homogeneous cell suspension was obtained. The cell suspension passing through the cell filter was then centrifuged at 400 × g for 10 minutes.
[0108] S2: Single-cell transcriptome library preparation: Single cells were resuspended in PBS containing 0.04% BSA to a final concentration of 500–1200 cells / ml. Approximately 10,000 cells were captured in droplets to generate milliliter-sized gel beads (GEMs). GEMs were then reverse transcribed in a C1000 Touch Thermal Cycler (Bio-Rad). After reverse transcription and cell barcoding, the emulsion was disrupted, and cDNA was isolated and purified using a Cleanup Mix containing DynaBeads and SPRIselect reagents, followed by PCR amplification. The cDNA to be sequenced was fragmented, end-repaired, base A added, sequencing adapters added, and PCR amplified. The amplified products underwent quality control.
[0109] S3: High-throughput sequencing: Single-cell RNA libraries were sequenced using an Illumina NovaSeq 6000 or BGISEQ DNBSEQ-T7 sequencer. The libraries were denatured, diluted, and mixed before sequencing, using the PE150 sequencing mode.
[0110] Following single-cell sequencing analysis, global cell type identification was performed. The single-cell transcriptome sequencing data of the samples were analyzed, and a secondary filtering process was used to remove low-quality cells and identify the global cell type, such as... Figure 3 As shown. Raw single-cell RNA sequence data was processed using the 10XGenomics cell Ranger toolkit (version 4.1.0), including demultiplexing FASTQ data, aligning it with the human reference genome (GRCh38, v3.0.0, from 10XGenamics), and generating a gene-cell unique molecular identifier (UMI) matrix using the “cellranger count” function.
[0111] Initial filtering and batch effect correction were performed: Data filtering, batch effect correction, dimensionality reduction, and UMI-based clustering analysis were conducted using the Seurat v4 R package. Initial filtering was performed on each sequencing sample: genes expressed in <3 cells were removed; cells expressing too many (>5000) or too few (<500) genes, or containing <400 or >25000 unique molecular identifiers (UMIs), were also discarded. Subsequently, batch effect correction and normalization were performed using the Seurat v4 function "IntegrateData" for ensemble analysis (RPCA). A new ensemble matrix was obtained, which was used solely for clustering and cell type classification.
[0112] Initial unsupervised clustering and global cell type identification are performed: After generating the ensemble matrix, the highly variable genes are calculated again based on "FindVariable Features" and then "Run PCA" is used for initial dimensionality reduction. "RunUMAP" is used to project individual cells into a two-dimensional space. "Find Neighbors" and "Find Clusters" are used to cluster and group individual cells, and UMAP is used to visualize the cell population. Each cluster was labeled with known markers (T cells: CD3D, CD3E, TRAC, TRBC1; B cells: CD79A, CD79B, MS4A1, TNFRSF17, MZB1; myeloid cells: CD14, CD68; epithelial cells: EPCAM, CD24; fibroblasts: COL1A2, COL3A1, MYH11, ACTA2; endothelial cells: VWF, PECAM1). Clusters that could not be labeled (high expression of mitochondrial genes), clusters with very low abundance, and conflicting clusters (co-expressing multiple broad markers) were deleted. The remaining global cell clusters of the same type were merged, resulting in six global cell types, such as... Figure 4 As shown.
[0113] Cells with high mitochondrial gene expression levels were individually removed because mitochondrial RNA content is highly cell type-dependent (significantly higher in epithelial cells). After obtaining six global cell types, the mitochondrial threshold for each global cell type was individually derived for further filtering. Epithelial cells were removed if the proportion of mitochondrial gene expression was greater than 75% (29.75% for epithelial cells). For the other five global cell types, mitochondrial gene expression levels were fitted using a median absolute deviation (MAD)-variance normal distribution, and cells with expression levels significantly higher than expected (determined by Bonferroni-corrected p<0.05) were removed at the following proportions: T cells 15%, B cells 10%, myeloid cells 12.84%, fibroblasts 8.11%, and endothelial cells 8.50%. The remaining 155,397 cells were used for re-clustering using Seurat. Ultimately, six high-quality global cell groups were obtained: epithelial cells, fibroblasts, endothelial cells, T cells, B cells, and myeloid cells.
[0114] Furthermore, the identified T cells, B cells, and myeloid cells were further clustered twice to remove confounding cells and cells with high mitochondrial expression, ultimately identifying 33 immune cell subclones. For example... Figure 5 As shown, it includes the following steps:
[0115] 1) Identification of T / NK subclones: A total of 29,033 cells were identified using T cell markers (CD3D, CD3E, CD2) and labeled as T cells. Re-clustering analysis revealed that these T cells were mixed with NK cells. In subsequent analysis, cell populations with high mitochondrial gene expression and confounding cells (10%) were excluded. The remaining cells (n=25,986) showed well-defined subpopulations annotated based on published characteristics, including the major T cell lineages (CD4+ T cells and CD8+ T cells) and three innate immune cell subpopulations, including γδ T cells expressing γδ receptor genes (TRDC, TRGC1, and TRGC2), CD3 gene-negative NK cells expressing NK cell genes (KLRD1, KLRC1, FCGR3A, and TYROBP), and a small number of native lymphocytes expressing KIT and KRT86. Furthermore, the proliferating T cells showed a mixed population of CD4+ and CD8+ T cells, which were further classified according to the normalized expression of CD4 and CD8A: CD4-Expr > CD8A-Expr, CD4+ prorif. T cells; CD4-Expr ≤ CD8A-Expr, CD8+ Trm-mitotic cells. Subpopulation results are as follows: Figure 6 As shown.
[0116] 2) Identification of B-cell subclones: A total of 20,229 B cells were identified using B-cell markers (MS4A1, CD79A, CD79B). After removing cells with high mitochondrial or ribosomal gene expression (16%), seven phenotypes were revealed. Cells were first classified into plasma cells (MZB1, XBP1, TNFRSF17A, IGHA1, IGHD1) and CD20+ B cells (MS4A1, CD79A, CD79B), and further annotated according to the marker genes. Subpopulation results are as follows: Figure 6 As shown.
[0117] 3) Identification of myeloid cell subclones: A total of 8469 myeloid cells were identified using the myeloid cell marker (CD14). After removing cells with high mitochondrial gene expression and mixed cells (14%), 10 subpopulations were revealed, including dendritic (DC) cell subpopulations, macrophage (Mac) subpopulations, and two mononuclear-like (Mono) subpopulations. The DC subpopulations included plasma cell DCs (pDC, pDC-LILR4A), mature DCs (mDC, mDC-LAMP3), and classical DCs (cDC-CD1C), characterized by high HLA-DR gene expression and low CD14 expression. Subclones with high CD68 or CD163 expression and negative expression of mononuclear cell markers (FCN1, S100A8, S100A12) were identified as macrophages. Cells expressing monocyte marker genes (FCN1, S100A8, S100A12) but not dendritic and macrophage markers were labeled as monocyte-like cells (Mono-FCN1, Mono-IL1B). Subpopulation results are as follows: Figure 6 As shown.
[0118] 4) Establishing subclonal marker gene tags: To characterize each cluster, the "FindAllMarkers" procedure in Seurat was used to identify marker gene tags for each subclone. This procedure uses the log-fold change (FC) of mean expression to identify markers and uses the Wilcoxon rank-sum test by default (min.pct = 0.25, logfc.threshold = 0.25). Genes with a fold change > 1.5 and an adjusted p-value < 0.05 were included as subclonal marker gene tags. A partial list of marker genes is shown below. Figure 7 As shown.
[0119] Furthermore, based on the analysis of identified cell subclones, their dynamic changes during immunotherapy are analyzed, i.e., a dynamic map of the single-cell transcriptome is constructed. For example... Figure 5 As shown, the proportion of each subclone in each sample was first calculated: subclone proportion = number of subclone cells / total number of cells. The Wilcoxon rank-sum test was used to compare the difference in subclone proportions between the pre-treatment and post-treatment complete remission (pCR) groups, with p < 0.05 considered statistically significant. Results are summarized as follows: Figure 8 As shown, in tumors exhibiting pCR response, the proportions of CD4 regulatory T cells (Tregs), proliferating CD8 tissue-resident memory cells (CD8+Trm mitotic), pro-inflammatory monocytes (Mono-IL1 B), IgA plasma cells, and proliferating CD4 (CD4 Prolif) significantly decreased after treatment, while the proportions of CD8 effector memory cells (CD8+Tem), CD4 helper T cells (CD4+Th), germinal center B cells (Bgc), and mature dendritic cells (mDC) significantly increased. The proportions of other cell types did not change significantly before and after treatment.
[0120] Example 4
[0121] This embodiment predicts complete response to immunotherapy and then performs transcriptome sequencing on pre-treatment samples to identify complete response profiles.
[0122] Tumor samples from 31 patients before treatment underwent routine transcriptome sequencing. After RNA extraction and library construction, the libraries were sequenced using a NovaSeq 6000 sequencing system (Illumina) in PE150 mode. The obtained RNA-seq reads were first checked using the FastP high-throughput sequencing data quality control tool. The reads were further processed using HISAT2 for filtering and alignment, and the TPM expression matrix was assembled using StringTie.
[0123] Based on the obtained marker gene tags and the obtained TPM expression matrix, immune cell abundance assessment and unsupervised clustering analysis were performed to identify drug resistance molecular subtypes. The steps are as follows: Figure 5 As shown, the ssGSEA (single-sample gene set enrichment analysis) algorithm was first used to quantify the relative abundance of each cell infiltration in the TME. The gene set used to label each type of infiltrating immune cell in the TME was obtained from the subclone-specific marker genes identified in Example 3. Genes with Foldchage > 1.5 and Padj < 0.05 in the subclones were considered as tag genes of that subclone. The enrichment score calculated by ssGSEA analysis was used to represent the relative abundance of 33 immune cell subclones infiltrating in each sample, including CD8 effector T cells (CD8 Tem), mature dendritic cells (mDC), SPP1 macrophages, natural killer T cells, and regulatory T cells. Subsequently, unsupervised clustering analysis was applied to identify different TME patterns and classify patients based on the relative abundance of the 33 infiltrating cells in the TME. For the NMF method, the standard "brunet" option was selected and 100 iterations were performed. The number of clusters, k, was set between 2 and 5, and the average silhouette width of the common membership matrix was determined using the R package "NMF". Common, dispersed, and silhouette indices were used to determine the optimal number of clusters; the optimal number of clusters chosen was 3. The results are as follows... Figure 8 and Figure 9 As shown, three molecular subtypes (IS-1, IS-2, IS-3) were obtained. IS-3 and IS-2 subtypes had a pCR rate of 100%, while IS-1 had the worst efficacy (pCR rate of 61%). IS-2 and IS-3 subtypes were both completely responsive to immunotherapy (subtype C), while IS-1 was incompletely responsive (subtype P). Patients predicted to be subtype C had better preoperative immunotherapy efficacy compared to subtype P (chi-square test, p = 0.034).
[0124] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. An immunotherapy prediction device, characterized in that, include: Sequencing module, quantification module, and clustering module; The sequencing module is used to perform transcriptome sequencing on the collected pre-treatment biopsy tissue to obtain a transcriptome sequencing expression matrix. The quantitative module is used to obtain subclonal marker gene tags of immune cells based on the dynamic single-cell transcriptome map, thereby obtaining a gene-normalized expression matrix, and to evaluate the abundance of subclonal immune cells based on the transcriptome sequencing expression matrix to obtain the relative abundance of each immune cell infiltration; wherein, the single cell is an immune cell; the dynamic single-cell transcriptome map is constructed using a dynamic single-cell transcriptome map construction method; The clustering module is used to identify and predict the response type of immunotherapy based on the relative abundance of each immune cell infiltration through unsupervised clustering analysis; wherein the response type of immunotherapy includes complete response and incomplete response. The method for constructing a dynamic transcriptome map of a single cell includes: Collect and construct a single-cell transcriptome library based on single-cell data; The single-cell transcriptome library was sequenced; The data obtained after sequencing are analyzed to construct a single-cell transcriptome atlas; Based on the single-cell transcriptome atlas, a dynamic single-cell transcriptome atlas is constructed. The construction of a dynamic single-cell transcriptome atlas based on the single-cell transcriptome atlas specifically includes: Based on the single-cell transcriptome map, identify the types of cell subclones; Calculate the proportion of each type of cell subclone in the overall annotation of cell types; Constructing cell subcloning marker gene tags; By selecting and comparing the differences in the proportion of cell subclones before and after treatment, a dynamic atlas of single-cell transcriptomes was constructed. The construction of the cell subcloning marker gene tag specifically includes: The marker genes of each cell subclone are identified by a preset procedure; wherein the preset procedure uses the logarithmic change in average expression to identify markers and uses a preset rank-sum test; The selection and comparison of the differences in the proportion of cell subclones before and after treatment, and the construction of a dynamic atlas of the single-cell transcriptome, specifically includes: Based on the preset rank-sum test, the difference in the proportion of cell subclones before and after treatment is compared. If the significance value of the difference in the proportion of cell subclones is less than the preset value, a dynamic map of the single-cell transcriptome is constructed.
2. The immunotherapy prediction device as described in claim 1, characterized in that, The single-cell transcriptome library includes: 10× Genomics 3' single-cell transcriptome library.
3. The immunotherapy prediction device as described in claim 1, characterized in that, The analysis of the data obtained after sequencing specifically includes: The sequencing data is visualized using dimensionality reduction algorithms; Global cells are classified and annotated using unsupervised clustering; Perform characteristic gene expression level analysis; Phenotypic proportion analysis; Mitochondrial characteristic gene analysis.
4. The immunotherapy prediction device as described in claim 3, characterized in that, The visualization of the sequenced data using dimensionality reduction algorithms specifically includes: Each sequencing sample underwent initial filtering, followed by batch processing effect correction and normalization. After calculating highly variable genes, dimensionality reduction is performed, and individual cells are projected into a two-dimensional space. Individual cells are clustered and grouped, and cell populations are visualized.
5. The immunotherapy prediction device as described in claim 3, characterized in that, The analysis of characteristic gene expression levels specifically includes: Detect characteristic genes of cell clusters and screen for characteristic genes and differentially expressed genes; Specifically, the expression level of characteristic genes is analyzed by classifying all cell markers using global cell marker genes. The global cell marker genes include: global T / NK cell marker genes, global B cell marker genes, global myeloid cell marker genes, global epithelial cell marker genes, global fibroblast marker genes, and global endothelial cell marker genes.
6. The immunotherapy prediction device as described in claim 3, characterized in that, The mitochondrial characteristic gene analysis specifically includes: After global cell type identification, a personalized analysis of the mitochondrial high expression threshold is performed, and cells with corresponding high mitochondrial gene expression levels are removed based on the threshold.
Citation Information
Patent Citations
Application of single cell sequencing in immune cell analysis
CN110819706A
A method to generate a cocktail of personalized cancer vaccines from tumor-derived genetic alterations for the treatment of cancer
WO2019036043A2