Immune typing method and device, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2024-08-23
- Publication Date
- 2026-04-24
AI Technical Summary
Existing breast cancer classification methods are mainly based on gene expression characteristics, lacking classification related to treatment response, and cannot effectively utilize tumor immune microenvironment information for treatment guidance.
By performing immune infiltration analysis on sample data, cell types related to treatment response were screened out, cluster analysis was performed to determine cell combinations for immunophenotyping, characteristic cell types were identified, and the association between the immune microenvironment and treatment response was established.
Significant correlation of treatment response based on immunophenotyping was achieved, which can more accurately guide the treatment of breast cancer and improve treatment outcomes.
Smart Images

Figure CN121925706A_ABST
Abstract
Description
Immune typing method and device, storage medium and program product TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to, but are not limited to, the field of biotechnology, and particularly relate to an immune typing method and device, a storage medium and a program product. BACKGROUND
[0002] Breast cancer can be divided into several main types according to the morphological, immunohistochemical markers and gene expression characteristics of the tumor.
[0003] Different breast cancer subtypes also have different treatment responses. The existing classification is mainly based on gene expression characteristics, and there are fewer subtypes related to treatment response. The immune microenvironment of the tumor regulates the survival, maintenance and growth of the tumor, and plays an important role in disease progression and treatment response, such as local drug resistance, immune escape and cancer metastasis. Analyzing the composition and function of different cell populations in the immune microenvironment of the tumor helps to explore more effective tumor treatment methods.
[0004] SUMMARY
[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0006] Embodiments of the present disclosure provide an immune typing method, comprising:
[0007] performing immune infiltration analysis on sample data to obtain the proportion of a plurality of cell types;
[0008] screening a plurality of cells related to treatment response from the plurality of cell types;
[0009] performing clustering analysis on a plurality of cell combinations, each cell combination including at least part or all of the plurality of cells screened out;
[0010] selecting one of the cell combinations as an immune typing cell combination according to the clustering analysis result, and determining the characteristic cell type corresponding to each immune typing group.
[0011] Embodiments of the present disclosure also provide an immune typing device, comprising a memory; and a processor connected to the memory, the memory being configured to store instructions, and the processor being configured to execute the steps of the immune typing method according to any one of the embodiments of the present disclosure based on the instructions stored in the memory.
[0012] Embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the immune typing method according to any one of the embodiments of the present disclosure.
[0013] The embodiments of the present disclosure further provide a program product comprising instructions, which, when executed by a computer, perform the immunophenotyping method according to any of the embodiments of the present disclosure.
[0014] The embodiments of the present disclosure further provide an immunophenotyping device comprising an immune infiltration module, a screening module, a clustering analysis module and a determination module, wherein:
[0015] The immune infiltration module is configured to perform immune infiltration analysis on sample data to obtain proportions of multiple cell types.
[0016] The screening module is configured to screen multiple cells related to treatment response from the multiple cell types.
[0017] The clustering analysis module is configured to perform clustering analysis on multiple cell combinations, each of which comprises at least part or all of the screened multiple cells.
[0018] The determination module is configured to select one of the cell combinations as an immunophenotyping cell combination according to the clustering analysis result, and determine the characteristic cell type corresponding to each immunophenotyping group.
[0019] The immunophenotyping method and device, storage medium and program product of the embodiments of the present disclosure perform immune infiltration analysis on sample data to obtain proportions of multiple cell types, screen multiple cells related to treatment response from the multiple cell types, perform clustering analysis on multiple cell combinations, each of which comprises at least part or all of the screened multiple cells, select one of the cell combinations as an immunophenotyping cell combination according to the clustering analysis result, and determine the characteristic cell type corresponding to each immunophenotyping group, thereby establishing an immunophenotyping related to the correlation between immune microenvironment and treatment response, and the immunophenotyping is significantly related to treatment response pCR.
[0020] Other features and advantages of the present disclosure will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present disclosure. Other advantages of the present disclosure will be realized and attained by the subject matter particularly pointed out in the specification and claims. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings are included to provide an understanding of the present technical solution, and constitute a part of the specification, and are used together with the embodiments of the present disclosure to explain the present technical solution, and do not constitute a limitation on the present technical solution.
[0022] FIG. 1 is a flowchart of an immunophenotyping method according to an exemplary embodiment of the present disclosure;
[0023] FIG. 2 is a graph of the clustering analysis results of 47 HER2-positive pre-treatment samples according to an example embodiment of the present disclosure;
[0024] FIG. 3 is a graph of the clustering analysis results of 47 HER2-positive post-treatment samples according to an example embodiment of the present disclosure;
[0025] FIG. 4 is a graph of the characteristic cell types of the immune subtyping groups determined for the 47 HER2-positive pre-treatment samples according to an example embodiment of the present disclosure;
[0026] FIG. 5 is a graph of the characteristic cell types of the immune subtyping groups determined for the 47 HER2-positive post-treatment samples according to an example embodiment of the present disclosure;
[0027] FIG. 6 is a graph of the correlation between the immune subtyping of pre-treatment samples and treatment response;
[0028] FIG. 7 is a graph of the correlation between the immune subtyping of post-treatment samples and treatment response;
[0029] FIG. 8 is a graph of the correlation between the combined pre-treatment immune subtyping of FIG. 6 and the post-treatment immune subtyping of FIG. 7 and treatment response;
[0030] FIG. 9 is a graph of the characteristic cell types of the immune subtyping groups in the TCGA validation data;
[0031] FIGS. 10A-10D are survival plots of the top 4 differential genes associated with survival for the IM1 immune subtyping group;
[0032] FIGS. 10E-10H are survival plots of the top 4 differential genes associated with survival for the IM2 immune subtyping group;
[0033] FIG. 11 is a structural diagram of an immune subtyping device according to an example embodiment of the present disclosure;
[0034] FIG. 12 is a structural diagram of another immune subtyping device according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] The present disclosure describes a plurality of embodiments, but the description is exemplary rather than limiting, and it will be apparent to those of ordinary skill in the art that more embodiments and implementations can be possible within the scope of the embodiments described in the present disclosure. Although a number of possible combinations of features are shown in the drawings and discussed in the specification, many other combinations of the disclosed features are possible. Unless specifically intended otherwise, any feature or element of any embodiment can be used with any other feature or element of any other embodiment, or in any other embodiment, whether or not that feature or element is specifically disclosed in combination with the other feature or element in any embodiment.
[0036] The present disclosure includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The presently disclosed embodiments, features and elements can also be combined with any conventional feature or element to form a unique application of the presently claimed invention that is defined by the claims. Any feature or element of any embodiment can also be combined with features or elements from other application schemes to form another unique application of the presently claimed invention that is defined by the claims. Accordingly, it should be understood that any feature shown and / or discussed in the present disclosure can be implemented alone or in any suitable combination. Embodiments, therefore, are not to be limited by other than in accordance with the claims as written and their equivalents. Moreover, various modifications and changes can be made within the scope of the claims.
[0037] Furthermore, in describing representative embodiments, the specification can have presented the method and / or process as a particular sequence of steps. However, to the extent that the method or process depends on the particular order of steps, this description should not be construed as limiting. Other steps can be performed in between described steps without departing from the scope of the present disclosure. As such, the particular order of the steps presented and / or discussed in the specification is not an limitation of the method and / or process. Moreover, the specification can present the steps of a method and / or process as an enumeration of particular chronological steps. However, the steps as presented and / or discussed in the specification should not be construed as being limited to the particular chronological order presented and / or discussed. Rather, the steps presented and / or discussed in the specification are used to illustrate the functionality of the present disclosure and should not be construed as limiting. The specification is to be construed as a limitation of the present disclosure only insofar as not otherwise falling within the scope of the claims.
[0038] As shown in FIG. 1, the embodiments of the present disclosure provide an immunophenotyping method, comprising:
[0039] Step 101, performing immune infiltration analysis on sample data to obtain the proportion of multiple cell types;
[0040] In some exemplary embodiments, in step 101, the sample data can be breast cancer disease sample data set. However, the present disclosure is not limited thereto.
[0041] Exemplarily, 94 HER2-positive sample data were selected from the Gene Expression Omnibus database as a training set, of which 47 samples were pre-treatment (baseline) and 47 samples were post-treatment (postoperative), as shown in Table 1.
[0042] Table 1
[0043] Wherein, Age is the age, Ki-67 is also called proliferation index, which represents the proliferation index of tumor cells, and is a protein existing in the cell nucleus. Estrogen receptor (ER) exists in the cell and can specifically bind with hormones to form a hormone-receptor complex, so that the hormone can exert its biological effect. The estrogen diffused to the cell nucleus binds with its nuclear receptor to cause the expression of downstream genes. Progesterone receptor (PR) plays a key role in Luminal A and Luminal B typing. Patients with PR>20% have longer disease-free survival. The human epidermal growth factor receptor family includes four members: HER1, HER2, HER3, and HER4. As transmembrane proteins, they can bind to growth factors on the cell surface, conduct signals, and regulate cell growth, division, and repair. Overexpression of HER2 can be treated with targeted drug Herceptin. The HER2 index can be divided into four grades: -, 1+, 2+, and 3+. P63: In situ carcinoma expression is negative, invasive carcinoma expression is positive, and part of the basal cell type invasive carcinoma can express positive. Molecular typing includes two types: HER2(+) HR(-) and HER2(+) HR(+), pathological classification includes breast invasive carcinoma and ductal carcinoma in situ, and response includes non-PCR (non-complete remission) and pCR (complete remission).
[0044] In some example embodiments, the sample data includes gene expression data and corresponding clinical information.
[0045] In the embodiments of the present disclosure, the sample data can be obtained by downloading from a public database. For example, the gene expression data and the corresponding clinical information can be obtained from the Gene Expression Omnibus (GEO). The clinical information includes survival time and survival status, etc. The GEO database is a gene expression database created and maintained by the National Center for Biotechnology Information (NCBI) of the United States. Since its establishment in 2000, GEO has collected gene expression data from research institutions around the world, covering multiple fields such as tumors, non-tumors, chips, next-generation sequencing (NGS), differential analysis, and molecular verification.
[0046] In some example embodiments, before step 101, the method further comprises:
[0047] The intersection of multiple gene expression data with survival time and survival state is taken, and the batch effect is removed to obtain a gene expression matrix denoted as E ∈ R g×n where g represents the number of intersection genes, and n represents the number of samples.
[0048] In multi-omics data, batch effect refers to the difference between samples caused by differences in experimental conditions, operators, instruments, and other factors. This difference may mask the real biological changes, thereby affecting the subsequent analysis results. There are various methods for processing batch effects. One method for removing batches is to use standardization techniques to normalize the differences between samples, thereby eliminating batch effects.
[0049] In some example embodiments, in step 101, the scores or proportions of multiple cell types are calculated by immune infiltration algorithms such as cibersort, xCell, and Kassandra, etc. deconvolution tools, denoted as a cell proportion matrix X ∈ R c×n where c represents the number of cell types.
[0050] Immunoinfiltration refers to the process by which immune cells (such as T cells, B cells, macrophages, etc.) cross the blood vessel walls of tissues or organs to enter the tissues or organs in the body to respond to infections, inflammation, tumors, or other diseases. Immunoinfiltration analysis is a method for studying the distribution and activity of immune cells in tissues or tumors. It helps to understand the response of the immune system under different disease states, especially tumors / cancers.
[0051] The tumor microenvironment (TME) regulates the survival, maintenance, growth, and immune surveillance of tumors, and plays an important role in disease progression and treatment response, such as local drug resistance, immune escape, and cancer metastasis. Immune cells, as an important component of the tumor microenvironment (TME), play a key role in the physiological function of tumors. Analyzing the composition and function of different cell populations in the TME through immune infiltration helps to explore more effective tumor treatment methods.
[0052] For example, still using the data in Table 1 as an example, the xCell deconvolution tool is used to calculate the cell proportions, and the xCell deconvolution tool obtains 64 cell types including B cells and their corresponding proportions.
[0053] Step 102, screening multiple cells related to treatment response from multiple cell types;
[0054] In some example embodiments, screening multiple cells related to treatment response from multiple cell types includes:
[0055] The statistical significance p-value (p_value) of each cell type in the treatment response is calculated using the T-test or F-test method, and N cell types with p_value < a preset first p-value threshold (such as 0.1) are screened out as N cells related to the treatment response, N is the number of screened cell types, and 1 < N < c.
[0056] For example, still taking the sample data in Table 1 as an example, 13 cell types related to the treatment response are screened out by the T-test or F-test method: B-cells, CD4+ naive T-cells, CD8+ T-cells, CD8+ Tcm, Class-switched memory B-cells, ly Endothelial cells, Mast cells, Monocytes, naive B-cells, NKT, Plasma cells, pro B-cells, and Tgd cells. The T-test studies whether there is a difference between two groups of data by comparing the mean values of different data. The F-test, also known as the variance ratio test or homogeneity test, is a test under the null hypothesis (H0) that the statistical value follows the F distribution. It is usually used to analyze statistical models with more than one parameter to determine whether all or part of the parameters in the model are suitable for estimating the whole.
[0057] In statistics, the p-value (p_value) is an important concept for testing hypotheses. The definition of the p-value is: the probability of obtaining equal or more extreme results than the actual observation under the assumption that the hypothesis is true. The p-value threshold is generally set to 0.05 or 0.1, and the result is considered statistically significant if the p-value is lower than the threshold, for example, p_value < 0.1, indicating that the difference shown is less than 10% likely to be due to chance, or in other words, the probability of obtaining opposite conclusions in the same conditions and repeated studies is less than 10%.
[0058] In some example embodiments, the method further comprises:
[0059] determining the number of immunophenotyping groups (i.e., the number of clusters) k.
[0060] In the embodiments of the present disclosure, the number of immunophenotyping groups (i.e., the number of clusters) k can be set as needed. For example, the number of immunophenotyping groups k can be between 2 and 5, however, the embodiments of the present disclosure do not limit this.
[0061] For example, the number of immunophenotyping groups for the pre-treatment sample and the post-treatment sample is set to 2. For the pre-treatment sample, the immunophenotyping groups include IM1 type and IM2 type, and the predicted treatment response of the samples in the IM1 type group is better than that of the samples in the IM2 type group. For the post-treatment sample, the immunophenotyping groups include TM1 type and TM2 type, and the treatment response of the samples in the TM1 type group is better than that of the samples in the TM2 type group.
[0062] Step 103, performing cluster analysis on a plurality of cell combinations, each cell combination including at least part or all of the cells in the screened plurality of cells;
[0063] In some example embodiments, before step 103, the method further includes: dividing the N cells into M fixed cells and (N-M) candidate cells, and each cell combination includes at least the M fixed cells and q candidate cells, 0≤M
[0064] In some example embodiments, dividing the N cells into M fixed cells and (N-M) candidate cells includes:
[0065] For each cell in the N cells, the following operations are performed:
[0066] Obtaining a treatment response related p-value and an odds ratio (OR) corresponding to the cell;
[0067] When the treatment response related p-value corresponding to the cell is less than a preset second p-value threshold and the absolute value of the odds ratio is greater than a preset odds ratio threshold, the cell is divided into a fixed cell;
[0068] Dividing the cells other than the fixed cells in the N cells into candidate cells.
[0069] The OR value, also known as the odds ratio, is used to evaluate the number of times the presence of a factor leads to a change in the result. When the OR value is greater than 1, it indicates that the exposure factor is a promoting factor for the occurrence of positive events. When the OR value is less than 1, it indicates that the exposure factor is a hindering factor for the occurrence of positive events. When the OR value is equal to 1, it indicates that the exposure factor has no effect on the occurrence of positive events, i.e., a particularly large or small OR value indicates a relatively obvious hindering or promoting effect on the occurrence of events.
[0070] In the embodiments of the present disclosure, the treatment response related p-value is a statistical significance p-value related to the treatment response group. For example, the preset second p-value threshold can be between 0.01 and 0.05, and the preset OR threshold can be greater than 1.25 or less than 0.75. For example, the preset second p-value threshold can be 0.05, and the preset OR threshold can be greater than 1.25 or less than 0.75. However, the embodiments of the present disclosure are not limited thereto.
[0071] Still taking the sample data in Table 1 as an example, since the number of obtained cell types is small, no fixed cell type is selected, i.e., M = 0 and q = (N-M) = 13.
[0072] In some example embodiments, in step 103, clustering analysis is performed on the plurality of cell combinations, including:
[0073] determining all possible cell combinations;
[0074] for each cell combination, performing the following operations: extracting the cell type matrix corresponding to the cell combination, performing non-negative matrix factorization on the cell type matrix to obtain a clustering result.
[0075] In the embodiments of the present disclosure, the clustering method used when performing clustering analysis on the plurality of cell combinations can also be other clustering algorithms, such as hierarchical clustering, K-means clustering, density clustering, and community detection algorithms, and the present disclosure does not limit this.
[0076] In the embodiments of the present disclosure, it is assumed that for each cell combination, the number of extracted candidate cells is m, since m cell types are selected from the q candidate cell type groups, there are combinations of cell types, therefore, all possible cell combinations include , where p is the minimum number of extracted cell types.
[0077] For example, p can be between 4 and 7. In the embodiments of the present disclosure, in order to ensure that the characteristic cell type number of each immunophenotyping group is greater than or equal to 3, the present disclosure sets the value of p to 6.
[0078] Taking the aforementioned 13 cell types as an example, the fixed cell number is 0 and the candidate cell number is 13, in all possible cell combinations, the value range of the extracted candidate cell number m is from p to N-M, i.e., from 6 to 13, therefore, all possible cell combinations include .
[0079] In some example embodiments, for each cell combination, the following operations are performed: extracting the cell type matrix of the fixed cell and the candidate cell from the original cell proportion matrix X ∈ R c×n , denoted as D ∈ R (M+m)×n ; using a non-negative matrix factorization algorithm to decompose the matrix D, D (M+m)×n = W (M+m)×k × H k×n , finding the row index corresponding to the maximum value in each column of the coefficient matrix H k×n , i.e., the clustering result, denoted as cluster = {arg max i=1,…,k (H k,1 ), arg maxi=1,…,k (H k,2 ),…,arg max i=1,…,k (H k,n )}。
[0080] Exemplarily, assuming that the currently selected candidate cell number m = 11, a cell type matrix of the candidate cells is extracted from the original cell proportion matrix X (the fixed number of extracted cells is 0), denoted as D ∈ R 11×47 ; using a non-negative matrix factorization algorithm, D 11×47 = W 11×2 × H 2×47 can be obtained; the row index corresponding to the maximum value in each column of the coefficient matrix H 2×47 is obtained, that is, the clustering result, denoted as cluster = {arg max i=1,…,2 (H 2,1 ), arg max i=1,…,2 (H 2,2 ), …, arg max i=1,…,2 (H 2,47 )}.
[0081] Step 104: selecting one of the cell combinations as the cell combination for immunophenotyping according to the clustering analysis result, and determining the characteristic cell type corresponding to each immunophenotyping group.
[0082] In some exemplary embodiments, selecting one of the cell combinations as the cell combination for immunophenotyping according to the clustering analysis result comprises:
[0083] For the clustering analysis result of each cell combination, the corresponding average silhouette coefficient and average treatment response proportion are calculated;
[0084] Selecting one of the cell combinations as the cell combination for immunophenotyping, the average silhouette coefficient corresponding to the clustering analysis result of the selected cell combination is greater than a preset silhouette coefficient threshold, and the average treatment response proportion corresponding to the clustering analysis result of the selected cell combination is greater than a preset average proportion threshold.
[0085] The silhouette coefficient is an index for evaluating the clustering effect, and its essence is to measure the ratio of the distance between each sample point and the sample in its cluster to the distance between its nearest cluster structure. The value range of the silhouette coefficient is [-1, 1], and the closer the silhouette coefficient is to 1, the better the clustering performance, and vice versa. The silhouette coefficient of the clustering result is the average value of the silhouette coefficients of each sample point.
[0086] Exemplarily, the preset profile coefficient threshold value can be between 0.5 and 0.8, for example, the preset profile coefficient threshold value can be 0.6, however, the embodiments of the present disclosure do not make any limitation in this regard.
[0087] In some exemplary embodiments, the clustering analysis result includes a first sample grouping and a second sample grouping, and a corresponding treatment response average proportion is calculated according to the following formula:
[0088] mean_proption_of_Response = mean(max(cluster1_pCR, cluster1_nonPCR) + max(cluster2_pCR, cluster2_nonPCR));
[0089] Wherein, mean_proption_of_Response is the treatment response average proportion, mean represents the mean operation, max represents the maximum operation, cluster1_pCR is the pCR sample proportion in the first sample grouping, cluster1_nonPCR is the nonPCR sample proportion in the first sample grouping, cluster2_pCR is the pCR sample proportion in the second sample grouping, and cluster2_nonPCR is the nonPCR sample proportion in the second sample grouping.
[0090] For example, assuming that the proportions of pCR samples and nonPCR samples in the first sample grouping cluster1 are 0.8:0.2, and the proportions of pCR samples and nonPCR samples in the second sample grouping cluster2 are 0.1:0.9, then the treatment response average proportion mean_proption_of_Response = mean(0.8+0.9) = 0.85. The greater the value of the treatment response average proportion mean_proption_of_Response, the more consistent the treatment effect of the samples in each sample grouping.
[0091] Exemplarily, the preset average proportion threshold value can be between 0.6 and 0.8, for example, the preset average proportion threshold value can be 0.70, however, the embodiments of the present disclosure do not make any limitation in this regard.
[0092] Exemplarily, taking the pre-treatment samples in the foregoing Table 1 as an example, as described above, since the number of obtained cell types is small, a fixed cell type is not selected, i.e., M = 0 and q = (N-M) = 13. The value range of the extracted candidate cell number m is from 6 to 13, therefore, all possible cell combinations include The number of clusters k is set to 2.
[0093] For each cell combination, the following operations are performed: a cell type matrix of the candidate group is extracted from the original cell proportion matrix X, denoted as D e R m×47 Using the NMF non-negative matrix factorization algorithm, D m×47 = W m×2 x H 2×47 can be obtained. The row index corresponding to the maximum value in each column of the coefficient matrix H 2×47 is the clustering result, and the average silhouette coefficient silhouette and the average proportion of response mean_proption_of_Response corresponding to the clustering result are calculated. The clustering analysis result graph shown in FIG. 2 is obtained, and the cell type combination with an average silhouette coefficient silhouette > 0.6 and an average proportion of response mean_proption_of_Response > 0.7 and the clustering result thereof (i.e., any one of all cell combinations in the gray area in FIG. 2) can be selected as the final immunophenotyping. In FIG. 2, each plus sign or plus sign with a circle represents a cell combination, the horizontal coordinate is the average silhouette coefficient silhouette, and the vertical coordinate is the average proportion of response.
[0094] For the post-treatment samples in Table 1, similar methods as before can be used to obtain the cell type combination and the clustering result of the immunophenotyping shown in FIG. 3, and the cell type combination with an average silhouette coefficient silhouette > 0.75 and an average proportion of response mean_proption_of_Response > 0.75 and the clustering result thereof (i.e., any one of all cell combinations in the gray area in FIG. 3) can be selected as the final immunophenotyping.
[0095] In some example embodiments, when there are multiple cell combinations whose clustering analysis result corresponds to an average silhouette coefficient greater than a preset silhouette coefficient threshold and an average proportion of response greater than a preset average proportion threshold, the selected cell combination is the cell combination with the largest average proportion of response among the multiple cell combinations that satisfy the corresponding average silhouette coefficient greater than the preset silhouette coefficient threshold and the corresponding average proportion of response greater than the preset average proportion threshold.
[0096] For the pre-treatment samples in Table 1, the cell combination selected in Figure 2 as the final immunophenotyping is the one that satisfies the conditions of having an average silhouette coefficient greater than the preset silhouette coefficient threshold, having an average proportion of treatment response greater than the preset average proportion threshold, and having the largest average proportion of treatment response. Specifically, the cell combination is (Tgd cells, naive B-cells, CD4+ naive T-cells, B-cells, Class-switched memory B-cells, ly Endothelial cells, CD8+ Tcm).
[0097] For the post-treatment samples in Table 1, the cell combination selected in Figure 3 as the final immunophenotyping is the one that satisfies the conditions of having an average silhouette coefficient greater than the preset silhouette coefficient threshold, having an average proportion of treatment response greater than the preset average proportion threshold, and having the largest average proportion of treatment response. Specifically, the cell combination is (Monocytes, naive B-cells, ly Endothelial cells, Class-switched memory B-cells, CD4+ naive T-cells, pro B-cells, NKT).
[0098] In some example embodiments, determining the characteristic cell type corresponding to each immunophenotyping group comprises:
[0099] For each cell type in the cell combination of the immunophenotyping, the following operations are performed: calculating the median of the cell type in each immunophenotyping group; determining the characteristic cell type of the immunophenotyping group corresponding to the cell type with the largest median.
[0100] In this embodiment, the characteristic cell type of each immunophenotyping group is determined by calculating the median of the proportion of each cell type in each immunophenotyping group. The cell type with the largest median in the immunophenotyping group is the characteristic cell type of the immunophenotyping group. For example, the medians of B cells in the three immunophenotyping groups are 0.05, 0.07, and 0.01, respectively. Therefore, B cells are the characteristic cell of the second immunophenotyping group.
[0101] As shown in FIG. 4, for the pre-treatment samples in Table 1, the characteristic cell types of the first immune typing group IM1 are (Tgd cells, naive B-cells, CD4+ naive T-cells, B-cells), and the characteristic cell types of the second immune typing group IM2 are (Class-switched memory B-cells, ly Endothelial cells, CD8+ Tcm).
[0102] As shown in FIG. 5, for the post-treatment samples in Table 1, the characteristic cell types of the first immune typing group TM1 are (Monocytes, naive B-cells, ly Endothelial cells, Class-switched memory B-cells, CD4+ naive T-cells, ), and the characteristic cell types of the second immune typing group TM2 are (pro B-cells, NKT).
[0103] FIG. 6 is a schematic diagram of the correlation between immune typing of pre-treatment samples and treatment response. FIG. 7 is a schematic diagram of the correlation between immune typing of post-treatment samples and treatment response. As shown in FIG. 8, the proportions of pCR of the treatment responses of the immune typing combinations of the pre-treatment and post-treatment samples, IM1+TM1, IM1+TM2, IM2+TM1, and TM2+TM2, from more to less, indicate that the immune typing before and after treatment is significantly correlated with the treatment response. Among them, IM1+TM1 means that the sample belongs to both IM1 and TM1, IM1+TM2 means that the sample belongs to both IM1 and TM2, IM2+TM1 means that the sample belongs to both IM2 and TM1, and IM2+TM2 means that the sample belongs to both IM2 and TM2.
[0104] In some example embodiments, the method further comprises determining the differentially expressed genes corresponding to each immune typing group, wherein the differentially expressed genes satisfy: a statistical significance p-value with respect to survival time is less than a preset third p-value threshold, a statistical significance p-value with respect to gene expression amount is less than a preset fourth p-value threshold, and log2 f c >1.
[0105] In this embodiment, the differentially expressed genes refer to genes that are differentially expressed in two or more groups of samples, that is, the expression levels of the genes are significantly different in different groups of samples.
[0106] In some examples, the screening conditions of the differential genes of each immune typing group are survival time related p_value < 0.05, log2 f c >1, and gene expression amount related p_value < 0.05.
[0107] In the embodiments of the present disclosure, the p_value related to the gene expression quantity, i.e., the statistical significance p-value related to the gene expression quantity, can be obtained by T-test or F-test method.
[0108] In some example embodiments, the method further comprises:
[0109] Obtaining a verification set, and verifying the verification set using the immunophenotyped cell combination and the differential genes thereof.
[0110] In the embodiments of the present disclosure, after obtaining the immunophenotyped cell combination, the verification of the verification set can be performed using the cell types and the differential genes determined by the training set immunophenotyping, and the results show that the verification set results are consistent with the training set.
[0111] In some example embodiments, the verification of the verification set using the immunophenotyped cell combination and the differential genes thereof comprises:
[0112] The scores or proportions of a plurality of cell types are calculated by the immune infiltration algorithm;
[0113] The immunophenotype of the sample is determined according to the characteristic cell types of each immunophenotype grouping obtained in the training set;
[0114] The gene expression is detected according to the differential genes of each immunophenotype grouping obtained in the training set.
[0115] For example, the gene expression data of 62 her2-positive breast cancer samples downloaded from the TCGA database are used as independent verification data, and the proportions of a plurality of cell types of the TCGA verification set are calculated by the immune infiltration algorithm and the same deconvolution tool xCell as the training set. The immunophenotype of the verification set sample is determined according to the characteristic cell types of the pre-treatment immunophenotype, as shown in FIG. 9. According to the immunophenotype of the TCGA verification set, the cox and limma toolkits of R language are used to find the differential genes related to survival, and the screening conditions are pvalue<0.05, log2fc>1, and differential expression pvalue<0.05. The top 4 differential genes related to survival of IM1 and IM2 are selected, as shown in FIGS. 10A to 10H. As can be seen from the results of FIGS. 10A to 10H, in the IM1 typing, the gene expressions of RAD21, BSPRY, TATDN1, and NUDCD1 are significantly up-regulated, and the prognosis of survival is good; in the IM2 typing, the gene expressions of ANAPC13, CAPN6, SLC25A29, and TUBAL3 are significantly up-regulated, and the prognosis of survival is poor.
[0116] The present disclosure provides an immune typing method based on immune infiltration analysis, which comprises immune infiltration analysis on sample data to obtain the proportion of a plurality of cell types; screening a plurality of cells related to treatment response from the plurality of cell types; performing cluster analysis on a plurality of cell combinations, each of which at least includes part or all of the plurality of cells screened; selecting one of the cell combinations as an immune typing cell combination according to the cluster analysis result, and determining the corresponding characteristic cell type of each immune typing group. The present disclosure divides the pre-treatment breast cancer immune typing into IM1 and IM2 types by comparing and analyzing the differences in treatment response and immune-related cell proportion of the pre-treatment and post-treatment transcriptome expression data, the characteristic cell types of IM1 are (Tgd cells, naive B-cells, CD4+ naive T-cells, B-cells), and the characteristic cell types of IM2 are (Class-switched memory B-cells, ly Endothelial cells, CD8+ Tcm), the post-treatment breast cancer immune typing is divided into TM1 and TM2 types, the characteristic cell types of TM1 are (Monocytes, naive B-cells, ly Endothelial cells, Class-switched memory B-cells, CD4+ naive T-cells), and the characteristic cell types of TM2 are (pro B-cells, NKT). The immune typing method of the present disclosure embodiment makes the proportion of treatment response pCR significantly related to the immune typing before and after treatment.
[0117] As shown in FIG. 11, the present disclosure embodiment also provides an immune typing device, comprising an immune infiltration module 1110, a screening module 1120, a cluster analysis module 1130 and a determination module 1140, wherein:
[0118] The immune infiltration module 1110 is configured to perform immune infiltration analysis on sample data to obtain the proportion of a plurality of cell types;
[0119] The screening module 1120 is configured to screen a plurality of cells related to treatment response from the plurality of cell types;
[0120] The cluster analysis module 1130 is configured to perform cluster analysis on a plurality of cell combinations, each of which at least includes part or all of the plurality of cells screened;
[0121] The determination module 1140 is configured to select one of the cell combinations as an immune typing cell combination according to the cluster analysis result, and determine the corresponding characteristic cell type of each immune typing group.
[0122] In some example embodiments, the determining module 1140 selects one of the cell combinations as the cell combination for the immune typing according to the clustering analysis results, including:
[0123] For the clustering analysis result of each cell combination, the corresponding average profile coefficient and average proportion of treatment response are calculated;
[0124] The selected cell combination has a clustering analysis result corresponding to an average profile coefficient greater than a preset profile coefficient threshold and an average proportion of treatment response greater than a preset average proportion threshold.
[0125] In some example embodiments, the clustering analysis result includes a first sample group and a second sample group, and the determining module 1140 calculates the corresponding average proportion of treatment response according to the following formula:
[0126] mean_proption_of_Response=mean(max(cluster1_pCR, cluster1_nonPCR)+max(cluster2_pCR, cluster2_nonPCR));
[0127] wherein mean_proption_of_Response is the average proportion of treatment response, mean represents the mean operation, max represents the maximum operation, cluster1_pCR is the proportion of pCR samples in the first sample group, cluster1_nonPCR is the proportion of nonPCR samples in the first sample group, cluster2_pCR is the proportion of pCR samples in the second sample group, and cluster2_nonPCR is the proportion of nonPCR samples in the second sample group.
[0128] In some example embodiments, when there are multiple cell combinations whose clustering analysis results correspond to an average profile coefficient greater than a preset profile coefficient threshold and an average proportion of treatment response greater than a preset average proportion threshold, the cell combination selected by the determining module 1140 is the cell combination with the largest average proportion of treatment response among the multiple cell combinations that satisfy the condition of the corresponding average profile coefficient being greater than the preset profile coefficient threshold and the corresponding average proportion of treatment response being greater than the preset average proportion threshold.
[0129] In some example embodiments, the determining module 1140 determines the characteristic cell type corresponding to each immune typing group, including:
[0130] For each cell type in the cell composition of the immune typing, the following operations are performed: the median of the cell type in each immune typing group is calculated; and the cell type is determined as a characteristic cell type corresponding to an immune typing group with the largest median.
[0131] In some example embodiments, the determining module 1140 is further configured to determine a differentially expressed gene corresponding to each immune typing group, the differentially expressed gene satisfying: a statistical significance p-value with respect to survival time being less than a preset third p-value threshold, a statistical significance p-value with respect to gene expression amount being less than a preset fourth p-value threshold, and log2 f c >1,
[0132] In some example embodiments, the clustering analysis module 1130 employs any one of the following clustering methods when performing clustering analysis on the plurality of cell compositions: non-negative matrix factorization, hierarchical clustering, K-means clustering, density clustering, and community detection.
[0133] In some example embodiments, the sample data comprises HER2-positive breast cancer pre-treatment samples.
[0134] The immune typing groups comprise IM1 type and IM2 type, the predicted treatment response of the IM1 type group samples is better than that of the IM2 type group samples, the characteristic cell types of the IM1 type group comprise Tgd cells, naive B-cells, CD4+ naive T-cells and B-cells, and the characteristic cell types of the IM2 type group comprise Class-switched memory B-cells, ly Endothelial cells and CD8+ Tcm.
[0135] In some example embodiments, the sample data comprises HER2-positive breast cancer post-treatment samples.
[0136] The immune typing groups comprise TM1 type and TM2 type, the treatment response of the TM1 type group samples is better than that of the TM2 type group samples, the characteristic cell types of the TM1 type group comprise Monocytes, naive B-cells, ly Endothelial cells, Class-switched memory B-cells and CD4+ naive T-cells, and the characteristic cell types of the TM2 type group comprise pro B-cells and NKT.
[0137] In some example embodiments, the screening module 1120 screens the plurality of cell types to obtain a plurality of cells associated with the treatment response from the plurality of cell types, including:
[0138] Using a T-test or F-test method, a statistical significance p-value of each cell type in the treatment response is calculated, and a plurality of cell types with a p-value less than a preset first p-value threshold are selected as the plurality of cells associated with the treatment response.
[0139] The embodiments of the present disclosure also provide an immunophenotyping device, including a memory; and a processor connected to the memory, the memory is configured to store instructions, and the processor is configured to execute the steps of the immunophenotyping method according to any of the embodiments of the present disclosure based on the instructions stored in the memory.
[0140] As shown in FIG. 12, in one example, the immunophenotyping device can include a processor 1210, a memory 1220, a bus system 1230, and a transceiver 1240, wherein the processor 1210, the memory 1220, and the transceiver 1240 are connected through the bus system 1230, the memory 1220 is configured to store instructions, and the processor 1210 is configured to execute the instructions stored in the memory 1220 to control the transceiver 1240 to transceive signals. Specifically, the transceiver 1240 can acquire sample data under the control of the processor 1210, the processor 1210 performs immune infiltration analysis on the sample data to obtain the proportion of the plurality of cell types; screens a plurality of cells associated with the treatment response from the plurality of cell types; performs clustering analysis on a plurality of cell combinations, each of which includes at least part or all of the screened plurality of cells; selects one of the cell combinations as the cell combination for immunophenotyping according to the clustering analysis result, and determines the characteristic cell type corresponding to each immunophenotyping group.
[0141] It should be understood that the processor 1210 can be a central processing unit (CPU), and the processor 1210 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0142] The memory 1220 can include read-only memory and random access memory, and provide instructions and data to the processor 1210. A part of the memory 1220 can also include non-volatile random access memory. For example, the memory 1220 can also store device type information.
[0143] The bus system 1230 can include a data bus, a power bus, a control bus, and a state signal bus, etc. in addition to the data bus. However, for the sake of clarity, all the buses are denoted as the bus system 1230 in FIG. 12.
[0144] In the implementation process, the processing performed by the processing device can be completed by the integrated logic circuit of the hardware in the processor 1210 or the instructions in the form of software. That is, the method steps of the embodiments of the present disclosure can be embodied as being completed by the hardware processor or being completed by the combination of the hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or the like. The storage medium is located in the memory 1220, and the processor 1210 reads the information in the memory 1220 and completes the steps of the above method in combination with the hardware. To avoid repetition, no longer detailed description is made here.
[0145] The embodiments of the present disclosure also provide a computer readable storage medium, which stores a computer program. The program is executed by a processor to implement the immunotyping method according to any of the embodiments of the present disclosure. The immunotyping method driven by the executable instructions is basically the same as the immunotyping method provided by the above embodiments of the present disclosure, and will not be described here.
[0146] In some possible implementation manners, various aspects of the immunotyping method provided by the embodiments of the present disclosure can also be implemented in the form of a program product, which includes program codes for causing a computer device to perform the steps of the immunotyping method according to various exemplary embodiments of the present disclosure described above in the specification when the program product is run on the computer device, for example, the computer device can execute the immunotyping method recorded in the embodiments of the present disclosure.
[0147] The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0148] Those of ordinary skill in the art will realize and understand that all or some of the steps in the methods disclosed above and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the components can be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer-readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Furthermore, it is common knowledge to those of ordinary skill in the art that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and can include any information delivery media.
[0149] It should be noted that the above-described embodiments or implementations are merely exemplary and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions of the form and details of implementation can be made without departing from the scope of the present disclosure.
Claims
1. An immune typing method, comprising: performing immune infiltration analysis on sample data to obtain proportions of a plurality of cell types; screening a plurality of cells associated with treatment response from the plurality of cell types; performing cluster analysis on a plurality of cell combinations, each cell combination comprising at least part or all of the screened plurality of cells; selecting one of the cell combinations as an immune typing cell combination according to the cluster analysis result, and determining a characteristic cell type corresponding to each immune typing group.
2. The method of claim 1, wherein, The selecting one of the cell combinations as an immune typing cell combination according to the cluster analysis result comprises: calculating a corresponding average silhouette coefficient and a treatment response average proportion for the cluster analysis result of each cell combination; selecting one of the cell combinations as an immune typing cell combination, the selected cell combination having a cluster analysis result corresponding to an average silhouette coefficient greater than a preset silhouette coefficient threshold and a treatment response average proportion greater than a preset average proportion threshold.
3. The method of claim 2, wherein, The cluster analysis result comprises a first sample group and a second sample group, and the treatment response average proportion is calculated according to the following formula: mean_proption_of_Response=mean(max(cluster1_pCR, cluster1_nonPCR)+max (cluster2_pCR, cluster2_nonPCR)); wherein, mean_proption_of_Response is the treatment response average proportion, mean represents the mean operation, max represents the maximum operation, cluster1_pCR is the proportion of pCR samples in the first sample group, cluster1_nonPCR is the proportion of nonPCR samples in the first sample group, cluster2_pCR is the proportion of pCR samples in the second sample group, and cluster2_nonPCR is the proportion of nonPCR samples in the second sample group.
4. The method of claim 2, wherein, When there are multiple cell combinations whose cluster analysis results correspond to an average silhouette coefficient greater than a preset silhouette coefficient threshold and a treatment response average proportion greater than a preset average proportion threshold, the selected cell combination is the cell combination with the largest treatment response average proportion among the multiple cell combinations.
5. The method of claim 1, wherein, The determining a characteristic cell type corresponding to each immune typing group comprises: for each cell type in the immune typing cell combination, performing the following operations: calculating the median of the cell type in each immune typing group; and determining the cell type as a characteristic cell type corresponding to the immune typing group with the largest median.
6. The method of claim 1, wherein, The screening a plurality of cells associated with treatment response from the plurality of cell types comprises: The statistical significance p-value of each cell type in the treatment response is calculated using a T-test or F-test method, and a plurality of cell types with a statistical significance p-value less than a preset first p-value threshold in the treatment response are selected as the plurality of cells related to the treatment response.
7. The method of claim 1, further comprising: determining differentially expressed genes corresponding to each of the immunotyping groups, the differentially expressed genes satisfying: a statistically significant p-value with respect to survival time being less than a preset third p-value threshold, a statistically significant p-value with respect to gene expression amount being less than a preset fourth p-value threshold, and 8. The method of claim 1, wherein, The clustering method used in the clustering analysis of the plurality of cell combinations is any one of the following: non-negative matrix factorization, hierarchical clustering, K-means clustering, density clustering, and community detection.
9. The method of claim 1, wherein, The sample data includes HER2-positive breast cancer pre-treatment samples. The immunophenotyping grouping includes IM1 type and IM2 type, the expected treatment response of the IM1 type grouped sample is better than that of the IM2 type grouped sample, the characteristic cell types of the IM1 type grouping include: Tgd cells, naive B-cells, CD4+ naive T-cells and B-cells, and the characteristic cell types of the IM2 type grouping include Class-switched memory B-cells, ly Endothelial cells and CD8+ Tcm.
10. The method of claim 1, wherein, The sample data includes HER2-positive breast cancer post-treatment samples. The immunophenotyping grouping includes TM1 type and TM2 type, the treatment response of the TM1 type grouped sample is better than that of the TM2 type grouped sample, the characteristic cell types of the TM1 type grouping include Monocytes, naive B-cells, ly Endothelial cells, Class-switched memory B-cells and CD4+ naive T-cells, and the characteristic cell types of the TM2 type grouping include pro B-cells and NKT.
11. An immunophenotyping device comprising a memory; and a processor connected to the memory, the memory being configured to store instructions, and the processor being configured to perform the steps of the immunophenotyping method according to any one of claims 1 to 10 based on the instructions stored in the memory.
12. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the immunophenotyping method according to any one of claims 1 to 10.
13. A computer program product comprising instructions which, when executed by a computer, perform the immunophenotyping method according to any one of claims 1 to 10.
14. An immunophenotyping device comprising an immune infiltration module, a screening module, a clustering analysis module and a determination module, wherein: The immune infiltration module is configured to perform immune infiltration analysis on sample data to obtain proportions of a plurality of cell types; The screening module is configured to screen a plurality of cells related to the treatment response from the plurality of cell types; The clustering analysis module is configured to perform clustering analysis on a plurality of cell combinations, each cell combination including at least some or all of the screened plurality of cells; and The determination module is configured to determine the plurality of cells related to the treatment response from the plurality of cell types based on the clustering analysis. The determining module is configured to select one of the cell combinations as the cell combination of the immunophenotype according to the clustering analysis result, and determine the characteristic cell type corresponding to each immunophenotype group.