Immunophenotyping method and device, storage medium, and program product

Through immune infiltration analysis and cluster analysis, characteristic cell combinations of DLBCL were screened out, which solved the problem of uncertain cell origin in existing DLBCL typing methods, achieved more accurate immune typing and biological heterogeneity interpretation, and is suitable for auxiliary analysis of hematological tumors and solid tumors.

WO2025200857A1PCT designated stage Publication Date: 2025-10-02BOE TECHNOLOGY GROUP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/077771
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-02-18
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing DLBCL typing methods cannot accurately determine the cell of origin, resulting in traditional typing being unable to meet clinical needs, and genotyping-based methods cannot predict undiscovered gene mutations or expression information.

Method used

The proportions of multiple cell types were obtained through immune infiltration analysis. Survival significance, variance, and correlation were used to screen N cell types, which were divided into M fixed cells and (NM) candidate cells. Cluster analysis was performed, and a cell combination was selected as the cell combination for immunophenotyping, and the characteristic cell types were determined.

Benefits of technology

It explains the differences in tumor treatment responses, discovers more differentially expressed genes, explains biological heterogeneity, provides more accurate immunophenotyping, and is suitable for auxiliary analysis of hematological tumors and solid tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025077771_02102025_PF_FP_ABST
    Figure CN2025077771_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An immunophenotyping method and device, a storage medium, and a program product. The method comprises: performing immune infiltration analysis on sample data to obtain a proportion of a plurality of cell types; on the basis of at least one of the following: survival significance, variance, and correlation, screening for N cells from the plurality of cell types, wherein N>1; dividing the N cells into M fixed cells and (N-M) candidate cells, and performing clustering analysis on a plurality of cell combinations, wherein each cell combination at least comprises M fixed cells and q candidate cells, M≥1, M<N, and 0≤q≤(N-M); and selecting one of the cell combinations as a cell combination for immunophenotyping, and determining a characteristic cell type corresponding to each immunophenotyping group.
Need to check novelty before this filing date? Find Prior Art

Description

Immunophenotyping method and device, storage medium and program product

[0001] This application claims priority to PCT patent application number PCT / CN2024 / 083837, filed with the State Intellectual Property Office of China on March 26, 2024, with invention name “Immunotyping method and device, storage medium and program product”, and priority to Chinese patent application number 202410509310.X, filed with the State Intellectual Property Office of China on April 25, 2024, with invention name “Immunotyping method and device, storage medium and program product”, the contents of which should be understood as incorporated into this application by reference. Technical Field

[0002] The embodiments of the present disclosure relate to, but are not limited to, the field of biotechnology, and in particular to an immunophenotyping method and apparatus, a storage medium, and a program product. Background Art

[0003] Diffuse large B-cell lymphoma (DLBCL) is an aggressive tumor that arises from mature B cells. DLBCL exhibits significant clinical heterogeneity, with nearly 20 subtypes listed in the 2016 WHO classification. This high degree of heterogeneity poses significant challenges to the diagnosis and treatment of DLBCL.

[0004] Existing research methods for DLBCL classification include: ① Cell of origin (COO) classification, which is based on immunohistochemistry. Although it can determine the patient's prognosis, its limitation is that the classification accuracy rate for the same genotype population is only 88%, and the cell of origin cannot be determined for some patients. Therefore, stratified treatment based on cell of origin classification can no longer meet clinical needs. ② Based on genotyping, using methods such as exome sequencing and transcriptome sequencing, it was found that different genetic subtypes have obvious differences in genotype (gene mutation), epigenetic and clinical characteristics, providing a potential pathological basis for the precision treatment strategy of DLBCL. However, these studies are based on genotyping standards set by information on a few known mutant genes and cannot accurately predict undiscovered gene mutations or gene expression information. Summary of the Invention

[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0006] The present disclosure provides an immunophenotyping method, comprising:

[0007] Perform immune infiltration analysis on sample data to obtain the proportions of various cell types;

[0008] Select N cell types from multiple cell types based on at least one of the following: survival significance, variance, and correlation, where N>1;

[0009] Dividing the N cells into M fixed cells and (NM) candidate cells, and performing cluster analysis on multiple cell combinations, each cell combination includes at least M fixed cells and q candidate cells, M≥1 and M<N, 0≤q≤(NM);

[0010] One of the cell combinations is selected as the cell combination for immunophenotyping, and the characteristic cell type corresponding to each immunophenotyping group is determined.

[0011] An embodiment of the present disclosure also provides an immunophenotyping device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the immunophenotyping method described in any embodiment of the present disclosure based on the instructions stored in the memory.

[0012] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the immunophenotyping method described in any embodiment of the present disclosure.

[0013] An embodiment of the present disclosure further provides a program product, comprising instructions. When the computer program product is executed by a computer, the instructions execute the immunophenotyping method as described in any embodiment of the present disclosure.

[0014] The present disclosure also provides an immunophenotyping device, comprising an immune infiltration module, a screening module, a cluster analysis module, and a selection module, wherein:

[0015] The immune infiltration module is configured to perform immune infiltration analysis on the sample data to obtain the proportions of multiple cell types;

[0016] The screening module is configured to screen N cells from a plurality of cell types according to at least one of the following: survival significance, variance, and correlation, where N>1;

[0017] The cluster analysis module is configured to divide the N types of cells into M types of fixed cells and (NM) types of candidate cells, and perform cluster analysis on multiple cell combinations, each cell combination including at least M types of fixed cells and q types of candidate cells, where M ≥ 1 and M < N, and 0 ≤ q ≤ (NM);

[0018] The selection module is configured to select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping grouping.

[0019] The immunophenotyping method and device, storage medium and program product of the embodiment of the present disclosure, by performing immune infiltration analysis on sample data, obtain the proportion of multiple cell types, according to at least one of the following: survival significance, variance and correlation, screen out N cells from multiple cell types, divide the N cells into M fixed cells and (NM) candidate cells, and perform cluster analysis on multiple cell combinations; select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping grouping, which can mine the immune cell composition in different typings and better explain the difference in treatment response of tumors; adopt multiple combinations of cell types to find the differences between samples to the greatest extent, thereby establishing immunophenotyping, and making immunophenotyping have significant differences in survival analysis. In addition, the present disclosure does not need to determine the origin of cells, which solves the problem that traditional typing cannot clearly type because it cannot determine the origin of cells; the present disclosure can find more differentially expressed genes, thereby better explaining the biological heterogeneity of tumors, while taking into account that immune cells play an important role in the physiological function of tumors. The present disclosure uses the screened cell types and differentially expressed genes to obtain consistent results in independent validation sets, which can be used for auxiliary analysis of various types of hematological tumors and solid tumors.

[0020] Other features and advantages of the present disclosure will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present disclosure. Other advantages of the present disclosure can be realized and obtained through the solutions described in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are used to provide an understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0022] FIG1 is a schematic flow chart of an immunophenotyping method provided by an exemplary embodiment of the present disclosure;

[0023] FIG2 is a graph showing the ratio scores of 64 cell types obtained using the xCell tool according to an exemplary embodiment of the present disclosure;

[0024] 3A to 3S are survival curves of 19 types of cells provided by an exemplary embodiment of the present disclosure;

[0025] FIG4 is a heat map showing the correlation of 19 types of cells according to an exemplary embodiment of the present disclosure;

[0026] FIG5 is a risk ratio diagram of 19 cell types provided by an exemplary embodiment of the present disclosure;

[0027] FIG6 is a diagram showing cluster analysis results according to an exemplary embodiment of the present disclosure;

[0028] FIG7 is a survival curve diagram of the immunophenotyping cell combination finally selected in FIG6 ;

[0029] FIG8 is a diagram showing cluster analysis results of another exemplary embodiment of the present disclosure;

[0030] FIG9 is a survival curve diagram of the immunophenotyping cell combination finally selected in FIG8 ;

[0031] FIG10 is a diagram of immunophenotyping and characteristic cell types provided by an exemplary embodiment of the present disclosure;

[0032] FIG11 is an immunophenotyping and differentially expressed gene graph provided by an exemplary embodiment of the present disclosure;

[0033] FIG12A is a diagram showing cluster analysis results according to another exemplary embodiment of the present disclosure;

[0034] FIG12B is a survival curve of the immunophenotyping cell combination finally selected in FIG12A;

[0035] FIG12C is a diagram of the immunophenotyping and characteristic cell types finally selected in FIG12A;

[0036] FIG13A is a diagram showing cluster analysis results according to another exemplary embodiment of the present disclosure;

[0037] FIG13B is a survival curve of the immunophenotyping cell combination finally selected in FIG13A;

[0038] FIG13C is a diagram of the immunophenotyping and characteristic cell types finally selected in FIG13A;

[0039] FIG14 is a schematic structural diagram of an immunophenotyping device provided by an exemplary embodiment of the present disclosure;

[0040] FIG15 is a schematic structural diagram of another immunophenotyping device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] The present disclosure describes a plurality of embodiments, but this description is exemplary rather than restrictive, and it will be apparent to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.

[0042] The present disclosure includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The disclosed embodiments, features, and elements of the present disclosure may also be combined with any conventional features or elements to form a unique inventive solution defined by the claims. Any features or elements of any embodiment may also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this disclosure may be implemented individually or in any appropriate combination. Therefore, the embodiments are not subject to other limitations except for the limitations set forth in the appended claims and their equivalents. In addition, various modifications and changes may be made within the scope of protection of the appended claims.

[0043] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as a limitation on the claims. In addition, the claims to the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can be changed and still remain within the spirit and scope of the disclosed embodiments.

[0044] As shown in FIG1 , the present disclosure provides an immunophenotyping method, comprising:

[0045] Step 101: Perform immune infiltration analysis on sample data to obtain the proportions of multiple cell types;

[0046] In some exemplary embodiments, in step 101, the sample data may be a DLBCL dataset. However, the present disclosure is not limited thereto, and the sample data may also be a mantle cell lymphoma (MCL) or any other type of disease dataset.

[0047] In some exemplary embodiments, the sample data includes gene expression data and corresponding clinical information.

[0048] In the disclosed embodiments, sample data can be downloaded from public databases. For example, gene expression data and corresponding clinical information can be obtained from the Gene Expression Omnibus (GEO), and the clinical information includes survival time and survival status. The GEO database is a gene expression database created and maintained by the National Center for Biotechnology Information (NCBI) of the United States. Since its establishment in 2000, GEO has collected gene expression data from research institutions around the world, covering multiple fields such as tumors, non-tumors, chips, next-generation sequencing (NGS), differential analysis and molecular verification.

[0049] In some exemplary embodiments, before step 101, the method further includes:

[0050] Take the intersection of genes for multiple gene expression data with survival time and survival status, remove the batch effect, and get the gene expression matrix denoted as E∈R g×n , where g represents the number of intersection genes and n represents the number of samples.

[0051] In multi-omics data, batch effects refer to differences between samples due to variations in experimental conditions, operators, instruments, and other factors. These differences can mask true biological variations, thereby affecting subsequent analytical results. There are various approaches to address batch effects. One approach to de-batch is to use normalization techniques to normalize differences between samples and thereby eliminate batch effects.

[0052] For example, 1,158 samples from nine DLBCL datasets were selected from the Gene Expression Comprehensive Database as the training set, of which 1,127 samples had survival time and survival status, as shown in Table 1. The intersection genes of these 1,127 samples were taken, resulting in a total of 10,696 intersection genes in this example. The combat function of the sva toolkit was then used to remove batch effects.

[0053] Table 1

[0054] Then, the scores or proportions of multiple cell types are calculated by using immune infiltration algorithms, such as cibersort, xCell and Kassandra deconvolution tools, and recorded as the cell proportion matrix X∈R c×n , where c represents the number of cell types.

[0055] Immune infiltration refers to the process by which immune cells (such as T cells, B cells, and macrophages) penetrate the blood vessels of tissues or organs and enter the body's tissues or organs to respond to infection, inflammation, tumors, or other diseases. Immune infiltration analysis is a method used to study the distribution and activity of immune cells in tissues or tumors. It helps understand the immune system's response to various disease states, particularly tumors and cancer.

[0056] The tumor microenvironment (TME) regulates tumor survival, maintenance, growth, and immune surveillance, playing a crucial role in disease progression and treatment response, including local drug resistance, immune escape, and cancer metastasis. Immune cells, as an essential component of the TME, play a key role in tumor physiology. Analyzing the composition and function of different cell populations in the TME through immune infiltration can help explore more effective tumor treatments.

[0057] For example, still taking the training set in Table 1 as an example, the xCell deconvolution tool is used to calculate the enrichment score of cells. The xCell deconvolution tool generates a total of 64 cell types including B cells, as shown in FIG2 .

[0058] Step 102: Screening N types of cells from the plurality of cell types based on at least one of the following: survival significance, variance, and correlation, where N is a natural number greater than 1;

[0059] In some exemplary embodiments, N cells are selected from a plurality of cell types based on survival significance, including:

[0060] The single-factor Cox survival analysis model was used to calculate the statistically significant p-value (p_value) of each cell type and the survival time and survival status, and N cell types with p_value < the preset p-value threshold (such as 0.05) were screened as N cell types significantly associated with prognosis survival, where N is the number of cell types screened, 1 <N<c。

[0061] Exemplary, still taking the training set of the aforementioned Table 1 as an example, 19 cell types significantly correlated with prognosis survival were screened out by the univariate cox survival analysis model. Figures 3A to 3S are survival curves of these 19 cell types, wherein the abscissa is the follow-up time (time, in days) and the ordinate is the survival probability (Survival probability). Each survival curve diagram includes a high expression curve, a low expression curve and a corresponding p value, wherein high expression refers to a relatively high expression level in the biological sample, and low expression refers to a relatively low expression level in the biological sample. In general, the greater the difference between the two curves of the high expression curve and the low expression curve, the greater the difference in the prognosis of the two groups of patients, and the more likely a statistical difference is to occur, and the smaller the p value, the more statistically significant the result is, i.e., the survival time of the two groups of patients is significantly different. As shown in Figures 3A to 3S, 19 cell types include B-cells, CD4+memory T-cells, CD4+naive T-cells, etc., and the difference is significant in the survival curve diagram.

[0062] In statistics, the p-value is an important concept used to test hypotheses. The p-value is defined as the probability of obtaining a result that is equal to or more extreme than the actual observed result, given that the hypothesis holds. The p-value threshold is typically set at 0.05 or 0.01; a p-value below this threshold is considered statistically significant. For example, a p-value < 0.05 indicates that there is less than a 5% chance that the difference in the results is due to chance, or that the probability of reaching the opposite conclusion if the same study is repeated under the same conditions is less than 5%.

[0063] In some exemplary embodiments, selecting N types of cells from a plurality of cell types based on variance comprises:

[0064] The variance of the proportion of each cell type is calculated, and N cell types whose variance is greater than a preset variance threshold are selected as the screened N cell types.

[0065] The larger the variance, the greater the data fluctuation; the smaller the variance, the smaller the data fluctuation. Therefore, based on the variance, cell types with large differences in proportion can be screened as the N screened cell types. In the disclosed embodiment, the preset variance threshold can be set based on the actual variance data.

[0066] In some exemplary embodiments, N cells are selected from a plurality of cell types based on correlation, including:

[0067] Calculate the correlation between each cell type and other cell types, count the number of cell types whose absolute value of the calculated correlation is greater than a preset correlation threshold, and select N cell types whose counted number is greater than the preset number threshold as the screened N cell types.

[0068] For example, the preset correlation threshold may be 0.6. The preset number threshold may be a positive integer greater than 2. However, the present disclosure does not limit this, and the preset correlation threshold and the preset number threshold may be set as needed.

[0069] In some exemplary embodiments, the method further comprises:

[0070] Calculate the correlation between N types of cells;

[0071] The number of immunophenotyping groups (ie, the number of clusters) k is determined based on the correlation between N types of cells.

[0072] In the disclosed embodiment, the correlation between N types of cells can be calculated using the Spearman correlation analysis method, where the Spearman correlation coefficient ranges from -1 to 1, with a value of -1 indicating a perfect negative correlation, a value of 1 indicating a perfect positive correlation, and a value of 0 indicating no correlation between the two variables.

[0073] For example, FIG4 is a heat map showing the correlation between the aforementioned 19 cell types. As shown in FIG4 , based on the correlation between the 19 cell types, the number k of immunophenotyping groups is determined to be 3. However, the present disclosure is not limited thereto, and the number k of immunophenotyping groups can be set as needed.

[0074] In the embodiment of the present disclosure, the number k of immunotyping groups may be between 2 and 5. However, the embodiment of the present disclosure is not limited thereto, and the number of immunotyping groups may also be any other natural number greater than or equal to 2. The following example is exemplified with the number of immunotyping groups being 3.

[0075] Step 103: Divide the N types of cells into M fixed cells and (NM) candidate cells, and perform cluster analysis on multiple cell combinations, where each cell combination includes at least M fixed cells and q candidate cells, where M is a natural number greater than or equal to 1 and M<N, and q is a natural number between 0 and (NM);

[0076] In some exemplary embodiments, dividing N cells into M fixed cells and (NM) candidate cells comprises:

[0077] For each of the N cells, perform the following operations:

[0078] Get the p-value and risk ratio corresponding to the cell;

[0079] When the p-value corresponding to the cell is less than the preset p-value threshold and the absolute value of the risk ratio is greater than the preset risk ratio threshold, the cell is classified as a fixed cell;

[0080] The cells other than the fixed cells among the N types of cells are classified as candidate cells.

[0081] HR values, also known as hazard ratios, are primarily used in survival analysis. HR values ​​assess the multiple change in the risk of an outcome event caused by the presence of a factor. An HR value greater than 1 indicates that the exposure factor promotes the occurrence of a positive event, an HR value less than 1 indicates that the exposure factor hinders the occurrence of a positive event, and an HR equal to 1 indicates that the exposure factor has no effect on the occurrence of a positive event. Specifically, extremely large or small HR values ​​indicate a significant inhibitory or promoting effect on the occurrence of the event. Figure 5 shows the HR values ​​for these 19 cell types. The dots represent the model coefficients corresponding to each predictor variable, and the lines represent the 95% confidence intervals. The figure shows the lower bound (lower.95) and upper bound (upper.95) of the 95% confidence intervals for the model coefficients in exponential form, indicating the uncertainty range of the estimate.

[0082] For example, the preset p-value threshold may be between 0.01 and 0.05, and the preset HR value threshold may be greater than 1.25 or less than 0.75. For example, the preset p-value threshold may be 0.05, and the preset HR value threshold may be greater than 1.25 or less than 0.75, however, the embodiments of the present disclosure are not limited thereto.

[0083] In the survival curve diagrams of Figures 3A to 3S, the p-values ​​corresponding to the six cell types, namely, Th1 cells, Memory B-cells, Class-switched memory B-cells, CD8+T-cells, CD8+naive T-cells, and Tregs, are less than the preset p-value threshold, and in the HR diagram of Figure 5, the absolute values ​​of the HR values ​​of these six cell types are greater than the preset HR value threshold. Therefore, the embodiment of the present disclosure selects the six cell types, namely, Th1 cells, Memory B-cells, Class-switched memory B-cells, CD8+T-cells, CD8+naive T-cells, and Tregs, as fixed cells.

[0084] In some exemplary embodiments, cluster analysis is performed on multiple cell combinations, comprising:

[0085] Identify all possible cell combinations;

[0086] For each cell combination, the following operations are performed: extract the cell type matrix corresponding to the cell combination, perform non-negative matrix decomposition on the cell type matrix to obtain the clustering results, and calculate the average silhouette coefficient and p-value corresponding to the cell combination based on the clustering results.

[0087] In the disclosed embodiment, the clustering method used for cluster analysis of multiple cell combinations may also be other clustering algorithms, such as hierarchical clustering, K-means clustering, density clustering, and community detection algorithms. After obtaining the clustering results, the average silhouette coefficient and p-value corresponding to the cell combination are calculated based on the clustering results.

[0088] In some exemplary embodiments, all possible cell combinations include kind.

[0089] Taking the aforementioned 19 cells as an example, excluding the 6 fixed cells that have been selected, there are 13 candidate cells left. From the q candidate cell groups, m candidate cells are selected, with a total of The value of m ranges from 1 to q, so all possible cell combinations include kind.

[0090] In some exemplary embodiments, for each cell combination, the following operations are performed: from the original cell ratio matrix X∈R c×n Extract the cell type matrix of fixed cells and candidate cells, denoted as D∈R (M+m)×n ; Use the non-negative matrix decomposition algorithm to decompose the matrix D, and you can get D (M+m)×n =W (M+m)×k ×H k×n , find the coefficient matrix H k×n The row index corresponding to the maximum value in each column is the clustering result, which is recorded as cluster = {arg max i=1,…,k (H k,1 ),arg max i=1,…,k (H k,2 ),…,arg max i=1,...,k (H k,n )}, calculate the average silhouette coefficient of the cluster, and calculate the significance p_value of the survival difference between samples according to the cluster grouping.

[0091] For example, assuming that the number of candidate cells currently selected is m=5, the cell type matrix of fixed cells and candidate cells is extracted from the original cell ratio matrix X, which is recorded as D∈R (6+5)×1158 ; Using the non-negative matrix decomposition algorithm, we can get D (6+5)×1158 =W (6+5)×3 ×H3×1158 ; Find the coefficient matrix H 3×1158 The row index corresponding to the maximum value in each column is the clustering result, which is recorded as cluster = {arg max i=1,…,3 (H 3,1 ),arg max i=1,…,3 (H 3,2 ),…,arg max i=1,…,3 (H 3,1158 )}, calculate the average silhouette coefficient of the cluster, and calculate the significance p_value of the survival difference between samples according to the cluster grouping, and obtain the cluster analysis result diagram shown in Figure 6.

[0092] Step 104: Select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping group.

[0093] In some exemplary embodiments, selecting one of the cell combinations as the cell combination for immunophenotyping comprises:

[0094] Based on the p-value and silhouette coefficient corresponding to each cell combination, one of the cell combinations is selected as the cell combination for immunophenotyping. The p-value and silhouette coefficient of the selected cell combination meet the following conditions 1 and 2:

[0095] Condition 1: The p-value is less than the preset p-value threshold;

[0096] Condition 2: The silhouette coefficient is greater than the preset silhouette coefficient threshold.

[0097] The Silhouette Coefficient (SCo) is a metric used to evaluate clustering effectiveness. Essentially, it measures the ratio of the distance from each sample point to its own cluster to the distance between the nearest cluster structure. The Silhouette Coefficient ranges from -1 to 1. A Silhouette Coefficient closer to 1 indicates better clustering performance, while a Silhouette Coefficient closer to -1 indicates poorer performance. The resulting Silhouette Coefficient is the average of the Silhouette Coefficients for each sample point.

[0098] Exemplarily, the preset silhouette coefficient threshold may be between 0.7 and 0.8. For example, the preset silhouette coefficient threshold may be 0.75. However, the embodiment of the present disclosure is not limited thereto.

[0099] Exemplary, taking the aforementioned 8192 cell combinations as an example, FIG6 is a cluster analysis result diagram of the 8192 cell combinations. As shown in FIG6 , a cell type combination with an average silhouette coefficient silhouette> 0.7 and -log10 (p_value)> 4 and its clustering result (i.e., any one of all cell combinations located in the gray area in FIG6 ) can be selected as the final immunophenotyping. In FIG6 , each plus sign or circled plus sign represents a cell combination, the horizontal axis is the average silhouette coefficient silhouette, and the vertical axis is the significance p_value taking the opposite number of log10. The better the significance, that is, the smaller the pvalue, the larger the -log10 (p_value).

[0100] In some exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0101] Condition 3: f1(p_value)-A*f2(silhouette)-b≥0, where f1(p_value) is a function of p_value, and the larger the p_value, the smaller the value of f1(p_value); f2(silhouette) is a function of silhouette, and the larger the silhouette, the larger f2(silhouette), A<0, b>0.

[0102] For example, f1(p_value) = -log10(p_value); f2(silhouette) = silhouette. However, the present disclosure is not limited to this, and the specific function forms of f1(p_value) and f2(silhouette) can be set as needed. The following description uses f1(p_value) = -log10(p_value); f2(silhouette) = silhouette as an example.

[0103] In this embodiment, A represents the slope of a descending line a, b is the intercept of the descending line a, and condition 3 indicates that the cell combination is located above the descending line a. As shown in Figure 6, the cell type combination located in the gray area and to the upper right of the descending line a and its clustering result (i.e., any one of all cell combinations represented by a circled plus sign in Figure 6) can be selected as the final immunophenotyping.

[0104] In some exemplary embodiments, Among them, y2=max(f1(p_value)), y1=f1(p_value)p_value<0.01, x2=f2(silhouette)silhouette∈[0.5,08], x1=max(f2(silhouette)), b=y2-A·x2, max means taking the maximum value.

[0105] Taking f1(p_value)=-log10(p_value) and f2(silhouette)=silhouette as an example, the values ​​of y2, y1, x2, and x1 can be: y2=max(-log10(p_value)), y1=4, x2=0.7, x1=max(silhouette). However, the present disclosure is not limited to this, and the values ​​of y2, y1, x2, and x1 can be set as needed. For example, in other examples, y1 can be a real number greater than 2.

[0106] In some other exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0107] Condition 3: f1(p_value)-A*f2(silhouette)-b≥c, where f1(p_value) is a function of p_value, and the larger the p_value, the smaller the value of f1(p_value); f2(silhouette) is a function of silhouette, and the larger the silhouette, the larger f2(silhouette), A<0, b>0, c>0.

[0108] In this embodiment, the cell combination with a larger c and a smaller p_value is more likely to be selected. The methods for determining the values ​​of f1 (p_value), f2 (silhouette), and A and b can be referred to above and will not be repeated here.

[0109] In some further exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0110] Condition 3: Among them, f1(p_value) is a function of p_value, and the larger p_value is, the smaller the value of f1(p_value) is; f2(silhouette) is a function of silhouette, and the larger the silhouette is, the larger f2(silhouette) is, A<0, b>0, c≥0.

[0111] In this embodiment, the larger the c, the easier it is to select a cell combination that satisfies both a smaller p_value and a larger silhouette. The methods for determining the values ​​of f1(p_value), f2(silhouette), and A and b can be referred to above and will not be repeated here.

[0112] In some exemplary embodiments, the selected cell combination is the cell combination with the smallest p-value among all cell combinations that simultaneously meet condition 1, condition 2, and condition 3.

[0113] Exemplary, the cell combination selected as final immunophenotyping in Fig. 6 is the cell type combination that satisfies condition 1, condition 2 and condition 3 simultaneously and significance p_value is minimum.Specifically, the cell type combination is (CD4+Tem, CD4+naive T-cells, CD8+naive T-cells, Macrophages M2, CD8+T-cells, Tregs, CD4+memory T-cells, Neutrophils, Fibroblasts, Th1 cells, Memory B-cells, Class-switched memory B-cells). Fig. 7 is the survival curve graph corresponding to the cell type combination.

[0114] In some exemplary embodiments, the selected cell combination is the cell combination with the smallest p-value and no intersection of survival curves among all cell combinations that simultaneously meet condition 1, condition 2, and condition 3.

[0115] Exemplary, the cell combination selected as final immunophenotyping in Fig. 8 is the cell type combination that satisfies condition 1, condition 2 and condition 3 simultaneously and survival curve does not cross and significance p_value is minimum.Specifically, this cell type combination is (CD4+Tem, CD8+naive T-cells, Macrophages M2, Macrophages M1, Th1cells, CD8+T-cells, Tregs, Neutrophils, Endothelial cells, Fibroblasts, Pericytes, naive B-cells, Memory B-cells, Class-switched memory B-cells). Fig. 9 is the survival curve figure corresponding to this cell type combination.

[0116] In some exemplary embodiments, selecting one of the cell combinations as the cell combination for immunophenotyping comprises:

[0117] Obtain all cell combinations that simultaneously meet conditions 1, 2, and 3, and record them as the first set;

[0118] Obtain a survival curve graph for each cell combination in the first set;

[0119] Based on the survival curve graph, detecting whether the survival probabilities of multiple immunophenotypes of each cell combination in the first set at multiple preset time points are ranked in the same order;

[0120] Determine a cell combination having the same ranking order of survival probability for multiple immunotypes at multiple preset time points, and record it as a second set;

[0121] The cell combination with the smallest p-value in the second set is selected as the final selected cell combination for immunophenotyping.

[0122] In this embodiment, multiple cell combinations that meet the above conditions 1, 2, and 3 are recorded as the first set G1, and the following operations are performed on each cell combination in the first set G1: a survival curve graph of the cell combination is obtained, and the survival probability ranking order of multiple immunotyping of the cell combination at multiple preset time points is detected to be the same. If they are the same, it is considered that the two curves in the survival curve graph of the cell combination do not intersect, and the cell combination in which all survival curve graphs in the first set G1 do not intersect is recorded as the second set G2, and the cell combination with the smallest p-value in the second set G2 is used as the final selected immunotyping cell combination.

[0123] Illustratively, the preset multiple time points may be 1 year, 5 years, 10 years, and 15 years, however, the present disclosure is not limited thereto.

[0124] Specifically, the following operations are performed on each cell combination in the first set G1: a survival curve diagram of the cell combination is obtained, the 1-year survival probability of each immunophenotyping group is sorted from high to low, the corresponding immunophenotyping groups are named cluster_1 to cluster_k, and the cell type combinations whose survival probabilities of cluster_1 to cluster_k decrease successively within the time range of time = (1 year, 5 years, 10 years, 15 years) are found. The found cell type combinations are added to the second set G2, and the cell type combination with the smallest significance p_value in the second set G2 is selected as the final immunophenotyping.

[0125] In some exemplary embodiments, determining the characteristic cell type corresponding to each immunophenotyping grouping comprises:

[0126] The following operations are performed for each cell type in the immunophenotyping cell combination: the median of the cell type in each immunophenotyping group is calculated; and the cell type is determined to be the characteristic cell type corresponding to the immunophenotyping group with the largest median.

[0127] In this embodiment, the characteristic cell type of each immunophenotyping grouping is determined by calculating the median proportion of each cell type in each immunophenotyping grouping. The cell type is the characteristic cell type of the immunophenotyping grouping in which the median proportion of the cell type is the largest. For example, if the medians of B cells in three immunophenotyping groups are 0.05, 0.07, and 0.01, respectively, then B cells are the characteristic cells of the second immunophenotyping grouping.

[0128] As shown in Figure 10, the characteristic cell types of the first immunophenotyping group IM1 are (CD4+Tem, CD8+naive T-cells, Macrophages M2, Macrophages M1, Th1 cells, CD8+T-cells, Tregs), the characteristic cell types of the second immunophenotyping group IM2 are (Neutrophils, Endothelial cells, Fibroblasts, Pericytes), and the characteristic cell types of the third immunophenotyping group IM3 are (naive B-cells, Memory B-cells, Class-switched memory B-cells).

[0129] In some exemplary embodiments, the method further comprises: determining the differentially expressed genes corresponding to each immunophenotyping group, wherein the differentially expressed genes satisfy: a p-value less than a preset p-value threshold and a log2 f c >1,

[0130] In this embodiment, differentially expressed genes refer to genes that are differentially expressed in two or more grouped samples, that is, their expression levels are significantly different in different grouped samples. In some examples, the screening condition for differentially expressed genes in each immunophenotyping group is log2f c >1 and p_value<0.05.

[0131] FIG11 is a heat map of differentially expressed genes in the immune typing groups shown in FIG10 , where the horizontal axis represents samples and the vertical axis represents genes. Only the names of genes related to immunity are marked in the figure.

[0132] In some exemplary embodiments, the method further comprises:

[0133] Obtain a validation set and use the immunophenotyping cell combination and its differentially expressed genes to validate the validation set.

[0134] In the disclosed embodiment, after obtaining the immunophenotyping cell combination, the cell types and differentially expressed genes determined by immunophenotyping of the training set can be used to verify the validation set. The results show that the validation set results are consistent with the training set.

[0135] In some exemplary embodiments, the validation set is validated using the immunophenotyping cell combination and its differentially expressed genes, including:

[0136] The scores or proportions of multiple cell types are calculated using the immune infiltration algorithm;

[0137] Determine the immunophenotyping of the sample based on the characteristic cell types of each immunophenotyping group obtained in the training set;

[0138] Based on the differentially expressed genes in each immunophenotyping group obtained in the training set, whether the gene expression is similar is detected.

[0139] Exemplarily, the validation set also comes from the GEO database. The present disclosure selects 9 DLBCL gene expression datasets independent of the training set as the validation set. The deconvolution tool xCell is used to calculate the proportions of multiple cell types in each validation set. The immunophenotyping of the sample is determined based on the characteristic cell types of each immunophenotyping group obtained in the training set. Based on the differential genes of each immunophenotyping group obtained in the training set, it is determined that the gene expression on each validation set is consistent with the training set results.

[0140] The immunophenotyping methods of the disclosed embodiments are applicable not only to the diagnosis and treatment of diffuse large B-cell lymphoma, but also to other diseases such as mantle cell lymphoma (MCL) and breast cancer. Mantle cell lymphoma is a subtype of non-Hodgkin lymphoma that arises from mature B cells. However, MCL is a heterogeneous tumor, with nearly 20% of cases exhibiting distinct clinicopathological, immunophenotypic, cytogenetic, and molecular features.

[0141] From the GEO database, two gene expression data sets of mantle cell lymphoma (MCL) were selected. According to the immunophenotyping method of the disclosed embodiment, 20 cell types significantly related to prognosis survival were screened out, 4 fixed cell types, and 16 candidate cell types. The clustering results of multiple cell combinations are shown in Figure 12A. The cell type combination and its clustering result with the minimum significance p_value of the average silhouette coefficient silhouette>0.7 and the survival difference were selected as the final immunophenotyping. The survival curve graph and cell ratio graph of the immunophenotyping are shown in Figures 12B and 12C.

[0142] Breast cancer expression data were downloaded from the TCGA database. According to the immunophenotyping method of the disclosed embodiment, 16 cell types significantly correlated with prognosis survival were screened out, 4 fixed cell types ("Class-switched memory B-cells", "CD8+Tcm", "CD8+T-cells", "cDC"), and 12 candidate cell types were selected. The clustering results of multiple cell combinations are shown in Figure 13A. The cell type combination with the smallest significance p_value of the average silhouette coefficient silhouette>0.7 and the survival difference was selected as the final immunophenotyping. The survival curve graph and cell ratio graph of this immunophenotyping are shown in Figures 13B and 13C.

[0143] The present disclosure provides an immunophenotyping method based on immune infiltration analysis, which includes calculating the scores or proportions of cell types through immune infiltration analysis of gene expression data, screening N cell types from multiple cell types based on at least one of the following: survival significance, variance and correlation, dividing the screened cell types into fixed cell types and candidate cell types, performing cluster analysis on multiple cell combinations, selecting one of the cell combinations as the final immunophenotyping cell combination based on the cluster analysis results, performing differential gene analysis on the immunophenotyping results, and using the screened cell types and differential genes to verify consistent results in independent data sets.

[0144] The present disclosure uses immune infiltration analysis to explore the immune cell composition of different typings, which can better explain the biological heterogeneity of tumors and differences in treatment response. It solves the problem that traditional typing cannot be performed due to the inability to determine the cell origin. By using multiple combinations of cell types, the present disclosure maximizes the differences between samples, thereby establishing immune typing and achieving significant differences in survival analysis. The present disclosure can discover more differentially expressed genes. Gene-based typing can better explain the biological heterogeneity of tumors, while taking into account the important role of immune cells in the physiological functions of tumors.

[0145] As shown in FIG14 , the embodiment of the present disclosure further provides an immunophenotyping device, comprising an immune infiltration module 1410 , a screening module 1420 , a cluster analysis module 1430 , and a selection module 1440 , wherein:

[0146] The immune infiltration module 1410 is configured to perform immune infiltration analysis on the sample data to obtain the proportions of multiple cell types;

[0147] The screening module 1420 is configured to screen N cells from a plurality of cell types according to at least one of the following: survival significance, variance, and correlation, where N>1;

[0148] a cluster analysis module 1430 configured to divide the N types of cells into M fixed cells and (NM) candidate cells, and perform cluster analysis on multiple cell combinations, each cell combination including at least M fixed cells and q candidate cells, where M ≥ 1 and M < N, and 0 ≤ q ≤ (NM);

[0149] The selection module 1440 is configured to select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping grouping.

[0150] In some exemplary embodiments, the clustering analysis module 1430 uses any one of the following clustering methods when performing cluster analysis on multiple cell combinations: non-negative matrix factorization, hierarchical clustering, K-means clustering, density clustering, and community detection.

[0151] In some exemplary embodiments, the selection module 1440 selects one of the cell combinations as the cell combination for immunophenotyping, including:

[0152] Based on the p-value and silhouette coefficient corresponding to each cell combination, one of the cell combinations is selected as the cell combination for immunophenotyping. The p-value and silhouette coefficient of the selected cell combination meet the following conditions 1 and 2:

[0153] Condition 1: The p-value is less than the preset p-value threshold;

[0154] Condition 2: The silhouette coefficient is greater than the preset silhouette coefficient threshold.

[0155] In some exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0156] Condition 3: f1(p_value)-A*f2(silhouette)-b≥c or Where p_value is the p-value, f1(p_value) is a function of p_value, and the larger the p_value, the smaller the value of f1(p_value); silhouette is the silhouette coefficient, f2(silhouette) is a function of silhouette, and the larger the silhouette, the larger f2(silhouette), A<0, b>0, c≥0.

[0157] For example, f1(p_value)=-log10(p_ value ); f2(silhouette)=silhouette.

[0158] In some exemplary embodiments, Among them, y2=max(f1(p_value)), y1=f1(p_value)p_value<0.01, x2=f2(silhouette)silhouette∈[0.5,08], x1=max(f2(silhouette)), b=y2-A·x2, max means taking the maximum value.

[0159] In some exemplary embodiments, the selected cell combination is the cell combination with the smallest p-value among all cell combinations that simultaneously meet Condition 1, Condition 2, and Condition 3.

[0160] In some exemplary embodiments, the selected cell combination is the cell combination with the smallest p-value and no intersection of survival curves among all cell combinations that simultaneously meet Condition 1, Condition 2, and Condition 3.

[0161] In some exemplary embodiments, the selection module 1440 selects one of the cell combinations as the cell combination for immunophenotyping, including:

[0162] All cell combinations that simultaneously meet conditions 1, 2, and 3 are recorded as the first set;

[0163] Obtain a survival curve graph for each cell combination in the first set;

[0164] Based on the survival curve graph, detecting whether the survival probabilities of multiple immunotypes of each cell combination in the first set at multiple preset time points are ranked in the same order;

[0165] The cell combinations with the same survival probability ranking order of multiple immunotypes at multiple preset time points are recorded as the second set;

[0166] The cell combination with the smallest p-value in the second set is selected as the final selected immunophenotyping cell combination.

[0167] In some exemplary embodiments, the selection module 1440 determines the characteristic cell type corresponding to each immunophenotyping group, including:

[0168] The following operations are performed for each cell type in the immunophenotyping cell combination: calculating the median of the cell type in each immunophenotyping group; and determining that the cell type is the characteristic cell type corresponding to the immunophenotyping group with the largest median.

[0169] In some exemplary embodiments, the screening module 1420 screens N cells from multiple cell types based on survival significance, including:

[0170] Using a univariate survival analysis model, the statistically significant p-value of each cell type and survival time and survival status was calculated, and N cell types with p-values ​​less than the preset p-value threshold were selected as N cells significantly associated with prognostic survival.

[0171] In some exemplary embodiments, the screening module 1420 screens N types of cells from multiple cell types based on variance, including:

[0172] The variance of the proportion of each cell type is calculated, and N cell types whose variance is greater than a preset variance threshold are selected as the screened N cell types.

[0173] In some exemplary embodiments, the screening module 1420 screens N types of cells from a plurality of cell types based on correlations, including:

[0174] Calculate the correlation between each cell type and other cell types, count the number of cell types whose absolute value of the calculated correlation is greater than a preset correlation threshold, and select N cell types whose counted number is greater than the preset number threshold as the screened N cell types.

[0175] In some exemplary embodiments, the clustering analysis module 1430 divides the N cells into M fixed cells and (NM) candidate cells, including:

[0176] For each of the N cells, perform the following operations:

[0177] Obtaining the p-value and risk ratio corresponding to the cell;

[0178] When the obtained p-value is less than a preset p-value threshold and the absolute value of the risk ratio value is greater than a preset risk ratio value threshold, classifying the cell as a fixed cell;

[0179] The cells other than the fixed cells among the N types of cells are classified as the candidate cells.

[0180] In some exemplary embodiments, the selection module 1440 is further configured to: determine the differentially expressed genes corresponding to each immunophenotyping group, wherein the differentially expressed genes satisfy: a p-value less than a preset p-value threshold and a log2f c >1,

[0181] In some exemplary embodiments, the immunophenotyping apparatus further comprises: a verification module configured to obtain a verification set and verify the verification set using the immunophenotyped cell combinations and their differential genes.

[0182] An embodiment of the present disclosure also provides an immunophenotyping device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the immunophenotyping method described in any embodiment of the present disclosure based on the instructions stored in the memory.

[0183] As shown in Figure 15, in one example, the immunophenotyping device may include: a processor 1510, a memory 1520, a bus system 1530 and a transceiver 1540, wherein the processor 1510, the memory 1520 and the transceiver 1540 are connected via the bus system 1530, the memory 1520 is used to store instructions, and the processor 1510 is used to execute the instructions stored in the memory 1520 to control the transceiver 1540 to send and receive signals. Specifically, the transceiver 1540 can obtain sample data under the control of the processor 1510, and the processor 1510 performs immune infiltration analysis on the sample data to obtain the proportion of multiple cell types; based on at least one of the following: survival significance, variance and correlation, N cells are screened out from multiple cell types, N>1; the N cells are divided into M fixed cells and (NM) candidate cells, and cluster analysis is performed on multiple cell combinations, each cell combination includes at least M fixed cells and q candidate cells, M≥1 and M<N, 0≤q≤(NM); one of the cell combinations is selected as the cell combination for immunophenotyping, and the characteristic cell type corresponding to each immunophenotyping group is determined.

[0184] It should be understood that the processor 1510 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0185] The memory 1520 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1510. A portion of the memory 1520 may also include a non-volatile random access memory. For example, the memory 1520 may also store information about the device type.

[0186] In addition to the data bus, the bus system 1530 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 1530 in FIG.

[0187] During implementation, the processing performed by the processing device can be completed by the hardware integrated logic circuit in the processor 1510 or by instructions in the form of software. That is, the method steps of the embodiment of the present disclosure can be embodied as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1520, and the processor 1510 reads the information in the memory 1520 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0188] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the immunophenotyping method described in any of the embodiments of the present disclosure. The immunophenotyping method driven by executing the executable instructions is substantially the same as the immunophenotyping method provided in the above embodiments of the present disclosure and is not further described here.

[0189] In some possible embodiments, various aspects of the immunophenotyping method provided by the present disclosure may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the immunophenotyping method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the computer device can execute the immunophenotyping method described in the embodiments of the present disclosure.

[0190] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0191] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0192] It should be noted that the above-described embodiments or implementations are merely illustrative and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementations without departing from the scope of the present disclosure.

Claims

1. An immunophenotyping method comprising: Perform immune infiltration analysis on sample data to obtain the proportions of various cell types; Select N cell types from multiple cell types based on at least one of the following: survival significance, variance, and correlation, where N>1; Dividing the N cells into M fixed cells and (NM) candidate cells, and performing cluster analysis on multiple cell combinations, each cell combination includes at least M fixed cells and q candidate cells, M≥1 and M<N, 0≤q≤(NM); One of the cell combinations is selected as the cell combination for immunophenotyping, and the characteristic cell type corresponding to each immunophenotyping group is determined.

2. The method according to claim 1, wherein The step of selecting one of the cell combinations as the cell combination for immunophenotyping comprises: Based on the p-value and silhouette coefficient corresponding to each cell combination, one of the cell combinations is selected as the cell combination for immunophenotyping. The p-value and silhouette coefficient of the selected cell combination meet the following conditions 1 and 2: Condition 1: The p-value is less than the preset p-value threshold; Condition 2: The silhouette coefficient is greater than the preset silhouette coefficient threshold.

3. The method according to claim 2, wherein: The p-value and silhouette coefficient of the selected cell combination also meet the following condition 3: Condition 3: f1(p_value)-A*f2(silhouette)-b≥c or Where p_value is the p-value, f1(p_value) is a function of p_value, and the larger the p_value, the smaller the value of f1(p_value); silhouette is the silhouette coefficient, f2(silhouette) is a function of silhouette, and the larger the silhouette, the larger f2(silhouette), A<0, b>0, c≥0.

4. The method according to claim 3, wherein: f1(p_value)=-log10(p_value); f2(silhouette)=silhouette.

5. The method according to claim 3, wherein Among them, y2=max(f1(p_value)), y1=f1(p_value)p_value<0.01, x2=f2(silhouette)silhouette∈[0.5,08], x1=max(f2(silhouette)), b=y2-A·x2, max means taking the maximum value.

6. The method according to claim 3, wherein: The selected cell combination is the cell combination with the smallest p value among all cell combinations that simultaneously meet the conditions 1, 2, and 3.

7. The method according to claim 3, wherein: The selected cell combination is the cell combination with the smallest p-value among all cell combinations that simultaneously meet the conditions 1, 2, and 3, and whose survival curves do not cross.

8. The method according to claim 7, wherein: The step of selecting one of the cell combinations as the cell combination for immunophenotyping comprises: All cell combinations that simultaneously meet conditions 1, 2, and 3 are recorded as the first set; Obtain a survival curve graph for each cell combination in the first set; Based on the survival curve graph, detecting whether the survival probabilities of multiple immunotypes of each cell combination in the first set at multiple preset time points are ranked in the same order; The cell combinations with the same survival probability ranking order of multiple immunotypes at multiple preset time points are recorded as the second set; The cell combination with the smallest p-value in the second set is selected as the final selected immunophenotyping cell combination.

9. The method according to claim 1, wherein: Determining the characteristic cell type corresponding to each immunophenotyping group includes: The following operations are performed for each cell type in the immunophenotyping cell combination: calculating the median of the cell type in each immunophenotyping group; and determining that the cell type is the characteristic cell type corresponding to the immunophenotyping group with the largest median.

10. The method according to claim 1, wherein The dividing the N types of cells into M types of fixed cells and (NM) types of candidate cells comprises: For each of the N cells, perform the following operations: Obtaining the p-value and risk ratio corresponding to the cell; When the obtained p-value is less than a preset p-value threshold and the absolute value of the risk ratio value is greater than a preset risk ratio value threshold, classifying the cell as a fixed cell; The cells other than the fixed cells among the N types of cells are classified as the candidate cells.

11. The method according to claim 1 , further comprising: Determine the differentially expressed genes corresponding to each immunophenotyping group, and the differentially expressed genes meet the following conditions: the p value is less than the preset p value threshold and the log2f c >1, 12. The method according to claim 1, wherein The clustering method used in the cluster analysis of the plurality of cell combinations is any one of the following: non-negative matrix decomposition, hierarchical clustering, K-means clustering, density clustering and community detection.

13. The method according to claim 1, wherein N cells are screened from a variety of cell types based on the survival significance, including: Using a univariate survival analysis model, the statistically significant p-value of each cell type and survival time and survival status was calculated, and N cell types with p-values ​​less than the preset p-value threshold were selected as N cells significantly associated with prognostic survival.

14. The method according to claim 1, wherein N types of cells are selected from a plurality of cell types according to the variance, including: The variance of the proportion of each cell type is calculated, and N cell types whose variance is greater than a preset variance threshold are selected as the screened N cell types.

15. The method according to claim 1, wherein N types of cells are screened from a plurality of cell types according to the correlation, including: Calculate the correlation between each cell type and other cell types, count the number of cell types whose absolute value of the calculated correlation is greater than a preset correlation threshold, and select N cell types whose counted number is greater than the preset number threshold as the screened N cell types.

16. An immunophenotyping device comprising a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the immunophenotyping method according to any one of claims 1 to 15 based on the instructions stored in the memory.

17. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the immunophenotyping method according to any one of claims 1 to 15 is implemented.

18. A computer program product comprising instructions, which, when executed by a computer, perform the immunophenotyping method according to any one of claims 1 to 15.

19. An immunophenotyping device comprising an immune infiltration module, a screening module, a cluster analysis module, and a selection module, wherein: The immune infiltration module is configured to perform immune infiltration analysis on the sample data to obtain the proportions of multiple cell types; The screening module is configured to screen N cells from a plurality of cell types according to at least one of the following: survival significance, variance, and correlation, where N>1; The cluster analysis module is configured to divide the N types of cells into M types of fixed cells and (NM) types of candidate cells, and perform cluster analysis on multiple cell combinations, each cell combination including at least M types of fixed cells and q types of candidate cells, where M ≥ 1 and M < N, and 0 ≤ q ≤ (NM); The selection module is configured to select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping grouping.

Citation Information

Patent Citations

  • Tumor immune subtype classification method and system

    CN112435714A

  • Immune typing method and device, storage medium and program product

    CN118447927A

  • Scoring model for evaluating tumor immune microenvironment and construction method therefor

    WO2023123005A1