Multi-omics analysis method and device based on immunophenotyping, and storage medium and program product

Through the multi-omics analysis method of immunophenotyping, cell types in DLBCL were screened and clustered, and differentially expressed genes were identified, which solved the problem of insufficient accuracy of existing DLBCL typing and provided higher subtype accuracy and clinical treatment guidance.

WO2025200452A1PCT designated stage Publication Date: 2025-10-02BOE TECHNOLOGY GROUP CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129452
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2024-11-01
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing DLBCL typing methods are inaccurate, especially the classification accuracy based on cell origin typing is low. Traditional methods lack multi-dimensional feature descriptions and cannot meet clinical needs.

Method used

A multi-omics analysis method based on immunophenotyping was used to obtain the cell type ratio through immune infiltration analysis, screen out cell types with high significance and correlation, perform cluster analysis, determine differentially expressed genes and omics characteristics, and establish a multi-dimensional description of disease characteristics.

Benefits of technology

Improves the accuracy of DLBCL subtyping, reveals new biomarkers, improves clinical management and treatment options, and is suitable for auxiliary analysis of hematological tumors and solid tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129452_02102025_PF_FP_ABST
    Figure CN2024129452_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A multi-omics analysis method and device based on immunophenotyping, and a storage medium and a program product. The method comprises: performing immune infiltration analysis on gene expression data to obtain the proportions of a plurality of cell types; on the basis of at least one of survival significance, variance and correlation, selecting N types of cells from the plurality of cell types, wherein N>1; dividing the N types of cells into M types of fixed cells and (N-M) types of candidate cells, and performing clustering analysis on a plurality of cell combinations, wherein each cell combination at least comprises M types of fixed cells and q types of candidate cells, wherein M≥1 and M<N, and 0≤q≤(N-M); selecting one of the cell combinations as a cell combination for immunophenotyping, and determining characteristic cell types corresponding to each immunophenotyping group; and determining differentially expressed genes corresponding to each immunophenotyping group, and determining at least one of the following omics features corresponding to each immunophenotyping group: differentially mutated genes, differential copy number variations and differentially methylated sites.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-omics analysis method and device based on immunophenotyping, storage medium and program product

[0001] This application claims priority to PCT patent application filed with the Patent Office of China on March 26, 2024, with application number PCT / CN2024 / 083837 and invention name “Immunotyping method and device, storage medium and program product”, and claims priority to Chinese patent application filed with the Patent Office of China on April 25, 2024, with application number 202410509310.X and invention name “Immunotyping method and device, storage medium and program product”, the contents of which should be understood as incorporated into this application by reference. Technical Field

[0002] The embodiments of the present disclosure relate to, but are not limited to, the field of biotechnology, and in particular to a multi-omics analysis method and device based on immunophenotyping, a storage medium, and a program product. Background Art

[0003] Diffuse large B-cell lymphoma (DLBCL) is a common type of non-Hodgkin's lymphoma. Its classification helps guide clinical treatment and prognosis assessment. The following are several common clinical DLBCL classification methods:

[0004] (1) Cell of origin (COO)

[0005] DLBCL can be divided into germinal center B cell-like (GCB) and non-germinal center B cell-like (ABC) according to the cell origin. Among them, GCB DLBCL is characterized by tumor cells showing gene expression characteristics of germinal center B cells and usually has a better prognosis; ABC DLBCL is characterized by tumor cells showing characteristics of non-germinal center B cells and is often associated with a poor prognosis.

[0006] (2) Immunophenotyping

[0007] Based on the expression of tumor cell surface markers, such as CD10, BCL6 (associated with germinal centers) and MUM1 / IRF4 (associated with non-germinal centers).

[0008] (3) Genetic typing

[0009] DLBCL can be divided into different types based on genetic typing, including chromosomal abnormalities and gene mutations. Chromosomal abnormalities, including certain specific chromosomal rearrangements (such as rearrangements of BCL2 and BCL6), are associated with prognosis; gene mutations, including mutations in TP53, MYC, and other related genes, can affect efficacy and prognosis.

[0010] (4) Clinical classification

[0011] The International Prognostic Index (IPI) assesses prognosis based on clinical information such as age, clinical stage, and lactate dehydrogenase (LDH) levels. DLBCL patients can be stratified based on the IPI score: 0-1 points are considered low risk; 2 points are considered low-intermediate risk; 3 points are considered high-intermediate risk; and 4-5 points are considered high risk.

[0012] However, these classification methods still have some problems, such as: ① Cell-of-origin classification, which is mainly based on immunohistochemistry. Although it can determine the patient's prognosis, its limitation is that the classification accuracy is only 88% for the same classification population, and the cell of origin cannot be determined for some patients. Therefore, stratified treatment based on cell-of-origin classification no longer meets clinical needs. ② Genotyping-based methods such as exome sequencing and transcriptome sequencing have found that different genetic subtypes have significant differences in genotype (gene mutation), epigenetic and clinical characteristics, providing potential pathological basis for precision treatment strategies for DLBCL. However, these studies are based on genotyping criteria set based on information from a few known mutant genes and cannot accurately predict undiscovered gene mutations or gene expression information. Moreover, the above methods are almost all based on single-omics differences and lack multi-dimensional (multi-omics data) feature descriptions.

[0013] Summary of the Invention

[0014] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0015] The present disclosure provides a multi-omics analysis method based on immunophenotyping, comprising:

[0016] performing immune infiltration analysis on first sample data to obtain ratios of multiple cell types, wherein the first sample data includes gene expression data;

[0017] Select N cell types from multiple cell types based on at least one of the following: survival significance, variance, and correlation, where N>1;

[0018] Dividing the N cells into M fixed cells and (NM) candidate cells, and performing cluster analysis on multiple cell combinations, each cell combination includes at least M fixed cells and q candidate cells, M≥1 and M<N, 0≤q≤(NM);

[0019] selecting one of the cell combinations as a cell combination for immunophenotyping, and determining a characteristic cell type corresponding to each immunophenotyping group;

[0020] Determine the differentially expressed genes corresponding to each of the immunophenotyping groups, and determine at least one of the following omics features corresponding to each of the immunophenotyping groups: differentially mutated genes, differential copy number variations, and differentially methylated sites.

[0021] An embodiment of the present disclosure also provides a multi-omics analysis device based on immunophenotyping, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the multi-omics analysis method based on immunophenotyping described in any embodiment of the present disclosure based on the instructions stored in the memory.

[0022] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-omics analysis method based on immunophenotyping as described in any embodiment of the present disclosure.

[0023] An embodiment of the present disclosure further provides a program product, comprising instructions. When the computer program product is executed by a computer, the instructions execute the multi-omics analysis method based on immunophenotyping as described in any embodiment of the present disclosure.

[0024] The present disclosure also provides a multi-omics analysis device based on immunophenotyping, including an immune infiltration module, a screening module, a cluster analysis module, a selection module, and an omics feature analysis module, wherein:

[0025] The immune infiltration module is configured to perform immune infiltration analysis on first sample data to obtain proportions of multiple cell types, wherein the first sample data includes gene expression data;

[0026] The screening module is configured to screen N cells from a plurality of cell types according to at least one of the following: survival significance, variance, and correlation, where N>1;

[0027] The cluster analysis module is configured to divide the N types of cells into M types of fixed cells and (NM) types of candidate cells, and perform cluster analysis on multiple cell combinations, each cell combination including at least M types of fixed cells and q types of candidate cells, where M ≥ 1 and M < N, and 0 ≤ q ≤ (NM);

[0028] The selection module is configured to select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping group;

[0029] The omics feature analysis module is configured to determine the differentially expressed genes corresponding to each of the immunophenotyping groups, and determine at least one of the following omics features corresponding to each of the immunophenotyping groups: differentially mutated genes, differential copy number variations, and differentially methylated sites.

[0030] The multi-omics analysis method and device, storage medium and program product based on immunophenotyping of the embodiments of the present disclosure obtain the proportion of multiple cell types by performing immune infiltration analysis on the first sample data, and screen N cells from the multiple cell types according to at least one of the following: survival significance, variance and correlation, divide the N cells into M fixed cells and (NM) candidate cells, and perform cluster analysis on the multiple cell combinations; select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping group, which can explore the immune cell composition in different types and better explain the differences in tumor treatment response; use multiple combinations of cell types to find the differences between samples to the greatest extent, thereby establishing immunophenotyping, and making the immunophenotyping have significant differences in survival analysis; by determining the differentially expressed genes corresponding to each immunophenotyping group, and determining at least one of the following omics features corresponding to each immunophenotyping group: differentially mutated genes, differential copy number variations, and differentially methylated sites, these different levels of omics data are integrated to provide a multi-dimensional disease feature description. This multidimensional analysis method has higher complexity and accuracy than previous studies on single cell types or gene expression; this method not only improves the accuracy of subtypes, but may also reveal new biomarkers, thereby improving clinical management and the selection of treatment options. In addition, the present disclosure does not require the determination of cell origin, which solves the problem that traditional typing cannot clearly classify cells due to the inability to determine cell origin; the present disclosure can discover more differentially expressed genes, thereby better explaining the biological heterogeneity of tumors, while taking into account the important role that immune cells play in the physiological functions of tumors. The present disclosure uses the screened cell types and differential genes to obtain consistent results in independent validation sets, which can be used for auxiliary analysis of various types of blood tumors and solid tumors.

[0031] Other features and advantages of the present disclosure will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present disclosure. Other advantages of the present disclosure can be realized and obtained through the solutions described in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings are used to provide an understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0033] FIG1 is a schematic diagram of a multi-omics analysis method based on immunophenotyping according to an exemplary embodiment of the present disclosure;

[0034] FIG2 is a graph showing the ratio scores of 64 cell types obtained using the xCell tool according to an exemplary embodiment of the present disclosure;

[0035] 3A to 3S are survival curves of 19 types of cells provided by an exemplary embodiment of the present disclosure;

[0036] FIG4 is a heat map showing the correlation of 19 types of cells according to an exemplary embodiment of the present disclosure;

[0037] FIG5 is a risk ratio diagram of 19 cell types provided by an exemplary embodiment of the present disclosure;

[0038] FIG6 is a diagram showing cluster analysis results according to an exemplary embodiment of the present disclosure;

[0039] FIG7 is a survival curve diagram of the immunophenotyping cell combination finally selected in FIG6 ;

[0040] FIG8 is a diagram showing cluster analysis results of another exemplary embodiment of the present disclosure;

[0041] FIG9 is a survival curve diagram of the immunophenotyping cell combination finally selected in FIG8 ;

[0042] FIG10A is a diagram of immunophenotyping and characteristic cell types provided by an exemplary embodiment of the present disclosure;

[0043] FIG10B is an immunophenotyping and differentially expressed gene map provided by an exemplary embodiment of the present disclosure;

[0044] FIG11A is a schematic diagram of the C-index value of the prognostic risk prediction model constructed using the training set and two validation sets according to an embodiment of the present disclosure;

[0045] FIG11B to FIG11D are schematic diagrams showing differences in survival curves of prognostic risk groups obtained using the prognostic risk prediction model constructed using the training set and two validation sets according to an embodiment of the present disclosure;

[0046] Figures 11E to 11F are schematic diagrams showing the correlation between the gene expression groups of the top two genes in importance and the prognostic risk groups;

[0047] FIG12A is a diagram showing cluster analysis results according to another exemplary embodiment of the present disclosure;

[0048] FIG12B is a survival curve of the immunophenotyping cell combination finally selected in FIG12A;

[0049] FIG12C is a diagram of the immunophenotyping and characteristic cell types finally selected in FIG12A;

[0050] FIG13A is a diagram showing cluster analysis results according to another exemplary embodiment of the present disclosure;

[0051] FIG13B is a survival curve of the immunophenotyping cell combination finally selected in FIG13A;

[0052] FIG13C is a diagram of the immunophenotyping and characteristic cell types finally selected in FIG13A;

[0053] FIG14A is a schematic diagram of differentially mutated genes in immunophenotyping groups according to an exemplary embodiment of the present disclosure;

[0054] 14B and 14C are schematic diagrams showing the verification of two differentially mutated genes in the immunophenotyping grouping according to an exemplary embodiment of the present disclosure;

[0055] 15A to 15G are schematic diagrams of differential CNAs of immunophenotyping groups according to an exemplary embodiment of the present disclosure;

[0056] Figures 16A to 16E, 16F to 16J, and 16K to 16O are schematic diagrams of the expression levels and methylation rates of the top five differentially methylated sites ranked by absolute values ​​of correlation coefficients in immunophenotyping groups B1, B2, and B3, respectively;

[0057] 17A to 17F are schematic diagrams showing the correlation between immunophenotyping groups and clinical characteristics according to an exemplary embodiment of the present disclosure;

[0058] FIG18 is a schematic structural diagram of a multi-omics analysis device based on immunophenotyping provided by an exemplary embodiment of the present disclosure;

[0059] FIG19 is a schematic structural diagram of another multi-omics analysis device based on immunophenotyping provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0060] The present disclosure describes a plurality of embodiments, but this description is exemplary rather than restrictive, and it will be apparent to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.

[0061] The present disclosure includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The disclosed embodiments, features, and elements of the present disclosure may also be combined with any conventional features or elements to form a unique inventive solution defined by the claims. Any features or elements of any embodiment may also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this disclosure may be implemented individually or in any appropriate combination. Therefore, the embodiments are not subject to other limitations except for the limitations set forth in the appended claims and their equivalents. In addition, various modifications and changes may be made within the scope of protection of the appended claims.

[0062] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as a limitation on the claims. In addition, the claims to the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can be changed and still remain within the spirit and scope of the disclosed embodiments.

[0063] As shown in FIG1 , the present disclosure provides a multi-omics analysis method based on immunophenotyping, comprising:

[0064] Step 101: performing immune infiltration analysis on first sample data to obtain ratios of multiple cell types, wherein the first sample data includes gene expression data;

[0065] In some exemplary embodiments, in step 101, the first sample data may be a DLBCL dataset. However, the present disclosure is not limited thereto, and the first sample data may also be a mantle cell lymphoma (MCL) or any other type of disease dataset.

[0066] In some exemplary embodiments, the first sample data includes not only gene expression data but also corresponding clinical information.

[0067] In the disclosed embodiment, the first sample data can be downloaded from a public database. For example, gene expression data and corresponding clinical information can be obtained from the Gene Expression Omnibus (GEO), and the clinical information includes survival time and survival status. The GEO database is a gene expression database created and maintained by the National Center for Biotechnology Information (NCBI) of the United States. Since its establishment in 2000, GEO has collected gene expression data from research institutions around the world, covering multiple fields such as tumors, non-tumors, chips, next-generation sequencing (NGS), differential analysis and molecular verification.

[0068] In some exemplary embodiments, before performing immune infiltration analysis on the first sample data, the method further comprises:

[0069] Take the intersection of genes for multiple gene expression data with survival time and survival status, remove the batch effect, and get the gene expression matrix denoted as E∈R g×n , where g represents the number of intersection genes and n represents the number of first samples.

[0070] In multi-omics data, batch effects refer to differences between samples due to variations in experimental conditions, operators, instruments, and other factors. These differences can mask true biological variations, thereby affecting subsequent analytical results. There are various approaches to address batch effects. One approach to de-batch is to use normalization techniques to normalize differences between samples and thereby eliminate batch effects.

[0071] For example, 1,158 samples from nine DLBCL datasets were selected from the Gene Expression Comprehensive Database as the training set, of which 1,127 samples had survival time and survival status, as shown in Table 1. The intersection genes of these 1,127 samples were taken, resulting in a total of 10,696 intersection genes in this example. The combat function of the sva toolkit was then used to remove batch effects.

[0072] Table 1

[0073] Then, the scores or proportions of multiple cell types are calculated by using immune infiltration algorithms, such as cibersort, xCell and Kassandra deconvolution tools, and recorded as the cell proportion matrix X∈R c×n , where c represents the number of cell types.

[0074] Immune infiltration refers to the process by which immune cells (such as T cells, B cells, and macrophages) penetrate the blood vessels of tissues or organs and enter the body's tissues or organs to respond to infection, inflammation, tumors, or other diseases. Immune infiltration analysis is a method used to study the distribution and activity of immune cells in tissues or tumors. It helps understand the immune system's response to various disease states, particularly tumors and cancer.

[0075] The tumor microenvironment (TME) regulates tumor survival, maintenance, growth, and immune surveillance, playing a crucial role in disease progression and treatment response, including local drug resistance, immune escape, and cancer metastasis. Immune cells, as an essential component of the TME, play a key role in tumor physiology. Analyzing the composition and function of different cell populations in the TME through immune infiltration can help explore more effective tumor treatments.

[0076] For example, using the training set in Table 1 above as an example, the xCell deconvolution tool was used to calculate the cell enrichment score. The xCell deconvolution tool generated a total of 64 cell types, including B cells, as shown in Figure 2. Using these 64 cell types, immune infiltration analysis allows for a more comprehensive exploration of the impact of the microenvironment on tumor subtypes.

[0077] Step 102: Screening N types of cells from the plurality of cell types based on at least one of the following: survival significance, variance, and correlation, where N is a natural number greater than 1;

[0078] In some exemplary embodiments, N cells are selected from a plurality of cell types based on survival significance, including:

[0079] The single-factor Cox survival analysis model was used to calculate the statistically significant p-value (p_value, hereinafter referred to as p-value) of each cell type and survival time and survival status, and N cell types with p-value < preset p-value threshold (such as 0.05) were screened as N cell types significantly associated with prognosis survival, where N is the number of cell types screened, 1 <N<c。

[0080] Exemplary, still taking the training set of aforementioned Table 1 as an example, 19 cell types significantly correlated with prognosis survival were screened out by the univariate cox survival analysis model. Figures 3A to 3S are survival curves of these 19 cell types, wherein the abscissa is the follow-up time (time, in days) and the ordinate is the survival probability (Survival probability). Each survival curve graph includes a high expression curve, a low expression curve and a corresponding p value, wherein high expression refers to its expression level in the biological sample is relatively high, and low expression refers to its expression level in the biological sample is relatively low. In general, the greater the difference between the two curves of the high expression curve and the low expression curve, the greater the difference in the prognosis of the two groups of patients, the easier it is to have a statistical difference, and the smaller the p value, the more statistically significant the result is, i.e., the survival time of the two groups of patients is significantly different. As shown in Figures 3A to 3S, 19 cell types include B-cells, CD4+ memory T-cells, CD4+ naive T-cells, etc., and the difference is significant in the survival curve graph. In the disclosed embodiments, prognosis mainly refers to the clinician's estimation of the treatment efficacy and survival time of the patient based on clinical data.

[0081] In statistics, the p-value is an important concept used to test hypotheses. The p-value is defined as the probability of obtaining a result that is equal to or more extreme than the actual observed result, given that the hypothesis holds. The p-value threshold is typically set at 0.05 or 0.01; a p-value below this threshold is considered statistically significant. For example, a p-value < 0.05 indicates that there is less than a 5% chance that the difference in the results is due to chance, or that the probability of reaching the opposite conclusion if the same study is repeated under the same conditions is less than 5%.

[0082] In some exemplary embodiments, selecting N types of cells from a plurality of cell types based on variance comprises:

[0083] The variance of the proportion of each cell type is calculated, and N cell types whose variance is greater than a preset variance threshold are selected as the screened N cell types.

[0084] The larger the variance, the greater the data fluctuation; the smaller the variance, the smaller the data fluctuation. Therefore, based on the variance, cell types with large differences in proportion can be screened as the N screened cell types. In the disclosed embodiment, the preset variance threshold can be set based on the actual variance data.

[0085] In some exemplary embodiments, N cells are selected from a plurality of cell types based on correlation, including:

[0086] Calculate the correlation between each cell type and other cell types, count the number of cell types whose absolute value of the calculated correlation is greater than a preset correlation threshold, and select N cell types whose counted number is greater than the preset number threshold as the screened N cell types.

[0087] For example, the preset correlation threshold may be 0.6. The preset number threshold may be a positive integer greater than 2. However, the present disclosure does not limit this, and the preset correlation threshold and the preset number threshold may be set as needed.

[0088] In some exemplary embodiments, the method further comprises:

[0089] Calculate the correlation between N types of cells;

[0090] The number of immunophenotyping groups (ie, the number of clusters) k is determined based on the correlation between N types of cells.

[0091] In the disclosed embodiment, the correlation between N types of cells can be calculated using the Spearman correlation analysis method, where the Spearman correlation coefficient ranges from -1 to 1, with a value of -1 indicating a perfect negative correlation, a value of 1 indicating a perfect positive correlation, and a value of 0 indicating no correlation between the two variables.

[0092] For example, FIG4 is a heat map showing the correlation between the aforementioned 19 cell types. As shown in FIG4 , based on the correlation between the 19 cell types, the number k of immunophenotyping groups is determined to be 3. However, the present disclosure is not limited thereto, and the number k of immunophenotyping groups can be set as needed.

[0093] In the embodiment of the present disclosure, the number k of immunotyping groups may be between 2 and 5. However, the embodiment of the present disclosure is not limited thereto, and the number of immunotyping groups may also be any other natural number greater than or equal to 2. The following example is exemplified with the number of immunotyping groups being 3.

[0094] Step 103: Divide the N types of cells into M fixed cells and (NM) candidate cells, and perform cluster analysis on multiple cell combinations, where each cell combination includes at least M fixed cells and q candidate cells, where M is a natural number greater than or equal to 1 and M<N, and q is a natural number between 0 and (NM);

[0095] In some exemplary embodiments, dividing N cells into M fixed cells and (NM) candidate cells comprises:

[0096] For each of the N cells, perform the following operations:

[0097] Get the p-value and risk ratio corresponding to the cell;

[0098] When the p-value corresponding to the cell is less than the preset p-value threshold and the absolute value of the risk ratio is greater than the preset risk ratio threshold, the cell is classified as a fixed cell;

[0099] Classify the cells other than the fixed cells among the N types of cells as candidate cells.

[0100] HR values, also known as hazard ratios, are primarily used in survival analysis. HR values ​​assess the multiple change in the risk of an outcome event caused by the presence of a factor. An HR value greater than 1 indicates that the exposure factor promotes the occurrence of a positive event, an HR value less than 1 indicates that the exposure factor hinders the occurrence of a positive event, and an HR equal to 1 indicates that the exposure factor has no effect on the occurrence of a positive event. Specifically, extremely large or small HR values ​​indicate a significant inhibitory or promoting effect on the occurrence of the event. Figure 5 shows the HR values ​​for these 19 cell types. The dots represent the model coefficients corresponding to each predictor variable, and the lines represent the 95% confidence intervals. The figure shows the lower bound (lower.95) and upper bound (upper.95) of the 95% confidence intervals for the model coefficients in exponential form, indicating the uncertainty range of the estimate.

[0101] For example, the preset p-value threshold may be between 0.01 and 0.05, and the preset HR value threshold may be greater than 1.25 or less than 0.75. For example, the preset p-value threshold may be 0.05, and the preset HR value threshold may be greater than 1.25 or less than 0.75, however, the embodiments of the present disclosure are not limited thereto.

[0102] In the survival curve diagrams of Figures 3A to 3S, the p-values ​​corresponding to the six cell types, Th1 cells, Memory B-cells, Class-switched memory B-cells, CD8+ T-cells, CD8+ naive T-cells, and Tregs, are less than the preset p-value threshold, and in the HR diagram of Figure 5, the absolute values ​​of the HR values ​​of these six cell types are greater than the preset HR value threshold. Therefore, the embodiment of the present disclosure selects the six cell types, Th1 cells, Memory B-cells, Class-switched memory B-cells, CD8+ T-cells, CD8+ naive T-cells, and Tregs, as fixed cells.

[0103] In some exemplary embodiments, cluster analysis is performed on multiple cell combinations, comprising:

[0104] Identify all possible cell combinations;

[0105] For each cell combination, the following operations are performed: extract the cell type matrix corresponding to the cell combination, perform non-negative matrix decomposition on the cell type matrix to obtain the clustering results, and calculate the average silhouette coefficient and p-value corresponding to the cell combination based on the clustering results.

[0106] In the disclosed embodiment, the clustering method used for cluster analysis of multiple cell combinations may also be other clustering algorithms, such as hierarchical clustering, K-means clustering, density clustering, and community detection algorithms. After obtaining the clustering results, the average silhouette coefficient and p-value corresponding to the cell combination are calculated based on the clustering results.

[0107] In some exemplary embodiments, all possible cell combinations include kind.

[0108] Taking the aforementioned 19 cells as an example, excluding the 6 fixed cells that have been selected, there are 13 candidate cells left. From the q candidate cell groups, m candidate cells are selected, with a total of The value of m ranges from 1 to q, so all possible cell combinations include kind.

[0109] In some exemplary embodiments, for each cell combination, the following operations are performed: from the original cell ratio matrix X∈R c×n Extract the cell type matrix of fixed cells and candidate cells, denoted as D∈R (M+m)×n ; Use the non-negative matrix decomposition algorithm to decompose the matrix D, and you can get D (M+m)×n =W (M+m)×k ×H k×n , find the coefficient matrix H k×n The row index corresponding to the maximum value in each column is the clustering result, which is recorded as cluster = {arg max i=1,…,k (H k,1 ),arg max i=1,…,k (H k,2 ),...,arg max i=1,...,k (H k,n )}, calculate the average silhouette coefficient of the cluster, and calculate the significance p_value of the survival difference between samples according to the cluster grouping.

[0110] For example, assuming that the number of candidate cells currently selected is m=5, the cell type matrix of fixed cells and candidate cells is extracted from the original cell ratio matrix X, which is recorded as D∈R (6+5)×1158 ; Using the non-negative matrix decomposition algorithm, we can get D (6+5)×1158 =W (6+5)×3 ×H3×1158 ; Find the coefficient matrix H 3×1158 The row index corresponding to the maximum value in each column is the clustering result, which is recorded as cluster = {arg max i=1,...,3 (H 3,1 ),arg max i=1,...,3 (H 3,2 ),...,arg max i=1,...,3 (H 3,1158 )}, calculate the average silhouette coefficient of the cluster, and calculate the significance p_value of the survival difference between samples according to the cluster grouping, and obtain the cluster analysis result diagram shown in Figure 6.

[0111] Step 104: Select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping group.

[0112] In some exemplary embodiments, selecting one of the cell combinations as the cell combination for immunophenotyping comprises:

[0113] Based on the p-value and silhouette coefficient corresponding to each cell combination, one of the cell combinations is selected as the cell combination for immunophenotyping. The p-value and silhouette coefficient of the selected cell combination meet the following conditions 1 and 2:

[0114] Condition 1: The p-value is less than the preset p-value threshold;

[0115] Condition 2: The silhouette coefficient is greater than the preset silhouette coefficient threshold.

[0116] The Silhouette Coefficient (SCo) is a metric used to evaluate clustering effectiveness. Essentially, it measures the ratio of the distance from each sample point to its own cluster to the distance between the nearest cluster structure. The Silhouette Coefficient ranges from -1 to 1. A Silhouette Coefficient closer to 1 indicates better clustering performance, while a Silhouette Coefficient closer to -1 indicates poorer performance. The resulting Silhouette Coefficient is the average of the Silhouette Coefficients for each sample point.

[0117] Exemplarily, the preset silhouette coefficient threshold may be between 0.7 and 0.8. For example, the preset silhouette coefficient threshold may be 0.75. However, the embodiment of the present disclosure is not limited thereto.

[0118] Exemplary, taking the aforementioned 8192 cell combinations as an example, FIG6 is a cluster analysis result diagram of the 8192 cell combinations. As shown in FIG6 , a cell type combination with an average silhouette coefficient silhouette> 0.7 and -log10 (p_value)> 4 and its clustering result (i.e., any one of all cell combinations located in the gray area in FIG6 ) can be selected as the final immunophenotyping. In FIG6 , each plus sign or circled plus sign represents a cell combination, the horizontal axis is the average silhouette coefficient silhouette, and the vertical axis is the significance p_value taking the opposite number of log10. The better the significance, that is, the smaller the pvalue, the larger the -log10 (p_value).

[0119] In some exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0120] Condition 3: f1(p_value)-A*f2(silhouette)-b≥0, where f1(p_value) is a function of p_value, and the larger the p_value, the smaller the value of f1(p_value); f2(silhouette) is a function of silhouette, and the larger the silhouette, the larger f2(silhouette), A<0, b>0.

[0121] For example, f1(p_value) = -log10(p_value); f2(silhouette) = silhouette. However, the present disclosure is not limited to this, and the specific function forms of f1(p_value) and f2(silhouette) can be set as needed. The following description uses f1(p_value) = -log10(p_value); f2(silhouette) = silhouette as an example.

[0122] In this embodiment, A represents the slope of a descending line a, b is the intercept of the descending line a, and condition 3 indicates that the cell combination is located above the descending line a. As shown in Figure 6, the cell type combination located in the gray area and to the upper right of the descending line a and its clustering result (i.e., any one of all cell combinations represented by a circled plus sign in Figure 6) can be selected as the final immunophenotyping.

[0123] In some exemplary embodiments, Among them, y2=max(f1(p_value)), y1=f1(p_value)p_value<0.01, x2=f2(silhouette)silhouette∈[0.5,08], x1=max(f2(silhouette)), b=y2-A·x2, max means taking the maximum value.

[0124] Taking f1(p_value)=-log10(p_value) and f2(silhouette)=silhouette as an example, the values ​​of y2, y1, x2, and x1 can be: y2=max(-log10(p_value)), y1=4, x2=0.7, x1=max(silhouette). However, the present disclosure is not limited to this, and the values ​​of y2, y1, x2, and x1 can be set as needed. For example, in other examples, y1 can be a real number greater than 2.

[0125] In some other exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0126] Condition 3: f1(p_value)-A*f2(silhouette)-b≥c, where f1(p_value) is a function of p_value, and the larger the p_value, the smaller the value of f1(p_value); f2(silhouette) is a function of silhouette, and the larger the silhouette, the larger f2(silhouette), A<0, b>0, c>0.

[0127] In this embodiment, the cell combination with a larger c and a smaller p_value is more likely to be selected. The methods for determining the values ​​of f1 (p_value), f2 (silhouette), and A and b can be referred to above and will not be repeated here.

[0128] In some further exemplary embodiments, the p-value and silhouette coefficient of the selected cell combination further satisfy the following condition 3:

[0129] Condition 3: Among them, f1(p_value) is a function of p_value, and the larger p_value is, the smaller the value of f1(p_value) is; f2(silhouette) is a function of silhouette, and the larger the silhouette is, the larger f2(silhouette) is, A<0, b>0, c≥0.

[0130] In this embodiment, the larger the c, the easier it is to select a cell combination that satisfies both a smaller p_value and a larger silhouette. The methods for determining the values ​​of f1(p_value), f2(silhouette), and A and b can be referred to above and will not be repeated here.

[0131] In some exemplary embodiments, the selected cell combination is the cell combination with the smallest p-value among all cell combinations that simultaneously meet condition 1, condition 2, and condition 3.

[0132] Exemplary, the cell combination selected as final immunophenotyping in Fig. 6 is the cell type combination that satisfies condition 1, condition 2 and condition 3 simultaneously and significance p_value is minimum.Specifically, the cell type combination is (CD4+ Tem, CD4+ naive T-cells, CD8+ naive T-cells, Macrophages M2, CD8+ T-cells, Tregs, CD4+ memory T-cells, Neutrophils, Fibroblasts, Th1cells, Memory B-cells, Class-switched memory B-cells). Fig. 7 is the survival curve graph corresponding to the cell type combination.

[0133] In some exemplary embodiments, the selected cell combination is the cell combination with the smallest p-value and no intersection of survival curves among all cell combinations that simultaneously meet condition 1, condition 2, and condition 3.

[0134] Exemplary, the cell combination selected as final immunophenotyping in Fig. 8 is the cell type combination that satisfies condition 1, condition 2 and condition 3 simultaneously and survival curve does not cross and significance p_value is minimum.Specifically, the cell type combination is (CD4+Tem, CD8+naive T-cells, Macrophages M2, Macrophages M1, Th1 cells, CD8+T-cells, Tregs, Neutrophils, Endothelial cells, Fibroblasts, Pericytes, naive B-cells, Memory B-cells, Class-switched memory B-cells). Fig. 9 is the survival curve figure corresponding to the cell type combination.

[0135] In some exemplary embodiments, selecting one of the cell combinations as the cell combination for immunophenotyping comprises:

[0136] Obtain all cell combinations that simultaneously meet conditions 1, 2, and 3, and record them as the first set;

[0137] Obtain a survival curve graph for each cell combination in the first set;

[0138] Based on the survival curve graph, detecting whether the survival probabilities of multiple immunophenotypes of each cell combination in the first set at multiple preset time points are ranked in the same order;

[0139] Determine a cell combination having the same ranking order of survival probability for multiple immunotypes at multiple preset time points, and record it as a second set;

[0140] The cell combination with the smallest p-value in the second set is selected as the final selected cell combination for immunophenotyping.

[0141] In this embodiment, multiple cell combinations that meet the above conditions 1, 2, and 3 are recorded as the first set G1, and the following operations are performed on each cell combination in the first set G1: a survival curve graph of the cell combination is obtained, and the survival probability ranking order of multiple immunotyping of the cell combination at multiple preset time points is detected to be the same. If they are the same, it is considered that the two curves in the survival curve graph of the cell combination do not intersect, and the cell combination in which all survival curve graphs in the first set G1 do not intersect is recorded as the second set G2, and the cell combination with the smallest p-value in the second set G2 is used as the final selected immunotyping cell combination.

[0142] Illustratively, the preset multiple time points may be 1 year, 5 years, 10 years, and 15 years, however, the present disclosure is not limited thereto.

[0143] Specifically, the following operations are performed on each cell combination in the first set G1: a survival curve diagram of the cell combination is obtained, the 1-year survival probability of each immunophenotyping group is sorted from high to low, the corresponding immunophenotyping groups are named cluster_1 to cluster_k, and the cell type combinations whose survival probabilities of cluster_1 to cluster_k decrease successively within the time range of time = (1 year, 5 years, 10 years, 15 years) are found. The found cell type combinations are added to the second set G2, and the cell type combination with the smallest significance p_value in the second set G2 is selected as the final immunophenotyping.

[0144] In some exemplary embodiments, determining the characteristic cell type corresponding to each immunophenotyping group comprises:

[0145] The following operations are performed for each cell type in the immunophenotyping cell combination: the median of the cell type in each immunophenotyping group is calculated; and the cell type is determined to be the characteristic cell type corresponding to the immunophenotyping group with the largest median.

[0146] In this embodiment, the characteristic cell type of each immunophenotyping grouping is determined by calculating the median proportion of each cell type in each immunophenotyping grouping. The cell type is the characteristic cell type of the immunophenotyping grouping in which the median proportion of the cell type is the largest. For example, if the medians of B cells in three immunophenotyping groups are 0.05, 0.07, and 0.01, respectively, then B cells are the characteristic cells of the second immunophenotyping grouping.

[0147] As shown in Figure 10A, the characteristic cell types of the first immunophenotyping group B1 are (CD4+Tem, CD8+naive T-cells, Macrophages M2, Macrophages M1, Th1cells, CD8+T-cells, Tregs), the characteristic cell types of the second immunophenotyping group B2 are (Neutrophils, Endothelial cells, Fibroblasts, Pericytes), and the characteristic cell types of the third immunophenotyping group B3 are (naive B-cells, Memory B-cells, Class-switched memory B-cells).

[0148] Step 105: Determine the differentially expressed genes (DEGs) corresponding to each immunophenotyping group, and determine at least one of the following omics features corresponding to each immunophenotyping group: differentially mutated genes, differential copy number variations, and differentially methylated positions (DMPs).

[0149] In some exemplary embodiments, determining the differentially expressed genes corresponding to each immunophenotyping group includes:

[0150] For each immunophenotyping group, the gene expression fold change and p-value of each gene between the current immunophenotyping group and the control group corresponding to the current immunophenotyping group were calculated, and genes with a calculated gene expression fold change greater than the preset second fold threshold and a p-value less than the preset first p-value threshold were selected as differentially expressed genes of the current immunophenotyping group.

[0151] In some exemplary embodiments, the number of immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

[0152] For example, the screening conditions for differentially expressed genes in each immunophenotyping group are that the gene expression fold change log2FC1>1 and p_value<0.05, That is, the preset second multiple threshold is 1, and the preset first p-value threshold is 0.05.

[0153] In this embodiment, differentially expressed genes refer to genes that are differentially expressed in two or more grouped samples, that is, their expression levels are significantly different in different grouped samples.

[0154] FIG10B is a heat map of differentially expressed genes in the immunophenotyping groups shown in FIG10A , where the horizontal axis represents samples and the vertical axis represents genes. Only the names of genes related to immunity are marked in the figure.

[0155] In some exemplary embodiments, the number of immunophenotyping groups is greater than or equal to 3, and determining the differentially expressed genes corresponding to each immunophenotyping group includes:

[0156] Determine that every two immunotyping groups are an experimental control combination, and set the experimental group and the control group in the experimental control combination (for example, one immunotyping group in the experimental control combination can be set as the experimental group, and the other immunotyping group can be set as the control group);

[0157] For each experimental-control combination, the fold change in gene expression and the p-value of each gene between the experimental and control groups were calculated;

[0158] The differentially expressed genes of each immunophenotyping group are determined according to the calculated gene expression fold change and the p-value. The differentially expressed genes of each immunophenotyping group are the intersection of multiple first differentially expressed genes and multiple second differentially expressed genes. The multiple first differentially expressed genes are genes whose calculated gene expression fold change is greater than a preset second fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the experimental group. The multiple second differentially expressed genes are genes whose calculated gene expression fold change is less than the opposite of the preset second fold threshold and whose p-value is less than the preset first p-value threshold when the immunophenotyping group is the control group.

[0159] For example, assuming there are three immunophenotyping groups: B1, B2, and B3, three experimental control combinations can be set up: B1 vs. B2, B1 vs. B3, and B2 vs. B3. Assuming that the differentially expressed genes in group B2 are required, in B2 vs. B3, group B2 is the experimental group. If the calculated gene expression fold change of a gene is greater than the preset second fold change threshold and the p-value is less than the preset first p-value threshold, the gene can be classified as the first differentially expressed gene. In B1 vs. B2, group B2 is the control group. If the calculated gene expression fold change of a gene is less than the inverse of the preset second fold change threshold and the p-value is less than the preset first p-value threshold, the gene can be classified as the second differentially expressed gene. The intersection of the first differentially expressed genes and the second differentially expressed genes is taken as the differentially expressed genes in group B2.

[0160] In some exemplary embodiments, the method further comprises:

[0161] Obtain a validation set and use the immunophenotyping cell combination and its differentially expressed genes to validate the validation set.

[0162] In the disclosed embodiment, after obtaining the immunophenotyping cell combination, the cell types and differentially expressed genes determined by the training set immunophenotyping can be used to verify the validation set. The results show that the validation set results are consistent with the training set.

[0163] In some exemplary embodiments, the validation set is validated using the immunophenotyping cell combination and its differentially expressed genes, including:

[0164] The scores or proportions of multiple cell types are calculated using the immune infiltration algorithm;

[0165] Determine the immunophenotyping of samples in the validation set based on the characteristic cell types of each immunophenotyping group obtained in the training set;

[0166] Detect whether the gene expression levels of the differentially expressed genes in the validation set are similar to the gene expression levels of the differentially expressed genes in each immunophenotyping group in the training set.

[0167] Exemplarily, the validation set also comes from the GEO database. The present disclosure selects 9 DLBCL gene expression datasets independent of the training set as the validation set. The deconvolution tool xCell is used to calculate the proportions of multiple cell types in each validation set. The immunophenotyping of the sample is determined based on the characteristic cell types of each immunophenotyping group obtained in the training set. Based on the differentially expressed genes of each immunophenotyping group obtained in the training set, it is determined that the gene expression on each validation set is consistent with the training set results.

[0168] In some exemplary embodiments, the method further comprises:

[0169] Screen the differentially expressed genes corresponding to each immunophenotyping group;

[0170] A prognostic risk prediction model is constructed using the screened differentially expressed genes, and the first sample data is grouped by prognostic risk according to the prognostic risk score output by the prognostic risk prediction model;

[0171] The differentially expressed genes that were significantly positively correlated with the prognostic risk groups were identified as important genes affecting the disease.

[0172] In the disclosed embodiments, when screening the differentially expressed genes corresponding to each immunophenotyping group, the following two screening methods can be used: (1) calculating the correlation between the differentially expressed genes and the prognosis survival, and retaining the differentially expressed genes that are significantly correlated with the prognosis survival and have a p-value less than a preset p-value threshold (such as 0.05). (2) using a machine learning algorithm to screen the differentially expressed genes corresponding to each immunophenotyping group. In the disclosed embodiments, the machine learning algorithm includes but is not limited to: L1 regularized regression (Lasso), elastic net (Elastic net), random forest survival analysis (randomForestSRC) and other algorithms.

[0173] In the disclosed embodiment, a prognostic risk prediction model can be constructed using machine learning models such as random forest survival analysis (randomForestSRC), gradient boosting machine (GBM), and support vector machine (SVM). The input of the prognostic risk prediction model is a differentially expressed gene matrix, the training label is the prognostic survival time and survival status of the sample, and the output of the prognostic risk prediction model is a prognostic risk score. The higher the prognostic risk score, the shorter the predicted survival time of the sample.

[0174] For example, the disclosed embodiment uses lasso to screen differentially expressed genes and GBM to construct a prognostic risk prediction model. The screened genes and their importance ranking in the model are shown in Table 2.

[0175] Table 2

[0176] In the embodiment of the present disclosure, the prediction accuracy of the prognostic risk prediction model can also be detected by calculating the consistency index (C-index). C-index is a commonly used statistical indicator for evaluating the predictive ability of the model, and is particularly widely used in survival analysis and risk prediction models. C-index measures whether the relative ranking of the risk score (or survival time) predicted by the model in a group of samples and a pair of randomly selected samples is consistent with the order of the events actually observed (such as death, recurrence, etc.). Exemplarily, for the validation set with prognostic survival information, the differentially expressed genes obtained by the above method are used to extract the gene expression matrix of the differentially expressed genes, which are input into the constructed prognostic risk prediction model to obtain the prognostic risk score, and the C-index value is calculated. As shown in Figure 11A, the C-index values ​​of the embodiment of the present disclosure on the training set and the validation set are higher than the results of the prognostic risk prediction model constructed by genes in other literature. The prognostic risk prediction model constructed by the embodiment of the present disclosure can help predict the survival of cancer patients and help doctors and patients have a better basis when making decisions.

[0177] In the disclosed embodiment, the prediction accuracy of the model can also be determined by detecting the correlation between the prognostic risk grouping and survival. Exemplarily, for the training set and the validation set, the median (or mean) of the prognostic risk score is used as a threshold to divide the samples into a high-risk group and a low-risk group. From Figures 11B to 11D, it can be seen that in the disclosed embodiment, the survival differences between the prognostic risk groups are significant, and the survival probability of patients in the high-risk group is lower than that of patients in the low-risk group, and the overall survival time is also relatively short.

[0178] In the disclosed embodiments, the important genes that affect the disease can be determined by determining the differentially expressed genes that are significantly positively correlated with the prognostic risk grouping. As previously mentioned, in the training set and the validation set, the median (or mean) of gene expression is used as the threshold, and the samples are divided into a high expression group and a low expression group according to the threshold. As can be seen from Figures 11E and 11F, the expression groups of the two genes ranked top two in importance in Table 2: MYC and ADH1B are significantly positively correlated with the prognostic risk grouping, indicating that these two genes have a negative effect on prognosis. Studies have shown that abnormal expression or gene rearrangement of the MYC gene is often associated with the aggressiveness and poor prognosis of the disease, and the ADH1B gene is of equal importance to MYC in the model, which may reveal new biomarkers.

[0179] The immunophenotyping-based multi-omics analysis method of the disclosed embodiments is applicable not only to the diagnosis and treatment of diffuse large B-cell lymphoma, but also to other diseases such as mantle cell lymphoma (MCL) and breast cancer. Mantle cell lymphoma is a subtype of non-Hodgkin lymphoma that arises from mature B cells. However, MCL is a heterogeneous tumor, with nearly 20% of cases exhibiting distinct clinicopathological, immunophenotypic, cytogenetic, and molecular features.

[0180] From the GEO database, two gene expression data sets of mantle cell lymphoma (MCL) were selected. According to the multi-omics analysis method based on immunophenotyping of the disclosed embodiment, 20 cell types significantly related to prognosis survival were screened out, 4 fixed cell types, and 16 candidate cell types. The clustering results of various cell combinations are shown in Figure 12A. The cell type combination and its clustering result with the minimum significance p_value of the average silhouette coefficient silhouette>0.7 and the survival difference were final immunophenotyping. The survival curve graph and cell ratio graph of the immunophenotyping are shown in Figures 12B and 12C.

[0181] Download the expression data of breast cancer from the TCGA database. According to the multi-omics analysis method based on immunophenotyping of the disclosed embodiment, 16 cell types significantly correlated with prognosis survival were screened out, 4 fixed cell types ("Class-switched memory B-cells", "CD8+Tcm", "CD8+T-cells", "cDC"), and 12 candidate cell types. The clustering results of multiple cell combinations are shown in Figure 13A. The cell type combination with the smallest significance p_value of the average silhouette coefficient silhouette>0.7 and the survival difference was selected as the final immunophenotyping. The survival curve graph and cell ratio graph of the immunophenotyping are shown in Figures 13B and 13C.

[0182] In some exemplary embodiments, in step 105, determining the differentially mutated genes corresponding to each immunophenotyping group includes:

[0183] Acquire second sample data, where the second sample data includes gene mutation data;

[0184] Encode the gene mutation data to obtain the gene mutation spectrum matrix;

[0185] For each immunophenotyping group, the mutation rate fold change and p-value of each gene between the current immunophenotyping group and the control group corresponding to the current immunophenotyping group are calculated according to the gene mutation spectrum matrix, and the genes with the calculated mutation rate fold change greater than the preset first fold threshold and the p-value less than the preset first p-value threshold are selected as the differential mutation genes of the current immunophenotyping group.

[0186] In some exemplary embodiments, the number of immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

[0187] In the disclosed embodiment, for a data set that has both gene expression data and gene mutation data, immunophenotyping groups can be obtained based on the gene expression data, and then the differential mutation genes can be determined based on the immunophenotyping groups. Specifically, the gene mutation data is first encoded (for example, one-hot encoding can be used), and the sample is recorded as 1 if there is a mutation on the gene, otherwise it is 0, and a gene mutation spectrum matrix with a row index of gene and a column index of sample is constructed. Then, according to the immunophenotyping grouping, the difference in mutation spectrum between one group and the remaining other groups is calculated, and the calculation method is as follows:

[0188] (1) Select one group as the experimental group and the remaining groups as the control group. Calculate the inter-group mutation rate fold change log2FC2 and p-value for each gene. The calculation formula of FC2 is as follows:

[0189] (2) One group in the immunophenotyping group was selected as the experimental group and the rest as the control group in turn. Genes with log2FC2>1 and p<0.05 were selected as the differentially mutated genes in the immunophenotyping group.

[0190] For example, the aforementioned steps are performed on a validation dataset with both gene expression data and gene mutation data to obtain differentially mutated genes for each immunophenotyping group. As shown in Figure 14A, the differentially mutated genes for immunophenotyping group B1 are (PKHD1L1, DNHD1, PIK3CD, FADD, LRRIQ1); the differentially mutated genes for immunophenotyping group B2 are (HIST1H1B, PCDH10, CIITAPOT1, PKD1, ETS1); and the differentially mutated genes for immunophenotyping group B3 are (IGLL5, IGLJ1, TNFAIP3, CD79B).

[0191] In another example, in another independent validation set, the above steps were performed to obtain immunophenotyping groups and corresponding differentially mutated genes. As shown in Figures 14B and 14C, it can be seen that the gene mutations of genes PCDH10 and CD79B also differ between the immunophenotyping groups.

[0192] In some exemplary embodiments, the number of immunophenotyping groups is greater than or equal to 3, and determining the differentially mutated genes corresponding to each immunophenotyping group includes:

[0193] Acquire second sample data, where the second sample data includes gene mutation data;

[0194] Encode the gene mutation data to obtain the gene mutation spectrum matrix;

[0195] Determine that every two immunotyping groups are an experimental control combination, and set the experimental group and the control group in the experimental control combination (for example, one immunotyping group in the experimental control combination can be set as the experimental group, and the other immunotyping group can be set as the control group);

[0196] For each experimental-control combination, the mutation rate fold change and p-value of each gene between the experimental group and the control group were calculated based on the gene mutation spectrum matrix;

[0197] The differentially mutated genes of each immunotyping group are determined based on the calculated mutation rate fold change and p-value. The differentially mutated genes of each immunotyping group are the intersection of multiple first differentially mutated genes and multiple second differentially mutated genes. The multiple first differentially mutated genes are genes whose calculated mutation rate fold change is greater than a preset first fold threshold and whose p-value is less than the preset first p-value threshold when the immunotyping group is the experimental group. The multiple second differentially mutated genes are genes whose calculated mutation rate fold change is less than the opposite of the preset first fold threshold and whose p-value is less than the preset first p-value threshold when the immunotyping group is the control group.

[0198] For example, assuming there are three immunophenotyping groups: B1, B2, and B3, three experimental control combinations can be set up: B1 vs. B2, B1 vs. B3, and B2 vs. B3. Assuming the differentially mutated genes in group B2 are required, in B2 vs. B3, group B2 is the experimental group. If the calculated mutation rate fold change for a gene is greater than a preset first fold change threshold and the p-value is less than the preset first p-value threshold, the gene can be classified as the first differentially mutated gene. In B1 vs. B2, group B2 is the control group. If the calculated mutation rate fold change for a gene is less than the inverse of the preset first fold change threshold and the p-value is less than the preset first p-value threshold, the gene can be classified as the second differentially mutated gene. The intersection of the first and second differentially mutated genes is taken as the differentially mutated genes in group B2.

[0199] In some exemplary embodiments, determining the differential copy number variation corresponding to each immunophenotyping group comprises:

[0200] Acquiring third sample data, where the third sample data includes genome copy number data;

[0201] For each immunophenotyping group, perform the following operations:

[0202] determining the copy number of genomic intervals on multiple chromosomes based on genomic copy number data;

[0203] For each chromosome in the multiple chromosomes, when the genomic interval copy number is greater than the preset copy number threshold, the immunotyping on the current chromosome is determined to include an increase in the genomic copy number and the proportion of the number of intervals with increased genomic copy number to the total number of intervals is counted; when the genomic interval copy number is less than the opposite of the preset copy number threshold, the immunotyping on the current chromosome is determined to include a loss of genomic copy number and the proportion of the number of intervals with lost genomic copy number to the total number of intervals is counted.

[0204] Copy Number Alteration (CNA) analysis is a technique used to study changes in the number of DNA copies within a genome. CNA typically refers to an increase (amplification) or decrease (deletion) in the number of DNA copies of a specific gene or genomic region within a cell or tissue sample. This analysis is particularly important in the study of cancer and other genetic diseases, as copy number changes can affect gene function and lead to pathological changes.

[0205] In some exemplary embodiments, the calculation formula for the genomic interval copy number is as follows:

[0206] Segment_Mean = log2(copynumber / 2), where Segment_Mean is the genomic interval copy number, copynumber is the copy number value within the interval, and the genomic interval copy number represents the change in the copy number value within the interval relative to the normal copy number (the normal copy number is 2).

[0207] In some exemplary embodiments, the preset copy number threshold may be a real number between 0.2 and 1. For example, the preset copy number threshold may be 0.2, 0.3, etc.

[0208] For example, it is assumed that the preset copy number threshold is set to 0.2. When the calculated Segment_Mean<-0.2, it is judged that the genome copy number is lost, and the number of intervals with copy number loss is counted as the ratio of the total number of intervals, which is recorded as FGL (Fraction of Genome Lost); when the calculated Segment_Mean>0.2, it is judged that the genome copy number is increased, and the number of intervals with copy number gain is counted as the ratio of the total number of intervals, which is recorded as FGG (Fraction of Genome Gained). As shown in Figures 15A to 15G, on chromosome 12, there are significant differences in copy number gain between the immunophenotyping groups, and on chromosomes 2, 8, 11, 15, 19, and 22, there are significant differences in copy number loss.

[0209] In some exemplary embodiments, determining the differentially methylated sites corresponding to each immunophenotyping group comprises:

[0210] Acquiring fourth sample data, the fourth sample data including methylation data, the methylation data including methylation rates of multiple sites;

[0211] Filtering the methylation data to determine the intersection sites in the fourth sample data;

[0212] For each immunophenotyping group, the differential methylation rate and p-value of each intersection site between the current immunophenotyping group and the control group corresponding to the current immunophenotyping group are calculated, and the intersection site with a calculated differential methylation rate greater than the preset first differential methylation rate threshold and a p-value less than the preset first p-value threshold is selected as the differential methylation site of the current immunophenotyping group.

[0213] For example, the preset first differential methylation rate threshold may be 10, 20, 30, etc., and the preset first p-value threshold may be 0.05. However, the present disclosure is not limited thereto.

[0214] In some exemplary embodiments, the number of immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

[0215] In the disclosed embodiments, for datasets with both gene expression data and methylation data, immunophenotyping groups can be obtained based on the gene expression data, and then differentially methylated sites can be determined based on the immunophenotyping groups. Specifically, the methylation data is first filtered, and sites with a sequencing depth greater than 10 are retained. The methylation rate of each site is calculated using the following formula:

[0216] All intersection sites between samples are retained.

[0217] Then, based on the immunophenotyping groups, the differential methylation sites between one group and the remaining groups were calculated as follows:

[0218] One group of immunotyping groups was selected in turn as the experimental group, and all other immunotyping groups were selected as the control group. The differential methylation rate (methydiff) of each site was calculated according to the following formula:

[0219] methydiff = average methylation rate (experimental group) - average methylation rate (control group);

[0220] The p-value of each site is calculated. Sites with a calculated differential methylation rate greater than a preset first differential methylation rate threshold (e.g., 20) and a p-value less than a preset first p-value threshold (e.g., 0.05) are retained as differential methylation sites for the immunophenotyping group.

[0221] In some exemplary embodiments, the number of immunophenotyping groups is greater than or equal to 3, and determining the differentially methylated sites corresponding to each immunophenotyping group includes:

[0222] Acquiring fourth sample data, the fourth sample data including methylation data, the methylation data including methylation rates of multiple sites;

[0223] Filtering the methylation data to determine the intersection sites in the fourth sample data;

[0224] Determine that every two immunotyping groups are an experimental control combination, and set the experimental group and the control group in the experimental control combination (for example, one immunotyping group in the experimental control combination can be set as the experimental group, and the other immunotyping group can be set as the control group);

[0225] For each experimental-control combination, the differential methylation rate and p-value of each intersection site between the experimental group and the control group were calculated;

[0226] The differential methylation sites of each immunophenotyping group are determined based on the calculated differential methylation rate and p-value. The differential methylation sites of the immunophenotyping group are the intersection of multiple first differential methylation sites and multiple second differential methylation sites. The multiple first differential methylation sites are the intersection sites where the calculated differential methylation rate is greater than the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold when the immunophenotyping group is the experimental group. The multiple second differential methylation sites are the intersection sites where the calculated differential methylation rate is less than the opposite of the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold when the immunophenotyping group is the control group.

[0227] For example, assuming there are three immunophenotyping groups: B1, B2, and B3, three experimental control combinations can be set up: B1 vs. B2, B1 vs. B3, and B2 vs. B3. Assuming the differentially methylated sites in group B2 are required, in B2 vs. B3, group B2 is the experimental group. If the calculated differential methylation rate for a site is greater than a preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold, the site is classified as the first differential methylation site. In B1 vs. B2, group B2 is the control group. If the calculated differential methylation rate for a site is less than the inverse of the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold, the site is classified as the second differential methylation site. The intersection of the first and second differential methylation sites is taken as the differential methylation sites in group B2.

[0228] In some exemplary embodiments, the method further comprises:

[0229] Determine the nearest gene to the differentially methylated site, calculate the correlation coefficient and p-value between the differentially methylated site and the nearest gene for each immunophenotyping group, retain the differentially methylated site with a p-value less than a preset second p-value threshold, and determine the nearest gene of the retained differentially methylated site as the associated gene of the differentially methylated site and the differentially expressed gene of the immunophenotyping group.

[0230] For example, the preset second p-value threshold may be 0.05, however, the present disclosure is not limited thereto.

[0231] In the disclosed embodiment, the nearest gene to the differentially methylated site is found, and for each immunophenotyping group, the correlation coefficient and p-value between the differentially methylated site and the nearest gene are calculated, and the differentially methylated site that is correlated with the nearest gene (i.e., the p-value is less than a preset second p-value threshold, for example, p-value <0.05) is retained, and the nearest gene of the retained differentially methylated site is determined to be the correlated gene of the differentially methylated site and the differentially expressed gene of the current immunophenotyping group.

[0232] In one example, for a validation set with gene expression data and methylation data, the aforementioned steps are performed to obtain immunophenotyping groups, the first differential methylation rate threshold methydiff is set to 20, the first p-value threshold is set to 0.01, the differential methylation sites are obtained by calculation, the nearest gene is found by the chipseeker tool, and the correlation coefficient between the differential methylation site and the nearest gene is calculated using person correlation. As shown in Table 2 and Figures 9-10, Table 3 shows the top 5 differential methylation sites ranked by the absolute value of the correlation coefficient (Correlation_r) in each immunophenotyping group, and Figures 16A to 16E, 16F to 16J, and 16K to 16O respectively show the expression levels and methylation rate schematic diagrams of the top 5 differential methylation sites ranked by the absolute value of the correlation coefficient in immunophenotyping groups B1, B2, and B3.

[0233] Table 3

[0234] In some exemplary embodiments, the method further comprises:

[0235] Acquiring fifth sample data, where the fifth sample data includes at least one clinical characteristic data;

[0236] The significant p-value between each clinical characteristic data and the immunophenotyping group was calculated, and the association or difference between the clinical characteristics and the immunophenotyping was determined based on the calculated significant p-value.

[0237] In the embodiments of the present disclosure, when the clinical characteristics are numerical variables, the significance p-value can be calculated by the F test or the Kruskal-Wallis test method; when the clinical characteristics are categorical variables, the significance p-value can be calculated by the chi-square test or the Fisher exact test method.

[0238] In one example, as shown in FIG17A to FIG17F , the significance p values ​​between multiple groups of clinical characteristic data and immunophenotyping groups were calculated. The calculation results showed that the significance p values ​​were all less than 0.05, indicating that the immunophenotyping groups were significantly correlated with the clinical characteristic data. Figure 17A shows the significant p-values ​​between age and immunophenotyping groups; Figure 17B shows the significant p-values ​​between lactate dehydrogenase (LDH) levels and immunophenotyping groups, with high LDH levels generally associated with a worse prognosis; Figure 17C shows the significant p-values ​​between the Ann Arbor stage (ann.arbor.stage) and immunophenotyping groups. This staging method is a recognized lymphoma classification standard both domestically and internationally, and is based on clinical manifestations, physical examination, B-ultrasound, CT, lower limb lymphangiography, and inferior vena cava angiography. Figure 17D shows the significant p-values ​​between cell of origin and immunophenotyping groups, with GCB generally having a better prognosis than ABC. Figure 17E shows the significant p-values ​​between the age-adjusted prognostic index (aaIPI) and immunophenotyping groups; and Figure 17F shows the significant p-values ​​between the prognostic index (IPI) and immunophenotyping groups, with a high IPI associated with a poor prognosis.

[0239] In some exemplary embodiments, the method further comprises:

[0240] Obtain a sample to be tested, determine one or more omics features of the sample to be tested, and determine the immunophenotyping group of the sample to be tested based on the determined omics features.

[0241] The disclosed embodiments further expand the description of immunophenotyping characteristics of related diseases (such as DLBCL) by analyzing one or more omics features (differentially expressed genes, differentially mutated genes, differential copy number variations, differential methylation sites) corresponding to each immunophenotyping group, and improve the characteristic spectrum of immunophenotyping in multiple omics, that is, the differences in multiple omics.

[0242] In summary, for DLBCL, the characteristic cell types of immunophenotyping group B1 are (CD4+Tem, CD8+naive T-cells, Macrophages M2, Macrophages M1, Th1cells, CD8+T-cells, Tregs), the characteristic cell types of immunophenotyping group B2 are (Neutrophils, Endothelial cells, Fibroblasts, Pericytes), and the characteristic cell types of immunophenotyping group B3 are (naive B-cells, Memory B-cells, Class-switched memory B-cells).

[0243] The omics characteristics of differentially mutated genes are: immunophenotyping group B1 (PKHD1L1, DNHD1, PIK3CD, FADD, LRRIQ1); immunophenotyping group B2 (HIST1H1B, PCDH10, CIITAPOT1, PKD1, ETS1); immunophenotyping group B3 (IGLL5, IGLJ1, TNFAIP3, CD79B).

[0244] The genomic characteristics of chromosome copy number variation are: copy number gain of chromosome 12, and copy number loss of chromosomes 2, 8, 11, 15, 19, and 22. Among them, immunotyping group B1 (the lowest proportion of copy number variation); immunotyping group B2 (copy number loss of chr2, chr15, chr19, and chr22); immunotyping group B3 (copy number gain of chr12, copy number loss of chr8 and chr11);

[0245] The omics characteristics of differentially methylated sites and related genes are as follows: differentially methylated sites and related genes in immunophenotyping group B1 include: chr10.13452301 (INPP5A), chr3.127749576. (RUVBL1), chr1.151169194 (PIP5K1A), chr3.48283854 (CDC25A), chr17.36805288 (MLLT6); differentially methylated sites and related genes in immunophenotyping group B2 include: chr3.71120196 (FOXP1), chr13.101317771 (TMTC4), chr7.73742163 (CLIP2), chr3.48757698 (IP6K2), chr19.50926533 (SPIB); The differentially methylated sites and related genes in immunophenotyping group B3 include: chr9.17135705 (CNTLN), chr14.61747027 (PRKCH), chr14.66974547 (GPHN), chr8.98290013 (TSPYL5), and chr7.3341596 (CARD11).

[0246] If a new sample does not have gene expression data but has other omics data, then immunophenotyping can be performed using these other omics data. That is, after perfecting the multi-omics profile of immunophenotyping, the disclosed embodiments can use any of these omics data, such as gene expression data, gene mutations, copy number variations, and methylation data, to determine immunophenotyping groups.

[0247] The present disclosure provides a multi-omics analysis method based on immunophenotyping, which includes analyzing gene expression data through immune infiltration, calculating the score or proportion of cell types, screening N cell types from multiple cell types based on at least one of the following: survival significance, variance and correlation, dividing the screened cell types into fixed cell types and candidate cell types, and performing cluster analysis on multiple cell combinations, selecting one of the cell combinations as the final immunophenotyping cell combination based on the cluster analysis results, and performing multi-omics feature analysis on the immunophenotyping results, integrating omics data at different levels, and providing a multi-dimensional disease feature description. This multi-dimensional analysis method has higher complexity and accuracy than previous studies on single cell types or gene expression; this method not only improves the accuracy of subtypes, but may also reveal new biomarkers, thereby improving clinical management and the selection of treatment options. In addition, the prognostic risk prediction model constructed by the present disclosure based on the differentially expressed genes after screening has higher accuracy than previous models.

[0248] The present disclosure uses immune infiltration analysis to explore the immune cell composition of different typings, which can better explain the biological heterogeneity of tumors and differences in treatment response. It solves the problem that traditional typing cannot be performed due to the inability to determine the cell origin. By using multiple combinations of cell types, the present disclosure maximizes the differences between samples, thereby establishing immune typing and achieving significant differences in survival analysis. The present disclosure can discover more differentially expressed genes. Gene-based typing can better explain the biological heterogeneity of tumors, while taking into account the important role of immune cells in the physiological functions of tumors.

[0249] As shown in FIG18 , the present disclosure also provides a multi-omics analysis device based on immunophenotyping, including an immune infiltration module 1810 , a screening module 1820 , a cluster analysis module 1830 , a selection module 1840 , and an omics feature analysis module 1850 , wherein:

[0250] an immune infiltration module 1810 configured to perform immune infiltration analysis on first sample data to obtain ratios of multiple cell types, wherein the first sample data includes gene expression data;

[0251] A screening module 1820 is configured to screen N cells from a plurality of cell types according to at least one of the following: survival significance, variance, and correlation, where N>1;

[0252] a cluster analysis module 1830 configured to divide the N types of cells into M fixed cells and (NM) candidate cells, and perform cluster analysis on multiple cell combinations, each cell combination including at least M fixed cells and q candidate cells, where M ≥ 1 and M < N, and 0 ≤ q ≤ (NM);

[0253] A selection module 1840 is configured to select one of the cell combinations as a cell combination for immunophenotyping, and determine a characteristic cell type corresponding to each immunophenotyping group;

[0254] The omics feature analysis module 1850 is configured to determine the differentially expressed genes corresponding to each immunophenotyping group, and determine at least one of the following omics features corresponding to each immunophenotyping group: differentially mutated genes, differential copy number variations, and differentially methylated sites.

[0255] In some exemplary embodiments, the omics feature analysis module 1850 is further configured to screen the differentially expressed genes corresponding to each of the immunotyping groups; use the screened differentially expressed genes to construct a prognostic risk prediction model, and group the first sample data into prognostic risk groups according to the prognostic risk score of the input first sample data according to the prognostic risk prediction model; and determine that the differentially expressed genes that are significantly positively correlated with the prognostic risk grouping are important genes that affect the disease.

[0256] In some exemplary embodiments, the omics feature analysis module 1850 determines the differentially mutated genes corresponding to each immunophenotyping group, including:

[0257] Acquire second sample data, where the second sample data includes gene mutation data;

[0258] Encoding the gene mutation data to obtain a gene mutation spectrum matrix;

[0259] For each of the immunotyping groups, the mutation rate fold change and p-value of each gene between the current immunotyping group and the control group corresponding to the current immunotyping group are calculated according to the gene mutation spectrum matrix, and genes with a calculated mutation rate fold change greater than a preset first fold threshold and a p-value less than a preset first p-value threshold are selected as the differentially mutated genes of the current immunotyping group.

[0260] In some exemplary embodiments, the number of the immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

[0261] In some exemplary embodiments, the number of immunophenotyping groups is greater than or equal to 3, and the omics feature analysis module 1850 determines the differentially mutated genes corresponding to each immunophenotyping group, including:

[0262] Acquire second sample data, where the second sample data includes gene mutation data;

[0263] Encoding the gene mutation data to obtain a gene mutation spectrum matrix;

[0264] Determining each of two immunophenotyping groups as an experimental control combination, and setting an experimental group and a control group in the experimental control combination;

[0265] For each experimental-control combination, the mutation rate fold change and p-value of each gene between the experimental group and the control group are calculated according to the gene mutation spectrum matrix;

[0266] The differential mutation genes of each immunophenotyping group are determined according to the calculated mutation rate fold change and p-value, and the differential mutation genes of each immunophenotyping group are the intersection of multiple first differential mutation genes and multiple second differential mutation genes. The multiple first differential mutation genes are genes whose calculated mutation rate fold change is greater than a preset first fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the experimental group. The multiple second differential mutation genes are genes whose calculated mutation rate fold change is less than the opposite of the preset first fold threshold and whose p-value is less than the preset first p-value threshold when the immunophenotyping group is the control group.

[0267] In some exemplary embodiments, the omics feature analysis module 1850 determines the differential copy number variation corresponding to each immunophenotyping group, including:

[0268] Acquiring third sample data, wherein the third sample data includes genome copy number data;

[0269] For each immunophenotyping group, perform the following operations:

[0270] determining the genomic interval copy numbers on a plurality of chromosomes based on the genomic copy number data;

[0271] For each of the multiple chromosomes, when the genomic interval copy number is greater than a preset copy number threshold, the immunotyping on the chromosome is determined to include an increase in the genomic copy number and the proportion of the number of intervals with increased genomic copy numbers to the total number of intervals is counted; when the genomic interval copy number is less than the opposite of the preset copy number threshold, the immunotyping on the chromosome is determined to include a loss of genomic copy number and the proportion of the number of intervals with lost genomic copy numbers to the total number of intervals is counted.

[0272] In some exemplary embodiments, the omics feature analysis module 1850 determines the differentially methylated sites corresponding to each of the immunophenotyping groups, including:

[0273] Acquiring fourth sample data, the fourth sample data including methylation data, the methylation data including methylation rates of multiple sites;

[0274] Filtering the methylation data to determine intersection sites in the fourth sample data;

[0275] For each of the immunotyping groups, the differential methylation rate and p-value of each intersection site between the current immunotyping group and the control group corresponding to the current immunotyping group are calculated, and the intersection site whose calculated differential methylation rate is greater than the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold is selected as the differential methylation site of the current immunotyping group.

[0276] In some exemplary embodiments, the number of immunophenotyping groups is greater than or equal to 3, and the omics feature analysis module 1850 determines the differentially methylated sites corresponding to each immunophenotyping group, including:

[0277] Acquiring fourth sample data, the fourth sample data including methylation data, the methylation data including methylation rates of multiple sites;

[0278] Filtering the methylation data to determine intersection sites in the fourth sample data;

[0279] Determining every two of the immunophenotyping groups as an experimental control combination, and setting an experimental group and a control group in the experimental control combination;

[0280] For each experimental-control combination, calculating the differential methylation rate and p-value of each intersection site between the experimental group and the control group;

[0281] The differential methylation sites of each immunotyping group are determined according to the calculated differential methylation rate and p-value, wherein the differential methylation sites of the immunotyping group are the intersection of multiple first differential methylation sites and multiple second differential methylation sites, and the multiple first differential methylation sites are the intersection sites whose calculated differential methylation rate is greater than the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold when the immunotyping group is the experimental group; the multiple second differential methylation sites are the intersection sites whose calculated differential methylation rate is less than the opposite of the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold when the immunotyping group is the control group.

[0282] In some exemplary embodiments, the omics feature analysis module 1850 is further configured to: determine the nearest gene to the differential methylation site, calculate the correlation coefficient and p-value between the differential methylation site and the nearest gene for each immunophenotyping group, retain the differential methylation site whose p-value is less than a preset second p-value threshold, and determine that the nearest gene of the retained differential methylation site is the correlated gene of the differential methylation site and the differentially expressed gene of the immunophenotyping group.

[0283] In some exemplary embodiments, the omics feature analysis module 1850 determines the differentially expressed genes corresponding to each of the immunophenotyping groups, including:

[0284] For each immunophenotyping group, the gene expression fold change and p-value of each gene between the current immunophenotyping group and the control group corresponding to the current immunophenotyping group are calculated, and genes whose calculated gene expression fold change is greater than a preset second fold threshold and whose p-value is less than a preset first p-value threshold are selected as differentially expressed genes of the current immunophenotyping group.

[0285] In some exemplary embodiments, the number of immunophenotyping groups is greater than or equal to 3, and the omics feature analysis module 1850 determines the differentially expressed genes corresponding to each immunophenotyping group, including:

[0286] Determining every two of the immunophenotyping groups as an experimental control combination, and setting an experimental group and a control group in the experimental control combination;

[0287] For each experimental control combination, the gene expression fold change and p-value of each gene between the experimental group and the control group were calculated;

[0288] The differentially expressed genes of each immunophenotyping group are determined according to the calculated gene expression fold change and the p-value, wherein the differentially expressed genes of the immunophenotyping group are the intersection of multiple first differentially expressed genes and multiple second differentially expressed genes, and the multiple first differentially expressed genes are genes whose calculated gene expression fold change is greater than a preset second fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the experimental group; and the multiple second differentially expressed genes are genes whose calculated gene expression fold change is less than the opposite of the preset second fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the control group.

[0289] In some exemplary embodiments, the omics feature analysis module 1850 is further configured to obtain fifth sample data, the fifth sample data including at least one clinical feature data; calculate a significant p-value between each clinical feature data and the immunophenotyping grouping, and determine the association or difference between the clinical feature and the immunophenotyping based on the calculated significant p-value. When the clinical feature is a numerical variable, the significant p-value can be calculated using an F test or a Kruskal-Wallis test; when the clinical feature is a categorical variable, the significant p-value can be calculated using a chi-square test or a Fisher's exact test.

[0290] In some exemplary embodiments, the omics feature analysis module 1850 is further configured to obtain a sample to be tested, determine one or more omics features of the sample to be tested, and determine the immunophenotyping group of the sample to be tested based on the determined omics features.

[0291] An embodiment of the present disclosure also provides a multi-omics analysis device based on immunophenotyping, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the multi-omics analysis method based on immunophenotyping as described in any embodiment of the present disclosure based on the instructions stored in the memory.

[0292] As shown in Figure 19, in one example, a multi-omics analysis device based on immunophenotyping may include: a processor 1910, a memory 1920, a bus system 1930 and a transceiver 1940, wherein the processor 1910, the memory 1920 and the transceiver 1940 are connected via the bus system 1930, the memory 1920 is used to store instructions, and the processor 1910 is used to execute the instructions stored in the memory 1920 to control the transceiver 1940 to send and receive signals. Specifically, the transceiver 1940 can obtain first sample data under the control of the processor 1910, and the first sample data includes gene expression data. The processor 1910 performs immune infiltration analysis on the first sample data to obtain the proportion of multiple cell types; based on at least one of the following: survival significance, variance and correlation, N cells are screened out from multiple cell types, N>1; the N cells are divided into M fixed cells and (NM) candidate cells, and a cluster analysis is performed on multiple cell combinations, each cell combination includes at least M fixed cells and q candidate cells, M≥1 and M<N, 0≤q≤(NM); one of the cell combinations is selected as the cell combination for immunotyping, and the characteristic cell type corresponding to each of the immunotyping groups is determined; the differentially expressed genes corresponding to each of the immunotyping groups are determined, and at least one of the following omics features corresponding to each of the immunotyping groups is determined: differentially mutated genes, differential copy number variations, and differentially methylated sites.

[0293] It should be understood that the processor 1910 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0294] The memory 1920 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1910. A portion of the memory 1920 may also include a non-volatile random access memory. For example, the memory 1920 may also store information on the device type.

[0295] In addition to the data bus, the bus system 1930 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 1930 in FIG.

[0296] During the implementation process, the processing performed by the processing device can be completed by the hardware integrated logic circuit in the processor 1910 or by instructions in the form of software. That is, the method steps of the embodiment of the present disclosure can be embodied as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1920, and the processor 1910 reads the information in the memory 1920 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0297] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-omic analysis method based on immunophenotyping as described in any embodiment of the present disclosure. The multi-omic analysis method based on immunophenotyping driven by executing the executable instructions is substantially the same as the multi-omic analysis method based on immunophenotyping provided in the above-mentioned embodiments of the present disclosure, and is not further described here.

[0298] In some possible embodiments, various aspects of the immunophenotyping-based multi-omics analysis method provided by the present disclosure can also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the immunophenotyping-based multi-omics analysis method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the computer device can execute the immunophenotyping-based multi-omics analysis method recorded in the embodiments of the present disclosure.

[0299] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0300] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium. It should be noted that the above-described embodiments or implementations are merely exemplary and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementation without departing from the scope of the present disclosure.

Claims

1. A multi-omics analysis method based on immunophenotyping, comprising: performing immune infiltration analysis on first sample data to obtain proportions of multiple cell types, wherein the first sample data includes gene expression data; Select N cell types from multiple cell types based on at least one of the following: survival significance, variance, and correlation, where N>1; Dividing the N cells into M fixed cells and (NM) candidate cells, and performing cluster analysis on multiple cell combinations, each cell combination includes at least M fixed cells and q candidate cells, M≥1 and M<N, 0≤q≤(NM); selecting one of the cell combinations as a cell combination for immunophenotyping, and determining a characteristic cell type corresponding to each immunophenotyping group; Determine the differentially expressed genes corresponding to each of the immunophenotyping groups, and determine at least one of the following omics features corresponding to each of the immunophenotyping groups: differentially mutated genes, differential copy number variations, and differentially methylated sites.

2. The method according to claim 1, further comprising: Screening the differentially expressed genes corresponding to each of the immunophenotyping groups; Constructing a prognostic risk prediction model using the screened differentially expressed genes, and grouping the first sample data into prognostic risk groups according to the prognostic risk score of the input first sample data obtained by the prognostic risk prediction model; The differentially expressed genes that are significantly positively correlated with the prognostic risk grouping are determined to be important genes affecting the disease.

3. The method according to claim 1, wherein Determining the differentially mutated genes corresponding to each of the immunophenotyping groups includes: Acquire second sample data, where the second sample data includes gene mutation data; Encoding the gene mutation data to obtain a gene mutation spectrum matrix; For each of the immunotyping groups, the mutation rate fold change and p-value of each gene between the current immunotyping group and the control group corresponding to the current immunotyping group are calculated according to the gene mutation spectrum matrix, and genes with a calculated mutation rate fold change greater than a preset first fold threshold and a p-value less than a preset first p-value threshold are selected as the differentially mutated genes of the current immunotyping group.

4. The method according to claim 3, wherein: The number of the immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

5. The method according to claim 1, wherein The number of the immunophenotyping groups is greater than or equal to 3, and determining the differentially mutated genes corresponding to each of the immunophenotyping groups comprises: Acquire second sample data, where the second sample data includes gene mutation data; Encoding the gene mutation data to obtain a gene mutation spectrum matrix; Determining each of two immunophenotyping groups as an experimental control combination, and setting an experimental group and a control group in the experimental control combination; For each experimental-control combination, the mutation rate fold change and p-value of each gene between the experimental group and the control group are calculated according to the gene mutation spectrum matrix; The differential mutation genes of each immunophenotyping group are determined according to the calculated mutation rate fold change and p-value, and the differential mutation genes of each immunophenotyping group are the intersection of multiple first differential mutation genes and multiple second differential mutation genes. The multiple first differential mutation genes are genes whose calculated mutation rate fold change is greater than a preset first fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the experimental group. The multiple second differential mutation genes are genes whose calculated mutation rate fold change is less than the opposite of the preset first fold threshold and whose p-value is less than the preset first p-value threshold when the immunophenotyping group is the control group.

6. The method according to claim 1, wherein Determining the differential copy number variation corresponding to each of the immunophenotyping groups includes: Acquiring third sample data, wherein the third sample data includes genome copy number data; For each immunophenotyping group, perform the following operations: determining the genomic interval copy numbers on a plurality of chromosomes based on the genomic copy number data; For each of the multiple chromosomes, when the genomic interval copy number is greater than a preset copy number threshold, the immunotyping on the chromosome is determined to include an increase in the genomic copy number and the proportion of the number of intervals with increased genomic copy numbers to the total number of intervals is counted; when the genomic interval copy number is less than the opposite of the preset copy number threshold, the immunotyping on the chromosome is determined to include a loss of genomic copy number and the proportion of the number of intervals with lost genomic copy numbers to the total number of intervals is counted.

7. The method according to claim 1, wherein Determining the differentially methylated sites corresponding to each of the immunophenotyping groups includes: Acquiring fourth sample data, the fourth sample data including methylation data, the methylation data including methylation rates of multiple sites; Filtering the methylation data to determine intersection sites in the fourth sample data; For each of the immunotyping groups, the differential methylation rate and p-value of each intersection site between the current immunotyping group and the control group corresponding to the current immunotyping group are calculated, and the intersection site whose calculated differential methylation rate is greater than the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold is selected as the differential methylation site of the current immunotyping group.

8. The method according to claim 7, wherein: The number of the immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

9. The method according to claim 1, wherein The number of the immunophenotyping groups is greater than or equal to 3, and determining the differentially methylated sites corresponding to each of the immunophenotyping groups comprises: Acquiring fourth sample data, the fourth sample data including methylation data, the methylation data including methylation rates of multiple sites; Filtering the methylation data to determine intersection sites in the fourth sample data; Determining every two of the immunophenotyping groups as an experimental control combination, and setting an experimental group and a control group in the experimental control combination; For each experimental-control combination, calculating the differential methylation rate and p-value of each intersection site between the experimental group and the control group; The differential methylation sites of each immunotyping group are determined according to the calculated differential methylation rate and p-value, wherein the differential methylation sites of the immunotyping group are the intersection of multiple first differential methylation sites and multiple second differential methylation sites, and the multiple first differential methylation sites are the intersection sites whose calculated differential methylation rate is greater than the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold when the immunotyping group is the experimental group; the multiple second differential methylation sites are the intersection sites whose calculated differential methylation rate is less than the opposite of the preset first differential methylation rate threshold and the p-value is less than the preset first p-value threshold when the immunotyping group is the control group.

10. The method according to any one of claims 7 to 9, wherein: The method further comprises: Determine the nearest gene to the differential methylation site, calculate the correlation coefficient and p-value between the differential methylation site and the nearest gene for each immunophenotyping group, retain the differential methylation site whose p-value is less than a preset second p-value threshold, and determine the nearest gene of the retained differential methylation site as the correlated gene of the differential methylation site and the differentially expressed gene of the immunophenotyping group.

11. The method according to claim 1, wherein Determining the differentially expressed genes corresponding to each of the immunophenotyping groups comprises: For each immunophenotyping group, the gene expression fold change and p-value of each gene between the current immunophenotyping group and the control group corresponding to the current immunophenotyping group are calculated, and genes whose calculated gene expression fold change is greater than a preset second fold threshold and whose p-value is less than a preset first p-value threshold are selected as differentially expressed genes of the current immunophenotyping group.

12. The method according to claim 11, wherein The number of the immunotyping groups is greater than or equal to 2, and the control group corresponding to the current immunotyping group is all remaining groups except the current immunotyping group.

13. The method according to claim 1, wherein The number of the immunophenotyping groups is greater than or equal to 3, and determining the differentially expressed genes corresponding to each of the immunophenotyping groups comprises: Determining every two of the immunophenotyping groups as an experimental control combination, and setting an experimental group and a control group in the experimental control combination; For each experimental control combination, the gene expression fold change and p-value of each gene between the experimental group and the control group were calculated; The differentially expressed genes of each immunophenotyping group are determined according to the calculated gene expression fold change and the p-value, wherein the differentially expressed genes of the immunophenotyping group are the intersection of multiple first differentially expressed genes and multiple second differentially expressed genes, and the multiple first differentially expressed genes are genes whose calculated gene expression fold change is greater than a preset second fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the experimental group; and the multiple second differentially expressed genes are genes whose calculated gene expression fold change is less than the opposite of the preset second fold threshold and whose p-value is less than a preset first p-value threshold when the immunophenotyping group is the control group.

14. The method according to claim 1, further comprising: Acquiring fifth sample data, wherein the fifth sample data includes at least one clinical characteristic data; Calculate the p-value between each clinical characteristic data and the immunophenotyping group, and determine the association or difference between the clinical characteristics and the immunophenotyping according to the calculated p-value.

15. The method according to claim 1, further comprising: Obtain a sample to be tested, determine one or more omics features of the sample to be tested, and determine the immunophenotyping group of the sample to be tested based on the determined omics features.

16. An immunophenotyping-based omics analysis device comprising a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the immunophenotyping-based omics analysis method according to any one of claims 1 to 15 based on the instructions stored in the memory. 17 . A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the computer-readable storage medium implements the immunophenotyping-based omics analysis method according to claim 1 . 18 . A computer program product comprising instructions, wherein when the computer program product is executed by a computer, the instructions execute the immunophenotyping-based omics analysis method according to claim 1 .

19. A multi-omics analysis device based on immunophenotyping, comprising an immune infiltration module, a screening module, a cluster analysis module, a selection module, and an omics feature analysis module, wherein: The immune infiltration module is configured to perform immune infiltration analysis on first sample data to obtain proportions of multiple cell types, wherein the first sample data includes gene expression data; The screening module is configured to screen N cells from a plurality of cell types according to at least one of the following: survival significance, variance, and correlation, where N>1; The cluster analysis module is configured to divide the N types of cells into M types of fixed cells and (NM) types of candidate cells, and perform cluster analysis on multiple cell combinations, each cell combination including at least M types of fixed cells and q types of candidate cells, where M ≥ 1 and M < N, and 0 ≤ q ≤ (NM); The selection module is configured to select one of the cell combinations as the cell combination for immunophenotyping, and determine the characteristic cell type corresponding to each immunophenotyping group; The omics feature analysis module is configured to determine the differentially expressed genes corresponding to each of the immunophenotyping groups, and to determine at least one of the following omics features corresponding to each of the immunophenotyping groups: Differentially mutated genes, differential copy number mutations, differentially methylated sites.

Citation Information

Patent Citations

  • Method for screening potential prognosis biomarker of lung adenocarcinoma based on tumor infiltrating immune cells

    CN113140258A

  • Immune typing method and device, storage medium and program product

    CN118447927A

  • Interfering with telomere maintenance in treatment of diseases

    US20050059077A1