Screening method of cancer prognosis marker
Patent Information
- Application Number
- CN202510896827.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-30
AI Technical Summary
然而,空间特征(尤其是巨噬细胞的活动)如何影响免疫细胞功能,仍未得到充分研究
本发明通过成像质谱流式细胞技术,提出了一种基于肿瘤微环境中细胞和空间组成筛选癌症预后标志物的方法。本发明的方法通过对胃癌肿瘤样本及预后状况进行分析,可得到空间属性对免疫细胞,尤其是巨噬细胞活动的影响。通过本发明的筛选方法得到了CD11chi巨噬细胞与I型胶原蛋白+成纤维细胞的空间互作是胃癌患者生存不良的主要决定因素。基于上述方法和结论,基于本发明筛选得到的标志物的预测模型,相比于依赖高维多组学数据的模型表现出更优越的预后准确性。
Smart Images

Figure CN120404541A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and particularly relates to a method for screening cancer prognosis markers. Background Art
[0002] Gastric cancer (GC) is a global health problem. In 2022, its incidence ranked fifth among malignant tumors globally, and the fatality rate ranked fifth. Currently, the main treatment methods for gastric cancer include surgery, chemotherapy, and radiotherapy, which have limited effects on advanced patients. Despite the continuous progress of clinical and basic research, the five-year survival rate of advanced gastric cancer patients is still less than 40%.
[0003] Given the poor prognosis of advanced gastric cancer patients, the tumor microenvironment (TME) has become a key factor affecting the treatment effect and the effectiveness of treatment regimens. The TME is a complex network composed of cellular and acellular components, which plays a decisive role in cancer dynamics and treatment effects. Macrophages play a key role therein, significantly affecting tumor behavior and treatment response. In recent years, studies have revealed the complex relationship between macrophages and gastric cancer, and emphasized new treatment targets and mechanisms driving tumor progression.
[0004] The formation and function of tertiary lymphoid structures (TLS) in the tumor microenvironment have also attracted attention. These structures are similar to secondary lymphoid organs but are located outside traditional lymphoid organs, and help activate naive T / B cells so that they can play a role in tumor immunity. In the research of various malignant tumors such as colorectal cancer, lung cancer, breast cancer, and malignant melanoma, it has been shown that the presence of TLS is associated with a good prognosis.
[0005] The immune microenvironment (especially TLS and specialized immune cell subsets) has a profound impact on gastric cancer prognosis and treatment outcomes. This understanding lays the foundation for the development of more personalized and efficient immunotherapies, and is expected to revolutionize the management strategies of malignant tumors. However, how spatial features (especially macrophage activities) affect immune cell functions has not been fully studied. To comprehensively analyze the complex relationship between the cellular and spatial components in the tumor microenvironment (TME), advanced analytical methods with spatial resolution capabilities are needed to effectively reveal spatial heterogeneity. By further combining histology and mass cytometry through imaging mass cytometry (IMC), 45 metal-labeled antibodies can be detected simultaneously on a single section, accurately presenting cell interactions and marker co-expression. Summary of the Invention
[0006] In order to overcome the problem in the prior art of lacking a method for screening markers based on the cellular and spatial composition in the tumor microenvironment (TME), the present invention proposes a method for screening cancer prognosis markers, and the above object is achieved through the following implementation manners of technical solutions: A method for screening cancer prognosis markers, comprising the following steps: Step 1) Obtain an analysis dataset of imaging mass cytometry, where the analysis dataset of imaging mass cytometry contains the expression intensity of markers to be screened in a sample and the corresponding prognosis of the sample; and divide different comparison groups according to the sample situation and prognosis; Step 2) Perform single-cell segmentation on the images in the analysis dataset of imaging mass cytometry, and annotate the cell types of single cells according to the expression intensity of the markers to be screened, to obtain the cell frequencies of different cell types; Step 3) Identify the regions where epithelial, immune, and fibrotic cells meet, and obtain the expression intensity of markers in the regions where epithelial, immune, and fibrotic cells meet; Step 4) Based on the cell frequencies and marker expression intensities of different cell types in the regions where epithelial, immune, and fibrotic cells meet, calculate the interaction analysis results between different cell types; Step 5) Perform regression analysis according to the cell frequencies, marker expression intensities, and cell interaction results in different comparison groups, and screen out prognostic markers.
[0007] Optionally, the cancer is gastric cancer.
[0008] Optionally, the specific process of dividing different comparison groups according to the sample situation and prognosis in Step 1) includes: Dividing into a tumor group and a peritumoral tissue group according to whether the sample belongs to a tumor or peritumoral tissue; Dividing into an early pathological stage group and a late pathological stage group according to whether the sample belongs to an early pathological stage or a late pathological stage; Dividing into an early clinical stage group and a late clinical stage group according to whether the sample belongs to an early clinical stage or a late clinical stage; Dividing into a long survival group and a short survival group according to the patient's prognosis.
[0009] Optionally, the markers to be screened in Step 1) include CD45, CD15, CD45RA, CD103, CD68, HLA-DR, CD45RO, B7H4, CD14, CD20, CD57, PD-L1, CD25, CD3, CD4, αSMA, CD8, CD7, FOXP3, KI67, CD163, CD16, CD69, vimentin, CD11b, CD11c, Caspase3, IL-1b, Ki-67, TNFα, type I collagen, LAG3, GATA3, PD-1, VISTA, CD31, granzyme B, IL-6, trypsin, E-Cadherin.
[0010] Optionally, the cell types include: Myofibroblasts, type I collagen + Fibroblasts, stromal cells, endothelial cells, epithelial cells, lymphocytes, myeloid cells; The epithelial cells include the following types: PD-L1 + Subsets, Ki-67 + Proliferative subsets, E-cadherin + Subsets, E-cadherin - Subsets; The lymphocytes include the following types: regulatory T cells, effector CD8 + T cells, memory CD4 + T cells, memory CD8 + T cells, effector CD4 + T cells, CD16 - Natural killer cells, CD16 + Natural killer cells, Lag3 + Natural killer cells, B cells; The myeloid cells include the following types: granulocytes, monocytes, Ki67 + Proliferating macrophages, CD11c low HLA-DR low Macrophages, CD11c hi Macrophages, HLA-DR hi Macrophages.
[0011] Optionally, the annotation of the cell types of the single cells in step ii) includes the following steps: The first round of clustering, identifying the main cell populations through the following markers: E-Cadherin, trypsin, CD3, CD4, CD8, CD20, CD7, CD57, CD45, CD68, CD15, CD14, CD16, CD31, αSMA, type I collagen, and vimentin; The second round of clustering, clustering myeloid cells using FOXP3, CD16, CD69, CD4, CD8, Caspase3, B7H4, VISTA, CD7, CD103, LAG3, CD20, granzyme B, PD-1, KI67, GATA3, CD45RA, CD3, TNFα, TL1β, CD45RO, CD57, CD25; Clustering lymphocytes using IL-6, CD14, CD16, Caspase3, CD163, PD-L1, CD11B, CD11C, CD15, Ki-67, HLA-DR, TNFα, and TL1β; Cluster epithelial cells by E-Cadherin, PDL1, Ki-67, and trypsin.
[0012] Optionally, the step of identifying the regions where epithelial, immune, and fibrotic cells meet in step (iii) includes the following steps: Based on spatial coordinates, use the k-nearest neighbor algorithm to calculate the 20 nearest neighbors of each cell, thereby constructing a local spatial adjacency graph for each cell; for the neighborhood of each cell, summarize the expression data of various markers of neighboring cells and aggregate according to cell type; The local subgraphs are screened to ensure that the central node of each subgraph corresponds to a specific cell type; the subgraphs are then connected through shared nodes to form a global connection graph, thereby generating spatial regions / patches; the spatial range of the constructed patches is expanded by adding n neighboring nodes to the existing patches, thereby expanding the spatial coverage of the constructed patches; Based on the spatial connection graph, obtain the immune region, fibroblast region, and epithelial region; the immune region includes lymphocytes and myeloid cells; the fibroblast region includes myofibroblasts and type I collagen + Fibroblasts; the epithelial region includes epithelial cells; Based on the immune region, fibroblast region, and epithelial region, the regions where epithelial, immune, and fibrotic cells meet can be obtained, and the regions where epithelial, immune, and fibrotic cells meet include: Immune-fibroblast junction; Immune-epithelial junction; Fibroblast-epithelial junction; In step (iv), by calculating the average number of surrounding cells in each region of interest, the number of interactions between the central cell and neighboring cells is statistically counted.
[0013] Optionally, in step (v), regression analysis is performed according to the cell frequencies, marker expression intensities, and results of cell-cell interactions in different comparison groups, specifically including the following steps: Variables with a zero value proportion lower than 50% are screened out, and cross-validation is performed by LASSO regression to identify non-zero coefficients; the screened variables are incorporated into a multivariable Cox regression model, and the proportional hazards assumption is evaluated by the cox.zph function; variables that meet the PH assumption are retained, and the association between the variables and survival is evaluated; the results are visualized by the ggforest function.
[0014] A method for establishing a cancer prognosis diagnosis model includes the following steps: Screen for cancer prognosis markers using the above method; The average expression intensity and prognosis of each biomarker in the cancer prognosis biomarker combination screened from the dataset analyzed by imaging mass cytometry are used to establish a cancer prognosis diagnosis model by machine learning methods.
[0015] A screening system for cancer prognosis biomarkers, comprising: A data acquisition module: used to acquire the dataset analyzed by imaging mass cytometry, where the dataset analyzed by imaging mass cytometry contains the expression intensity of the biomarkers to be screened in the sample and the corresponding prognosis of the sample; and different comparison groups are divided according to the sample situation and prognosis. A cell type annotation module: used to perform single-cell segmentation on the images in the dataset analyzed by imaging mass cytometry, and annotate the cell types of single cells according to the expression intensity of the biomarkers to be screened, so as to obtain the cell subset frequencies of different cell types. A regional expression intensity analysis module: used to identify the regions where epithelial, immune, and fibrotic cells intersect, and obtain the biomarker expression intensity in the regions where epithelial, immune, and fibrotic cells intersect. An interaction analysis module, used to calculate the interaction analysis results between different cell types based on the cell frequencies and biomarker expression intensities of different cell types in the regions where epithelial, immune, and fibrotic cells intersect. A regression analysis module, which performs regression analysis according to the cell frequencies, biomarker expression intensities, and cell interaction results in different comparison groups to screen out prognosis biomarkers.
[0016] The present invention has the following beneficial effects: Through imaging mass cytometry, the present invention proposes a method for screening cancer prognosis biomarkers based on the cell and spatial composition in the tumor microenvironment. The method of the present invention can obtain the influence of spatial attributes on immune cells, especially macrophage activity, by analyzing gastric cancer tumor samples and prognosis. Through the screening method of the present invention, CD11c hi Macrophages and type I collagen + The spatial interaction between fibroblasts is the main determinant of poor survival in gastric cancer patients. Based on the above methods and conclusions, the prediction model based on the biomarkers screened by the present invention shows more superior prognosis accuracy compared with the model relying on high-dimensional multi-omics data. Description of the Drawings
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a graph showing the analysis results of the cell proportions of different comparison groups.
[0019] Figure 2 It is a box plot showing the distribution of the proportions (percentage of total cells) of selected cell types in the clinical group among different groups.
[0020] Figure 3 It is the difference in the expression of functional markers within specific cell types in different tissue regions.
[0021] Figure 4 It is the average frequency of cell-cell interactions within the fibroblast-immune interface region.
[0022] Figure 5 It is the result of summarizing the output of a multivariable survival model.
[0023] Figure 6 It is a graph showing the results of survival analysis during the model validation process.
[0024] Figure 7 It is a comparison graph of the evaluation results between the model in the embodiment and the existing prognosis evaluation models. Specific Embodiments
[0025] Now, various exemplary embodiments of the present invention will be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, characteristics, and implementation schemes of the present invention. It should be understood that the terms described in the present invention are only for describing specific embodiments and are not used to limit the present invention.
[0026] In addition, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded from the range.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although only preferred methods and materials are described herein, any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of this invention.
[0028] As used herein, terms such as "comprising," "including," "having," "containing," and the like are open-ended terms, meaning including but not limited to.
[0029] The present invention discloses a method for screening cancer prognosis markers, comprising the following steps: Step 1): Obtain an imaging mass cytometry analysis dataset, which contains the expression intensity of the markers to be screened in the sample and the corresponding prognosis of the sample; and divide different comparison groups according to the sample conditions and prognosis; Step 2): Perform single-cell segmentation on the images in the imaging mass cytometry analysis dataset, and annotate the cell types of single cells according to the expression intensity of the markers to be screened, to obtain the cell subset frequencies of different cell types; Step 3): Identify the regions where epithelial, immune, and fibrotic cells intersect, and obtain the marker expression intensities in the regions where epithelial, immune, and fibrotic cells intersect; Step 4): Based on the cell frequencies and marker expression intensities, obtain the results of cell interaction analysis; Step 5): Perform regression analysis modeling according to the cell frequencies, marker expression intensities, and cell interaction results in different comparison groups, and screen out the prognosis markers.
[0030] Taking the screening of gastric cancer markers as an example, the specific operation steps are as follows: In step 1) above, obtain an imaging mass cytometry analysis dataset, which contains the expression intensity of the markers to be screened in the sample and the corresponding prognosis of the sample; and divide different comparison groups according to the sample conditions and prognosis, specifically including: To identify regional factors influencing the prognosis and clinical characteristics of gastric cancer (GC) patients, a detailed spatial multi-omics analysis was performed on 205 GC patients. Imaging mass cytometry (IMC) analysis was conducted on 205 tumors and 39 regions of interest (ROIs) around the tumors from 205 GC patients. This IMC-based single-cell spatial proteomics technology can identify different regions in the gastric cancer tumor microenvironment (TME). In addition, the technology also helped to reveal other spatial features, including cell-cell interactions, such as tertiary lymphoid structures (TLS). By integrating these spatial features with clinical data, such as prognosis, the most influential spatial determinants can be obtained.
[0031] To explore the complex spatial heterogeneity in the gastric cancer (GC) microenvironment while preserving the structural integrity of the GC ecosystem, we designed an imaging mass cytometry (IMC) panel containing 40 markers (Table S2). The panel was specifically tailored for GC patients and optimized from established protocols. The markers we selected cover a wide range of cell types, including epithelial cells, endothelial cells, and stromal cells, as well as various immune cells. In addition, key cytokines (such as IL-1β, TNF-α, and IL-6), the proliferation marker Ki-67, and the apoptosis indicator cleaved caspase-3 are also included. Specifically, they are: CD45, CD15, CD45RA, CD103, CD68, HLA-DR, CD45RO, B7H4, CD14, CD20, CD57, PD-L1, CD25, CD3, CD4, αSMA, CD8, CD7, FOXP3, KI67, CD163, CD16, CD69, vimentin, CD11b, CD11c, Caspase3, IL-1b, Ki-67, TNFα, type I collagen, LAG3, GATA3, PD-1, VISTA, CD31, granzyme B, IL-6, trypsin, E-Cadherin.
[0032] Before IMC use, each antibody was rigorously validated by immunohistochemistry (IHC) to ensure specificity and then conjugated with a unique heavy metal. After conjugation, each metal-labeled antibody was re-validated for its performance by IMC.
[0033] After staining and IMC scanning, the acquired images undergo a series of preprocessing steps to improve image quality. These steps include compensation, noise reduction, and contrast enhancement. Noise reduction is achieved through median filtering to effectively suppress background noise; contrast enhancement is applied using a linear regression model. To accurately segment and analyze individual cells or components in different channels of IMC images, we adopted a deep learning-based segmentation method. This method enables precise division and quantification of the expression levels of each marker, and subsequently, the data is quantified into a matrix format for further analysis.
[0034] Imaging Mass Cytometry (IMC) single-cell segmentation: To perform single-cell segmentation on IMC images, a pre-trained DeepCell model was adopted. The DeepCell model requires two imaging channels for segmentation: the first channel is the nuclear staining channel (usually DAPI staining) to identify the location of cell nuclei; the second channel is the cell membrane or cytoplasmic region to define the cell boundaries. In this study, the membrane staining channel was selected as the second imaging channel for segmentation analysis.
[0035] We established four comparison groups based on key clinical and pathological parameters, including tumor vs. peritumoral tissue (Group 1), early vs. late pathological stage (Stage I vs. Stage II, Group 2), early vs. late clinical stage (Stage I vs. Stage II, Group 3), and patients were divided into long survival group and short survival group according to overall survival (OS) (Group 4). Patients with an overall survival of more than five years were classified into the long survival group, while patients with an overall survival of less than one year were classified into the short survival group (Group 4). The comparison results of cell frequencies between these groups are as Figure 1 shown.
[0036] Performing single-cell segmentation on the images in the imaging mass cytometry analysis dataset in step (ii) above, and annotating the cell types of single cells according to the expression intensity of the to-be-screened markers to obtain the cell subset frequencies of different cell types; specifically including: IMC preprocessing and cell type identification: First, the marker expression was transformed using the hyperbolic arcsine function, and the signal values of each channel were restricted within the range of 1% and 99% to set the minimum and maximum values. Subsequently, Min-Max Normalization was applied to each channel separately to standardize the data. To address possible batch effects, the R package Harmony (version 1.2.3) was used for alignment. Subsequently, cell clustering was completed by FastPG (version 0.0.8), setting the number of nearest neighbors to 100 and performing two rounds of clustering. In the first round of clustering, the main cell populations were identified by the following lineage-specific markers: E-Cadherin, trypsin, CD3, CD4, CD8, CD20, CD7, CD57, CD45, CD68, CD15, CD14, CD16, CD31, αSMA, type I collagen, vimentin. To distinguish which of myofibroblasts, type I collagen + fibroblasts, stromal cells, endothelial cells, epithelial cells, lymphocytes, and myeloid cells it belongs to. In the second round of clustering, the myeloid cell subsets were reclustered using additional markers FOXP3, CD16, CD69, CD4, CD8, Caspase3, B7H4, VISTA, CD7, CD103, LAG3, CD20, granzyme B, PD-1, KI67, GATA3, CD45RA, CD3, TNFα, TL1β, CD45RO, CD57, CD25. Lymphocytes were reclustered using markers such as IL-6, CD14, CD16, Caspase3, CD163, PD-L1, CD11B, CD11C, CD15, Ki-67, HLA-DR, TNFα, and TL1β. Epithelial cells were analyzed by clustering using markers such as E-Cadherin, PDL1, KI67, and trypsin. Finally, the average expression levels of the markers in each cluster were visualized by a heatmap, and this information was used to annotate the corresponding cell types (CT). Here, the expression situation refers to whether there is a relatively higher expression level among all cell types being clustered: divided into yin and yang according to expression (+) and no expression (-), where expression (+) can be further subdivided into high expression (hi) and low expression (low). For example, in a heatmap with red~blue corresponding to high~low, blue represents no expression (-), red represents expression (+), high expression (hi) is bright red, and low expression (low) is orange or light red.
[0037] After batch correction, we successfully analyzed 1,562,075 single cells and classified them into seven cell clusters. These cell clusters were mainly annotated by classical cell type markers. This analysis identified including E-cadherin+ Epithelial cells, CD3 + or CD20 + Lymphocytes, CD68 + or CD15 + Myeloid cells, CD31 + Endothelial cells, α-SMA + Myofibroblasts, type I collagen + Fibroblasts and vimentin + Multiple cell populations including stromal cells. Each cluster was then further subdivided based on specific functional markers, providing a map of the cellular heterogeneity within the dataset. For example, in our analysis, epithelial cells were classified as PD-L1 + , Ki-67 + proliferative subsets and E-cadherin positive or negative subsets. Lymphocytes were classified into several functionally distinct types: Foxp3 + regulatory T cells, granzyme B + effector CD8 + T cells, CD45RO + memory CD4 + and CD8 + T cells, effector CD4 + T cells, and CD16 - , CD16 + , Lag3 + CD57 + natural killer cells (NK cells) and CD20 + B cells. In addition, myeloid cells were subdivided into CD15 + granulocytes, CD14 + CD16 + monocytes, Ki67 + proliferating macrophages, CD11c low HLA-DR low macrophages, CD11c hi macrophages and HLA-DR hi macrophages.
[0038] The distribution of cell subset frequencies under different clinical and pathological conditions is shown in Figure 2 . The analysis results showed that for gastric cancer (GC), there were significant differences in the cell population frequencies between tumor tissue and peritumoral tissue ( Figure 1 and 2 ). In the tumor area, multiple macrophage subsets as well as CD4 + and CD8 +The proportion of T cell subsets increased significantly, with the only exception being regulatory T cells (Tregs). Contrary to expectations, the abundance of Tregs was lower at the tumor site, a finding that differs from the common immunosuppressive characteristics in the tumor microenvironment (TME) as Tregs are generally considered key immunosuppressive cellular components ( Figure 1 and Figure 2 ). In addition, type I collagen + -producing fibroblasts were more common in the tumor area, which is consistent with the generally enhanced fibrotic properties at the tumor site ( Figure 1 and 2 ). In contrast, natural killer (NK) cell subsets were significantly reduced at the tumor site, indicating a weakened innate anti-tumor immune response ( Figure 1 and Figure 2 ). Similarly noteworthy is that the number of epithelial cell subsets was greater in the tumor area, which is consistent with the nature that most tumor cells in these areas are mainly derived from the epithelial lineage ( Figure 1 and Figure 2 ).
[0039] A comparative analysis was performed on gastric cancer (GC) at different clinical stages. Notably, higher levels of CD4 + and CD8 + memory T cells were observed at the advanced stage of the disease, indicating an increase in the memory T cell phenotype with the progression of gastric cancer. More importantly, regarding survival outcomes, I found that the proportion of type I collagen + -producing fibroblasts was higher in the short-survival group, suggesting its adverse effect on the prognosis of gastric cancer. In contrast, stromal cells were less common in short-surviving patients, indicating its possible protective role in disease progression ( Figure 1 and Figure 2 ).
[0040] The frequency distribution of cell subsets in the microenvironment of gastric cancer (GC) was comprehensively studied by imaging mass cytometry (IMC), providing important insights into their dynamic interactions in the tumor environment. We identified multiple key subsets, including fibroblasts, multiple macrophage subsets, and T cell subsets, which may significantly affect the staging and prognosis of gastric cancer patients.
[0041] After studying cell frequencies under different clinical conditions, we further utilized spatial proteomics data to analyze the key spatial features provided by imaging mass cytometry (IMC) at single-cell spatial resolution. Our initial study focused on tertiary lymphoid structures (TLS). TLS consists of dense aggregations of B and T cells and is thought to have important implications for treatment outcomes. To test this hypothesis, we employed a modified version of the patch algorithm, which is specifically designed to detect "tissue patch"-like clusters characterized by a significant B cell concentration.
[0042] No statistically significant differences in TLS size were observed between different clinical or survival groups. Therefore, we further delved into the cell frequencies of different cell subsets within TLS to evaluate how changes within TLS affect gastric cancer (GC). As expected, the majority of cells within TLS were lymphocytes, predominantly B cells and CD4 + T cells, which also validated the accuracy of our analysis algorithm. In addition to lymphocytes, we also observed a significant population of stromal cells within TLS, including mesenchymal and endothelial cells. Furthermore, we detected HLA-DR hi macrophages within TLS.
[0043] Subsequently, we performed a comparative analysis of cell frequencies under different conditions. The differences between the tumor site and the peritumoral tissue were consistently significant. In the TLS of tumor tissue, we noted a significant increase in multiple macrophage subsets, including HLA-DR hi and CD11c hi macrophages. This finding highlights the key role of macrophages in TLS dynamics. Unexpectedly, we also detected granulocytes within TLS. This may imply a certain degree of immunosuppressive function within TLS. Additionally, in tumor tissue, the epithelial cell subset within TLS increased. This phenomenon may reflect a higher concentration of epithelial lineage cells at the tumor site rather than a specific adaptation feature of the tumor microenvironment. Furthermore, we observed that the presence of HLA-DR hi macrophages was significantly increased in the tertiary lymphoid structures (TLS) of gastric cancer (GC) tumor tissue, corresponding to a higher pathological grade. Similarly, granulocytes and CD11c hi macrophages showed higher levels in more advanced clinical stages, indicating that immunosuppressive cell activity may intensify as gastric cancer progresses. However, when analyzing cell frequencies, we did not observe statistically significant differences between the long-survival and short-survival groups. This suggests that cell composition alone may not be sufficient to distinguish long- and short-term survival outcomes in gastric cancer patients. This finding emphasizes the complexity of the tumor microenvironment and implies that other factors may influence the prognosis of gastric cancer.
[0044] Therefore, we extended our analysis to the expression of functional markers within TLS in different comparison groups. Although no significant differences were observed between the short-survival and long-survival groups in terms of the frequencies of cell subsets within TLS, we found some interesting changes in functional markers. In the short-survival group, the expression of Ki-67 in CD11c hi macrophages was significantly upregulated, which supported its role in immunosuppression in TLS. This finding was consistent with our previous observation that the frequency of CD11c hi macrophages in TLS increased in advanced clinical stages. In contrast, in the TLS of the long-survival group, effector CD4 + T cells showed higher levels of expression of residency markers, including CD69 and CD103. This indicated that the residency status of CD4 + T cells was associated with a better prognosis.
[0045] When evaluating other comparison groups, we observed elevated levels of immune checkpoints in TLS of tumor tissues, including PD-1, PD-L1, and LAG3. In addition, the levels of pro-inflammatory cytokines (such as TNF-α, IL-6, and IL-1β) within these TLS also increased. This phenomenon was partially reproduced in the pathological stage of advanced gastric cancer, where the level of TNF-α was also significantly elevated. Similarly, in advanced clinical stages, the levels of LAG3 and IL-6 also increased. In contrast, in the early clinical stages of gastric cancer, functional markers related to cytotoxic activity (such as granzyme B and GATA3) were more prominent.
[0046] In further studies, we focused on elucidating the specific patterns of cell interactions within TLS. Given that B cells are the major cell subset in these structures, significant interactions involving B cells and other cell subsets were expected to be observed. Notably, in the short-survival group, the interaction between type I collagen + fibroblasts and E-cadherin⁻ epithelial cells increased significantly, indicating that these non-immune cells play a role in mediating immunosuppression within TLS. In addition, in this group, the interaction frequency between B cells and CD11c low HLA-DR low macrophages was higher, further indicating its contribution to the immunosuppressive environment. In contrast, in the long-survival group, the interactions between HLA-DR hi macrophages and monocytes, as well as between effector CD8 + T cells and effector CD4 + T cells increased significantly. These interactions suggest that there may be a more robust and coordinated immune response between effector CD4 + and CD8 + T cells in TLS.
[0047] In summary, the dynamic changes in TLS composition, particularly the frequency of lymphocytes and their functional markers, as well as cell-cell interactions within TLS, highlight the complex interactions between tumor microenvironments.
[0048] To understand the broader spatial features beyond TLS, the tumor microenvironment (TME) was explored, which includes different regions such as epithelium, immune, and fibrosis. Previous studies have shown that these regions exhibit different interactions and responses during treatment. Therefore, the analysis scope was extended to these specific regions, which are characterized by continuous cell extensions. To improve the analysis accuracy, the k-nearest neighbor (KNN) algorithm was adopted, which aims to identify homogeneous clusters of cell types. This algorithm helps to delineate different regions in tissue samples. Subsequently, the overlapping regions were precisely mapped, especially the regions where epithelial, immune, and fibrotic cells meet, such as the epithelial-immune, epithelial-fibrosis, and fibroblast-immune junctions. This approach enables a deeper understanding of cell interactions in these complex microenvironments.
[0049] Step (iii) above: identifying the regions where epithelial, immune, and fibrotic cells meet, and obtaining the expression intensity of markers in the regions where epithelial, immune, and fibrotic cells meet; specifically including: Identification of spatial regions: First, based on spatial coordinates, the 20 nearest neighbors of each cell were calculated using the k-nearest neighbor (KNN) algorithm to construct a local spatial adjacency graph for each cell. Next, for the neighborhood of each cell, various marker expression data of neighboring cells were aggregated and aggregated according to cell type. Subsequently, the local subgraphs were screened to ensure that the central node of each subgraph corresponds to a specific cell type (such as B cells). These subgraphs were then connected through shared nodes to form a global connection graph, thereby generating spatial regions / patches. Finally, the spatial range of the existing patches was expanded by adding n neighboring nodes to expand the spatial coverage of the constructed patches.
[0050] Downstream spatial analysis: Based on the spatial connection graph, three main regions were defined: Immune region: including "lymphocytes" and "myeloid cells"; Fibroblast region: including "myofibroblasts" and "type I collagen + fibroblasts"; Epithelial region: including "epithelial cells".
[0051] Based on the above regions, three overlapping regions were defined: the immune-fibroblast junction, the immune-epithelial junction, and the fibroblast-epithelial junction. By calculating the average number of surrounding cells within each region of interest (ROI), the number of interactions between the central cell and neighboring cells was statistically analyzed.
[0052] Step (iv) Obtain the results of cell interaction analysis based on cell frequency and marker expression intensity; specifically including: Differential analysis was performed on cell frequency, marker expression, and cell-cell interaction within the defined regions. And patients in the short-term survival group and the long-term survival group were comparatively evaluated and ranked based on the frequency of cell subsets. It is worth noting that in the long survival group, effector CD4 + T cells, B cells, and NK cell subsets showed a significant increase in frequency in the immune region or immune junction region. This enhanced infiltration of effector cells was positively correlated with a better prognosis. In contrast, lag3 + NK cells, Ki-67 + macrophages, regulatory T cells (Tregs), and type I collagen + fibroblasts showed a significant increase in frequency, indicating that cell exhaustion and the proliferation of immunosuppressive cells may have an adverse effect on patient prognosis.
[0053] Furthermore, a detailed comparison of the expression of functional markers for each cell subset within these regions was performed, and the markers with the most significant expression changes were identified ( Figure 3 ). At the fibrotic-epithelial junction, significant changes in marker expression occurred, indicating that the tumor fibrotic environment significantly affected the dynamic changes of markers. In particular, exhaustion markers (such as PD-1) and pro-inflammatory markers (such as TNF-α) were significantly expressed in CD4 + effector cells, lag3 + NK cells, granulocytes, CD8 + T cells, and type I collagen + fibroblasts. This pattern indicates that the presence of exhaustion and inflammatory markers in effector cells, memory cells, myeloid cells, and fibroblasts may be associated with a poor prognosis.
[0054] Conversely, activation markers (such as GATA3 and HLA-DR), as well as the proliferation marker Ki-67, were significantly expressed in effector CD4 + and CD8 + T cells in the long-term survival group ( Figure 3 ). These findings indicate that the activation and proliferation of effector T cells are associated with a better prognosis. This enhanced effector function and expansion in the tumor microenvironment may play a key role in mediating a more effective anti-tumor immune response. Based on the comparison of cell frequency and marker expression, further cell interaction analysis was performed (Figure 4 ). Figure 4 In which, the x-axis represents the starting (central) cell type, and the y-axis represents the interacting adjacent cell types. Statistically significant differences in interaction frequencies (p < 0.05) between the long-survival group and the short-survival group are emphasized with larger markers. It is observed that in the long-term survival group, HLA-DR hi macrophages and PD-L1 + epithelial cells, NK cells, B cells, and effector CD4 + T cells interact more frequently. This indicates that HLA-DR hi macrophages play a key role in promoting the interaction between tumor cells and immune effector cells, which may contribute to the prolongation of the patient's survival period.
[0055] In contrast, in the short-survival group, CD11c hi macrophages and type I collagen + fibroblasts showed a significant increase in interaction ( Figure 4 ). This pattern indicates that the interaction mediated by CD11c hi macrophages and type I collagen + fibroblasts may promote tumor activity and is associated with poor prognosis. Therefore, the different roles of macrophage subsets in tumor-immune interactions seem to be crucial in determining the clinical progression and prognosis of gastric cancer.
[0056] Through the comparative analysis of regional cell frequencies, marker expressions, and cell interactions, different mechanisms related to the patient's prognosis can be statistically obtained: in the long-survival group, the frequency and activation status of effector cells increased significantly, while in the short-survival group, it mainly showed the depletion state of immune effector cells and type I collagen + fibroblasts. Most notably, HLA-DR hi macrophages can promote better interaction with PD-L1 + epithelial cells and other immune effector cells. On the contrary, in the short-survival group, the interaction between CD11c hi macrophages and type I collagen + fibroblasts promoted the tumor. These findings emphasize the key role of specific macrophage subsets in regulating tumor-immune dynamics, and these findings are important factors in determining the clinical prognosis of gastric cancer.
[0057] The step (v) performs regression analysis modeling based on the results of cell frequencies, marker expression intensities, and cell interactions in different comparison groups, and the screened prognostic markers specifically include: Multivariate survival analysis: After completing the spatial analysis, functional marker expression, cell type frequencies, and cell-cell interactions were extracted and integrated for multivariate survival analysis. Features with a zero-value proportion below 50% were screened, and cross-validation was performed via LASSO regression to identify non-zero coefficients. The selected variables were then incorporated into a multivariate Cox regression model, and the proportional hazards assumption (PH assumption) was evaluated using the cox.zph() function. Variables meeting the PH assumption were retained, and their association with survival was evaluated. The results were visualized using the ggforest() function.
[0058] In the above process, after integrating all spatial features (including cell frequencies, functional marker expression of regional subpopulations, and cell-cell interactions), a multivariate Cox regression analysis was performed. First, variables were constructed based on the marker expression levels, cell interactions, and cell frequencies of various cell types in different regions. Subsequently, Lasso regression was used for variable selection by introducing a penalty term, extracting representative variables (non-zero coefficients), and significantly reducing the model complexity (from 4807 variables to 19), thus avoiding overfitting. Then, a proportional hazards (PH) assumption test was conducted to ensure that the hazard ratio in the model remained constant over time. In addition, a multicollinearity assessment was performed to confirm the independence of the selected variables, thereby enhancing the reliability and interpretability of the model. Finally, 16 spatial risk factors were identified, such as Figure 5 As shown, among the 16 finally identified spatial risk factors, HLA-DR hi macrophages mediated spatial features associated with favorable prognosis, and HLA-DR hi macrophages were able to promote better interactions between PD-L1 + epithelial cells and other immune effector cells. This was consistent with the previous regional analysis results. In contrast, CD11c hi macrophages and type I collagen + fibroblasts interaction was found to be the primary spatial high-risk factor highly associated with poor prognosis ( Figure 5 ), and this result was consistent with the above interaction research results.
[0059] After single-cell spatial deconvolution, we evaluated the key spatial features identified by IMC analysis at each point, including HLA-DR hi macrophage-mediated positive points, and CD11c hi macrophage and type I collagen + fibroblast interaction points. This approach enabled us to accurately assess the presence and distribution of these key spatial features in tissue sections and provided insights into the mechanisms of this spatial organization in the tumor microenvironment.
[0060] Based on the above analysis results, 4 key cell types related to gastric cancer prognosis were screened out, namely PD-L1 + epithelial cells, HLA-DR hi macrophages, CD11c hi macrophages and type I collagen + fibroblasts. The markers used for clustering analysis of the above cells are the key markers that can be used for predicting gastric cancer prognosis.
[0061] Based on the markers identified by the above imaging mass cytometry (IMC) technology analysis, 7 markers were screened out in this example, namely CD68, PD-L1, PAN-CK, HLA-DR, CD45, CD11c, and type I collagen. This is because the expression of these 7 markers was used to cluster and identify the above four key cell types in the above method: epithelial cells, HLA-DR hi macrophages, CD11c hi macrophages and type I collagen + The interactions between fibroblasts were used to extract independent images from each region of interest (ROI), generate masks corresponding to these cell types, and then establish a survival prediction model.
[0062] The method for establishing the model is as Figure 1 shown, including the following steps: From the imaging mass cytometry (IMC), two types of integrated image features were designed: marker intensity features and spatial similarity features.
[0063] The marker intensity features are the measurement results of the average expression intensity of 7 markers in the IMC image data, CD68, PD-L1, PAN-CK, HLA-DR, CD45, CD11c, and type I collagen.
[0064] The spatial similarity features were calculated by the structural similarity index (SSIM) and were used to evaluate the pairwise spatial relationships of epithelial cells, HLA-DR hi macrophages, CD11c hi macrophages and type I collagen + fibroblasts in the IMC image data, reflecting the pairwise spatial relationships of epithelial cells, HLA-DR + macrophages, CD11c + macrophages and type I collagen +Interactions between fibroblasts. The structural similarity index SSIM is a metric for measuring the similarity between two images. Different from traditional pixel-difference-based evaluation methods such as mean squared error (MSE) or peak signal-to-noise ratio (PSNR), SSIM takes more into account the brightness, contrast, and structural information of images, and thus can more accurately reflect the perception of image quality by the human visual system. These features, along with the prognostic survival status of patients in the dataset, are input into a deep neural network, and the trained model can be used for the prognostic clinical prediction of gastric cancer patients.
[0065] The tissue samples of 27 patients were used to validate the above-mentioned markers and model. The imaging mass cytometry (IMC) image results of the tissue samples of the 27 patients were obtained and input into the traditional survival prediction model based on all IMC markers and the cancer prognostic risk assessment model trained in each space based on 7 key markers in the present invention. Combining the marker expression intensity features and spatial similarity features can obtain obvious advantages. The above validation data were analyzed using existing prognostic analysis models, including RF (Random Forest), GCN (Graph Convolutional Network), and DeepSurv (Deep Survival Model). The results showed that the model of the present invention showed significant performance improvement in comparison with existing studies ( Figure 6 , Figure 7 ).
[0066] Based on the above screening method, the present invention also proposes a screening system for cancer prognostic markers, including: Data acquisition module: used to acquire the imaging mass cytometry analysis dataset, where the imaging mass cytometry analysis dataset contains the expression intensity of the markers to be screened in the samples and the corresponding prognostic conditions of the samples; and different comparison groups are divided according to the sample conditions and prognostic conditions; Cell type annotation module: used to perform single-cell segmentation on the images in the imaging mass cytometry analysis dataset, and annotate the cell types of single cells according to the expression intensity of the markers to be screened, to obtain the cell subpopulation frequencies of different cell types; Regional expression intensity analysis module: used to identify the regions where epithelial, immune, and fibrotic cells meet, and obtain the marker expression intensity in the regions where epithelial, immune, and fibrotic cells meet; Interaction analysis module: used to obtain the cell interaction analysis results based on cell frequencies and marker expression intensities; The regression analysis module performs regression analysis modeling based on the cell frequencies, marker expression intensities, and results of cell-cell interactions in different comparison groups, and screens out prognostic markers.
[0067] The above system can run on a computer device, which includes: one or more processors, a memory, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface).
[0068] The processor can be a central processing unit, a network processor, or a combination thereof. Among them, the processor can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0069] Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the method shown in the above embodiments.
[0070] The memory can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory can include a memory remotely set relative to the processor, and these remote memories can be connected to the computer device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0071] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A screening method for a cancer prognosis marker, characterized in that, It includes the following steps: Step 1): Obtain an analysis dataset of imaging mass cytometry, where the analysis dataset of imaging mass cytometry contains the expression intensity of markers to be screened in a sample and the corresponding prognosis of the sample; and divide different comparison groups according to the sample situation and prognosis; Step 2): Perform single-cell segmentation on the images in the analysis dataset of imaging mass cytometry, and annotate the cell types of single cells according to the expression intensity of the markers to be screened to obtain the cell frequencies of different cell types; Step 3): Identify the regions where epithelial, immune, and fibrotic cells meet, and obtain the expression intensity of markers in the regions where epithelial, immune, and fibrotic cells meet; Step 4): Based on the cell frequencies and marker expression intensities of different cell types in the regions where epithelial, immune, and fibrotic cells meet, calculate the interaction analysis results between different cell types; Step 5): Perform regression analysis according to the cell frequencies, marker expression intensities, and cell interaction results in different comparison groups to screen out prognostic markers.
2. The screening method of the cancer prognosis marker according to claim 1, characterized in that, The cancer is gastric cancer.
3. The screening method of the cancer prognosis marker according to claim 2, wherein In step 1), dividing different comparison groups according to the sample situation and prognosis specifically includes: Dividing into a tumor group and a peritumoral tissue group according to whether the sample belongs to tumor or peritumoral tissue; Dividing into an early pathological stage group and a late pathological stage group according to whether the sample belongs to an early pathological stage or a late pathological stage; Dividing into an early clinical stage group and a late clinical stage group according to whether the sample belongs to an early clinical stage or a late clinical stage; Dividing into a long survival group and a short survival group according to the patient's prognosis; 4. The screening method of the cancer prognosis marker according to claim 2, wherein The markers to be screened in step 1) include CD45, CD15, CD45RA, CD103, CD68, HLA-DR, CD45RO, B7H4, CD14, CD20, CD57, PD-L1, CD25, CD3, CD4, αSMA, CD8, CD7, FOXP3, KI67, CD163, CD16, CD69, vimentin, CD11b, CD11c, Caspase3, IL-1b, Ki-67, TNFα, type I collagen, LAG3, GATA3, PD-1, VISTA, CD31, granzyme B, IL-6, trypsin, E-Cadherin.
5. The screening method of the cancer prognosis marker according to claim 1, characterized in that, The cell types include: Myofibroblasts, type I collagen⁺ fibroblasts, stromal cells, endothelial cells, epithelial cells, lymphocytes, myeloid cells; The epithelial cells include the following types: PD-L1 + subpopulation, Ki-67 + proliferation subpopulation, E-cadherin + subpopulation, E-cadherin - subpopulation; The lymphocytes include the following types: regulatory T cells, effector CD8 + T cells, memory CD4 + T cells, memory CD8 + T cells, effector CD4 + T cells, CD16 - natural killer cells, CD16 + natural killer cells, Lag3 + natural killer cells, B cells; The myeloid cells include the following types: granulocytes, monocytes, Ki67 + Proliferating macrophages, CD11c low HLA-DR low Macrophages, CD11c hi Macrophages, HLA-DR hi Macrophages.
6. The screening method of the cancer prognosis marker according to claim 5, wherein In step 2), annotating the cell types of single cells includes the following steps: The first round of clustering, identifying the main cell populations through the following markers: E-Cadherin, trypsin, CD3, CD4, CD8, CD20, CD7, CD57, CD45, CD68, CD15, CD14, CD16, CD31, αSMA, type I collagen, and vimentin; For the second round of clustering, cluster myeloid cells using FOXP3, CD16, CD69, CD4, CD8, Caspase3, B7H4, VISTA, CD7, CD103, LAG3, CD20, granzyme B, PD-1, KI67, GATA3, CD45RA, CD3, TNFα, TL1β, CD45RO, CD57, CD25; Cluster lymphocytes using IL-6, CD14, CD16, Caspase3, CD163, PD-L1, CD11B, CD11C, CD15, Ki-67, HLA-DR, TNFα, and TL1β; Cluster epithelial cells using E-Cadherin, PDL1, Ki-67, and trypsin.
7. The screening method of the cancer prognosis marker according to claim 1, wherein The steps for identifying the regions where epithelial, immune, and fibrotic cells intersect in step iii) include the following: Based on spatial coordinates, use the k-nearest neighbor algorithm to calculate the 20 nearest neighbors of each cell, thereby constructing a local spatial adjacency graph for each cell; for the neighborhood of each cell, aggregate the expression data of various markers of neighboring cells and aggregate according to cell type; The local subgraphs are screened to ensure that the central node of each subgraph corresponds to a specific cell type; the subgraphs are then connected through shared nodes to form a global connection graph, thereby generating spatial regions / patches; the spatial range of the existing patches is expanded by adding n neighboring nodes to the existing patches, thereby expanding the spatial coverage of the constructed patches; Based on the spatial connection graph, an immune region, a fibroblast region, and an epithelial region are obtained; the immune region includes lymphocytes and myeloid cells; the fibroblast region includes myofibroblasts and type I collagen; + fibroblasts; the epithelial region includes epithelial cells; Based on the immune region, fibroblast region, and epithelial region, the regions where epithelial, immune, and fibrotic cells intersect can be obtained. The regions where epithelial, immune, and fibrotic cells intersect include: Immune-fibroblast junction; Immune-epithelial junction; Fibroblast-epithelial junction; In step iv), the average number of surrounding cells in each region of interest is calculated to count the number of interactions between the central cell and neighboring cells.
8. The screening method of the cancer prognosis marker according to claim 1, characterized in that In step v), regression analysis is performed based on the cell frequencies, marker expression intensities, and results of cell-cell interactions in different comparison groups. Specifically, it includes the following steps: Variables with a zero value proportion lower than 50% are screened out, and cross-validation is performed through LASSO regression to identify non-zero coefficients; the screened variables are incorporated into a multivariable Cox regression model, and the proportional hazards assumption is evaluated through the cox.zph function; variables that meet the PH assumption are retained, and the association between the variables and survival is evaluated; the results are visualized through the ggforest function.
9. A method for establishing a cancer prognosis diagnosis model, characterized in that, Include the following steps: Use the method in claim 1 to screen for cancer prognosis markers; Analyze the average expression intensity of each marker in the cancer prognosis marker combination screened from the dataset using imaging mass cytometry, and establish a cancer prognosis diagnosis model using machine learning methods for the prognosis.
10. A screening system for cancer prognosis markers, characterized in that, Include: Data acquisition module: used to acquire the analysis dataset of imaging mass cytometry, where the analysis dataset of imaging mass cytometry contains the expression intensity of the markers to be screened in the sample and the corresponding prognosis of the sample; and different comparison groups are divided according to the sample situation and the prognosis situation; Cell type annotation module: used to perform single-cell segmentation on the images in the analysis dataset of imaging mass cytometry, and annotate the cell types of single cells according to the expression intensity of the markers to be screened, so as to obtain the cell subpopulation frequencies of different cell types; Regional expression intensity analysis module: used to identify the regions where epithelial, immune, and fibrotic cells meet, and obtain the expression intensity of the markers in the regions where epithelial, immune, and fibrotic cells meet; Interaction analysis module, used to calculate the interaction analysis results between different cell types based on the cell frequencies and marker expression intensities of different cell types in the regions where epithelial, immune, and fibrotic cells meet; Regression analysis module, performs regression analysis according to the cell frequencies, marker expression intensities, and cell interaction results in different comparison groups, and screens out the prognostic markers.
Citation Information
Patent Citations
System for predicting prognosis of gastric cancer in subjects
CN111724903A
Immune microenvironment score prediction method, system and equipment based on tissue mass spectrum imaging
CN114487075A
Prognosis survival prediction model construction method and prediction method
CN119007201A
Method for analyzing tumor microenvironment of esophageal squamous cell carcinoma
CN119517167A
Cancer status prediction method and uses thereof
US20220260576A1
Cited By
Immunohistochemical staining machine data acquisition and analysis method and system
CN121207795A