A patient stratification method based on breast cancer single-cell spatial proteome multi-cell enrichment patterns

By segmenting cells and constructing networks from single-cell spatial proteomics images of breast cancer, detecting cell communities, and extracting and characterizing the distribution patterns of the tumor microenvironment in patients, the problem of inaccurate stratification of breast cancer patients in existing technologies is solved, and more accurate cancer prognosis assessment is achieved.

CN119560168BActive Publication Date: 2025-11-18ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411557830.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-11-18
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

Current technologies lack effective methods to analyze the spatial structure and heterogeneity of the breast cancer tumor microenvironment, resulting in insufficient accuracy in cancer prognostic assessment, especially in the lack of methods for stratifying patients in multi-cell enrichment patterns.

Method used

Deep learning methods and image processing techniques were used to segment and identify cell types in single-cell spatial proteomics images of breast cancer, construct cell network graphs, detect tightly connected cell communities, and extract and characterize the distribution patterns of the tumor microenvironment of patients using clustering algorithms, and stratify patients using spatial topological information.

Benefits of technology

It enables precise stratification of breast cancer patients, makes full use of spatial topological information, reveals the structural characteristics of the tumor microenvironment, and improves the accuracy of cancer prognostic assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
  • Figure HDA0005117024440000011
    Figure HDA0005117024440000011
Patent Text Reader

Abstract

The present application belongs to the technical field of breast cancer treatment, and particularly relates to a patient stratification method based on breast cancer single-cell spatial proteomics multi-cell enrichment mode, comprising the following steps: S1: breast cancer single-cell spatial proteomics image data cell segmentation and cell type identification; S2: constructing a cell network graph and simulating a tumor microenvironment; S3: detecting closely connected cell communities and extracting cell community features; S4: clustering the cell community features and representing a patient tumor microenvironment distribution mode; and S5: calculating the similarity of the patient tumor microenvironment distribution mode and stratifying the patient. The method can stratify the patient, and the prognosis of the patient groups is significantly different.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of breast cancer treatment technology, and in particular relates to a patient stratification method based on the multi-cell enrichment pattern of single-cell spatial proteome in breast cancer. Background Technology

[0002] Cancer is a major global public health problem. Currently, the prognostic assessment systems for breast cancer mainly include the following four types: (1) TNM (Tumor Node Metastasis) staging system for breast cancer; (2) histological grading of breast cancer; (3) molecular subtyping of breast cancer, based on gene expression information of estrogen receptor ER, progesterone receptor PR, and human epidermal growth factor receptor 2 (HER2); and (4) gene detection methods for breast cancer. These methods mainly focus on the morphology, proliferation capacity, and molecular expression of tumor cells in breast cancer. However, the occurrence and development of breast cancer are not only related to the malignant proliferation caused by gene mutations in cancer cells, but also closely related to the tumor microenvironment (TME). The TME is highly heterogeneous, stemming from differences in cell phenotype, intercellular interactions, and cell spatial location, which affect tumor development and prognosis. Therefore, accurate analysis of the TME structure and its correlation with tumor development and prognosis is of great significance.

[0003] To accurately analyze the structure of the tumor microenvironment (TME) and establish a tumor single-cell atlas, single-cell analysis has entered the era of multi-omics. The molecular information set of cells is used as a parameter set to depict cellular phenotype and function. For example, single-cell sequencing technology detects the transcriptome information of individual cells in dissected tissues to analyze cell phenotype and function and predict cell-cell interactions, but it lacks spatial information of individual cells, cell-cell interaction relationships, and tissue structural characteristics. Spatial transcriptome sequencing technology obtains gene expression information of individual cells or neighboring cell populations in tissue samples and their relationship with their spatial location within the tissue, but this currently popular spatial transcriptome technology cannot achieve single-cell resolution. Spatial metabolome technology can obtain metabolome information within single cells, analyze the TME at the spatial metabolic level, discover new cellular metabolic phenotypes, and elucidate the mechanisms by which metabolic pathways remodel the microenvironment, but it currently cannot accurately identify cell phenotypes. Spatial proteome technology detects protein expression in individual cells and obtains the spatial location information of individual cells; simultaneously, protein biomarkers can precisely define cell phenotype and function. Among transcriptomics, metabolomics, and proteomics, proteins are the main carriers of cellular life activities, and their abnormalities are often closely related to diseases. Moreover, tissue space is the research level closest to biological reality. Therefore, obtaining the protein localization and expression profiles of single cells at the tissue spatial level is of great value for analyzing the tissue microenvironment, diagnosis and prognosis, and precision medicine.

[0004] Currently, single-cell spatial proteomics is an emerging technological field, and there is no unified method for analyzing this data. In particular, the analysis of TME structure based on multi-cell enrichment patterns does not fully utilize spatial topological information, resulting in insufficient research on the relationship between cancer spatial dimension analysis and prognosis. Moreover, there are few methods for stratifying patients based on TME structure analysis using multi-cell enrichment patterns. Therefore, this invention, with breast cancer as the background, proposes a patient stratification method based on multi-cell enrichment patterns in single-cell spatial proteomics, which can achieve patient stratification and obtain some new biological insights. Summary of the Invention

[0005] To comprehensively address the aforementioned problems, especially the shortcomings of existing technologies, this invention provides a patient stratification method based on the multi-cell enrichment pattern of single-cell spatial proteome in breast cancer, which can stratify breast cancer patients and obtain some new biological insights.

[0006] In view of this, the present invention provides a patient stratification method based on the multi-cell enrichment pattern of single-cell spatial proteome in breast cancer, comprising the following steps:

[0007] S1: Spatial proteomics image data of single breast cancer cells for cell segmentation and cell type identification

[0008] The image data is segmented using deep learning methods such as Mesmer or CellProfiler combined with Ilastik software to obtain cell mask images. The cell mask images are then overlaid on each protein channel image to obtain the expression of each protein in each cell, resulting in a cell protein expression matrix. Based on the cell protein expression matrix, a clustering algorithm is used to cluster the cells to obtain a cell population protein expression heatmap. The cell phenotype is then identified based on the protein expression of the cell population.

[0009] S2: Constructing a cell network diagram to simulate the tumor microenvironment.

[0010] Using the cell mask image mentioned in S1, the coordinates of each cell centroid are extracted using the image processing software package skimage or OpenCV. The distance between each cell centroid and the centroids of other cells is calculated. A radius value is set, and cells within this value are considered as the neighbor of the cell. That is, an edge is connected between these two cells. Cells beyond the radius are not connected by an edge. The cell network graph is constructed.

[0011] S3: Detect tightly connected cellular communities and extract cellular community features.

[0012] Using the cell network graph mentioned in S2, a community detection algorithm is used to detect tightly connected cell communities. To make full use of spatial topological information, the number of edges connecting each cell type is extracted for each cell community. The number of edges connecting each cell type represents the importance of that cell type in the cell community and can better reflect the arrangement of that cell in the cell community. The number of edges connecting each cell type is used as the feature representation of that cell community.

[0013] S4: Cell community feature clustering, characterizing the distribution pattern of the patient's tumor microenvironment.

[0014] Using the cell community features extracted by S3, a clustering algorithm was used to obtain k tumor microenvironment feature structures. The frequency of each tumor microenvironment feature structure in each patient was counted as a representation of the distribution pattern of the tumor microenvironment in that patient.

[0015] S5: Calculate the similarity of tumor microenvironment distribution patterns in patients and stratify them.

[0016] The tumor microenvironment distribution pattern of each patient was obtained using S4, and the similarity of the tumor microenvironment distribution patterns between each pair of patients was calculated. Then, a graph was constructed using patients as nodes and the similarity between patients as the weights of the edges connecting patients. Finally, a community detection algorithm was used to classify patients into different groups.

[0017] The beneficial effects of this invention are: based on a large dataset, this invention can make full use of spatial topological information to extract multi-cell enrichment patterns to analyze TME structural features, and characterize the TME structural distribution pattern of the patient's tumor microenvironment, so as to accurately stratify patients with prognostic differences. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the overall operation of the present invention;

[0019] Figure 2 This is a schematic diagram of the cell community network constructed in this invention;

[0020] Figure 3 This is a schematic diagram illustrating the clustering of cell community features extracted in this invention;

[0021] Figure 4 This is a schematic diagram illustrating the patient stratification based on the distribution pattern of the tumor microenvironment in the present invention;

[0022] Figure 5 This is a schematic diagram of the patient stratification results of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] In the description of this application, it should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. For ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0025] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and are not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0026] It should be noted that in the description of this application, the directional terms such as "front, back, up, down, left, right", "horizontal, vertical, horizontal" and "top, bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description. Unless otherwise stated, these directional terms do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the scope of protection of this application. The directional terms "inner" and "outer" refer to the inner and outer contours relative to the outline of each component itself.

[0027] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0028] Example 1:

[0029] like Figures 1-5 As shown, this embodiment provides a patient stratification method based on the multi-cell enrichment pattern of single-cell spatial proteome in breast cancer, including the following steps:

[0030] S1: Spatial proteomics image data of single breast cancer cells for cell segmentation and cell type identification

[0031] For the Mesmer deep learning method, input images of the nuclear channel and the cell membrane or cytoplasmic channel are required. For single-cell spatial proteomics images, there are often multiple images of the nuclear channel and multiple images of the cell membrane or cytoplasmic channel. Pixel values ​​of multiple nuclear channel images, cell membrane images, or cytoplasmic channel images are averaged to obtain fused nuclear channel, cell membrane, or cytoplasmic channel images, which are then input into the Mesmer method to output a cell mask image. For the CellProfiler software combined with the Ilastik software method, the cell nucleus is first detected using Ilastik software, and then the nuclear region is expanded using a propagation method to cover cell membrane and cytoplasmic signals. The probability map is generated by stopping at the cell edge based on the signal intensity gradient, and then input into the CellProfiler software to generate a cell mask image.

[0032] After obtaining the cell mask, the expression value of a certain protein on a single cell is defined as the average pixel value of the image region enclosed by the cell mask. A cell protein expression matrix is ​​obtained in this way. Then, the FlowSOM clustering algorithm is used to cluster cells based on their protein expression, resulting in a cell population protein expression matrix. Finally, the cell population type is identified based on the protein expression of each cell population.

[0033] S2: Constructing a cell network diagram to simulate the tumor microenvironment.

[0034] Using the cell mask image obtained from S1, the coordinates of the centroid of each cell are extracted using the Python package skimage or OpenCV. The distance between the centroid of each cell and the centroids of other cells is calculated. A radius value is set, and cells within this value are considered "neighbors" of the cell. That is, an edge is connected between these two cells. Cells beyond the radius are not connected by an edge. In other words, cells are treated as nodes, and if cells are neighbors, they are connected by an edge. A cell network graph is constructed.

[0035] S3: Detect tightly connected cellular communities and extract cellular community features.

[0036] Using the cell network graph obtained from S2, the random walktrap community detection algorithm is used to identify tightly connected cell communities. In order to better utilize the spatial topological information of cells, the number of edges connecting cell types in each cell community is extracted to characterize the cell community, i.e., the multi-cell enrichment pattern.

[0037] S4: Cell community feature clustering, characterizing the distribution pattern of the patient's tumor microenvironment.

[0038] Using the representation vector of each multicellular enrichment pattern obtained from S3, hierarchical clustering is performed using Canberra distance as the distance metric to group similar multicellular enrichment patterns together, resulting in k tumor microenvironment feature structures. The frequency of occurrence of feature microenvironment structures for each patient is statistically analyzed to characterize the distribution pattern of the patient's tumor microenvironment, i.e.:

[0039]

[0040] S5: Calculate the similarity of tumor microenvironment distribution patterns in patients and stratify them.

[0041] Using the tumor microenvironment distribution pattern obtained for each patient from S4, the cosine similarity of the tumor microenvironment distribution patterns between each pair of patients is calculated, i.e.:

[0042]

[0043] The similarity matrix is ​​obtained as follows:

[0044]

[0045] Based on similarity values, the k nearest neighbors of each patient were obtained, and the Intersection over Union (IOU) matrix was calculated, which is the intersection number of the k nearest neighbors between two patients divided by the union number. A patient-to-patient network graph was constructed using the IOU matrix, and then the Louvain community detection algorithm was used to group patients. Finally, combining the overall survival of patients and the occurrence of death events, the multivariate log-rank test was used to compare the survival outcomes of each group. The results showed statistically significant differences between groups. Combined with clinical information, it was found that patients could be accurately stratified with prognostic differences.

[0046] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A patient stratification method based on multi-cell enrichment patterns of single-cell spatial proteome in breast cancer, characterized in that, Includes the following steps: S1. Cell segmentation and cell type identification from spatial proteomics image data of single-cell breast cancer cells; S2. Construct a cell network diagram to simulate the tumor microenvironment; S3. Detect tightly connected cellular communities and extract cellular community features; S4. Cell community feature clustering to characterize the distribution pattern of the patient's tumor microenvironment; S5. Calculate the similarity of the distribution patterns of the tumor microenvironment in patients and stratify the patients accordingly; S3 specifically includes the following steps: Using the cell network graph mentioned in S2, a community detection algorithm is used to detect tightly connected cell communities. For each cell community, the number of edges connecting cell types is extracted. The number of edges connecting cell types represents the importance of that cell type in the cell community and can better reflect the arrangement of that cell type in the cell community. The number of edges connecting each cell type is used as the feature representation of that cell community. S4 specifically includes the following steps: Using the cell community features extracted by S3, a clustering algorithm was used to obtain k tumor microenvironment feature structures. The frequency of each tumor microenvironment feature structure in each patient was counted as a representation of the distribution pattern of the tumor microenvironment in that patient. S5 specifically includes the following steps: Using the tumor microenvironment distribution pattern obtained by S4 for each patient, the similarity of the tumor microenvironment distribution patterns between each pair of patients is calculated. Then, a graph is constructed with patients as nodes and the similarity between patients as the weights of the edges connecting patients. Finally, a community detection algorithm is used to classify patients into different groups.

2. The patient stratification method based on the multi-cell enrichment pattern of single-cell spatial proteome in breast cancer according to claim 1, characterized in that: S1 specifically includes the following steps: The image data is segmented into cell mask images using deep learning methods such as Mesmer or CellProfiler combined with Ilastik software. The cell mask images are then overlaid on each protein channel image to obtain the expression of each protein in each cell, resulting in a cell protein expression matrix. Based on the cell protein expression matrix, a clustering algorithm is used to cluster the cells to obtain a cell population protein expression heatmap. The cell type is then identified based on the protein expression of the cell population.

3. The patient stratification method based on the multi-cell enrichment pattern of single-cell spatial proteome in breast cancer according to claim 2, characterized in that: S2 specifically includes the following steps: Using the cell mask image mentioned in S1, the coordinates of each cell centroid are extracted using an image processing software package. The distance between each cell centroid and the centroids of other cells is calculated. A radius value is set, and cells within this value are considered as the neighbor of that cell. That is, an edge is connected between these two cells. Cells beyond the radius are not connected by an edge. A cell network graph is constructed.

Citation Information

Patent Citations

  • Prediction of brcaness / homologous recombination deficiency of breast tumors on digitalized slides

    CA3226033A1

  • Breast cancer pathological image analysis method and device based on convolutional neural network

    CN113838558A