A method and system for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells
By constructing a cell topological network and spatial weight matrix, combined with Moran index analysis, the problems of low efficiency and insufficient accuracy in traditional tumor microenvironment analysis are solved, and accurate analysis and dynamic change prediction of tumor microenvironment are achieved.
Patent Information
- Application Number
- CN202510309219.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Traditional tumor microenvironment analysis methods rely on manual evaluation, have low efficiency and large errors, making it difficult to comprehensively analyze the spatial relationship and dynamic effects between cells, and lack recognition accuracy when the nucleus is densely stacked.
The quantitative analysis method of cell spatial distribution characteristics of tumor microenvironment was used to obtain cell segmentation and classification information through pre-trained cell recognition models, and the cell topology network and spatial weight matrix were constructed. The cell adjacency relationship and spatial distribution characteristics were analyzed using the Delaune triangle network and the Voronoi diagram, combining with the Moran index to quantify the spatial autocorrelation and interaction of cell types.
Accurate analysis of the tumor microenvironment is achieved, cell segmentation accuracy is improved, and the adjacency relationship and spatial structure between cells is fully revealed, providing more accurate prediction of cell morphology and dynamic changes, and enhancing the spatial dimension and accuracy of tumor microenvironment analysis.
Smart Images

Figure CN119915815B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly relates to a method and system for quantitatively analyzing the spatial distribution characteristics of cells in the tumor microenvironment. Background Art
[0002] The tumor microenvironment refers to the area between tumor cells and the surrounding normal tissues, which is composed of tumor cells, extracellular matrix of tumor cells, tumor-associated fibroblasts, immune cells, etc.; the tumor microenvironment is closely related to the growth, metastasis, and immune escape of tumors.
[0003] The tumor microenvironment is usually applied in the diagnosis and treatment of tumors. Traditional tumor microenvironment analysis mainly relies on manual evaluation. However, due to the complexity and heterogeneity of tumor tissues, it is easily affected by subjectivity, resulting in problems such as low efficiency and large errors. In the prior art, whole-slide image scanning technology is mostly used, which can provide relatively accurate cell localization information. However, the analysis of the tumor microenvironment is not comprehensive enough, easily ignores the spatial relationship between cells, and is also difficult to capture the mutual dynamic interactions of cells in the tumor microenvironment in advance. Moreover, there is still a problem of insufficient recognition accuracy in the case of densely stacked cell nuclei, ultimately affecting the accuracy of tumor microenvironment analysis.
[0004] Therefore, there is a need for a solution to solve at least one of the above problems. Summary of the Invention
[0005] In view of the defects in the prior art, the present application proposes a method and system for quantitatively analyzing the spatial distribution characteristics of cells in the tumor microenvironment. To solve at least one of the above technical problems, the technical solution of the present application is as follows: A method for quantitatively analyzing the spatial distribution characteristics of cells in the tumor microenvironment, including: obtaining a pathological section image, and obtaining cell segmentation and classification information in the pathological section image through a pre-trained cell recognition model; based on the cell segmentation and classification information, obtaining the cell spatial distribution information of the pathological section image; according to the cell spatial distribution information, constructing a cell topological network to obtain the first cell spatial distribution characteristic of the tumor microenvironment; according to the cell spatial distribution information, constructing a spatial weight matrix to obtain the second cell spatial distribution characteristic of the tumor microenvironment; based on the first cell spatial distribution characteristic and the second cell spatial distribution characteristic, judging the tumor state to realize the typing of the target object.
[0006] In a specific embodiment, the cell topological network includes: a first cell spatial topological network and / or a second cell spatial topological network; the first cell spatial topological network is constructed based on the relevant principles of the Delaunay triangular network, and the second cell spatial topological network is constructed through the Voronoi diagram;
[0007] The first cell spatial distribution features obtained through the first cell spatial topology network include one or any combination of the proportion of each type of cell link, the chain edge distance, and the link frequency; the first cell spatial distribution features obtained through the second cell spatial topology network include one or any combination of the area of each cell pixel, the area variance, the mean, the median, and the clutter degree.
[0008] In a specific embodiment, the first cell spatial topology network includes simplices, and the union of all simplices constitutes the convex hull of cell distribution. The triangulation of any cell position contains no more than triangles; where represents the number of cells, represents the spatial dimension, represents the number of cells in the convex hull; in the second cell spatial topology network, the Voronoi cell associated with the cell center is defined as: ;
[0009] where is a metric space, d is the distance function, represents the cell center, represents the Voronoi cell associated with the cell center.
[0010] In a specific embodiment, the acquisition of the second cell spatial distribution features specifically includes:
[0011] Based on the spatial weight matrix, obtain the grid adjacency weights after meshing the tumor microenvironment; obtain the number of cells in each grid as the grid cell number information; combine the grid adjacency weights and the grid cell number information to analyze the spatial autocorrelation of the same type of cells in the tumor microenvironment and the spatial autocorrelation between different types of cells in the tumor microenvironment.
[0012] In a specific embodiment, the analysis of the spatial autocorrelation of the same type of cells in the tumor microenvironment is realized based on the univariate Moran index; the calculation formula of the univariate Moran index is:
[0013] ;
[0014] where represents the number of grids after meshing the tumor microenvironment, represents the adjacency weight between grid i and grid j, and respectively represent the number of a certain type of cells in grid and grid , then represents the mean number of a certain type of cells in all grids;
[0015] Analyze the spatial autocorrelation among different types of cells in the tumor microenvironment, which is achieved based on the bivariate Moran's index; the calculation formula of the bivariate Moran's index is:
[0016] ;
[0017] wherein, represents the number of grids after the gridification of the tumor microenvironment, represents the number of the first type of cells in the th grid, represents the number of the second type of cells in the th grid, is an element in the weight matrix, representing the adjacency weight between grids, represents the mean value of the first type of cells, ; represents the mean value of the second type of cells, ; represents the spatial lag value, .
[0018] In a specific embodiment, based on the first cell spatial distribution feature and the second cell spatial distribution feature, the tumor state is judged to type the target object, which specifically includes: respectively selecting preset typing feature parameters from the first cell spatial distribution feature and the second cell spatial distribution feature; the preset typing feature parameters include: the chain edge distance between tumor cells and immune cells, the standard deviation of the chain edge distance between stromal cells and stromal cells, the variance of the chain edge distance between stromal cells and stromal cells, the number of chain edges between tumor cells and immune cells, the number of triangles in the triangular network, the number of polygons, the standard deviation of the chain edge distance between tumor cells and immune cells, the variance of the chain edge distance between tumor cells and immune cells, one or any combination of them; obtaining the clinical information of the target object; combining the association established by the clinical information of the target object, the typing feature parameters and the survival period of the target object to judge the tumor state so as to type the target object; the tumor state includes: one or any combination of in situ, local expansion and metastasis.
[0019] In a specific embodiment, the selection of the preset parameter features specifically includes: performing KM survival analysis and Cox proportional hazards analysis based on the overall survival period and the progression-free survival period of the target object; judging whether each feature parameter in the first cell spatial distribution feature and the second cell spatial distribution feature meets the preset conditions; the preset conditions include the value ranges of one or any combination of the logrank-p value and the Cox-p value corresponding to each of the feature parameters for the overall survival period and the progression-free survival period; if it meets the preset conditions, it is used as the typing feature parameter.
[0020] In a specific embodiment, the acquisition of cell segmentation and classification information specifically includes: segmenting the pathological section image according to a preset specification to obtain sub-images of the pathological section; identifying the cells in each of the sub-images of the pathological section through the pre-trained cell recognition model to obtain the cell segmentation and classification information; the acquisition of cell spatial distribution information specifically includes: restoring the pathological section image by in-situ backtracking of each of the sub-images of the pathological section; and obtaining the cell spatial distribution information of the pathological section image based on the restored pathological section image and the cell segmentation and classification information.
[0021] In a specific embodiment, the cell recognition model is trained based on the STARDIST model, which specifically includes: acquiring and preprocessing a training sample set, where the sample set includes multiple training samples; the input data includes: training images in multiple of the training samples; the output data includes: the classification of each cell in the corresponding training image and the number of cells of each type; using the input data and the output data to perform deep learning model training and optimize the parameter settings of the deep learning model. If the classification of each cell in the training image output by the deep learning model and the number of cells of each type meet the preset target, the training is terminated to obtain the cell recognition model.
[0022] A quantitative analysis system for the spatial distribution characteristics of tumor microenvironment cells, which is applied to the quantitative analysis method for the spatial distribution characteristics of tumor microenvironment cells in any one of the foregoing, includes: a cell recognition and segmentation module: used to acquire a pathological section image and obtain the information about cell segmentation and classification in the pathological section image through a pre-trained cell recognition model; a distribution information acquisition module: used to acquire the cell spatial distribution information of the pathological section image based on the cell segmentation and classification information; a first feature acquisition module: used to construct a cell topological network according to the cell spatial distribution information and obtain the first cell spatial distribution characteristic of the tumor microenvironment; a second feature acquisition module: used to construct a spatial weight matrix according to the cell spatial distribution information and obtain the second cell spatial distribution characteristic of the tumor microenvironment; a judgment and typing module: used to judge the tumor state based on the first cell spatial distribution characteristic and the second cell spatial distribution characteristic to realize the typing of the target object.
[0023] Beneficial effects:
[0024] The present invention provides a method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells, which has good cell segmentation and classification accuracy, fully takes into account the adjacency relationship, interaction intensity and spatial structure between cells, conducts quantitative analysis of multi-dimensional spatial distribution characteristics, achieves good tumor microenvironment cell analysis results, can accurately realize the typing of target objects, and has good versatility. Specifically, by using a preset cell recognition model to recognize, segment and classify cells, it can effectively handle the overlap between cells, accurately locate the position of each cell, and provide a more accurate cell morphology, improving the cell segmentation accuracy. By constructing a cell topology network to obtain the first cell spatial distribution characteristics, it can accurately model the adjacency relationship and interaction between cells in the tumor microenvironment at the cell scale, and then comprehensively reveal the spatial structure of the tumor microenvironment, increasing the spatial dimension of the analysis of the tumor microenvironment. By constructing a spatial weight matrix, combining the univariate Moran index and the bivariate Moran index to obtain the second cell spatial distribution characteristics, it quantifies the spatial distribution pattern between different cell types in the tumor microenvironment, reveals the spatial aggregation or dispersion trend between cell types. The univariate Moran index can be used to analyze the spatial autocorrelation of the same type of cells, while the bivariate Moran index helps to analyze the spatial interaction between different cell types, providing a new perspective for the functional research of the tumor microenvironment, and thus helping to capture cell dynamic changes in advance. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0026] Figure 1 It is a flowchart of the method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells in the embodiment;
[0027] Figure 2 It is a composition diagram of the system for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells in the embodiment;
[0028] Figure 3 It is an example diagram of the first cell spatial topology network in the embodiment;
[0029] Figure 4 It is an example diagram of the second cell spatial topology network in the embodiment;
[0030] Figure 5 It is a KM curve diagram of the standard deviation of the edge distance between tumor cells and immune cells in the embodiment;
[0031] Figure 6KM curve graph of the standard deviation of the chain-side distance between stromal cells in the embodiment;
[0032] Figure 7 Recognition effect of the cell recognition model in the embodiment under different thresholds Figure 1 ;
[0033] Figure 8 Recognition effect of the cell recognition model in the embodiment under different thresholds Figure 2 ;
[0034] Figure 9 Graph of the results of the multivariate Cox regression analysis in this embodiment;
[0035] Figure 10 Example diagram of the training samples input into the cell recognition model in this embodiment;
[0036] Figure 11 Recognition effect diagram of the cell output of the cell recognition model in this embodiment.
[0037] Reference numerals:
[0038] 1 - Cell recognition and segmentation module; 2 - Distribution information acquisition module; 3 - First feature acquisition module; 4 - Second feature acquisition module; 5 - Judgment typing module. Detailed implementation manners
[0039] In the following, various embodiments of the present disclosure will be described more fully. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions that fall within the spirit and scope of the various embodiments of the present disclosure. In the following, the term "comprising" or "may comprise" used in the various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. Further, as used in the various embodiments of the present disclosure, the terms "comprising," "having," and their cognates are only intended to indicate a specific feature, number, step, operation, element, component, or combination of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items or precluding the possibility of adding one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items. In the various embodiments of the present disclosure, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the listed words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B. Expressions (such as "first," "second," etc.) used in the various embodiments of the present disclosure may modify various constituent elements in the various embodiments, but do not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only used for the purpose of distinguishing one element from other elements. For example, a first user device and a second user device indicate different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of the present disclosure, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element. It should be noted that if it is described that one constituent element is "connected" to another constituent element, the first constituent element may be directly connected to the second constituent element, and a third constituent element may be "connected" between the first constituent element and the second constituent element. Conversely, when one constituent element is "directly connected" to another constituent element, it can be understood that there is no third constituent element between the first constituent element and the second constituent element. The term "user" used in the various embodiments of the present disclosure may indicate a person who uses an electronic device or a device that uses an electronic device (e.g., an artificial intelligence electronic device). The terms used in the various embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the various embodiments of the present disclosure. As used herein, the singular form is intended to also include the plural form unless the context clearly indicates otherwise. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the present disclosure pertain.The terms (such as those defined in a commonly used dictionary) will be interpreted to have the same meaning as their contextual meaning in the relevant technical field and will not be interpreted to have an idealized meaning or an overly formal meaning, unless clearly defined in various embodiments of the present disclosure.
[0040] Embodiment
[0041] An embodiment of the present application provides a method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells, as Figure 1 shown including:
[0042] S100: Obtain a pathological section image, and through a pre-trained cell recognition model, obtain cell segmentation and classification information in the pathological section image;
[0043] Optionally, the cell recognition model can be trained based on the VGGNet model architecture, InceptionNet model architecture, HD-Yolo model architecture, Mask-RCNN model architecture, or Transformer model architecture, or can be trained based on one or any combination of the FCN model architecture or RetinaNet model architecture; of course, no specific limitation is imposed on the specific model architecture, as long as it meets cell segmentation and classification recognition; specifically, in some embodiments of the present application, the cell recognition model is trained based on the STARDIST deep learning model, abandoning the traditional boundary box-based segmentation method, and using star-shaped convex polygons parameterized by radial distance to represent the nucleus shape, thereby improving the accuracy of cell classification and segmentation.
[0044] S200: Based on the cell segmentation and classification information, obtain the cell spatial distribution information of the pathological section image; in some embodiments of the present application, the cell segmentation and classification information mainly includes, after cell recognition pattern recognition, the morphology, range division or segmentation of each cell, and the category judgment of the divided or segmented cells, such as tumor cells, stromal cells, etc.
[0045] S300: According to the cell spatial distribution information, construct a cell topological network to obtain the first cell spatial distribution characteristic of the tumor microenvironment; in some embodiments of the present application, the cell spatial distribution information mainly refers to the number, distribution position, relative distance, etc. of cells in space.
[0046] S400: According to the cell spatial distribution information, construct a spatial weight matrix to obtain the second cell spatial distribution characteristic of the tumor microenvironment;
[0047] Among them, the first cell spatial distribution feature is obtained based on the cell topological network, focusing on the adjacency relationship, interaction intensity and spatial topological structure between local individual cells. The cell spatial distribution in the tumor microenvironment is extracted as the cell spatial distribution feature in a numerically quantitative manner, revealing the spatial organization pattern between cells, quantifying the geometric relationship and interaction intensity between cells, and making up for the deficiencies in the analysis of the spatial relationship between cells in the existing technology; the second cell spatial distribution feature is obtained based on the spatial weight matrix, focusing on the spatial position distribution relationship of local overall cells in the tumor microenvironment, such as the spatial aggregation or dispersion trend between cells of the same type, and the spatial interaction relationship between different types of cells. The spatial position distribution relationship of cells is reflected in a numerical form, realizing an effective characterization of the spatial distribution patterns of different cell types in the tumor microenvironment, and revealing the relationship between cell spatial distribution and tumor progression and treatment response. The cell types may specifically include: tumor cells, immune cells and stromal cells.
[0048] S500: Based on the first cell spatial distribution feature and the second cell spatial distribution feature, determine the tumor state to achieve the typing of the target object.
[0049] Combining the first cell spatial distribution feature and the second cell spatial distribution feature enriches the analysis dimension. On the basis of the original cell morphology, quantity and position, the adjacency relationship, action intensity and distribution trend between cells are quantitatively processed to accurately quantify the geometric relationship and spatial interaction between cells, realizing multi-dimensional quantitative analysis of the cell spatial distribution in the tumor microenvironment, helping to obtain more comprehensive cell distribution information, enhancing the understanding of the tumor microenvironment, facilitating understanding the role of the interaction between cells in the tumor microenvironment in tumor development, immune escape and metastasis, improving the accuracy of the analysis results, helping to predict the dynamic role and development of cells in the tumor microenvironment, and further improving the accuracy of the typing of the target object. Specifically, in some embodiments of the present application, the pathological section image is from the target object, and the target object may be a patient.
[0050] Furthermore, the cell topological network includes: the first cell spatial topological network and / or the second cell spatial topological network;
[0051] The first cell space topological network is constructed based on the relevant principles of the Delaunay triangulation network, and the second cell space topological network is constructed through the Voronoi diagram. Specifically, the relevant principles of the Delaunay triangulation network include: the Delaunay criterion, which uses geometric conditions to determine whether a pair of adjacent triangles represents an optimal connection choice. If a pair of triangles does not meet this criterion, the connection method will be readjusted; Delaunay triangulation, which divides its convex hull into several triangles, requiring that the circumcircles of these triangles do not contain any other points. This process maximizes the size of the smallest angle in each triangle and tends to avoid the appearance of overly elongated triangles; the Delaunay triangulation is not unique, and the two possible triangulations of dividing a quadrilateral into two triangles both satisfy the "Delaunay condition", that is, the interiors of the circumcircles of all triangles are empty. More specifically, the specific process of constructing the first cell space topological network includes: by obtaining the centroid of each cell nucleus, a spatial topological network is constructed using the Delaunay triangulation algorithm, where each cell nucleus is connected to its adjacent cell nuclei by edges; through the relevant principles of the Delaunay triangulation network, an optimal network structure is provided for the interaction and transmission between cells, as Figure 3 shown.
[0052] Specifically, the Voronoi diagram divides the plane into regions close to a given set of objects, and it can also be classified as a kind of tessellation; these objects are only a finite number of points in the plane (called seeds, sites, or generators); for each seed, there is a corresponding region, called a Voronoi cell, which contains all the points in the plane that are closer to this seed than to other seeds. That is, the Voronoi diagram divides the space into multiple regions according to the nearest attribute of the objects, and the distance from any point of the convex polygon to the object point in the region it contains is less than the distance to any other object point; by regarding cells as discrete points in space, the Voronoi diagram can effectively divide the cell space, making the proximity and relative position of each cell clear, as Figure 4 shown.
[0053] The first cell space distribution characteristics obtained through the first cell space topological network include one or any combination of the proportion of each type of cell link, the chain edge distance, and the link frequency;
[0054] The first cell space distribution characteristics obtained through the second cell space topological network include one or any combination of the area, area variance, mean, median, and clutter degree of each cell pixel.
[0055] By constructing the cell topological network respectively based on the relevant principles of the Delaunay triangulation network and the Voronoi diagram, the differences in the characteristics of the Delaunay triangulation network and the Voronoi diagram are fully utilized to realize the separate acquisition of different characteristic parameters such as cell adjacency relationship, interaction intensity, and spatial topological structure;
[0056] For example, the Delaunay triangulation network can accurately represent the adjacency relationships of cells, optimize the connections between cells, provide an efficient communication structure and mechanical framework, simplify the complex cell spatial structure, and also provide a basis for in-depth research on cell population behavior, cell communication, cell mechanics, etc. It is applicable to the analysis of the pathways of material diffusion between cells and the links of signal propagation. Another example is that the Voronoi diagram can help understand the distribution of cells in space, and can also be used to calculate the range or influence area of cell interactions, and is applicable to the analysis of the tumor immune escape mechanism; on the basis of identifying the aggregation and dispersion patterns of cells, the functional roles and ecological niches of cells can be inferred, and spatial characteristic information such as the area, variance, mean, median, etc. of each cell pixel polygon can be calculated. Among them, the area fluctuation is the difference value of the polygon area here.
[0057] Furthermore, the first cell space topological network includes simplices, and the union of all simplices constitutes the convex hull of cell distribution. The triangulation of any cell position contains no more than triangles;
[0058] Among them, represents the number of cells, represents the spatial dimension, represents the number of cells in the convex hull;
[0059] A simplex is a geometric and mathematical concept, usually used to describe a geometric shape formed by several linearly independent points; a convex hull refers to a given set of points, and is the smallest convex polygon containing this set of points. Specifically, if the obtained spatial information is in a two-dimensional plane ( ), and there are cells on the convex hull, then the triangulation of any cell position contains at most triangles, and additionally contains an external face. If the cell positions are distributed according to a Poisson process with a constant intensity, then each cell has an average of six adjacent triangles; more generally, in -dimensional space, the same Poisson process makes the average number of neighbors of each cell depend only on .
[0060] In the cell network, neighboring cells often have a higher probability of interaction, such as signal transmission or material exchange, etc.; at the same time, the first cell space topological network follows the criterion of "maximizing the minimum angle", avoiding long and narrow triangles, thus forming a geometrically optimal connection structure; in the cell space network, this structure minimizes unnatural long-distance connections and largely ensures that the connections between cells reflect the true topological relationships.
[0061] In the second cell space topological network, the Voronoi cell associated with the cell center is defined as:
[0062] ;
[0063] wherein, is a metric space, d is the distance function, K represents the index set defining the centers of a group of cells, that is, the index set of the sites, represents a group of non-empty cells in the space X, represents the cell center, represents the Voronoi cell associated with the cell center.
[0064] The Voronoi cell contains all those points that are closest to the cell center, that is, all points in the plane whose distance to the cell is less than or equal to the distance to other cell centers. The Voronoi cell of each cell defines a unique region in the space, and all points within this region are more likely to correspond to this cell type; this structure of the second cell space topology network can not only help understand the distribution of cells in the space, but also be used to calculate the range or influence area of cell interactions, and thus can clearly characterize the interactions and spatial distributions between cells.
[0065] Furthermore, the acquisition of the second cell spatial distribution characteristics specifically includes:
[0066] Based on the spatial weight matrix, obtain the grid adjacency weights after meshing the tumor microenvironment;
[0067] Among them, specifically, in some embodiments of the present application, the construction steps of the spatial weight matrix are as follows: Grid division, divide each 2560×2560 pixel pathological image block into 40×40 uniform grids, each grid being 64×64 pixels. Adjacency definition, if two grids share an edge or a corner, that is, the Queen adjacency rule, then they are defined as spatially adjacent, and the weight = 1, otherwise = 0. Row normalization, perform row normalization on the weight matrix so that the sum of the adjacency weights of each grid is 1, that is, = 1.
[0068] Obtain the number of cells in each grid as the grid cell number information;
[0069] Combine the grid adjacency weights with the grid cell number information, analyze the spatial autocorrelation of the same type of cells in the tumor microenvironment and the spatial autocorrelation between different types of cells in the tumor microenvironment, and thereby quantify the spatial interactions between cell types.
[0070] By quantitatively analyzing the spatial autocorrelation of the same type of cells and the spatial autocorrelation between different types of cells, the spatial distribution trend of cells in the tumor microenvironment can be revealed, thereby analyzing the spatial interaction between different types of cells, such as the mutual relationship between tumor cells and immune cells, stromal cells, and further improving the accuracy of describing the spatial heterogeneity and cell interaction pattern in the tumor microenvironment.
[0071] Furthermore, the spatial autocorrelation of the same type of cells in the tumor microenvironment is analyzed based on the univariate Moran's index; the calculation formula of the univariate Moran's index is:
[0072] ;
[0073] Among them, represents the number of grids after grid processing of the tumor microenvironment, represents the adjacency weight between grid i and grid j, and respectively represent the number of a certain type of cells in grid and grid ; then represents the average value of the number of a certain type of cells in all grids; specifically, in some embodiments of the present application, the value of
[0074] includes "0" and "1", "0" means non-adjacent, and "1" means adjacent. This index reflects the degree of aggregation or dispersion of a certain type of cells in space. The cell types include the aforementioned tumor cells, stromal cells, and immune cells; if the value of the index is close to 1, it means that this type of cells has a high degree of aggregation in space; if the value of the index is close to -1, it means that this type of cells is highly dispersed in space; and if the value of the index
[0075] is close to 0, it means that the distribution of this type of cells is random.
[0076] ;
[0077] Among them, represents the number of grids after grid processing of the tumor microenvironment, represents the number of the first type of cells in the th grid, represents the number of the second type of cells in the th grid, is an element in the weight matrix, representing the adjacency weight between grids, represents the mean of the first type of cells, ; represents the mean of the second type of cells, ; represents the spatial lag value, .
[0078] If is equal to 1, then grids and grid are adjacent. If is equal to 0, then grids and grid are not adjacent. The bivariate Moran's I is used to measure the degree of aggregation or dispersion of two types of cells in the tumor microenvironment space. The spatial lag value represents the weighted average of the number of neighborhood cell types of a certain grid. For the cross-correlation between the first type of cells (which can be tumor cells) and the second type of cells (which can be immune cells), the spatial lag value is the weighted average of the number of the third type of cells (which can be stromal cells); the basis for weighting is the spatial adjacency relationship between grids, using the Queen adjacency method, that is, when two grids share an edge or a corner, they are considered adjacent.
[0079] When calculating the lag value, the weight matrix weights the neighborhood of each grid. The weight of adjacent grids in the weight matrix is 1, and the weight of non-adjacent grids is 0; based on the foregoing settings, for each grid, the spatial lag value of the third type of cells is calculated, that is, the weighted average of the number of stromal cells in all grids around the grid. Specifically, in some embodiments of the present application, the Moran's I (autocorrelation of like and unlike cells) of each grid is obtained through grid division; and the results of all valid grids are aggregated to calculate the average Moran's I, that is, the average value of the Moran's I of each grid is taken, and invalid values are skipped, that is, cell-free grids or background grids.
[0080] If the bivariate Moran's I is >0 and close to 1, it indicates that the first type of cells and the second type of cells are spatially aggregated, that is, these cell types tend to cluster together in certain regions, and thus there may be an interaction between the first type of cells and the second type of cells; if the bivariate Moran's I is <0 and close to -1: it indicates that the first type of cells and the second type of cells are spatially dispersed, that is, these cell types tend to be far away from each other; if the bivariate Moran's I is close to 0: it indicates that the first type of cells and the second type of cells are randomly distributed in space, without a significant spatial pattern.
[0081] Based on the above conclusions, for example, when tumor cells and stromal cells are spatially aggregated, it may be due to the stromal support required for tumor growth; for another example, when tumor cells and stromal cells are spatially dispersed, it may be due to immune responses and tumor suppression; thus, it is used for the preliminary judgment of tumor status and tumor microenvironment.
[0082] Furthermore, based on the first cell spatial distribution feature and the second cell spatial distribution feature, the tumor status is judged to classify the target object, which specifically includes:
[0083] Preset classification feature parameters are respectively selected from the first cell spatial distribution feature and the second cell spatial distribution feature;
[0084] Specifically, in some embodiments of the present application, a total of 115 feature parameters are set for the first cell spatial distribution feature and the second cell spatial distribution feature to capture the spatial features of the spatial arrangement and structural features of the tumor microenvironment; there are 109 in the first cell spatial distribution feature. For example, the margins of the topological network obtained based on the spatial positions (i.e., cell positions) of the graph nodes of the first cell spatial topological network and the second cell spatial topological network in the graph measurement total 35, which are used to describe the edge distances of the cell chains in the first cell spatial topological network; the frequency of the edges total 6, which are used to describe the proportion of the cell chain edges in the first cell spatial topological network; the absolute cell number total 11, which are used to describe the absolute number of cell chain edges and the absolute number of cells; the coefficient of variation of the cell polygon area (CellCV) total 4, which are used to describe the coefficient of variation of the polygon regions in the second cell spatial topological network; the cell neighborhood range total 16, which are used to describe the average distance (mean, median, variance, standard deviation) between a specific type of cell and all its neighbors; the cell polygon area total 18, which are used to describe the polygon regions of the second cell spatial topological network graph; the cell density total 3, which are used to describe the cell density; and the cell diversity (CellDiv) around the cells total 16, which are used to describe the cell diversity of all types of cells.
[0085] The second cell spatial distribution feature includes a total of 6 slice Moran indices; among them, the above slices are arbitrarily selected from the WSI (whole-slide digital sections) of each patient according to a preset quantity, so the topological features of each patient are calculated based on the corresponding values of all the slices selected from the same WSI. Of course, no limitation is made on the selection of feature parameters here.
[0086] For another example, the calculation formula for the degree of cell diversity is as follows:
[0087] ;
[0088] is the target cell of the number of cell types of the is the target cell number of nearest neighbor cells.
[0089] When the target cell and the nearest neighbor cell belong to different types, = 1, otherwise 0. The cell diversity heterogeneity being 0 (minimum value) indicates that there are no other types of cells as neighbors, while the cell diversity heterogeneity being 1 (maximum value) indicates that there are no cells of the same type as neighbors; the cell diversity of each cell type is calculated in this way.
[0090] The calculation formula for the variance of cell influence is as follows: ;
[0091] wherein, represents the standard deviation of the area of the Voronoi diagram polygon, represents the average area of the Voronoi diagram polygon of a specific cell type.
[0092] When the spatial distribution of cells and their neighbors is relatively consistent, the area change of the Voronoi diagram polygon is small and the CellCV value is also low, and vice versa. When the point set is randomly distributed, the CellCV value is between 33% and 64%, indicating a random distribution; when the point set is aggregated, the CellCV value is greater than 64%, indicating an aggregated distribution; when the CellCV value is less than 33%, the point set is uniformly distributed; the variance of cell influence of each cell type is calculated in this way.
[0093] The preset typing feature parameters include: the chain edge distance between tumor cells and immune cells, the standard deviation of the chain edge distance between stromal cells and stromal cells, the variance of the chain edge distance between stromal cells and stromal cells, the number of chain edges between tumor cells and immune cells, the number of triangles in the triangular network, the number of polygons, the standard deviation of the chain edge distance between tumor cells and immune cells, the variance of the chain edge distance between tumor cells and immune cells, one or any combination of them;
[0094] Obtain the clinical information of the target object, and the clinical information of the target object can include age grouping information and the overall cancer staging information; of course, no restrictions are imposed on the specific content of the clinical information of the target object. Exemplarily, it can also include the emotions and physical conditions of the target object, etc.; the overall cancer staging includes: the first stage, Stage I, early cancer, usually local and not spread; the second stage, Stage II and the third stage, Stage III, the cancer is relatively local or has spread to adjacent lymph nodes or tissues; the fourth stage, Stage IV: the cancer has metastasized distantly, usually indicating advanced cancer.
[0095] Establish an association by combining the clinical information of the target object, the classification feature parameters, and the survival period of the target object, and judge the tumor status to achieve the classification of the target object; Exemplarily, judge the tumor status to achieve the classification of the target object; This can be achieved by constructing a classification model for the target object;
[0096] Specifically, based on the influence of the classification feature parameters and the clinical information of the target object on the survival period of the target object, establish connections with the survival period of the target object and the tumor status respectively; Exemplarily, for example, the feature "standard deviation of the edge distance between tumor and immunity" has an almost positive correlation with the survival period of the target object, and the survival period of the target object with a high standard deviation is significantly higher than that of the target object with a low standard deviation; Another example is that the feature "standard deviation of the edge distance between stroma and stroma" also has an almost positive correlation with the survival period of the target object; At the same time, a large number of training samples are given, and the influence weights and scores of the classification feature parameters are set and optimized. For example, a positive correlation can be set as a positive value, and a negative correlation can be set as a negative value;
[0097] More specifically, through a large number of training samples, adjust the influence weights and scores of the classification feature parameters until the target accuracy is met. Calculate the disease score based on all the selected classification feature parameters combined with the clinical information of the target object, and set and optimize the scoring threshold, establish the connection between the scoring threshold and the survival period and tumor status of the target object, and then, based on the tumor status, achieve the classification of the target object; Further, based on the historical changes of the disease score of the target object, judge the dynamic state changes of the tumor; The dynamic changes of the tumor include: one or any combination of tumor development, deterioration and metastasis, and static stability; Exemplarily, a connection can be established between the change of the disease score and the dynamic change of the tumor. If the change of the disease score of the target object is small, it indicates that the tumor status is relatively stable; Another example is that if the disease score of the target object drops sharply, it indicates that the tumor status has deteriorated.
[0098] Furthermore, the selection of the preset parameter features specifically includes:
[0099] Perform KM survival analysis and Cox proportional hazards analysis based on the overall survival and progression-free survival of the target object; among them, KM survival analysis (Kaplan-Meier survival analysis) is a statistical method commonly used in biomedical and clinical research, mainly used to estimate and compare the survival rates or survival times of different groups; Cox proportional hazards model (Cox Proportional Hazards Model) is a regression analysis method commonly used in medicine and biostatistics, mainly used to evaluate the impact of different variables (such as treatment methods, clinical characteristics, gene mutations, etc.) on the survival of cancer patients; the Cox model is particularly suitable for survival data analysis and can model the relationship between the survival time of patients and a series of predictive factors without assuming the data distribution.
[0100] Specifically, in some embodiments of the present application, the analysis of the relationship between 115 spatial parameter characteristics and the classification of the target object is implemented based on the following method: sort all the values of a certain type of spatial characteristic parameter according to the numerical value, and select a cut-off value to group the patients corresponding to each value, which can be specifically divided into a high group and a low group, and perform KM survival analysis and Cox proportional hazards analysis in combination with the clinical follow-up information of TCGA patients (including survival period, survival status, etc.) to obtain the relationship between the spatial characteristic parameter and the patient analysis or clinical prognosis prediction evaluation, so as to facilitate risk prediction and survival outcome prediction.
[0101] Judge whether each characteristic parameter in the first cell spatial distribution characteristic and the second cell spatial distribution characteristic meets the preset conditions; the preset conditions include the value ranges of one or any combination of the logrank-p value and the Cox-p value corresponding to the overall survival and progression-free survival of each characteristic parameter; the Log-rank test is a test method used to compare whether there are significant differences in the survival curves of two or more groups, mainly used to compare whether there are significant differences in the survival rates of different groups (such as treatment group and control group), the smaller the p value, the more significant the difference in the survival curves between groups; the Cox regression model (i.e., the Cox proportional hazards model) is a regression analysis method used to analyze the impact of multiple factors on the survival time, which can evaluate how one or more covariates affect the survival time, and the smaller the p value, the more significant the impact of the variable on the survival time; thus, judge whether there is an impact of the characteristic parameter on the target object and the size of the impact.
[0102] If it meets the preset conditions, use it as a classification characteristic parameter to obtain a more significant analysis effect in KM survival analysis and Cox proportional hazards analysis after grouping the target object under the optimal cut-off, and still obtain relatively good significance after correcting the p values obtained by KM and Cox through FDR correction.
[0103] Specifically, in some embodiments of the present application, there are four specific preset conditions: for the overall survival analysis, the log-rank p-value is less than 0.05 and the Cox-p value is less than 0.05; for the progression-free survival analysis, the log-rank p-value is less than 0.05 and the Cox-p value is less than 0.05. Of course, there is no limitation on the setting of the preset conditions. After performing FDR correction on each feature parameter, as shown in Table 1, the feature parameters that meet the preset conditions are respectively the edge distance between tumor cells and immune cells, the standard deviation of the edge distance between stromal cells and stromal cells, the variance of the edge distance between stromal cells and stromal cells, the number of edges between tumor cells and immune cells, the number of triangles in the triangular network, the number of polygons, the standard deviation of the edge distance between tumor cells and immune cells, and the variance of the edge distance between tumor cells and immune cells. Among them, FDR correction is a method used in multiple hypothesis testing to control the false positive rate (i.e., the proportion of wrongly rejecting the null hypothesis), aiming to reduce the number of false positive errors when multiple tests are conducted simultaneously.
[0104] Table 1 KM and Cox analysis results of classification feature parameters
[0105]
[0106] The close distance between immune cells and cancer cells may be a signal of a high level of immune cell infiltration in the tumor microenvironment, which means that there is a large amount of immune cell infiltration around the tumor, usually indicating an active interaction between the tumor and the immune system. A high degree of immune infiltration is associated with a good anti-tumor response and a better prognosis, especially in patients for whom immunotherapy is effective.
[0107] Exemplarily, as Figure 5 shown, Figure 5 The content is a survival curve of the "standard deviation of the edge distance between tumor and immune cells" based on the Kapan-Meier survival analysis. In the survival curve of the "standard deviation of the edge distance between tumor and immune cells", patients with a high standard deviation tend to have a higher long-term survival probability. A high standard deviation means a large difference in the distance between immune cells and tumor cells, which may reflect a more complex and heterogeneous immune response in the tumor microenvironment, such as a multi-level immune infiltration pattern; a low standard deviation indicates that the distribution of immune cells is more concentrated and uniform, possibly indicating that immune cells are aggregated in a specific area of the tumor and cannot effectively infiltrate the entire tumor tissue.
[0108] Exemplarily, as Figure 6 shown, Figure 6The content is a survival curve graph of "standard deviation of the edge-to-edge distance between stroma - stroma" based on Kaplan - Meier survival analysis. In the survival curve of "standard deviation of the edge-to-edge distance between stroma - stroma", it shows that the survival probability of the high standard deviation group is significantly higher than that of the low standard deviation group; a high standard deviation means that the distance differences between stromal cells are large, the distribution is uneven, the cell density is high in some areas, while it is sparse in other areas. This high spatial heterogeneity of stromal cells limits the overall progression of tumors. On the contrary, a low standard deviation indicates that the distance differences between stromal cells are small and the cell distribution is relatively uniform. This uniform cell distribution may provide a more stable supportive environment for tumors, promote their progression, and result in a lower survival probability. Among them, the standard deviation of the distance between stromal cells is closely related to the death risk of the target object. The distribution and behavior of stromal cells can affect the progression of tumors and the prognosis of the target object. In addition, the applicant also systematically evaluated whether the classification characteristic parameters can be independently used for the classification of the target object based on the multi-factor Cox regression analysis. During the experiment, the "gender", "age", "tumor stage", "tumor type", etc. in the clinical follow-up texts corresponding to 158 TCGA patients were combined with the spatial characteristics corresponding to the classification characteristic parameters for multivariate analysis. Among them, the p-value of the variance (standard deviation) of the edge-to-edge distance between stroma - stroma is less than 0.1, and it has a relatively good effect for the classification of the target object. The experimental results are as Figure 9 shown.
[0109] In addition, the applicant also grouped 72 patients in a cancer prevention and treatment center of a university using the cut-off values of the spatial characteristics corresponding to the classification characteristic parameters calculated and screened from TCGA data. The experimental results also show that the above-mentioned classification characteristic parameters have a significant effect on the prognosis of patients.
[0110] Furthermore, the acquisition of cell segmentation and classification information specifically includes: segmenting the pathological section image according to a preset specification to obtain sub-images of the pathological section; specifically, in some embodiments of the present application, the specification of the pathological section image is 2560*2560, and the specification of the sub-image of the pathological section is 256*256; of course, no limitation is imposed on the size of the image rules. Identifying the cells in each sub-image of the pathological section through a pre-trained cell recognition model to obtain the information of cell segmentation and classification; the acquisition of cell spatial distribution information specifically includes: restoring the pathological section image by in-situ backtracking of each sub-image of the pathological section; based on the restored pathological section image and the information of cell segmentation and classification, obtaining the cell spatial distribution information of the pathological section image.
[0111] Specifically, in some embodiments of the present application, some patients with a last follow-up time exceeding 3 years were selected as the analysis objects, a total of 158 cases; and the pathological sections of each patient were segmented into a number of pathological section images of 2560*2560. Then, 10 pathological section images of 2560*2560 were randomly selected from each patient for image cutting and recognition, segmented into 100 pathological section sub-images of 256*256. The STARDIST model was used to recognize each pathological section sub-image, and then the spatial information of the cells in each pathological section image was sorted out by the in-situ backtracking method to obtain the cell spatial distribution information of the pathological section image; based on this, a cell topology network was constructed. Finally, the cell spatial distribution information of the randomly selected 10 pathological section images was averaged to obtain the calculated value of the final cell spatial distribution information of the target object.
[0112] Further, the cell recognition model is trained based on the STARDIST model, specifically including:
[0113] Obtain and preprocess the training sample set, which includes multiple training samples; specifically, in some embodiments of the present application, the Lizard dataset is used, a total of 4981 non-overlapping image patches with a size of 256×256, containing 535063 labeled cell nuclei. 90% of the data is divided into the training set, and 10% of the data is divided into the validation set. The training set contains 4482 images, and the validation set contains 499 images;
[0114] Since the number of different cell types varies greatly, for example, the epithelial cell category accounts for 64.35%, while the neutrophil category only accounts for 0.79%. This means that if the model does not perform class balancing processing, the training may tend to predict the majority class (such as epithelial cells), thus ignoring the minority classes (such as neutrophils and eosinophils). Therefore, the applicant adopted the method of oversampling to deal with the imbalance problem in the dataset. The optimized training cell distribution is shown in Table 2. Of course, no limitation is made on the specific dataset type and data sample size.
[0115] Table 2 Cell distribution in the training set after oversampling
[0116]
[0117] The input data includes: training images in multiple training samples;
[0118] The output data includes: the classification of each cell and the number of cells of each type in the corresponding training image;
[0119] Using the input data and output data, perform deep learning model training and optimize the parameter settings of the deep learning model. If the cell classification and the number of each type of cell in the training image output by the deep learning model meet the preset goals, end the training to obtain a cell recognition model.
[0120] Specifically, in some embodiments of the present application, the cell recognition model includes 4 encoder and decoder layers, each layer applies two convolution operations, the convolution kernel size is 3×3, the initial number of convolution kernels is 64, and each layer is downsampled through a 2×2 pooling layer; the activation function uses ReLU, and a convolution layer with 256 convolution kernels is added after the UNet for further processing. Specifically, in some embodiments of the present application, optimizing the parameter settings of the deep learning model can be a threshold, that is, a decision boundary, used to convert the predicted output of the model into the final classification label or decision. The applicant performs 80 iterations of training on the deep learning model, and optimizes the threshold of the model based on the metrics commonly used in machine learning model evaluation and performance analysis, such as total loss, probability loss, classification loss, mean square error, etc. The candidate thresholds are 0.1, 0.2, 0.3. Compare the values in the probability map with the current threshold to generate a binary mask. For each binary mask, apply non-maximum suppression to remove overlapping predictions; this can reduce duplicate detections and keep the most likely bounding box or segmentation region; return the optimal threshold configuration as prob_thresh = 0.499797 and nms_thresh = 0.2. When the prediction probability of the model for a certain pixel or target is greater than or equal to 0.499797, this pixel or target will be regarded as a positive class (for example, the target exists), otherwise it will be regarded as a negative class (for example, the target does not exist); when the IoU of two bounding boxes is greater than this threshold, they are considered overlapping and one of them is retained.
[0121] More specifically, the cell recognition effect of the cell recognition model is referred to Figure 10 and Figure 11 as shown.
[0122] In addition, the applicant also selected 400 CRC images of 256*256 from other datasets as the test set, and paired the ground truth and prediction results of the 400 image sets to visualize the relationship between the IoU threshold and the performance metrics ( Figure 2-3), it can be found that on the test set, when the IoU threshold (nms_thresh) is set to 0.2, the model can still perform quite well in recognition ability. Among them, TP (True Positive, true positive example) represents that the IoU is greater than the threshold, and the predicted box overlaps with the ground truth box, indicating that the target is correctly detected; FP (False Positive, false positive example) represents that the IoU is less than or equal to the threshold, and the predicted box exists, but there is no corresponding ground truth, indicating a false detection (detected but there is actually no target); FN (False Negative, false negative example) represents that the IoU is less than the threshold, and there is a ground truth box, but there is no corresponding predicted box, indicating a missed detection (the target actually exists but is not detected). The specific recognition effect of the cell recognition model at different thresholds is as Figure 7 and Figure 8 shown.
[0123] The cell recognition model is trained based on the STARDIST model architecture. The shape of the cell nucleus is represented by a star-shaped convex polygon parameterized by the radial distance, which can effectively handle the overlap between cells, accurately locate the position of each cell, and provide a more accurate cell morphology; it overcomes the problem of large segmentation errors when dealing with high-density cell populations, making the cell segmentation results in a complex tumor microenvironment more accurate, and can also avoid the influence brought by subjective factors to a certain extent, effectively improving the reliability, consistency and objectivity of the recognition results, and at the same time significantly improving the recognition speed and recognition efficiency.
[0124] The embodiments of the present application at least have the following beneficial effects:
[0125] The embodiment of the present application provides a method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells, which has good cell segmentation and classification accuracy, fully pays attention to the adjacency relationship, interaction intensity and spatial structure between cells, comprehensively analyzes the quantitative spatial distribution characteristics in multiple dimensions, realizes good analysis results of tumor microenvironment cells, is used to accurately classify the target object, and has good versatility; specifically, the preset cell recognition model can effectively handle the overlap between cells when recognizing, segmenting and classifying cells, accurately locate the position of each cell, and provide a more accurate cell morphology, improving the cell segmentation accuracy; constructing a cell topological network to obtain the first cell spatial distribution characteristics can accurately model the adjacency relationship and interaction between cells in the tumor microenvironment at the cell scale, and then comprehensively reveal the spatial structure of the tumor microenvironment, increasing the spatial dimension of the analysis of the tumor microenvironment; constructing a spatial weight matrix, combining the univariate Moran index and the bivariate Moran index, to obtain the second cell spatial distribution characteristics, quantifying the spatial distribution pattern between different cell types in the tumor microenvironment, revealing the spatial aggregation or dispersion trend between cell types, the univariate Moran index can be used to analyze the spatial autocorrelation of the same type of cells, while the bivariate Moran index helps to analyze the spatial interaction between different cell types, providing a new perspective for the functional research of the tumor microenvironment, and then helping to capture the dynamic changes of cells in advance.
[0126] Embodiment 2
[0127] The embodiment of the present application provides a system for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells, which is applied to the method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells in Embodiment 1, as Figure 2 shown, including:
[0128] Cell recognition and segmentation module 1: used to obtain the pathological section image and, through the pre-trained cell recognition model, obtain the information about cell segmentation and classification in the pathological section image;
[0129] Distribution information acquisition module 2: used to obtain the cell spatial distribution information of the pathological section image based on the information about cell segmentation and classification;
[0130] First feature acquisition module 3: used to construct a cell topological network according to the cell spatial distribution information and obtain the first cell spatial distribution characteristics of the tumor microenvironment;
[0131] Second feature acquisition module 4: used to construct a spatial weight matrix according to the cell spatial distribution information and obtain the second cell spatial distribution characteristics of the tumor microenvironment;
[0132] Judgment and classification module 5: used to judge the tumor state based on the first cell spatial distribution characteristics and the second cell spatial distribution characteristics to realize the classification of the target object.
[0133] Further, the cell topological network includes: a first cell space topological network and / or a second cell space topological network; the first cell space topological network is constructed based on the relevant principles of the Delaunay triangulation network, and the second cell space topological network is constructed by the Voronoi diagram; the first cell space distribution characteristics obtained through the first cell space topological network include: one or any combination of the proportion of each type of cell link, the chain edge distance, and the link frequency; the first cell space distribution characteristics obtained through the second cell space topological network include: one or any combination of the area of each cell pixel, the area variance, the mean, the median, and the clutter degree.
[0134] Further, the first cell space topological network includes simplices, and the union of all simplices constitutes the convex hull of the cell distribution, and the triangulation of any cell position contains no more than triangles; where represents the number of cells, represents the spatial dimension, represents the number of cells in the convex hull; in the second cell space topological network, the Voronoi cell associated with the cell center is defined as: ; where is a metric space, d is a distance function, K represents the index set defining the centers of a group of cells, that is, the index set of the sites, represents a set of non-empty cells in the space X, represents the cell center, represents the Voronoi cell associated with the cell center.
[0135] Further, the acquisition of the second cell space distribution characteristics specifically includes: based on the spatial weight matrix, obtaining the grid adjacency weights after the tumor microenvironment is gridded; obtaining the number of cells in each grid as the grid cell number information; combining the grid adjacency weights and the grid cell number information to analyze the spatial autocorrelation of the same type of cells in the tumor microenvironment and the spatial autocorrelation between different types of cells in the tumor microenvironment.
[0136] Further, the analysis of the spatial autocorrelation of the same type of cells in the tumor microenvironment is realized based on the univariate Moran index; the calculation formula of the univariate Moran index is:
[0137] ;
[0138] where represents the number of grids after the tumor microenvironment is gridded, represents the adjacency weight between grid i and grid j, and respectively represent the grids and the grid the number of a certain type of cell in then represents the mean value of the number of a certain type of cell in all grids;
[0139] Analyze the spatial autocorrelation between different types of cells in the tumor microenvironment, which is achieved based on the bivariate Moran's index; the calculation formula of the bivariate Moran's index is:
[0140] ;
[0141] Among them, represents the number of grids after the meshing process of the tumor microenvironment, represents the number of the first type of cells in the th grid, represents the number of the second type of cells in the th grid, is an element in the weight matrix, representing the adjacency weight between grids, ; represents the mean value of the first type of cells, ; represents the spatial lag value, .
[0142] Furthermore, based on the spatial distribution characteristics of the first type of cells and the spatial distribution characteristics of the second type of cells, judge the tumor state to classify the target object, specifically including: respectively select the preset classification characteristic parameters from the spatial distribution characteristics of the first type of cells and the spatial distribution characteristics of the second type of cells; the preset classification characteristic parameters include: the chain edge distance between tumor cells and immune cells, the standard deviation of the chain edge distance between stromal cells and stromal cells, the variance of the chain edge distance between stromal cells and stromal cells, the number of chain edges between tumor cells and immune cells, the number of triangles in the triangular network, the number of polygons, the standard deviation of the chain edge distance between tumor cells and immune cells, the variance of the chain edge distance between tumor cells and immune cells, one or any combination of them; obtain the clinical information of the target object; combine the clinical information of the target object, the classification characteristic parameters and the association established with the survival period of the target object to judge the tumor state, so as to realize the classification of the target object; the tumor state includes: one or any combination of being in situ, local expansion and metastasis.
[0143] Further, the selection of the preset parameter features specifically includes: performing KM survival analysis and Cox proportional hazards analysis based on the overall survival time and progression-free survival time of the target object; determining whether each feature parameter in the first cell spatial distribution feature and the second cell spatial distribution feature meets the preset conditions; the preset conditions include the value ranges of one or any combination of the logrank-p value and Cox-p value corresponding to the overall survival time and progression-free survival time of each feature parameter; if it meets the preset conditions, it is used as the classification feature parameter.
[0144] Further, the acquisition of cell segmentation and classification information specifically includes: segmenting the pathological section image according to a preset specification to obtain sub-images of the pathological section; identifying the cells in each sub-image of the pathological section through a pre-trained cell recognition model to obtain cell segmentation and classification information; the acquisition of cell spatial distribution information specifically includes: restoring the pathological section image by in-situ backtracking of each sub-image of the pathological section; based on the restored pathological section image and the cell segmentation and classification information, obtaining the cell spatial distribution information of the pathological section image.
[0145] Further, the cell recognition model is trained based on the STARDIST model, specifically including:
[0146] Obtaining and preprocessing a training sample set, where the sample set includes multiple training samples; the input data includes: training images in multiple training samples; the output data includes: the classification of each cell in the corresponding training image and the number of cells of each type; using the input data and output data to perform deep learning model training and optimizing the parameter settings of the deep learning model. If the classification of each cell in the training image output by the deep learning model and the number of cells of each type meet the preset target, the training is ended to obtain the cell recognition model.
[0147] Since the technical problem to be solved in this embodiment, the technical solution used to solve this technical problem, and the technical effects achieved are all the same as those of the method described in Embodiment 1. For the technical details of the system in this embodiment, reference can be made to Embodiment 1, which will not be elaborated here.
[0148] Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the implementation scenario description, or can be correspondingly changed and located in one or more devices different from this implementation scenario.
[0149] The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules. The serial numbers of the present invention above are only for description and do not represent the advantages or disadvantages of the implementation scenario.
[0150] The above are only several specific implementation scenarios of the present invention. However, the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art shall fall within the protection scope of the present invention.
Claims
1. A method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells, characterized in that, Including: Obtain a pathological section image, and through a pre-trained cell recognition model, obtain cell segmentation and classification information in the pathological section image; The acquisition of cell segmentation and classification information specifically includes: Segment the pathological section image according to a preset specification to obtain a pathological section sub-image; Through the pre-trained cell recognition model, identify the cells in each pathological section sub-image to obtain the cell segmentation and classification information; Based on the cell segmentation and classification information, obtain the cell spatial distribution information of the pathological section image; The acquisition of cell spatial distribution information specifically includes: Restore the pathological section image by in-situ backtracking of each pathological section sub-image; Based on the restored pathological section image and the cell segmentation and classification information, obtain the cell spatial distribution information of the pathological section image; According to the cell spatial distribution information, construct a cell topological network to obtain the first cell spatial distribution feature of the tumor microenvironment; The cell topological network includes: a first cell spatial topological network and / or a second cell spatial topological network; The first cell spatial topological network is constructed based on the relevant principles of the Delaunay triangulation network, and the second cell spatial topological network is constructed through the Voronoi diagram; The first cell spatial distribution feature obtained through the first cell spatial topological network includes: one or any combination of the proportion of each type of cell link, the chain edge distance, and the link frequency; The first cell spatial distribution feature obtained through the second cell spatial topological network includes: one or any combination of the area of each cell pixel, the area variance, the mean value, the median value, and the mixing degree; According to the cell spatial distribution information, construct a spatial weight matrix to obtain the second cell spatial distribution feature of the tumor microenvironment; The acquisition of the second cell spatial distribution feature specifically includes: Based on the spatial weight matrix, obtain the grid adjacency weight after grid processing of the tumor microenvironment; Obtain the number of cells in each grid as grid cell number information; Combine the grid adjacency weight and the grid cell number information to analyze the spatial autocorrelation of the same type of cells in the tumor microenvironment and the spatial autocorrelation between different types of cells in the tumor microenvironment; Based on the first cell spatial distribution feature and the second cell spatial distribution feature, judge the tumor state to achieve the typing of the target object; Based on the first cell spatial distribution feature and the second cell spatial distribution feature, judge the tumor state to achieve the typing of the target object, specifically including: Select preset typing feature parameters from the first cell spatial distribution feature and the second cell spatial distribution feature respectively; The preset typing feature parameters include: one or any combination of the chain edge distance between tumor cells and immune cells, the standard deviation of the chain edge distance between stromal cells and stromal cells, the variance of the chain edge distance between stromal cells and stromal cells, the number of chain edges between tumor cells and immune cells, the number of triangles in the triangular network, the number of polygons, the standard deviation of the chain edge distance between tumor cells and immune cells, and the variance of the chain edge distance between tumor cells and immune cells; Obtain the clinical information of the target object; Based on the association established between the clinical information of the target object, the classification characteristic parameters, and the survival period of the target object, judge the tumor status to achieve the classification of the target object; The tumor status includes: being in situ, locally extended, and / or metastatic.
2. The method for quantitatively analyzing the spatial distribution characteristics of cells in the tumor microenvironment according to claim 1, characterized in that The first cellular spatial topological network includes simplices, and the union of all simplices forms the convex hull of cell distribution. The triangulation of any cell position contains no more than triangles; Among them, represents the number of cells, represents the spatial dimension, represents the number of cells in the convex hull; In the second cell spatial topological network, the Voronoi unit associated with the cell center is defined as: ; Among them, is a metric space, and d is the distance function, represents the cell center, represents the Voronoi cell associated with the cell center.
3. The quantitative analysis method for the spatial distribution characteristics of tumor microenvironment cells according to claim 1, wherein Analyze the spatial autocorrelation of the same type of cells in the tumor microenvironment, which is achieved based on the univariate Moran index; The calculation formula of the univariate Moran index is: ; Among them, represents the number of grids after meshing the tumor microenvironment, represents the adjacency weight between grid i and grid j, and respectively represent the number of a certain type of cells in grid and grid ; represents the mean value of the number of a certain type of cells in all grids. Analyze the spatial autocorrelation between different types of cells in the tumor microenvironment, which is achieved based on the bivariate Moran index; The calculation formula of the bivariate Moran index is: ; Among them, represents the number of grids after the meshing of the tumor microenvironment, represents the th number of the first type of cells in the grid, represents the th number of the second type of cells in the grid, is an element in the weight matrix, representing the adjacency weight between grids, represents the mean value of the first type of cells, ; represents the mean value of the second type of cells, ; represents the spatial lag value, .
4. A method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells according to claim 1, characterized in that The selection of the preset parameter characteristics specifically includes: Perform KM survival analysis and Cox proportional hazards analysis based on the overall survival period and progression-free survival period of the target object; Judge whether each characteristic parameter in the first cell spatial distribution characteristic and the second cell spatial distribution characteristic meets the preset conditions; The preset conditions include the value ranges of one or any combination of the logrank-p value and Cox-p value corresponding to each of the characteristic parameters for the overall survival period and the progression-free survival period; If it meets the preset conditions, use it as the classification characteristic parameter.
5. A method for quantitatively analyzing the spatial distribution characteristics of tumor microenvironment cells according to claim 1, characterized in that, The cell recognition model is trained based on the STARDIST model, specifically including: Obtain and preprocess the training sample set, and the sample set includes multiple training samples; The input data includes: training images in multiple of the training samples; The output data includes: the classification of each cell in the training image and the number of cells of each type; Use the input data and the output data to train the deep learning model and optimize the parameter settings of the deep learning model. If the classification of each cell in the training image and the number of cells of each type output by the deep learning model meet the preset target, end the training to obtain the cell recognition model.
6. A quantitative analysis system for the spatial distribution characteristics of tumor microenvironment cells, which is applied to the quantitative analysis method for the spatial distribution characteristics of tumor microenvironment cells according to any one of claims 1 to 5, and is characterized in that, It includes: Cell recognition and segmentation module: used to obtain the pathological section image and, through the pre-trained cell recognition model, obtain the information on cell segmentation and classification in the pathological section image; The acquisition of the cell segmentation and classification information specifically includes: Segment the pathological section image according to the preset specifications to obtain pathological section sub-images; Through the pre-trained cell recognition model, identify the cells in each of the pathological section sub-images to obtain the cell segmentation and classification information; Distribution information acquisition module: used to obtain the cell spatial distribution information of the pathological section image based on the cell segmentation and classification information; The acquisition of the cell spatial distribution information specifically includes: Restore the pathological section image by in-situ backtracking of each of the pathological section sub-images; Based on the restored pathological section image and the cell segmentation and classification information, obtain the cell spatial distribution information of the pathological section image; First feature acquisition module: used to construct a cell topological network based on the cell spatial distribution information and obtain the first cell spatial distribution characteristic of the tumor microenvironment; The cell topological network includes: a first cell spatial topological network and / or a second cell spatial topological network; The first cell spatial topological network is constructed based on the relevant principles of the Delaunay triangulation network, and the second cell spatial topological network is constructed through the Voronoi diagram; The first cell spatial distribution features obtained through the first cell spatial topological network include: one or any combination of the proportion of each type of cell link, the edge distance of the chain, and the link frequency; The first cell spatial distribution features obtained through the second cell spatial topological network include: one or any combination of the area of each cell pixel, the area variance, the mean value, the median value, and the mixture degree; The second feature acquisition module: used to construct a spatial weight matrix based on the cell spatial distribution information and obtain the second cell spatial distribution features of the tumor microenvironment; The acquisition of the second cell spatial distribution features specifically includes: Based on the spatial weight matrix, obtaining the grid adjacency weights after grid processing of the tumor microenvironment; Obtaining the number of cells in each grid as the grid cell number information; Combining the grid adjacency weights and the grid cell number information to analyze the spatial autocorrelation of the same type of cells in the tumor microenvironment and the spatial autocorrelation between different types of cells in the tumor microenvironment; The classification judgment module: used to judge the tumor state based on the first cell spatial distribution features and the second cell spatial distribution features to achieve the classification of the target object; Judging the tumor state based on the first cell spatial distribution features and the second cell spatial distribution features to achieve the classification of the target object specifically includes: Respectively selecting preset classification feature parameters from the first cell spatial distribution features and the second cell spatial distribution features; The preset classification feature parameters include: one or any combination of the edge distance between tumor cells and immune cells, the standard deviation of the edge distance between stromal cells and stromal cells, the variance of the edge distance between stromal cells and stromal cells, the number of edges between tumor cells and immune cells, the number of triangles in the triangular network, the number of polygons, the standard deviation of the edge distance between tumor cells and immune cells, and the variance of the edge distance between tumor cells and immune cells; Obtaining the clinical information of the target object; Combining the clinical information of the target object, the classification feature parameters, and the established association with the survival period of the target object to judge the tumor state to achieve the classification of the target object; The tumor state includes: one or any combination of in situ, local expansion, and metastasis.
Citation Information
Patent Citations
Tumor microenvironment spatial relationship modeling system and method based on digital pathological image
CN114565919A
Tumor microenvironment analysis method integrating spatial multi-omics data
CN118866107A