Colon cancer lymph node metastasis risk prediction method and system based on image analysis

By using image analysis methods, combined with pathological sections and immunohistochemical images, adaptive clustering and multi-parameter fusion evaluation are performed to construct a distribution map of metastatic branch paths. This solves the problems of invasiveness and time consumption of traditional methods and achieves accurate prediction of lymph node metastasis risk.

CN121506482AInactive Publication Date: 2026-02-10SHENZHEN BAOAN DISTRICT PEOPLES HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511635253.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-02-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional methods for assessing lymph node metastasis in colon cancer rely on pathological examination and histological diagnosis after surgical resection. These methods are highly invasive, time-consuming, and may cause patients to miss the optimal treatment window. There is an urgent need for a non-invasive, rapid, and accurate predictive method.

Method used

Image-based analysis methods were employed, including acquiring pathological slide scan images for layer-by-layer downsampling and feature reconstruction, performing adaptive clustering analysis and multi-parameter fusion evaluation, combining immunohistochemical staining images for single-cell localization and counting, constructing a metastatic branch path distribution map, and conducting multi-regional lymphatic metastasis simulation and risk prediction.

Benefits of technology

It preserves pathological image details at different resolutions, enhances the global and local details of the analysis, provides more accurate lymph node metastasis risk assessment, can identify high-risk areas at an early stage, and helps to develop individualized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506482A_ABST
    Figure CN121506482A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image analysis, in particular to a colon cancer lymph node metastasis risk prediction method and system based on image analysis. The method comprises the following steps: acquiring a pathological section scanning image, carrying out layer-by-layer downsampling and feature reconstruction, and generating pathological section topological features; carrying out adaptive clustering analysis and multi-parameter fusion evaluation on the basis of the pathological section topological features to obtain independent migration evaluation values of individual tumor cells; performing tumor metastasis branch analysis according to the independent migration evaluation value, and constructing a metastasis branch path distribution diagram; extracting an immunohistochemical staining image of a patient to obtain immune cell space distribution data; and performing multi-region lymphatic metastasis simulation and metastasis risk situation prediction based on the immune cell spatial distribution data and the metastasis branch path distribution diagram to obtain a final metastasis risk assessment result. According to the method, the lymph node metastasis risk is accurately and efficiently analyzed, and the credibility of a pathological analysis result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image analysis, and in particular to a method and system for predicting the risk of lymph node metastasis in colon cancer based on image analysis. Background Technology

[0002] With the continuous advancement of medical imaging technology, especially the development of technologies such as computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound imaging, image analysis has gradually become an important research direction in the medical field for tumor diagnosis and treatment. Colorectal cancer, as one of the most prevalent malignant tumors worldwide, requires early diagnosis and precise treatment to improve patient survival rates. Lymph node metastasis in colorectal cancer is a key factor in determining cancer stage and assessing patient prognosis; therefore, accurately predicting the risk of lymph node metastasis in colorectal cancer is of significant practical importance for improving clinical treatment outcomes and optimizing treatment plans.

[0003] Traditional methods for assessing lymph node metastasis in colorectal cancer primarily rely on pathological examination and histological diagnosis after surgical resection. While these methods offer high accuracy, they have limitations: firstly, pathological examination typically requires invasive procedures to obtain samples, which are time-consuming and expensive; secondly, performing lymph node metastasis examination after surgical resection may not only increase the patient's physical burden but also potentially cause them to miss the optimal treatment window. Therefore, finding a non-invasive, rapid, and accurate method to predict lymph node metastasis in colorectal cancer has become a pressing issue in the field of tumor diagnosis. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a method and system for predicting the risk of lymph node metastasis in colon cancer based on image analysis, thereby resolving at least one of the aforementioned technical problems.

[0005] To achieve the above objectives, this invention provides a method for predicting the risk of lymph node metastasis in colorectal cancer based on image analysis, comprising the following steps: Step S1: Obtain the scanned images of pathological sections, perform layer-by-layer downsampling and feature reconstruction, and generate the topological features of the pathological sections; Step S2: Based on the topological features of pathological sections, perform adaptive clustering analysis and multi-parameter fusion evaluation to obtain the independent migration evaluation value of individual tumor cells; Step S3: Perform tumor metastasis branch analysis based on the independent migration assessment values ​​to construct a metastasis branch path distribution map; Step S4: Extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain spatial distribution data of immune cells; Step S5: Based on the spatial distribution data of immune cells and the distribution map of metastatic branch paths, perform multi-regional lymphatic metastasis simulation and predict the metastasis risk situation to obtain the final metastasis risk assessment results.

[0006] This specification provides an image analysis-based system for predicting the risk of lymph node metastasis in colorectal cancer, used to perform the image analysis-based method for predicting the risk of lymph node metastasis in colorectal cancer as described above, including: The image layering and parsing module is used to acquire scanned images of pathological slides, perform layer-by-layer downsampling and feature reconstruction, and generate topological features of pathological slides. The migration assessment module is used to perform adaptive clustering analysis and multi-parameter fusion assessment based on the topological features of pathological sections to obtain independent migration assessment values ​​for individual tumor cells. The metastasis path distribution module is used to perform tumor metastasis branch analysis based on the independent migration assessment value and construct a metastasis branch path distribution map. The immune cell calculation module is used to extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain the spatial distribution data of immune cells; The metastasis prediction module is used to simulate multi-regional lymphatic metastasis and predict metastasis risk based on the spatial distribution data of immune cells and the distribution map of metastasis branch paths, so as to obtain the final metastasis risk assessment results.

[0007] The beneficial effects of this invention are as follows: Through layer-by-layer downsampling and feature reconstruction, we can preserve the details of pathological images at different resolutions, especially key features such as cell morphology, arrangement, and tumor boundaries. Multi-scale processing ensures that various features from microscopic cells to tissue structures are fully captured. Low-resolution layers extract information on a large scale of tissue structures, while high-resolution layers focus on microscopic features at the cell level, enhancing the global and local details of the analysis. By combining multi-scale image features, we generate topological features of pathological sections, including indicators such as cell clustering, boundary irregularity, and spatial connectivity. These topological features can effectively reflect the spatial heterogeneity of tumors, providing a solid data foundation for subsequent metastasis risk prediction. Adaptive clustering analysis can perform more refined clustering of tumor regions based on the topological features of pathological sections, effectively distinguishing different tumor subregions and identifying potential high-risk migration areas. Each tumor cell can obtain an independent migration assessment value through clustering analysis, thus providing a quantitative basis for the migration ability of individual cells. By fusing multiple topological features (such as cell aggregation degree, boundary irregularity, and stromal invasion depth), this step can provide a more comprehensive and accurate assessment. Compared to single-feature assessments, fusion evaluations better reflect the complexity of tumor tissue, especially how tumor cells interact with their surrounding environment, thus predicting their potential metastatic potential. Tumor metastasis branch analysis can clearly track and depict the migration trajectory of tumor cells and their expansion paths within the body. The constructed metastasis branch path distribution map not only shows the direction of tumor spread but also reveals the spatial connectivity between different regions, providing an intuitive basis for further metastasis risk assessment. The metastasis path distribution map helps physicians identify potential metastatic hotspots and take timely intervention measures. For those paths with high metastasis risk, targeted treatment or monitoring can be implemented to prevent metastasis. Immunohistochemical staining technology can accurately label immune cells in tumor tissue, such as T cells, B cells, and macrophages. Through precise single-cell-level localization, the location of each immune cell can be efficiently identified, thereby extracting the spatial distribution pattern of immune cells. This helps reveal changes in the tumor immune microenvironment and immune escape mechanisms. Automated counting technology reduces errors from manual labeling and counting, providing efficient and accurate assessment of immune cell numbers. By calculating the density and distribution patterns of immune cells, quantitative data on the tumor's immune response can be provided, thus providing auxiliary information for subsequent metastasis risk prediction. By combining spatial distribution data of immune cells with metastatic pathway maps, multi-regional lymphatic metastasis simulations can be performed. This process simulates the spread of tumor cells within the immune microenvironment, revealing how immunosuppression in different regions affects the risk of tumor metastasis. This not only helps predict tumor metastasis pathways but also identifies regions with insufficient immune responses, providing a reference for treatment plans.By comprehensively analyzing the distribution of immune cells, tumor cell migration characteristics, and metastatic pathways, personalized metastasis risk assessments can be derived. This risk prediction considers the interaction of multiple factors and can provide more accurate early warning results than traditional methods, helping to detect potential lymph node metastasis risks at an early stage. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the steps in the method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a flowchart illustrating the detailed implementation steps of step S2. Detailed Implementation

[0009] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0010] This application provides a method and system for predicting the risk of lymph node metastasis in colorectal cancer based on image analysis. The executing entities of the method and system include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, and network upload devices that can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio-visual management system, an information management system, and a cloud-based data management system.

[0011] Please see Figures 1 to 3 This invention provides a method for predicting the risk of lymph node metastasis in colorectal cancer based on image analysis, comprising the following steps: Step S1: Obtain the scanned images of pathological sections, perform layer-by-layer downsampling and feature reconstruction, and generate the topological features of the pathological sections; Step S2: Based on the topological features of pathological sections, perform adaptive clustering analysis and multi-parameter fusion evaluation to obtain the independent migration evaluation value of individual tumor cells; Step S3: Perform tumor metastasis branch analysis based on the independent migration assessment values ​​to construct a metastasis branch path distribution map; Step S4: Extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain spatial distribution data of immune cells; Step S5: Based on the spatial distribution data of immune cells and the distribution map of metastatic branch paths, perform multi-regional lymphatic metastasis simulation and predict the metastasis risk situation to obtain the final metastasis risk assessment results.

[0012] In the embodiments of the present invention, see Figure 1This is a flowchart illustrating the steps of an image analysis-based method for predicting the risk of lymph node metastasis in colorectal cancer according to the present invention. In this example, the steps of the image analysis-based method for predicting the risk of lymph node metastasis in colorectal cancer include: Step S1: Obtain the scanned images of pathological sections, perform layer-by-layer downsampling and feature reconstruction, and generate the topological features of the pathological sections; In this embodiment, high-resolution full-field pathological section scanning of colon cancer tissue samples was performed at a magnification of 40×, a spatial resolution of 0.25 μm / pixel, and an output format of TIFF. To facilitate multi-scale structural analysis, the images underwent multi-level downsampling to generate a four-level pyramidal resolution sequence (1×, 1 / 2×, 1 / 4×, 1 / 8×). Bilinear interpolation was used during downsampling to maintain structural continuity, and adaptive contrast enhancement (CLAHE algorithm) was employed to correct brightness differences between different levels. Subsequently, morphological features, texture features, and color distribution features were extracted from each image layer, including nuclear density, cytoplasmic ratio, histogram of oriented gradients (HOG) statistics, and gray-level co-occurrence matrix (GLCM) features. The extraction window size was 128×128 pixels, with a stride of 64 pixels. To obtain spatial topological relationships, the features were mapped to a two-dimensional coordinate network, where nodes represent cell center locations, edges connect adjacent cell centroids, and edge weights are generated by weighting inter-cell distance and morphological similarity. The topological feature matrix is ​​obtained by calculating the node degree distribution, local clustering coefficient, and topological entropy, with a dimension of approximately (N×10). 4 The topological features of the pathological sections, multiplied by 15, are used to characterize the structural complexity and cell connectivity patterns within the pathological tissue. These final output features serve as the basis for subsequent adaptive clustering and migration assessment.

[0013] Step S2: Based on the topological features of pathological sections, perform adaptive clustering analysis and multi-parameter fusion evaluation to obtain the independent migration evaluation value of individual tumor cells; In this embodiment, after obtaining topological features, an adaptive clustering algorithm is used to classify the tumor cell population based on spatial and morphological features. First, a density-based clustering algorithm (DBSCAN) is used for preliminary grouping, with a minimum sample size of 15 and a distance threshold ε of 20 pixels to identify different cell aggregation regions. Then, the topological difference index for each cluster is calculated, including local aggregation difference (ΔC), neighborhood diversity index (NDI), and topological entropy change rate (ΔH). Local aggregation difference measures the change in the degree of spatial aggregation of cells in a local area, NDI reflects the level of heterogeneity, and the topological entropy change rate characterizes the trend of structural order. These three factors are normalized and then linearly fused with weights of 0.4, 0.3, and 0.3 to obtain a comprehensive topological difference value. This difference value, along with cell morphological features (roundness, nuclear atypia), is input into a support vector regression (SVR) model for multi-parameter fitting, outputting an independent migration assessment value for each cell, ranging from 0 to 1. Cells with high assessment values ​​indicate strong spatial independent migration potential and are typically concentrated at tumor margins or fracture structures. This step provides quantitative migration data for subsequent modeling of tumor invasion direction and metastasis pathways.

[0014] Step S3: Perform tumor metastasis branch analysis based on the independent migration assessment values ​​to construct a metastasis branch path distribution map; In this embodiment, cell-level independent migration assessment values ​​are mapped to the entire slice space, and a tumor metastasis path network is constructed using a migration potential field. Specifically, a potential function is constructed using the cell centroid as a node and the lines connecting adjacent high-migrating cells as edges, based on the migration assessment values. Higher potential energy indicates greater cell migration potential. The shortest path search algorithm (Dijkstra's method) is used to identify the optimal migration path from the high-potential region (tumor core) to the low-potential edge (mesenchymal junction). To ensure spatial continuity, only paths with a connection length exceeding 50 micrometers and an average migration potential greater than 0.6 are retained as valid branches. A metastasis branch topology is generated based on the number, direction, and convergence of paths, where each branch node includes attributes such as path length, angle direction, and local curvature. A three-dimensional reconstruction algorithm (Marching Cubes) is used to extend the two-dimensional paths into three-dimensional space, forming a tumor metastasis branch path distribution map. This distribution map reflects the potential spatial migration direction of tumor cells and the potential direction of lymphatic metastasis, laying a spatial framework for risk prediction combined with immune information.

[0015] Step S4: Extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain spatial distribution data of immune cells; In this embodiment, immunohistochemical staining images corresponding to the aforementioned pathological sections are acquired, typically including T cell markers such as CD3 and CD8, or immunosuppression-related markers such as PD-L1. After scanning, the image resolution is set to 0.25 μm / pixel, and a color deconvolution algorithm is used to separate positive and background signals. The dye channel absorption matrices are set to DAB=[0.65,0.70,0.29], hematoxylin=[0.27,0.57,0.78], and a threshold of 0.02. After obtaining the separated positive signal image, Gaussian smoothing (σ=1.2) and Otsu thresholding are applied to the image to extract immunopositive cells. A watershed algorithm based on morphological reconstruction is used to separate overlapping cells, with a minimum cell radius of 4 pixels. Subsequently, the centroid coordinates, area, and staining intensity of each immune cell are calculated and recorded in a coordinate matrix. After spatial registration, the immune cell positions are mapped to the pathological section coordinate system, thus forming spatial distribution data of immune cells. The data format is a two-dimensional coordinate point set, with each point carrying a type identifier and staining intensity information, which can be used to assess the density and distribution pattern of immune cells.

[0016] Step S5: Based on the spatial distribution data of immune cells and the distribution map of metastatic branch paths, perform multi-regional lymphatic metastasis simulation and predict the metastasis risk situation to obtain the final metastasis risk assessment results.

[0017] In this embodiment, a multi-regional lymphatic metastasis simulation model is established by combining the spatial distribution of immune cells with the branching pathways of tumor metastasis. First, the pathological sections are divided into three regions: the tumor center, the invasion front region, and the surrounding stroma region. The immune cell density (cells / mm²) of each region is calculated based on the immune cell distribution, and the immune defense level is defined using the immune type diversity index (Shannon value). Then, the immune escape trend index is mapped to each metastatic pathway, and the pathway risk weights (weight coefficients of 0.5, 0.3, and 0.2, respectively) are calculated based on migration assessment values ​​and immune density. During the simulation, a time step of 0.5 units is set to simulate the dynamic changes in tumor cell migration and spread along the pathways. The risk value of each pathway is dynamically updated according to the degree of immunosuppression, ultimately forming a multi-time-point risk heat map. Risk evolution curves are obtained through time series analysis, and the peak, rate of increase, and stable phase of metastatic risk are further calculated. The final output metastatic risk assessment results include: an overall risk index (0–1), the proportion of high-risk pathways, and a spatially visualized metastatic trend map. These results comprehensively reflect the lymph node metastasis potential and temporal trend in colorectal cancer patients, providing a quantitative basis for individualized prognostic prediction.

[0018] In this embodiment, see Figure 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Obtain pathological scan images of the patient's colon cancer tumor tissue; The pathological slide scan images are subjected to multi-region color normalization and local contrast enhancement to obtain standardized pathological images. A multi-scale resolution hierarchical system is defined, and standardized pathological images are downsampled and reconstructed layer by layer to extract cell features at each level. The cell features at each level include cell morphological atypia indicators, nucleoplasmic distribution characteristics, and cell arrangement disorder. Based on the cell characteristics at each level, the cell cluster aggregation coefficient, boundary irregularity, interstitial infiltration depth, spatial connectivity path length, and topological complexity are calculated to obtain the spatial heterogeneity of cell structure. Deep convolutional encoding is performed based on the spatial heterogeneity of cell structure to generate topological features of pathological slides.

[0019] In this embodiment, after obtaining authorization from the patient and the hospital, the patient's colon cancer tumor tissue was digitally scanned using a whole-slide imaging (WSI) scanner. The resolution was typically 0.25 μm / pixel (40× magnification) to ensure that cellular structural details were fully preserved in the image. The resulting pathological images were ultra-high resolution multi-level TIFF or SVS format files with multiple scaling levels for subsequent image processing and analysis.

[0020] Due to variations in scanning equipment, staining batches, and sample processing conditions, color differences in HE images can occur. To improve the model's generalization ability and the stability of feature extraction, color normalization of the slice images is necessary. A multi-region color normalization method based on the Macenko algorithm is employed. First, several representative regions (including tumor, glandular, and stromal regions) are selected from the image, and the RGB space is mapped to the OD space using optical density (OD) transformation. Next, principal component analysis (PCA) is used to estimate the dye vector for each region, and a linear regression model is used to map the sample staining distribution to the staining distribution of the reference template slice, achieving color matching and normalization across the entire image. After color normalization, Contrast Limited Adaptive Histogram Equalization (CLAHE) is used to enhance the contrast of local regions, improving the structural discriminability between the cell nucleus and cytoplasm. The CLAHE window size is set to 16×16 pixels, and the clipping threshold is 0.01 to prevent amplification of local noise. A multi-scale resolution hierarchy is constructed. The original pathological image was set as the highest resolution layer (Level 0), and then the resolution was progressively reduced layer by layer using a pyramid downsampling method (e.g., a downsampling ratio of 2× per layer, up to Level 4, corresponding to 1 / 16 of the original image resolution). At each resolution level, morphological and arrangement features at the cell level were extracted using a feature extraction method based on kernel segmentation. Kernel segmentation could employ an improved U-Net or Hover-Net model, trained with pixel-level annotations to achieve kernel segmentation and classification. For each detected cell nucleus, cell morphological heterogeneity indicators (such as nuclear area, aspect ratio, shape roundness, and fractal dimension of nuclear contour), nucleoplasmic distribution features (nuclear / cytoplasmic area ratio, mean and standard deviation of nuclear staining intensity distribution), and cell arrangement disorder (neighborhood orientation variance, Voronoi neighborhood area variation coefficient, etc.) were extracted. At different levels, these features can reflect multi-scale information from local cell structure to overall tissue morphology. Through multi-level feature fusion, hierarchical integration of multi-scale cell features was achieved, establishing a data foundation for spatial topology analysis.

[0021] A spatial relationship graph (CellGraph) was constructed using Delaunay triangulation and Voronoi diagram models to determine the interconnections between cell clusters. Based on this spatial network, the cluster coefficient, boundary irregularity (based on contour curvature variance), invasion depth (measured by the average distance from the tumor-stromal interface to the tumor interior), spatial connectivity path length (based on the average length of the shortest path in the graph), and topological complexity (calculated using the graph's Betti number and fractal dimension) were calculated. These indicators reflect the growth patterns and structural disorder of tumor cells at the microscale. For example, samples with high boundary irregularity and high topological complexity often correspond to more aggressive tumors with a higher risk of metastasis. To ensure statistical stability, spatial features were calculated using a 256×256 μm sliding window (50% overlap), and the entire graph was scanned and features were summarized to obtain a stable distribution of spatial heterogeneity features. A Deep Convolutional Encoder Network (DCEN) structure is employed to feature map the spatial heterogeneity matrix. The input is a two-dimensional feature map containing metrics such as clustering coefficients, connected path lengths, and topological complexity. Spatial patterns are extracted layer by layer through multi-layer convolutions (kernel size = 3×3, stride = 1) and pooling operations. The output of the encoding network is compressed into a 128-dimensional topological feature vector through fully connected layers to capture the high-order spatial relationships of the organizational structure. Furthermore, to enhance the model's ability to perceive heterogeneity at different scales, a multi-scale convolution module (dilated convolution with dilation rates of 1, 2, and 4) is introduced into the network to achieve cross-scale feature fusion. The model is trained using the Adam optimizer with a learning rate of 1×10⁻⁶. The batch size was 32, and unsupervised learning was performed using the Mean Squared Error as the loss function. After training, the resulting encoded vectors became the topological representations of the pathological slides, which could be further input into subsequent classification or risk prediction models (such as Transformer-based graph convolutional prediction networks) to determine the patient's lymph node metastasis risk level.

[0022] In this embodiment, see Figure 3 The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Adaptive clustering analysis based on the topological features of pathological sections is used to obtain image cell classification; Based on the image cell classification, multi-location cell topological differences are calculated to obtain topological difference indices at different locations, including local clustering differences, neighborhood diversity indices, and topological entropy change rates. Based on the aforementioned topological difference values, tumor cells are precisely identified, and tumor cell nuclei are labeled. Calculate the spatial coordinate information of the tumor cell nuclei; perform image frame segmentation based on the spatial coordinate information, and extract multiple tumor cell nuclei image frames; Based on multiple tumor cell nucleus image frames, the leading edge region is identified, and the local edge curvature change rate, leading edge direction vector offset, and structural breakpoint distribution density are calculated to generate a tumor cell dispersion index. Multi-parameter fusion evaluation was performed based on tumor cell dispersion index to obtain independent migration evaluation values ​​for individual tumor cells.

[0023] In this embodiment, the topological feature matrix generated by deep convolutional encoding (each cell nucleus corresponds to a set of high-dimensional topological feature vectors) is standardized, and Z-score normalization is used to eliminate dimensional differences. Subsequently, t-SNE (t-Distributed Stochastic Neighbor Embedding) or UMAP (Uniform Manifold Approximation and Projection) is used for nonlinear dimensionality reduction, compressing the high-dimensional feature space into a two-dimensional embedding space for cluster visualization and density analysis. To avoid bias caused by a preset number of clusters, a density-based adaptive clustering algorithm (such as DBSCAN or OPTICS) is selected to adaptively identify different cell subpopulations based on the local density of the samples. In the experiment, the minimum number of samples (MinPts) for DBSCAN was set to 10, and the neighborhood radius ε was set to 0.5; the parameters were automatically optimized using SilhouetteCoefficient. The clustering results typically identify different subpopulations, such as highly atypically atypical tumor cells, mesenchymal fibroblasts, immune cells, and glandular epithelial cells. Spatial mapping of clustering results allows for a visual representation of the spatial distribution patterns of different cell types on pathological slides, laying the foundation for subsequent multi-location topological difference analysis. After cell clustering and classification, the entire pathological image is divided into several analysis unit regions with fixed pixel sizes, and the topological difference characteristics of cell distribution are calculated for each region. This calculation process includes three indicators: local clustering difference, neighborhood diversity index, and topological entropy change rate. Local clustering difference is expressed based on Ripley's K-function or L-function, used to measure the degree of cell aggregation or dispersion in local space; the neighborhood diversity index is calculated using the Shannon diversity index, reflecting the composition ratio and heterogeneity level of various cell types within a local region; the topological entropy change rate is calculated using the scale derivative of the node degree distribution entropy, characterizing the stability of the cell connectivity graph under spatial scale changes. The topological difference feature values ​​of each region are standardized and mapped to a topological heterogeneity map of the tissue space.

[0024] Based on the spatial distribution of topological difference indices, a multidimensional feature fusion model was constructed for the precise identification of tumor cell nuclei. Input features included topological difference values, cell morphological features (such as nuclear area, roundness, aspect ratio, and nucleocytoplasmic ratio), and local density information. The model determined cell categories through supervised learning, outputting the probability that each cell belonged to a tumor cell. Cells with probabilities higher than a set threshold were labeled as tumor cells and explicitly distinguished in the image. This topological and morphological fusion-based identification mechanism reduces the possibility of misclassification of non-tumor cells while maintaining identification accuracy. The final result is a tissue image containing precisely labeled tumor cell nuclei, serving as the basis for subsequent spatial coordinate extraction and image box generation. After identifying and labeling all tumor cell nuclei, the spatial coordinate information of each nucleus was extracted, and a global spatial index was established in the form of centroid coordinates (x, y). A spatial connectivity graph between cells was constructed using Delaunay triangulation or Voronoi partitioning to achieve a global characterization of the distribution relationships of the tumor cell population. Subsequently, fixed-size image frames are generated centered on each cell nucleus, covering the range of its neighboring cells to ensure complete local microenvironment information. To avoid redundant overlap between image frames, spatial distance constraints are applied to the frame generation to ensure independence and regional representativeness. Each image frame records its corresponding spatial coordinates, cell type, and topological attributes, providing accurate data support for subsequent frontier region identification and dispersion calculation.

[0025] The location of the tumor invasion front in the tissue is determined by edge detection and region localization algorithms, and the image bounding box is spatially matched with the distance to the front. Cells located within a preset distance threshold are defined as cells in the front region. For these regions, indices such as the rate of change of local edge curvature, the offset of the front direction vector, and the distribution density of structural breakpoints are calculated. The rate of change of local edge curvature is obtained from the variance of the curvature derivative of the front curve, which is used to characterize the complexity of the boundary morphology; the offset of the front direction vector is calculated based on the angle between the cell alignment direction and the boundary normal, which is used to measure the consistency of the cell alignment direction; the distribution density of structural breakpoints quantifies the degree of disruption of tissue continuity by analyzing the number and distribution pattern of breakpoints in the intercellular connectivity. After standardization and weighted integration, the three parameters form a tumor cell dispersion index, which reflects the degree of disorder in the local structure and the tendency of cells to detach from the population. The dispersion index, along with topological difference parameters, morphological features, and spatial distribution features, are incorporated into a multi-parameter fusion model to establish a tumor cell independent migration risk assessment system. The model uses multi-layer nonlinear regression or neural network as the core structure, and optimizes the correlation between different parameters through feature weighted learning. Input parameters include local aggregation differences, topological entropy change rate, dispersion index, nuclear area, and edge curvature change rate. The output is an independent migration assessment value for each tumor cell. Higher assessment values ​​represent the cell's potential for autonomous migration; higher values ​​indicate a greater likelihood of the cell detaching from the group and metastasizing within the spatial structure. A migration risk heatmap is generated by analyzing the distribution of assessment values ​​across the entire map, providing a quantitative basis for identifying potentially invasive regions in colorectal cancer tissue and predicting lymph node metastasis risk.

[0026] In this embodiment, step S3 includes the following steps: The pathological slide scan images are subjected to intelligent anatomical structure recognition to extract the anatomical structures of the slides, including lymphatic vessels, nerve bundles and blood vessels; Based on the analysis of the boundary morphology of the anatomical structure of the slice, and the three-dimensional geometric reconstruction, a three-dimensional geometric model of the pathological structure is constructed. Tumor invasion path tracing and extraction based on a three-dimensional geometric model of pathological structure; By analyzing the lymphatic drainage pathways of tumor invasion, potential invasion path directions can be obtained. Based on the independent migration assessment values, the risk intensity of potential invasion path directions is predicted, and the distribution of transfer branches is visualized to construct a transfer branch path distribution map.

[0027] In this embodiment, anatomical structure-level recognition is performed on digital colorectal cancer pathology slide scan images. The image resolution is set to 0.25 micrometers / pixel, and full-field scanning data is acquired using 40x optical magnification. After color correction and smoothing / denoising, the images are input into a structure recognition model based on a convolutional neural network (CNN) for tissue-level segmentation. The model structure adopts a multi-scale UNet network, with an input image block size of 512×512 pixels, a stride of 256 pixels, and output channels including five types of structures: epithelial tissue, stromal region, blood vessels, nerve bundles, and lymphatic vessels. To enhance the recognition of different tissues, a channel attention mechanism (channel weights ranging from 0.1 to 1.0) is introduced during the model training phase, and boundary-weighted loss is applied to blood vessel and lymphatic vessel regions to enhance the accuracy of structural boundary recognition. Through morphological filtering and connectivity analysis, the identified target regions are post-processed to remove noise blocks smaller than 200 pixels, preserving complete and continuous anatomical structure regions. The output is a pathological slice with multi-structure mask annotations, which clearly distinguishes the spatial distribution of lymphatic vessels, nerve bundles, and blood vessels, laying the foundation for subsequent 3D geometric modeling. After completing the 2D structure recognition, sequence registration and geometric reconstruction are performed on the multi-layer pathological slices. The spatial interval of each slice is set to 3 micrometers, and the slices are precisely aligned using a registration algorithm based on feature point matching (SIFT + affine transformation). The alignment error is controlled within 2 pixels. Subsequently, the boundary contours of each anatomical structure are extracted, and spline curves are used for smooth fitting to eliminate edge discontinuities caused by slice cutting. For tubular structures such as lymphatic vessels, blood vessels, and nerve bundles, a method combining centerline tracing and surface reconstruction is used for 3D geometric reconstruction: first, center path points (2-pixel spacing) are extracted, and then the 3D surface is reconstructed through radial dilation. The reconstruction mesh resolution is set to 1 micrometer, and the mesh smoothness control parameter is 0.05 to ensure a smooth surface and maintain the realistic morphology. The final generated three-dimensional geometric model of the pathological structure is represented in the form of a multi-layer STL file, which includes spatial coordinates, volume information and boundary normal vectors. It can intuitively show the spatial layout of blood vessels, lymphatic vessels and nerve channels inside the colon tissue, and provide accurate spatial reference for subsequent tumor invasion path analysis.

[0028] Using an established three-dimensional geometric model of the pathological structure, combined with the spatial distribution coordinates of tumor cells, the spatial trajectory of tumor cell populations spreading outward was tracked. First, the tumor region boundary was defined in the three-dimensional model, with the invasion initiation layer set as the main tumor boundary (Z=0). A layer-by-layer search was performed along the normal direction with a step size of 1 micrometer to identify the shortest connection path for tumor cell migration. Path tracking was implemented using a graph search algorithm, combining three-dimensional A* search with a topologically constrained shortest path algorithm. The search weights comprehensively considered cell density gradient, dispersion index, and tissue permeability parameters (weight ratio of 0.4:0.4:0.2). To avoid pseudo-path generation, the path continuity threshold was set to 0.85, and path angle changes were limited to within 45°. Finally, multiple spatial invasion trajectories extending from the main tumor to surrounding tissues (especially blood vessels and lymphatic vessels) were extracted, and the path length, direction vector, and tissue penetration level were recorded. These trajectories represent the potential tumor invasion paths. After extracting the tumor invasion pathways, the spatial relationship between tumor cell invasion and lymphatic drainage direction is analyzed by combining the geometric orientation and spatial connectivity of lymphatic vessel structures in the 3D model. First, the endpoint coordinates of all invasion pathways are spatially matched with the lymphatic vessel centerline, with a matching radius of 10 micrometers. If the endpoint distance from the lymphatic vessel centerline is less than this threshold, the pathway is defined as having a potential connection with the lymphatic system. Then, based on the matched pathway set, the lymphatic drainage direction distribution is calculated, and the main drainage directions are extracted using a principal direction vector aggregation algorithm. The direction vector resolution is set to a 15° angle interval, and the clustering threshold is 0.9 to ensure the stability of the direction statistics. To further quantify the path flow characteristics, the concentration index (range 0 to 1) and diffusion angle range (difference between the maximum and minimum direction angles) of the path directions are calculated. Higher concentration and smaller angles indicate greater consistency and directionality of the invasion pathways.

[0029] The cell-independent migration assessment values ​​calculated in the previous stage are fused with the lymphatic drainage direction characteristics in this stage to predict the risk intensity of potential metastatic pathways of tumor cells. First, the migration assessment values ​​of all cells along each invasion pathway are averaged to obtain the overall migration potential value of the pathway. Then, a comprehensive risk intensity scoring model is constructed by combining pathway length, directional stability, and connection strength with lymphatic vessels (defined as the reciprocal of the distance from the pathway endpoint to the lymphatic centerline). The scoring weight ratio is set as follows: migration assessment value 0.5, directional stability 0.3, and connection strength 0.2. The output risk intensity ranges from 0 to 1; pathways exceeding 0.7 are identified as high-risk metastatic channels. Subsequently, a 3D visualization module is used to graphically represent the pathway branches: high-risk branches are displayed in red, medium-risk branches in yellow, and low-risk branches in blue. Branch thickness is linearly mapped to risk intensity, and the size of the node sphere corresponds to the cell density at the endpoint of the invasion pathway. In this embodiment, step S4 includes the following steps: Extract immunohistochemical staining images from the patient; The immunohistochemical staining image is deconvolved with color space to obtain positive staining signal, background tissue signal, and staining image after signal separation. Subpixel-level spatial registration is performed on the stained image and pathological slide scan image after signal separation to obtain a spatially registered group image. Based on the positive staining signal, the spatially registered histochemical images are precisely located and automatically counted at the single-cell level to obtain the spatial distribution data of immune cells.

[0030] In this embodiment, immunohistochemical (IHC) staining image data matching the corresponding pathological sections were acquired. The staining targets were generally key immune markers in the tumor microenvironment, such as CD3, CD8, CD20, or PD-L1. Image acquisition was performed using a full-field digital slide scanning system with a magnification of 40x, a resolution of 0.25 micrometers / pixel, a color depth of 24 bits, and TIFF image format. To ensure color consistency, brightness normalization and white balance correction were performed on the scanned images. The correction parameters were automatically calculated using a reference blank tissue area to control the staining deviation between different batches within ±3%. Subsequently, image region cropping and background removal were performed, removing tissueless areas and blank edge portions, retaining effective tissue areas for analysis. Each immunohistochemical image was saved at approximately 2GB of its original resolution, and a number mapping relationship was established with the corresponding H&E pathological sections to provide a consistent basis for subsequent registration and signal separation steps. For the acquired immunohistochemical images, a color deconvolution method based on optical density matrix decomposition was used to spatially decouple the staining signals. The image is first converted from RGB space to optical density space, and then different dye channels are separated using a non-negative matrix factorization algorithm. Taking the common DAB-hematoxylin staining as an example, the absorption matrix of the DAB channel is set to [0.65, 0.70, 0.29], and the hematoxylin channel is set to [0.27, 0.57, 0.78]. The matrix factorization threshold is set to 0.02 to ensure that the separated signals are free from cross-contamination. The separation results include three main outputs: a positive staining signal image (reflecting DAB characteristics), a background tissue signal image (mainly hematoxylin nuclear staining), and a clean staining image after signal separation. To suppress noise and enhance the edges of positive signals, the separated image is Gaussian smoothed (σ=1.2), and the positive region is extracted using the Otsu adaptive thresholding algorithm.

[0031] To achieve a one-to-one correspondence between immunohistochemical staining results and the morphological structure of H&E pathological sections, subpixel-level spatial registration was performed between the signal-separated IHC images and the corresponding pathological section scan images. First, feature point information was extracted from both types of images, and the Scale Invariant Feature Transform (SIFT) algorithm was used to detect key points, with a feature matching distance threshold set to 0.75. Then, the Random Sample Consensus Algorithm (RANSAC) was used to remove incorrect matches, with a maximum of 1000 iterations and matching accuracy controlled within ±1 pixel. The matching results were used to estimate the affine transformation matrix (including rotation, translation, and scaling parameters), and bicubic interpolation was used for image geometric correction. To further improve accuracy, a refinement adjustment was performed after the initial registration, using a local registration algorithm based on mutual information to reduce the registration error to within 0.3 pixels. After spatial registration, single-cell-level localization and automatic counting of immune cells were performed using positive staining signals. First, threshold segmentation and connected component analysis were performed on the registered positive signal image to identify independent cell regions. To eliminate errors caused by signal overlap, a watershed segmentation method based on morphological reconstruction was used for cell separation, with a minimum nuclear radius of 4 pixels and a maximum nuclear radius of 10 pixels. The centroid coordinates of each cell were then extracted as spatial positioning points, and their area, staining intensity, and roundness were recorded. To ensure positioning accuracy, all positive regions had to simultaneously meet the conditions of an intensity threshold > 0.25 and an area threshold > 20 pixels². An automatic counting module normalized the number of positive cells to the tissue area, obtaining the immune cell density value per square millimeter. The final generated spatial distribution data of immune cells was stored in the form of a coordinate matrix, containing the location, type, and staining intensity information of each immune cell.

[0032] In this embodiment, the specific steps of step S5 are as follows: The pathological slide scan images are classified into regions based on multiple tumor cell nucleus image frames to obtain tumor invasion classification regions, including the tumor central region, invasion front region, and surrounding stroma region. Based on the spatial distribution data of immune cells, the immune cell density of tumor invasion classification regions is calculated to obtain the immune cell density of multiple regions. Spatial interaction distribution analysis of immune cell density in multiple regions was performed to construct a tumor-immune microenvironment interaction network; Immune escape analysis was performed based on the tumor-immune microenvironment interaction network to generate immune escape trend indices for different regions; Multi-regional lymphatic metastasis simulation was performed based on the immune escape trend index of different regions to analyze the distribution map of metastatic branch paths, and the risk change trend was predicted to construct the final metastasis risk assessment result.

[0033] In this embodiment, a region classification standard is established using cell density gradient and dispersion index (value range 0–1): when the cell nucleus density is higher than 450 / m², the classification standard is established. Furthermore, a dispersion below 0.3 was defined as the tumor center region; a density between 150–450 cells / mm² with a dispersion between 0.3–0.6 was defined as the invasion front region; and a density below 150 cells / mm² with a dispersion greater than 0.6 was defined as the surrounding stroma region. Region segmentation was achieved through spatial clustering based on kernel density estimation (KDE), with a bandwidth parameter set to 50 pixels and a sliding window overlap rate of 0.4 to ensure smooth boundaries. Morphological closing operations were performed on the results after segmentation to eliminate isolated pixels. The final classification results were represented by a three-channel labeled image: red areas represented the tumor center region, orange areas the invasion front region, and green areas the stroma region, thus achieving precise spatial delineation of different functional levels of tumor tissue. Based on the region segmentation results, the spatial distribution data of immune cells obtained from immunohistochemical images were mapped to the same coordinate system to ensure a one-to-one correspondence with the tumor classification regions. For each region, the spatial density index of immune cells, i.e., the number of immune cells per square millimeter, was calculated. The calculation employed a sliding window statistical method, with a window size of 200×200 pixels and a step size of 100 pixels to smooth out local density variations. This was done to differentiate between immune cell types (such as CD3). + T cells, CD8 + Cell types (such as B cells, etc.) are retained in the input data, and density values ​​for each type are calculated separately. Results include total immune cell density and subtype density (unit: cells / mm²), typically ranging from 200 to 3000 cells / mm². For ease of subsequent analysis, density values ​​are Z-score normalized to ensure comparability between different sections. Finally, three types of immune cell density parameter matrices are obtained: tumor center, frontal zone, and stroma.

[0034] Based on the results of multi-regional immune cell density analysis, the spatial interaction and distribution relationship between tumor cells and immune cells was further analyzed. First, a heterogeneous network model was constructed with tumor cells and immune cells as nodes, with a total number of nodes approximately 10^4–10^5. The edge weights between nodes were defined as the product of spatial proximity and functional interaction strength. Spatial proximity was calculated based on inter-cell distance (within a threshold of 30 pixels considered a neighborhood connection), and functional interaction strength was set according to cell type correlation (e.g., CD8). +The interaction weights between T cells and tumor cells were set to 1.0, and those between B cells and tumor cells to 0.6. This method forms a weighted undirected graph structure. To analyze the immune activity interactions between regions, subgraphs of each region were further extracted, and node degree centrality, aggregation coefficient, and betweenness centrality were calculated. Node degree centrality reflects the activity level of cell participation in the interaction, the aggregation coefficient represents the local immune aggregation level, and betweenness centrality is used to identify potential immunosuppressive channels. Finally, a tumor-immune microenvironment interaction network graph was constructed, showing the strength of immune connections and spatial topological characteristics in different regions. Based on the established interaction network, the immune escape characteristics of each region were quantitatively analyzed. First, the interaction density decrease ratio between immune cell nodes and tumor cell nodes was calculated, i.e., the rate of change of the average connectivity of tumor cell nodes relative to the connectivity of immune cell nodes. When this ratio is below 0.6, it is considered that the immune interaction is weakened. Then, the three parameters of local immune cell density, immune cell diversity, and network aggregation coefficient were weighted and fused to form an immune escape trend index. The weights of each parameter were set to 0.4, 0.35, and 0.25, respectively. The index ranges from 0 to 1, with higher values ​​indicating a greater likelihood of immune escape. To improve spatial resolution, the index distribution was Gaussian smoothed (σ = 20 pixels) to obtain a continuous immune escape trend map. The results show that the tumor center typically exhibits a low escape index (<0.3), the invasion front zone a moderate one (0.3–0.6), while the surrounding stroma may show high escape values ​​(>0.6), reflecting that tumor cells are more likely to evade immune recognition in the peripheral areas. This trend index provides quantitative parameters for subsequent metastasis risk simulation.

[0035] A multi-regional lymph node metastasis simulation model was established by fusing the immune escape trend index with the previously generated metastatic branch path distribution map. First, each metastatic path was mapped to the region it traverses, and the average immune escape index for each path was calculated. Then, based on the path risk calculation model, the metastasis risk intensity of each path was updated. The update formula used the escape index as the main weight (weight coefficient 0.6), combined with the original path migration potential (0.3) and directional stability (0.1) for comprehensive calculation. The risk intensity results were remapped to three-dimensional space to form a multi-regional risk gradient field. During the simulation, the time step was set to 0.5 units to simulate the spread trend of tumor cells under the influence of immune escape, with 100 iterations until the risk distribution stabilized. The final output risk profile was represented by a color gradient: red for high-risk metastasis pathways, yellow for medium-risk, and blue for low-risk. The system simultaneously generated an overall metastasis risk index (range 0–1) to quantify the patient's overall lymph node metastasis risk level, providing a visual decision-making basis for clinical staging and prognostic assessment of colorectal cancer.

[0036] In this embodiment, the specific steps for simulating multi-regional lymphatic metastasis based on the distribution map of metastatic branch paths using the immune escape trend index of different regions, predicting risk changes, and constructing the final metastasis risk assessment result are as follows: Multi-regional lymphatic metastasis simulation was performed on the distribution map of metastatic branch paths based on the immune escape trend index of different regions, generating dynamic simulation data. Calculate the time window for immune escape based on dynamic simulation data; The attack path advancement analysis was performed on the dynamic simulation data, and the risk transfer advancement speed was calculated; Based on the time window of immune escape occurrence and the speed of risk transfer, a time-series risk evolution analysis was performed to generate a risk transfer evolution curve; Predict the risk change trend of the risk evolution curve and construct the final risk assessment result.

[0037] In this embodiment, a temporal evolution mechanism is first introduced based on the three-dimensional metastatic branch path distribution map to construct a space-time simulation platform to generate dynamic simulation data. The simulation employs a hybrid approach: a field-based diffusion-migration model is used macroscopically to characterize changes in cell population density, while an agent-based model is used microscopically to describe the migration behavior of individual tumor cells or cell clusters along the path. The spatial grid resolution is set to 5 micrometers, the time step to 0.5 time units, and the total simulation duration is defaulted to 100 time units. Each identified metastatic path is assigned inherent path properties at the initial moment: initial migration potential (weighted mapping by migration assessment values, ranging from 0 to 1), path directional stability, and connection strength with lymphatic vessels. The immune escape tendency index is mapped to the local immunosuppression coefficient (range 0 to 1) at each grid point on the path, which affects the cell's survival probability and migration rate at that location. The randomness of migration behavior is controlled by a noise term with an amplitude of 0.05 to simulate microenvironmental heterogeneity. To assess uncertainty, the simulation employed a Monte Carlo sampling strategy, running 500 independent cycles. Each cycle recorded the evolution of cell density over time along each path, the path occupancy probability, and the temporal distribution of arrival at key nodes (such as lymphatic confluences). Outputs included a time-series heatmap, a path occupancy probability matrix, and a three-dimensional risk field at each time point. Using the dynamic simulation data generated in the first step, the time window for immune escape was defined and identified as the critical time interval for risk evolution. First, a local immune escape threshold event was calculated for each path and the spatial grid points it covered: when the immune escape trend index at a grid point continuously exceeded 0.6 (threshold adjustable) and this state was maintained for more than 3 time units, it was determined to be a "local escape initiation event." At the path scale, if more than 30% of the key grid points on the same path triggered a local escape initiation event, the path was considered to have entered the "global immune escape phase." The immune escape occurrence time window was defined as the interval from the time point when the overall escape phase of the path was first triggered to the time point when immune suppression on the path fell back below 0.4. To avoid misjudgments caused by single noise events, the triggering event must occur repeatedly in at least two independent Monte Carlo runs. The length, start and end times, and frequency of the time window are output as a statistical distribution, including the median, interquartile range, and confidence interval (95%). In addition, the spatial distribution properties of the window (i.e., which path segments enter the escape window earlier) are calculated and stored in the form of a time-space matrix.

[0038] Based on existing time window and path occupancy data, a quantitative analysis of the invasion path's advancement is conducted to quantify the risk propagation speed. Advancement analysis uses a predetermined threshold of cell density along the path (e.g., 0.2 relative units per grid point) as the marker for an advancement event. The timing of the event is recorded along the path from start to finish, and the advancement time difference between adjacent nodes is calculated. Advancement speed is defined as the physical distance between nodes (in micrometers) divided by the advancement time difference, expressed in micrometers per unit of time. To enhance robustness, speed statistics are expressed using the mean and standard deviation of all Monte Carlo runs, and outliers (exceeding ±3 standard deviations of the full sample mean) are removed. In the advancement speed calculation, the immune escape trend index is used as a moderating factor: in high escape regions, the advancement speed is multiplied by a coefficient of 1.2–1.5 to reflect faster advancement potential; in low escape regions, it is multiplied by 0.6–0.9 to reflect obstructed advancement. Further cluster analysis of the advancement speed identifies "fast advancement segments," "medium-speed advancement segments," and "slow advancement segments," with the proportion and spatial location of each segment used to depict the risk diffusion pattern. Advance velocity data is exported as time-series tables and path velocity curves, and can be used to correct subsequent risk evolution models. Using the immune escape time window and advance velocity data as driving variables, a time-series risk evolution model is constructed to generate the transfer risk evolution curves for each path and the overall tissue. Risk value updates on the time axis follow a phased strategy: before the escape window begins, risk increases slowly over time; after entering the escape window, risk increases rapidly with increasing advance velocity; when the path reaches the lymphatic confluence node or immunosuppression declines, the risk level tends to stabilize or decrease. Specifically, a state machine-based sequence model is used to fit different stages piecewise, and the curves are denoised using exponential smoothing and moving average filters, with the time resolution consistent with the simulation step size (0.5 time units). To represent uncertainty, the risk evolution curves are output as a mean curve plus a confidence band (e.g., 95% confidence interval), and quantile curves (10%, 50%, 90% quantiles) are also provided to show the tail risk probability. Each curve includes key time point annotations (escape initiation time, peak propulsion velocity time, and time to lymph node arrival), and supports grouping and comparison by region (central region, frontal region, and interstitial region) and by pathway type. Results are saved in the form of graphical curves, tables, and interactive time slice data for easy clinical interpretation and decision support.

[0039] After obtaining the historical and current status of the risk evolution curve, future trend predictions are performed, and the results are synthesized to form the final metastasis risk assessment. The prediction employs a combination of a short-term predictor based on Bayesian updates and a long-term trend predictor based on regression trees: the short-term predictor uses the latest observed risk curve and the rate of advancement to progressively update and predict risk changes over several time steps (e.g., the next 10 time units); the long-term predictor estimates the overall metastasis probability distribution and the expected value of the potential time to critical metastasis events based on path attributes, statistical characteristics of the immune escape window, and clustering results of the rate of advancement. The prediction results are presented in a multi-dimensional output format: the future risk curve for each path, the probability of reaching lymph nodes (e.g., probability thresholds for 72 hours, 1 week, and 1 month), and a risk heatmap and overall quantitative risk index (0–1) for the entire map. To facilitate clinical application, the final metastasis risk assessment also includes recommended action levels (e.g., low, medium, high) and corresponding suggested time windows to support the prioritization of short-term follow-up plans or intraoperative / postoperative intervention strategies for patients. All results are accompanied by uncertainty measures (confidence intervals, tail risk probabilities) to clearly indicate the confidence level of the predictions, and are exported in the form of printable reports and interactive visualization panels, forming a complete time-transition risk assessment output.

[0040] In this embodiment, an image analysis-based system for predicting the risk of lymph node metastasis in colorectal cancer is provided, for performing the image analysis-based method for predicting the risk of lymph node metastasis in colorectal cancer as described above, including: The image layering and parsing module is used to acquire scanned images of pathological slides, perform layer-by-layer downsampling and feature reconstruction, and generate topological features of pathological slides. The migration assessment module is used to perform adaptive clustering analysis and multi-parameter fusion assessment based on the topological features of pathological sections to obtain independent migration assessment values ​​for individual tumor cells. The metastasis path distribution module is used to perform tumor metastasis branch analysis based on the independent migration assessment value and construct a metastasis branch path distribution map. The immune cell calculation module is used to extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain the spatial distribution data of immune cells; The metastasis prediction module is used to simulate multi-regional lymphatic metastasis and predict metastasis risk based on the spatial distribution data of immune cells and the distribution map of metastasis branch paths, so as to obtain the final metastasis risk assessment results.

[0041] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0042] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for predicting the risk of lymph node metastasis in colon cancer based on image analysis, characterized in that, Includes the following steps: Step S1: Acquire scanned images of pathological slides, perform layer-by-layer downsampling and feature reconstruction, and generate topological features of pathological slides; Step S2: Based on the topological features of pathological sections, perform adaptive clustering analysis and multi-parameter fusion evaluation to obtain independent migration assessment values ​​for individual tumor cells; Step S3: Perform tumor metastasis branch analysis based on the independent migration assessment values ​​and construct a metastasis branch path distribution map; Step S4: Extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain spatial distribution data of immune cells; Step S5: Based on the spatial distribution data of immune cells and the distribution map of metastatic branch paths, perform multi-regional lymphatic metastasis simulation and predict the metastasis risk situation to obtain the final metastasis risk assessment results.

2. The method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to claim 1, characterized in that, The specific steps of step S1 are as follows: Obtain pathological scan images of the patient's colon cancer tumor tissue; The pathological slide scan images are subjected to multi-region color normalization and local contrast enhancement to obtain standardized pathological images. A multi-scale resolution hierarchical system is defined, and standardized pathological images are downsampled and reconstructed layer by layer to extract cell features at each level. The cell features at each level include cell morphological atypia indicators, nucleoplasmic distribution characteristics, and cell arrangement disorder. Based on the cell characteristics at each level, the cell cluster aggregation coefficient, boundary irregularity, interstitial infiltration depth, spatial connectivity path length, and topological complexity are calculated to obtain the spatial heterogeneity of cell structure. Deep convolutional encoding is performed based on the spatial heterogeneity of cell structure to generate topological features of pathological slides.

3. The method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to claim 1, characterized in that, The specific steps of step S2 are as follows: Adaptive clustering analysis based on the topological features of pathological sections is used to obtain image cell classification; Based on the image cell classification, multi-location cell topological differences are calculated to obtain topological difference indices at different locations, including local clustering differences, neighborhood diversity indices, and topological entropy change rates. Based on the aforementioned topological difference values, tumor cells are precisely identified, and tumor cell nuclei are labeled. Calculate the spatial coordinate information of the tumor cell nuclei; Image frame segmentation is performed based on the spatial coordinate information to extract multiple tumor cell nucleus image frames; Based on multiple tumor cell nucleus image frames, the leading edge region is identified, and the local edge curvature change rate, leading edge direction vector offset, and structural breakpoint distribution density are calculated to generate a tumor cell dispersion index. Multi-parameter fusion evaluation was performed based on tumor cell dispersion index to obtain independent migration evaluation values ​​for individual tumor cells.

4. The method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to claim 1, characterized in that, Step S3 is as follows: The pathological slide scan images are subjected to intelligent anatomical structure recognition to extract the anatomical structures of the slides, including lymphatic vessels, nerve bundles and blood vessels; Based on the analysis of the boundary morphology of the anatomical structure of the slice, and the three-dimensional geometric reconstruction, a three-dimensional geometric model of the pathological structure is constructed. Tumor invasion path tracing and extraction based on a three-dimensional geometric model of pathological structure; By analyzing the lymphatic drainage pathways of tumor invasion, potential invasion path directions can be obtained. Based on the independent migration assessment values, the risk intensity of potential invasion path directions is predicted, and the distribution of transfer branches is visualized to construct a transfer branch path distribution map.

5. The method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to claim 1, characterized in that, The specific steps of step S4 are as follows: Extract immunohistochemical staining images from the patient; The immunohistochemical staining image is deconvolved with color space to obtain positive staining signal, background tissue signal, and staining image after signal separation. Subpixel-level spatial registration is performed on the stained image and pathological slide scan image after signal separation to obtain a spatially registered group image. Based on the positive staining signal, the spatially registered histochemical images are precisely located and automatically counted at the single-cell level to obtain the spatial distribution data of immune cells.

6. The method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to claim 1, characterized in that, The specific steps of step S5 are as follows: The pathological slide scan images are classified into regions based on multiple tumor cell nucleus image frames to obtain tumor invasion classification regions, including the tumor central region, invasion front region, and surrounding stroma region. Based on the spatial distribution data of immune cells, the immune cell density of tumor invasion classification regions is calculated to obtain the immune cell density of multiple regions. Spatial interaction distribution analysis of immune cell density in multiple regions was performed to construct a tumor-immune microenvironment interaction network; Immune escape analysis was performed based on the tumor-immune microenvironment interaction network to generate immune escape trend indices for different regions; Multi-regional lymphatic metastasis simulation was performed based on the immune escape trend index of different regions to analyze the distribution map of metastatic branch paths, and the risk change trend was predicted to construct the final metastasis risk assessment result.

7. The method for predicting the risk of lymph node metastasis in colon cancer based on image analysis according to claim 6, characterized in that, The specific steps for simulating multi-regional lymphatic metastasis based on the distribution map of metastatic branch pathways using the immune escape trend index of different regions, predicting risk changes, and constructing the final metastasis risk assessment result are as follows: Multi-regional lymphatic metastasis simulation was performed on the distribution map of metastatic branch paths based on the immune escape trend index of different regions, generating dynamic simulation data. Calculate the time window for immune escape based on dynamic simulation data; The attack path advancement analysis was performed on the dynamic simulation data, and the risk transfer advancement speed was calculated; Based on the time window of immune escape occurrence and the speed of risk transfer, a time-series risk evolution analysis was performed to generate a risk transfer evolution curve; Predict the risk change trend of the risk evolution curve and construct the final risk assessment result.

8. A colorectal cancer lymph node metastasis risk prediction system based on image analysis, characterized in that, A method for performing image analysis-based prediction of colorectal cancer lymph node metastasis risk as described in claim 1, comprising: The image layering and parsing module is used to acquire scanned images of pathological slides, perform layer-by-layer downsampling and feature reconstruction, and generate topological features of pathological slides. The migration assessment module is used to perform adaptive clustering analysis and multi-parameter fusion assessment based on the topological features of pathological sections to obtain independent migration assessment values ​​for individual tumor cells. The metastasis path distribution module is used to perform tumor metastasis branch analysis based on the independent migration assessment value and construct a metastasis branch path distribution map. The immune cell calculation module is used to extract the patient's immunohistochemical staining images, perform precise single-cell localization and automatic counting, and obtain the spatial distribution data of immune cells; The metastasis prediction module is used to simulate multi-regional lymphatic metastasis and predict metastasis risk based on the spatial distribution data of immune cells and the distribution map of metastasis branch paths, so as to obtain the final metastasis risk assessment results.

Citation Information

Cited By

  • Method for converting non-paired pathological image into pseudo DIC image and application thereof

    CN122199252A

  • Methods and applications for converting unpaired pathological images to pseudo-DIC images

    CN122199252B