Resolution Adaptive Seamless Semantic Segmentation Method for Digital Pathology Whole Slide Images
Through adaptive indexing and self-supervised learning methods, the problem of semantic segmentation accuracy reduction and splicing noise gaps caused by inconsistent resolution in digital pathological panoramic slices is solved, and efficient and accurate seamless semantic segmentation is achieved. It is suitable for a variety of slice formats and staining types, improving the flexibility and accuracy of the model.
Patent Information
- Application Number
- CN202510475225.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When processing digital pathological panoramic slices with different resolutions, the prior art has problems such as degradation in the accuracy of semantic segmentation models and noise gaps in splicing results, especially when the slice resolution is inconsistent in multi-center and large-scale queues.
Through adaptive indexing and mask-based self-supervised learning methods, a semantic segmentation model is built, and the sampling size and edge cropping are adjusted using physical resolution information to achieve seamless splicing, reduce dependence on labeled data, and improve model generalization capabilities.
High-quality seamless semantic segmentation at different resolutions is achieved, which improves the applicability and accuracy of the model, reduces data processing time, is suitable for a variety of slice formats and staining types, and enhances the understanding of complex tissue structures.
Smart Images

Figure CN119992552B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for segmenting digital pathology panoramic slices, specifically a resolution adaptive seamless semantic segmentation method for digital pathology panoramic slices, belonging to the technical field of pathological image segmentation. Background Art
[0002] Digital pathology slices refer to images formed by digitizing glass slides using a dedicated digital slide scanner, and the data can be read using a database or after format conversion. Nevertheless, due to differences in scanner manufacturers, selected components such as eyepieces, objective lenses, optical paths, etc., and scanning configurations, the magnification of the images may vary, and this phenomenon is particularly significant in multi-center, large-scale cohorts. In digital pathology images, the magnification information of the image is recorded at the resolution, usually in units of micrometers per pixel (μm / px), and the common resolution range is 0.227μm / px to 0.552μm / px.
[0003] Computational pathology uses technologies such as image analysis and deep learning to identify regions of interest or targets of interest in pathological images. Computational pathology models are usually developed in a single cohort, extracting images at a fixed level such as level 0, and combining physician annotations to generate image patches for training image classification or semantic segmentation tasks.
[0004] However, when applying the trained model to subsequent test slices and hoping to obtain a complete segmentation map of the whole slice, two severe challenges will be encountered:
[0005] First, the actual physical resolution of the slice is inconsistent with the resolution of the images used for model training, and it is impossible to find a close resolution level. Directly applying the model will cause a serious decline in accuracy; when the resolutions between different slices are inconsistent, a more complex manual selection and processing process will be faced;
[0006] Second, when processing pathological panoramic slices, non-overlapping sliding window sampling is used, and the model is used to process local images separately, and then the results are stitched together. This solution will introduce noise at the edges due to zero padding in the local image convolution operation, resulting in obvious stitching seams composed of noise in the stitching result and reducing the result quality. Summary of the Invention
[0007] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide a resolution adaptive seamless semantic segmentation method for digital pathology panoramic slices.
[0008] Technical Solution: The resolution adaptive seamless semantic segmentation method for digital pathology panoramic slices of the present invention includes the following steps:
[0009] Step 1, obtaining a pathological panoramic slice image to be processed;
[0010] Step 2: Perform adaptive indexing on the pathological panoramic slice image based on the physical resolution information to obtain the target image block;
[0011] Step 3: Construct a semantic segmentation model, and based on the dataset with the specified resolution, use the mask-based self-supervised learning method to train the semantic segmentation model to obtain the target semantic segmentation model;
[0012] Step 4: Use the target semantic segmentation model to perform semantic segmentation on the pathological panoramic slice with any physical resolution.
[0013] Furthermore, Step 2 includes:
[0014] Calculate the physical resolution corresponding to each level according to the original physical resolution of the panoramic pathological slice and the downsampling ratio of each level;
[0015] Traverse the physical resolutions of all levels to find the most suitable level as the target level;
[0016] Calculate the deviation ratio coefficient according to the target resolution and the physical resolution of the target level, and adjust the sampling size according to the deviation ratio coefficient to obtain the new sampling size;
[0017] Read the corresponding local image block from the panoramic pathological slice through the target level, position, and new sampling size, and adjust the local image block to the specified target size to obtain the target image block.
[0018] Furthermore, using the mask-based self-supervised learning method to train the semantic segmentation model includes:
[0019] For each slice in the dataset, randomly select a proportion of the image area of the slice for masking, and the remaining image area is used as the unmasked area, where the value range of
[0020] is from 0 to 1;
[0021] Use the feature extraction sub-module of the semantic segmentation model to extract the features of the unmasked area;
[0022] Use the reconstruction sub-module of the semantic segmentation model to predict the pixel values of the masked area.
[0023] ,
[0024] In the formula, is the loss of the self-supervised task, is the loss of the supervised task, is the loss weight for the self-supervised task, is the loss weight for the supervised task, and .
[0025] Further, step 4 includes:
[0026] Calculate the difference coefficient r between the specified resolution Q of the semantic segmentation model and the actual resolution s of the pathological whole-slide image, with the formula: ;
[0027] According to the size of the original pathological whole-slide image and the difference coefficient r of the resolution, calculate the size of the segmentation result map equivalent to the resolution Q, with the formulas:
[0028] , ,
[0029] In the formula, H represents the height of the original pathological whole-slide image at the highest resolution, h represents the height of the segmentation result map equivalent to the resolution Q, W represents the width of the original pathological whole-slide image at the highest resolution, and w represents the width of the segmentation result map equivalent to the resolution Q;
[0030] Initialize the semantic segmentation label map M according to the calculated size of the segmentation result map;
[0031] Perform a sliding window sampling loop with a step size of s on the semantic segmentation label map M, and fill the local image blocks obtained in each loop into the corresponding areas of the semantic segmentation label map M;
[0032] After completing the loop, obtain the complete semantic segmentation label map M.
[0033] Further, performing a sliding window sampling loop with a step size of s on the semantic segmentation label map M includes:
[0034] Determine the position of the sampling starting point in the original pathological whole-slide image, with the formula:
[0035] ,
[0036] ,
[0037] In the formula, , respectively represent the horizontal and vertical coordinates of the sampling starting point on the original pathological whole-slide image; , respectively represent the horizontal and vertical coordinates of the current sliding window position on the segmentation result map, which are loop variables used to traverse the entire segmentation result map;
[0038] Read the target image patch from the original pathological panoramic section according to the position of the sampling starting point, and obtain a semantic segmentation map with a size of ;
[0039] Crop at the center of the semantic segmentation map to obtain a valid part R with a size of ;
[0040] Fill the valid part R into the area of the semantic segmentation label map M.
[0041] Further, finding the most suitable level as the target level includes:
[0042] When the target resolution of a certain level is greater than or equal to the current level's physical resolution multiplied by the resolution tolerance ratio, then this level is regarded as the most suitable level for sampling.
[0043] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are:
[0044] 1. The present invention finds the most suitable level for sampling by calculating the physical resolution of each level of the section, ensuring that any semantic segmentation model can run at the correct image resolution level, and solving the problem in the prior art that the accuracy of the semantic segmentation model decreases due to inconsistent section resolutions;
[0045] 2. The present invention proposes an edge cropping and stitching scheme with universal applicability. During the sliding window sampling process, only the central valid area of the segmentation result is used for filling, avoiding the influence of edge noise, forming a high-quality seamless panoramic segmentation result, improving the stitching accuracy, and avoiding the problems of stitching seams and inaccurate segmentation results caused by edge noise in the prior art;
[0046] 3. The present invention is applicable to any digital slice format with physical resolution information and a data reading interface, as well as digital pathological slice analysis of different organs and different staining types;
[0047] 4. The present invention supports various model input and output sizes, which are achieved by adjusting the padding mode of the convolutional layer, meeting the applicable conditions of different models, and improving the flexibility and applicability of the models;
[0048] 5. The present invention has efficient data reading ability. Through the adaptive resolution image indexing algorithm, it can quickly and accurately obtain the required local image patch, reducing the time for data reading and processing;
[0049] 6. The present invention reduces the dependence on labeled data by introducing self-supervised learning methods, while enhancing the generalization ability of the semantic segmentation model and the understanding ability of complex tissue structures; through the masking process and the reconstruction module, the model can learn the global structure and context information of the image without labeled data, thus showing higher accuracy and robustness when dealing with complex tissue structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of a resolution adaptive seamless semantic segmentation method for digital pathology panoramic slides;
[0051] Figure 2 It is a structural diagram of a digital pathology image;
[0052] Figure 3 It is a flowchart of an adaptive indexing of whole slide images based on physical resolution;
[0053] Figure 4 It is a flowchart of the working process of a semantic segmentation model;
[0054] Figure 5 It is a flowchart of constructing a semantic segmentation label map M;
[0055] Figure 6 It is an effect diagram of lung cancer area detection in a panoramic pathological slide;
[0056] Figure 7 It is a schematic diagram of multi-class tissue segmentation results;
[0057] Figure 8 It is a graph of the overall training loss;
[0058] Figure 9 It is a graph of the overall Dice coefficient. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0060] The resolution adaptive seamless semantic segmentation method for digital pathology panoramic slides described in this embodiment has a flowchart as Figure 1 shown, and this method includes the following steps:
[0061] Step 1, obtain a pathological panoramic slide image to be processed.
[0062] To facilitate improving the response speed in user interaction, digital pathology images are usually stored in the format of multi-resolution images, with different levels representing images at different zoom levels. Level 0 stores the image with the original size and the highest resolution. The higher the level, the higher the downsampling ratio, and the smaller the image size. For example, Figure 2 as shown
[0063] Step 2: Perform adaptive indexing on the panoramic pathological section image based on the physical resolution information to obtain the target image block.
[0064] Further, Step 2 includes:
[0065] Query location:
[0066] According to the original physical resolution of the panoramic pathological section and the downsampling ratio of each level , calculate the physical resolution corresponding to each level . The calculation formula is:
[0067] ,
[0068] Traverse the physical resolutions of all levels and find the most suitable level as the target level;
[0069] Calculate the sampling size:
[0070] Calculate the deviation ratio coefficient according to the target resolution and the physical resolution of the target level, and adjust the sampling size according to the deviation ratio coefficient to obtain the new sampling size; among them, the calculation method of adjusting the sampling size is to multiply the original sampling size by the deviation ratio coefficient to obtain the new sampling size.
[0071] Sampling by level, post-scaling processing:
[0072] Read the corresponding local image block from the panoramic pathological section through the target level, location, and new sampling size, and adjust the local image block to the specified target size to obtain the target image block.
[0073] Among them, the obtained local image block is scaled, and the scaling ratio is the reciprocal of the deviation ratio coefficient, and finally the target image block that meets the requirements is obtained.
[0074] Further, finding the most suitable level as the target level includes:
[0075] When the target resolution of a certain level is greater than or equal to the physical resolution of the current level multiplied by the resolution tolerance ratio , then this level is used as the most suitable level for sampling.
[0076] In one example, as the value can be 0.85.
[0077] The adaptive index described in this embodiment is applicable to any digital slice format with physical resolution information and a data reading interface, and can handle any physical resolution required by the user. The flowchart of the adaptive index of the whole-slide image based on physical resolution is as Figure 3 shown. The adaptive resolution image index algorithm described in this example can further establish a whole-slide inference framework, implement semantic segmentation of each sub-region of the whole-slide and perform in-situ stitching.
[0078] Step 3: Construct a semantic segmentation model, and train the semantic segmentation model using a mask-based self-supervised learning method based on a dataset with a specified resolution to obtain a target semantic segmentation model.
[0079] Among them, the specified resolution is the resolution adapted in the training stage and is the resolution specified by the user. The dataset can be any pathological image dataset with the same resolution. For example, the dataset of The Cancer Genome Atlas can be selected.
[0080] Furthermore, training the semantic segmentation model using a mask-based self-supervised learning method includes:
[0081] For each slice in the dataset, randomly select a proportion of the image regions of the slice for masking, and the remaining image regions are used as unmasked regions, where the value range of is from 0 to 1;
[0082] Use the feature extraction sub-module of the semantic segmentation model to extract the features of the unmasked regions. The features of the unmasked regions provide context information for the model to help it better predict the content of the masked regions;
[0083] Use the reconstruction sub-module of the semantic segmentation model to predict the pixel values of the masked regions. The pixel values of the masked regions are the targets that the model needs to predict. By predicting these pixel values, the model can learn the global structure and context information of the image.
[0084] Combined with Figure 4 shown, furthermore, step 4 includes:
[0085] Calculate the difference coefficient r between the specified resolution Q of the semantic segmentation model and the actual resolution s of the pathological whole-slide slice. The formula is: ;
[0086] According to the size of the original pathological whole-slide slice and the difference coefficient r of the resolution, calculate the size of the segmentation result map equivalent to the resolution Q. The formulas are respectively:
[0087] , ,
[0088] In the formula, H represents the height of the original pathological whole-slide image at the highest resolution, h represents the height of the segmentation result map equivalent to the resolution Q, W represents the width of the original pathological whole-slide image at the highest resolution, and w represents the width of the segmentation result map equivalent to the resolution Q;
[0089] Initialize the semantic segmentation label map M according to the calculated size of the segmentation result map;
[0090] Among them, the label map means that each pixel represents a category, and initialize a zero matrix of all zeros, calculate the size of the segmentation result map to obtain h and w, and then create an empty label map M according to h and w for storing the final semantic segmentation result;
[0091] Perform a sliding window sampling loop with a step size of s on the semantic segmentation label map M, and fill the local image block obtained in each loop into the corresponding area of the semantic segmentation label map M;
[0092] After completing the loop, obtain the complete semantic segmentation label map M.
[0093] Combined with Figure 5 As shown, further, perform a sliding window sampling loop with a step size of s on the semantic segmentation label map M, including:
[0094] Determine the position of the sampling starting point in the original pathological whole-slide image, and the formula is:
[0095] ,
[0096] ,
[0097] In the formula, , respectively represent the horizontal and vertical coordinates of the sampling starting point on the original pathological whole-slide image; , respectively represent the horizontal and vertical coordinates of the current sliding window position on the segmentation result map, which are loop variables used to traverse the entire segmentation result map; S represents the input size of the semantic segmentation model, and s represents the effective output size of the semantic segmentation model;
[0098] Read the target image block from the original pathological whole-slide image according to the position of the sampling starting point, and obtain a semantic segmentation map with a size of ;
[0099] Crop at the center of the semantic segmentation map to obtain an effective part R with a size of ;
[0100] Fill the valid part R into the region of the semantic segmentation label map M.
[0101] Improve the stitching accuracy of the segmentation result through step 3. Only the result of the central size will be finally adopted. However, due to the abandonment of the edges, this factor must be considered during stitching and appropriate adjustments should be made to form a seamless stitching effect.
[0102] Step 4: Use the target semantic segmentation model to perform semantic segmentation on a pathological whole-slide image with any physical resolution.
[0103] Using the trained target semantic segmentation model, semantic segmentation can be performed on a pathological whole-slide image with any physical resolution to obtain a segmentation result with a specified resolution.
[0104] If the output size of the semantic segmentation model is a downsampling of the input size, the result can be upsampled proportionally, and the above process can still be used. For a fully convolutional network using a non-padding mode, its convolutional layer should be adjusted to a padding-to-the-same-size mode to meet the applicable conditions of the algorithm.
[0105] Figure 6 The figure shows the recognition result map of the lung cancer region in the whole-slide pathological section implemented by using the method proposed in this embodiment. In the figure, "image" represents the original pathological image, and "cutoff" represents the cropping width. In this example, four different cropping widths of 0, 32, 64, and 128 are selected respectively. The part pointed by the black arrow in the figure is the segmentation gap. By configuring an appropriate cropping size and cooperating with the whole-slide segmentation algorithm, it can be clearly seen that the gap is eliminated and a high-quality segmentation result is obtained.
[0106] Figure 7 The figure shows the multi-class tissue segmentation result obtained by using the target semantic segmentation model described in this embodiment. The main classes described in the figure include: epithelial-like tissue corresponding to cyan in the figure, stroma corresponding to orange, immune infiltration corresponding to blue, and microvessels corresponding to yellow in the figure. To establish this multi-class tissue segmentation model, 669 regions of interest (rois) were extracted from 5 cancer genome atlas cohorts. When extracting the regions of interest, the most appropriate image level was first obtained. The average area of these rois is 1.991 mm2, and the standard deviation is , and the rois are saved in TIF format with a fixed resolution of .
[0107] The annotations of these ROIs were completed using the Automatic Slide Analysis Platform (ASAP) software and saved in XML format. After the annotation stage, the XML file was converted to a TIF mask using the Python interface of ASAP. Then the ROIs were aligned with the mask and jointly cropped into patches for model training and validation. Specifically, the sliding window method was used to extract patches of pixels from the tissue images and masks, with a stride of 448 pixels. Overlapping sampling was adopted to avoid under-training of pixels located at the edges. A total of 17,159 patches were generated through image cropping and then randomly divided into a training set and a validation set at a ratio of 7:3. Sampling was then performed, which avoided edge effects, made full use of the image information, and reduced errors during the stitching process. Patches of pixels with a stride of 448 pixels were extracted from the tissue images and masks using the sliding window method. Overlapping sampling was used to avoid under-training of pixels at the edges. A total of 17,159 patches were generated by image cropping and then randomly divided into a training set and a validation set at a ratio of 7:3. Sampling was then carried out, avoiding edge effects, making full use of image information, and reducing errors during the stitching process.
[0108] To further improve the generalization ability of the model, a mask-based self-supervised learning method was introduced during the training process. During the data preparation stage, some image regions were randomly selected for masking, such as randomly setting 30% of the image regions to zero, and then letting the model predict the pixel values of these masked regions. In this way, the model can learn the global structure and context information of the image without labeled data, thereby enhancing its understanding ability of complex tissue structures. For the architecture of the semantic segmentation model, it can be adjusted based on the existing U-Net model. The encoder serves as a feature extraction sub-module, responsible for extracting features from unmasked regions, while the decoder serves as a reconstruction sub-module for predicting the pixel values of masked regions. To balance the influence of these two tasks of self-supervised and supervised tasks, a weighted hybrid loss function was adopted, where the loss weight of the self-supervised task can be set to 0.3 and the loss weight of the supervised task can be set to 0.7. In this way, the semantic segmentation model can simultaneously learn the supervised information in the labeled data and the context information in the unlabeled data.
[0109] Exemplarily, an existing semantic segmentation model can be used and appropriately adjusted and optimized during the training process. The generated slices are used to train the semantic segmentation model. For example, a U-Net model with a VGG-19 encoding branch is used as the semantic segmentation model for training, and the channels in the decoder path are adjusted, and a lightweight coordinate attention mechanism is introduced. During the training process, the input size of the model is set to pixels, and mild color augmentation is introduced to improve the generalization ability of the model. Through the resolution-adaptive seamless semantic segmentation method described in this embodiment, significant results have been achieved. Through the fixed-resolution image input, the optimization of the model structure, and the application of the sliding window method, the model performs excellently in the segmentation of various tissue types, such as Figures 8 to 9As shown, the Dice coefficients of most categories exceed 0.8, and even for the more challenging small blood vessel category, a score exceeding 0.6 can be achieved in the later stage of training. The loss function is a 1:1 mixture of cross-entropy and dice loss, using the Adam optimizer with an initial learning rate of , and after 60 times, the learning rate is multiplied by 0.9 every 10 times, and the training stops after 160 times.
Claims
1. A resolution - adaptive seamless semantic segmentation method for digital pathology whole - slide images, characterized in that It includes the following steps: Step 1: Obtain the pathological whole-slide image to be processed; Step 2: Perform adaptive indexing on the pathological whole-slide image based on the physical resolution information to obtain the target image patch; Step 3: Construct a semantic segmentation model, and train the semantic segmentation model using a mask-based self-supervised learning method based on a dataset with a specified resolution to obtain the target semantic segmentation model; Step 4: Use the target semantic segmentation model to perform semantic segmentation on pathological whole-slide images with any physical resolution; Step 2 includes: According to the original physical resolution of the whole-pathology slide and the downsampling magnification of each level, calculate the physical resolution corresponding to each level; Traverse the physical resolutions of all levels, and find the most suitable level as the target level; Calculate the deviation ratio coefficient according to the target resolution and the physical resolution of the target level, and adjust the sampling size according to the deviation ratio coefficient to obtain a new sampling size; Read the corresponding local image patch from the whole-pathology slide through the target level, position, and new sampling size, and adjust the local image patch to the specified target size, where the obtained local image patch is scaled, and the scaling ratio is the reciprocal of the deviation ratio coefficient, to obtain the target image patch; Finding the most suitable level as the target level includes: When the target resolution of a certain level is greater than or equal to the current level's physical resolution multiplied by the resolution tolerance ratio, then this level is used as the most suitable sampling level.
2. The resolution adaptive seamless semantic segmentation method for digital pathology whole slide according to claim 1, characterized in that Training the semantic segmentation model using a mask-based self-supervised learning method includes: For each slice in the dataset, randomly select γ proportion of the image regions of the slice for masking, and the remaining image regions are used as unmasked regions, where the value range of γ is from 0 to 1; Use the feature extraction sub-module of the semantic segmentation model to extract the features of the unmasked regions; Use the reconstruction sub-module of the semantic segmentation model to predict the pixel values of the masked regions.
3. The resolution adaptive seamless semantic segmentation method for digital pathology panoramic slides according to claim 2, characterized in that The semantic segmentation model adopts a weighted hybrid loss function, and the formula is: , Wherein, is the loss of the self-supervised task, is the loss of the supervised task, is the loss weight of the self-supervised task, is the loss weight of the supervised task, and .
4. The resolution adaptive seamless semantic segmentation method for digital pathology panoramic slides according to claim 3, characterized in that Step 4 includes: Calculate the difference coefficient r between the specified resolution Q of the computational semantic segmentation model and the actual resolution s of the pathological whole-slide section, with the formula: ; According to the size of the original pathological whole slide and the difference coefficient r of the resolution, calculate the size of the segmentation result map equivalent to the resolution Q. The formulas are as follows: , , In the formula, H represents the height of the original pathological whole-slide at the highest resolution, h represents the height of the segmentation result map equivalent to the resolution Q, W represents the width of the original pathological whole-slide at the highest resolution, and w represents the width of the segmentation result map equivalent to the resolution Q; Initialize the semantic segmentation label map M according to the calculated size of the segmentation result map; Perform a sliding window sampling loop with a step size of s on the semantic segmentation label map M, and fill the local image patch obtained in each loop into the corresponding area of the semantic segmentation label map M; After completing the loop, obtain the complete semantic segmentation label map M.
5. The resolution - adaptive seamless semantic segmentation method for digital pathology whole - slide according to claim 4, characterized in that, Performing a sliding window sampling loop with a step size of s on the semantic segmentation label map M includes: Determine the position of the sampling starting point in the original pathological whole-slide, and the formula is: , , Wherein, and respectively represent the abscissa and ordinate of the sampling starting point on the original pathological whole slide; and respectively represent the abscissa and ordinate of the current sliding window position on the segmentation result map, which are loop variables used to traverse the entire segmentation result map; Read the target image patch from the original pathological panoramic section according to the position of the sampling starting point, and obtain a semantic segmentation map with a size of ; Crop at the center of the semantic segmentation map to obtain the effective part R with a size of ; Fill the valid part R into the region of the semantic segmentation label map M.
Citation Information
Patent Citations
Multi-scale fusion remote sensing image semantic segmentation method and system
CN115512103A
WSI digital slice segmentation refining method, device, equipment, medium and product
CN118212413A
Cited By
Tumor nerve infiltration quantification method and device based on digital pathological panoramic section and medium
CN121998985A
A method and device for quantifying tumor neural infiltration based on digital pathology whole slide images, and a medium
CN121998985B