A method and device for segmenting cancerous regions in breast tissue sections

By introducing residual attention module and multi-scale feature fusion module in the U-Net++ model, combined with the fully connected CRF model, the complexity problem in breast pathological section segmentation is solved, and the segmentation accuracy of cancerous areas and image detail recovery effect is improved.

CN115439493BActive Publication Date: 2025-07-18SHANGHAI PAIYING MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211111864.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-07-18
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

The existing deep learning-based breast pathological section segmentation method is difficult to effectively deal with the complexity of breast tissue pathological images, especially the high density distribution of tissue primitives, overlapping and winding of tissues, small differences in the characteristics at the junction of normal tissue areas and abnormal tissue areas, and many hollows in the image, resulting in poor segmentation effect.

Method used

U-Net++ is used as the benchmark model, and the original convolutional block is replaced by the residual attention module, combining the dual attention module and the multi-scale feature fusion module to enhance the feature extraction capability, and image post-processing is performed through the fully connected conditional random field CRF model to improve the accuracy of cancerous region segmentation.

Benefits of technology

The boundary segmentation effect of the lesion area and the small target area is improved, the accuracy of segmentation of cancerous areas is improved, the feature extraction ability of pathological images is enhanced, and the detailed parts of the image are restored.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115439493B_ABST
    Figure CN115439493B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for segmenting cancerous regions in breast tissue sections, including: making a training set and a test set based on the whole cancerous slide and the annotation of the cancerous regions therein; designing a segmentation model based on multi-scale fusion and attention mechanism, and training the segmentation model with the training set to obtain a trained segmentation model; inputting the whole slide image to be detected into the trained segmentation model to obtain a segmentation result in the form of a heat map; applying a fully connected conditional random field (CRF) model to refine the segmentation result to obtain image patches; and stitching the obtained image patches to form a segmentation result of the whole slide. By applying the embodiments of the present invention, it is intended to realize automatic segmentation of cancerous regions in breast pathological images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of in-line resistance slicing, and particularly to a method and device for segmenting cancerous regions in breast tissue slices. Background Art

[0002] The method for making digital pathological slices of breast cancer first stains tissue slides with hematoxylin and eosin (H&E), and then forms digital pathological images through electron microscope scanning. Pathological slices usually contain millions of cells, and pathologists usually need to analyze multiple slides in a day, which is a very time-consuming and laborious task. The characteristics of pathological images are complexity and diversity, and extremely strict requirements are imposed on the professionalism and experience of pathologists.

[0003] The research on classification methods for breast pathological slices based on deep learning has become mature, and the classification accuracy is close to or even exceeds the accuracy of doctor diagnosis. However, in the clinical diagnosis process, accurately locating the lesion area is more intuitive than simply giving the classification result of the image. Manual annotation of medical images requires high costs and time. Therefore, semantic segmentation of medical images using artificial intelligence has always been a popular topic in the field of medical image analysis.

[0004] Segmentation techniques based on deep learning have demonstrated their superior performance in many computer vision tasks. Due to their strong feature extraction ability and easy training and optimization, they have now become a complete and robust image segmentation method. In recent years, segmentation methods based on deep learning are mainly segmentation methods based on fully convolutional models, and representative ones include Fully Convolutional Networks (FCN), U-Net network, and U-Net++ network, etc. For example, Gu et al. proposed a context encoder network (CE-Net) to capture context information and retain spatial information, and applied it to 2D medical image segmentation tasks. Qu [8] et al. proposed a fully convolutional neural network to achieve instance segmentation of cell nuclei and glandular regions by maintaining full-resolution feature maps. Feng et al. used a new context pyramid fusion network (CPFNet) to fuse global and multi-scale context information by combining two pyramid modules.

[0005] Although extensive research has been conducted on medical image segmentation, histopathological images have the following problems compared to other medical images: First, the tissue primitives in histopathological images are distributed with high density, and there are overlaps and entanglements between tissues. Second, the feature differences at the junction between normal and abnormal tissue regions in histopathological images are relatively small, making it difficult to accurately segment. Finally, due to the special tissue structure of breast tissue pathological images, there are many holes in the images, which cause certain interference to the learning of the segmentation model. As a result, segmentation models that perform excellently in other tasks are difficult to achieve good segmentation results in the task of segmenting pathological cancerous tissue regions. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and device for segmenting cancerous regions in breast tissue sections, aiming to use U-Net++ as a benchmark model, replacing the original convolutional blocks with residual attention modules to deepen the depth and complexity of the network, and using dual attention modules to adaptively adjust features in the spatial and channel dimensions respectively to enhance the feature extraction ability for pathological images. Finally, a multi-scale feature fusion module is added in the middle layer to obtain a larger receptive field and context information, improving the boundary segmentation effect of the lesion region and small target region. Finally, a fully connected CRF is used for image post-processing to restore the detailed part of the image and further improve the accuracy of cancerous region segmentation.

[0007] To achieve the above purpose, the present invention provides a method for segmenting cancerous regions in breast tissue sections, characterized in that the method includes:

[0008] Based on the whole cancerous slide and the annotation of the cancerous region therein, a training set and a test set are made;

[0009] A segmentation model based on multi-scale fusion and attention mechanism is designed, and the training set is used to train the segmentation model to obtain a trained segmentation model;

[0010] The whole slide image to be detected is input into the trained segmentation model to obtain a segmentation result in the form of a heat map;

[0011] The fully connected conditional random field CRF model is applied to refine the segmentation result to obtain image patches;

[0012] The obtained image patches are stitched together to form a segmentation result of the whole slide.

[0013] In one implementation, the step of making a training set and a test set based on the whole cancerous slide and the annotation of the cancerous region therein includes:

[0014] Based on a pre-provided dataset, obtain the non-cancerous regions, cancerous regions, and corresponding annotations of the cancerous regions for each whole cancer slide, and convert each whole cancer slide into a corresponding cancer region mask image. In the mask image, the pixel value of the cancerous region is 255, and the pixel value of the non-cancerous region is 0;

[0015] Use the Ostu segmentation method to convert the whole cancer slide into a grayscale image, and find the grayscale value when the variance between the tissue region and the background is the largest according to the pixel distribution. Convert the extracted tissue region into a tissue region mask file corresponding to the slide according to the obtained grayscale value;

[0016] Randomly sample image patches of 256×256 pixel size in the non-cancerous regions and cancerous regions of the whole cancer slide respectively, and perform segmentation operations on the corresponding positions of the mask image of the cancer slide to extract mask patches of the same size; the proportion of the tissue region contained in the extracted image patch in the whole image patch should be not less than 50%, and at the same time, the normal image patch does not contain cancer pixels, and the tumor image patch should contain at least one cancer pixel;

[0017] Use the Reinhard color transfer method to normalize the color of the pathological tissue image patches, and use the histogram equalization technique to enhance the image patches to more prominently show the pathological cell-level features of the breast tissue. Finally, divide the dataset into a training set and a test set according to a preset ratio.

[0018] In one implementation, the steps of designing a segmentation model based on multi-scale fusion and attention mechanism include:

[0019] Use the U-Net++ network model as the backbone network of the segmentation model, and perform operations on the decoder part of the original U-Net++ network model to reduce a large number of skip structures, so that each node in the decoder only saves one skip structure and is directly connected to the preset side nodes;

[0020] Use residual attention blocks to replace the convolutional blocks of the original U-Net++ network model to obtain a segmentation model based on multi-scale fusion and attention mechanism;

[0021] For a given intermediate feature map input F, pass through two convolutional blocks, and then input the features into the spatial attention module and the channel attention module respectively. Fuse the features output by the two modules, and use the identity mapping of the input feature and the fused dual-attention feature as the output, where each convolutional block contains a normalization layer BN, a ReLU layer, and a convolutional layer with a convolutional kernel of 3×3;

[0022] Add a multi-scale feature fusion module to the intermediate layer between the encoder and the decoder. The multi-scale feature fusion module consists of three dilated convolutions with different dilation rates and skip connections, used to obtain features with three receptive fields, capture multi-scale image attributes and global shape features of cancerous tissues;

[0023] The multi-scale feature fusion module consists of three dilated convolutions with dilation rates of 6, 12, and 18 respectively. The convolutional kernel of the dilated convolution is 3×3, the stride is 1, and a ReLU layer follows each dilated convolution; the features of three scales and the input features after upsampling operation are fused in the channel dimension using skip connections; after fusing the features of different scales, use 1×1 convolution to convert them into feature maps of a fixed size.

[0024] In one implementation, training the segmentation model using the training set to obtain a trained segmentation model includes:

[0025] Use flipping, rotation, and cropping methods to perform data augmentation on the image patches and mask maps in the training set. Use the Adam optimization algorithm with a relatively fast convergence speed. During training, use the cross-entropy loss function as the loss function in the initial stage. When the model iterates to a certain degree of convergence, use the Dice loss function to adjust the model parameters.

[0026] In one implementation, the step of inputting the whole slide image to be detected into the trained segmentation model to obtain a segmentation result in the form of a heat map includes:

[0027] Make a segmentation dataset for the whole slide image to be detected. Use a sliding window to extract image patches to be detected with a size of 256×256 pixels in the tissue area of the pathological image, and perform color normalization on the image patches to be detected;

[0028] Input the image patches to be detected into the trained segmentation model to predict the cancerous area; in the prediction result, each pixel in the image patch to be detected can obtain a cancer probability;

[0029] According to the cancer probability values of all pixels, draw a heat map of the cancerous area of the image patch to be detected, and display the area where the malignant tumor is located in a set color area. This segmentation result is used to predict the existence and location of the cancerous area.

[0030] In one implementation, the step of refining the segmentation result using the fully connected conditional random field (CRF) model to obtain an image patch includes:

[0031] The structure of the fully connected conditional random field (CRF) model includes:

[0032] Use the Gibbs energy distribution function to represent the label distribution E(x|I) in the image:

[0033]

[0034] Among them, the unary term energy function θ i (x i ) = -logP(x i ), where P(x i ) represents the predicted cancer probability value of the i-th pixel in the image patch. The larger the predicted cancer probability value, the smaller the energy and the more accurate the prediction.

[0035] The binary term energy function θ ij (x i , x j ) is specifically expressed as:

[0036] θ ij (x i , x j ) = μ(x i , x j )(ω1k1, ω2k2)

[0037]

[0038]

[0039] Among them, k1 and k2 are the Gaussian kernels of the binary term energy function; σ α , σ β , σ γ are the standard deviation parameters of the Gaussian kernel; ω1 and ω2 are the weights of the Gaussian function; p i , p j and I i , I j respectively represent the position information and color information of pixel points i and j, and μ(x i , x j ) represents the compatibility measure between two labels.

[0040] In one implementation, the step of splicing the obtained image patches to form the segmentation result of the whole slice includes:

[0041] Convert the image patches in the whole slice that have undergone image post-processing operations into binary images through a set threshold, where cancer pixels are 1 and normal pixels are 0.

[0042] Splice the processed segmentation maps into an image with the same size as the original whole slice to form the segmentation result of the entire whole slice.

[0043] In addition, the present invention also discloses a device for segmenting cancerous regions in breast tissue sections, including:

[0044] A production module, configured to produce a training set and a test set based on a whole cancer slide and the annotation of the cancerous regions therein.

[0045] A training module, configured to design a segmentation model based on multi-scale fusion and attention mechanism, and train the segmentation model with the training set to obtain a trained segmentation model.

[0046] A segmentation module, configured to input a whole slide image to be detected into the trained segmentation model to obtain a segmentation result in the form of a heat map.

[0047] A refinement module, configured to apply a fully connected conditional random field (CRF) model to refine the segmentation result to obtain image patches.

[0048] A stitching module, configured to stitch the obtained image patches to form a segmentation result of the whole slide.

[0049] Applying the method and device for segmenting cancerous regions of breast tissue sections provided by the embodiments of the present invention, using U-Net++ as a benchmark model, replacing the original convolutional blocks with residual attention modules to deepen the depth and complexity of the network, and using dual attention modules to adaptively adjust features in the spatial dimension and channel dimension respectively to enhance the feature extraction ability for pathological images. Finally, a multi-scale feature fusion module is added in the middle layer to obtain a larger receptive field and context information, improving the boundary segmentation effect of the lesion region and small target region. Finally, a fully connected CRF is used for image post-processing to restore the detailed part of the image, further improving the accuracy of cancerous region segmentation. Description of the Drawings

[0050] Figure 1 is a schematic flowchart of a method for segmenting cancerous regions of breast tissue sections according to an embodiment of the present invention.

[0051] Figure 2 is a schematic structural diagram of a residual attention module according to an embodiment of the present invention.

[0052] Figure 3 is a schematic structural diagram of a multi-scale fusion module according to an embodiment of the present invention.

[0053] Figure 4 is a structural diagram of a segmentation network according to an embodiment of the present invention.

[0054] Figure 5 is a structural diagram of a segmentation network of the prior art. Detailed Embodiments

[0055] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0056] Please refer to Figures 1-4 . It should be noted that the diagrams provided in this embodiment only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0057] Regarding the first problem, researchers often adopt the method of extracting deeper image features by increasing the width and depth of the segmentation network to improve the segmentation accuracy. For example, Ibtehaz et al. proposed MultiResUNet, which uses the MultiRes module to replace each unit in the encoder of U-Net, increasing the network learning ability while performing multi-resolution analysis. Zhou et al. added dense connections to the U-Net network, thus introducing the idea of deep supervision, and integrating U-Net structures of different sizes into a network through a redesigned skip connection path. However, these models have more parameters, consume more memory during training, and increase the training difficulty.

[0058] For the following two problems, researchers usually adopt the method of fusing image features at different magnification levels to improve the segmentation effect in the transition region. Xu et al. used multi-scale features combined with a deep convolutional neural network to segment cancerous regions in pathological image patches. Rijthoven et al. proposed a segmentation network called HookNet, which fuses the context and detail features of tissue regions through multiple encoder and decoder branches to achieve high-resolution semantic segmentation. Schmitz et al. introduced model paths at different spatial scales to implement a multi-encoder fully convolutional neural network with deep fusion while maintaining spatial relationships. However, these methods have high requirements for the dataset, requiring pathological images and corresponding annotations at different magnification levels, and since the network contains multiple branches, the number of network parameters is large, and a large number of training samples are required to prevent the model from overfitting. In addition, Zhang et al. used the idea of the ASPP (Atrous spatial pyramid pooling) module, using dilated convolution to extract multi-scale features and fusing them to obtain more context information. However, when the dilation rate used in this module increases, it is easy to lose local information.

[0059] Therefore, due to problems such as the complexity of breast pathological section images and the too small difference between cancerous cells and normal cells, traditional segmentation models cannot achieve ideal segmentation effects. The present invention designs a method for segmenting cancerous regions of breast tissue sections based on multi-scale fusion and attention mechanism. Using U-Net++ as the baseline model, the original convolutional blocks are replaced with residual attention modules to deepen the depth and complexity of the network, and the dual attention module can adaptively adjust the features in the spatial dimension and channel dimension respectively to enhance the feature extraction ability of pathological images. Finally, a multi-scale feature fusion module is added to the intermediate layer to obtain a larger receptive field and context information, improving the boundary segmentation effect of the lesion region and small target region. Finally, a fully connected CRF is used for image post-processing to restore the detailed part of the image and further improve the accuracy of cancerous region segmentation.

[0060] Such as Figure 1 The present invention provides a method for segmenting cancerous regions of breast tissue sections, including:

[0061] S101, based on the whole cancerous slide and the annotation of the cancerous region therein, making a training set and a test set.

[0062] In the embodiment of the present invention, the dataset provides xml format annotations for the cancerous regions of each whole cancerous slide. Combining the original pathological section (i.e., the whole cancerous slide) and the region annotation, it is converted into a mask image corresponding to the cancerous region of the section. In the mask image, the pixel value of the cancerous tissue region is 255, and the pixel value of the normal tissue region is 0;

[0063] Then the Ostu segmentation method is used to segment the tissue region. Specifically: first, the whole cancerous slide is converted into a grayscale image, and then the grayscale value when the variance between the tissue region and the background is the largest is found according to the pixel distribution; the extracted tissue region is converted into a mask file corresponding to the tissue region of the section. It should be noted that the grayscale is an image parameter extracted by the Ostu segmentation method, and the background region and non-background region are divided according to the change of the grayscale value.

[0064] For the normal tissue region and the cancerous region of the whole cancerous slide respectively, image blocks with a size of 256×256 pixels are randomly sampled. At the same time, the same segmentation operation is performed on the corresponding positions of the mask image corresponding to the section to extract mask blocks of the same size as the pixel-level labels of the image blocks. The proportion of the tissue region contained in the extracted image block in the whole image block needs to be greater than 50%, and at the same time, the normal image block does not contain cancer pixels, and at least one cancer pixel needs to be contained in the tumor image block.

[0065] The numerous obtained image blocks are the original images of the test set and the training set.

[0066] Use the Reinhard color transfer method to normalize the color of the pathological tissue image patches extracted in the above steps, and use histogram equalization technology to enhance the image patches (for example, perform statistical analysis on the three color channels of each screened image patch, calculate their pixel means and standard variances respectively, and then perform linear transformation on each pixel value in the channel), so as to more prominently show the pathological cell-level features of breast tissue. Finally, divide the data set into a training set and a test set in a ratio of 8:2.

[0067] S102, design a segmentation model based on multi-scale fusion and attention mechanism, and use the training set to train the segmentation model to obtain a trained segmentation model.

[0068] Use the U-Net++ network as the backbone network of the segmentation model, and in the decoder part of the original network structure, reduce the cross-scale dense skip structure, so that each node in the decoder only saves the skip structure once and is directly connected to the node on the right. This approach can effectively reduce the number of parameters of the model and make the network pay more attention to the feature extraction part, improving the recognition accuracy of the segmentation target.

[0069] Use residual attention blocks to replace the convolutional blocks in the model to deepen the network depth, thereby improving the network's ability to distinguish between normal tissue regions, abnormal tissue regions, and the background, and improving the segmentation effect of the transition region between normal and abnormal tissues, obtaining the residual attention module as shown in Figure 2 Figure.

[0070] In a specific embodiment, the residual attention block structure includes:

[0071] For a given intermediate feature map input F, first pass through two convolutional blocks, then input the features into the spatial attention module and the channel attention module respectively, then fuse the features output by the two modules, and finally use the input feature and the double-attention feature after fusion as the output through identity mapping. Each convolutional block contains a normalization layer (BatchNormalization, BN), a ReLU layer, and a convolutional layer with a convolution kernel of 3×3. Compared with the original convolutional block, the residual attention block deletes the original max pooling operation, and uses a convolution with a stride of 2 in the first convolutional block to reduce the size of the feature map to half of the original to achieve downsampling, thereby improving the calculation speed of the model. At the same time, each convolutional layer uses zero padding to ensure the consistency of the input size.

[0072] In Figure 2 the intermediate feature map input F passes through BN and a 1*1 convolution all the way, and the other way gets X through two cascaded convolutional blocks, and then gets X SA and the channel attention mechanism module X CA, the combined X undergoes matrix multiplication and element-wise addition to obtain X'. The convolution output of X' and the 1*1 convolution is obtained by element-wise addition to get the output result F'.

[0073] The spatial attention mechanism includes: max pooling and average pooling layers. After channel stacking, they are input into a 3*3 convolution, and after BN and Sigmoid, X is output. SA 。

[0074] The channel attention mechanism includes: max pooling and average pooling layers. After passing through a shared MLP, they are respectively input into two ReLU layers. After element-wise addition, they are input into Sigmoid.

[0075] Add a multi-scale feature fusion module to the intermediate layer between the encoder and the decoder. The multi-scale feature fusion module consists of three dilated convolutions with different dilation rates and skip connections, used to obtain features with 3 receptive fields, capture multi-scale image attributes and the global shape features of cancerous tissues.

[0076] In a specific embodiment, such as Figure 3 In, the structure of the multi-scale feature fusion module includes:

[0077] Such as Figure 3 As shown, the multi-scale feature fusion module consists of three dilated convolutions with dilation rates of 6, 12, and 18 respectively. The convolution kernel of the dilated convolution is 3×3, the stride is 1, and a ReLU layer follows each dilated convolution. The features of 3 scales and the input features after upsampling operation are fused in the channel dimension using skip connections. After fusing the features of different scales, they are converted into a feature map of a fixed size using 1×1 convolution, thereby improving the network's ability to adjust the feature weights of different receptive fields and promoting the fusion of multi-scale features.

[0078] Use methods such as flipping, rotating, and cropping to perform data augmentation on the image patches and mask images in the training set to increase the diversity of the training set. Use the Adam algorithm with a relatively fast convergence speed as the training algorithm (use the Adam optimization algorithm with a relatively fast convergence speed for gradient descent). During the training process, use the cross-entropy loss function as the loss function in the initial stage. When the model iterates to a certain degree of convergence, use the Dice loss function to fine-tune the model parameters.

[0079] See Figure 4, which is provided by an embodiment of the present invention, is a segmentation network structure diagram based on multi-scale fusion and attention mechanism. The first layer is 5 RABs (residual attention modules) with skip connections, the second layer is 3 RABs with skip connections, and the first 3 RABs in the first layer and the first 2 RABs in the second layer are downsampling and upsampling in sequence, with a total of 2 upsamplings and 2 downsamplings. The third RAB in the second layer and the fourth RAB in the first layer are upsampled, and the fourth RAB in the second layer and the fifth RAB in the first layer are upsampled.

[0080] The third layer is that RAB, MFB (multi-scale feature fusion module), and RAB are connected in sequence with skip connections. And the first 3 RABs in the second layer and the RAB and MFB in the third layer are downsampling and upsampling in sequence, with a total of 2 upsamplings and 2 downsamplings; the second RAB in the third layer and the fourth RAB in the second layer are upsampled.

[0081] The fourth layer is 2 RABs with skip connections. The first RAB in the third layer downsamples the first RAB in the fourth layer. The first RAB in the fourth layer is connected with the MFB in the third layer for upsampling, and the second RAB in the fourth layer is connected with the second RAB in the third layer for upsampling.

[0082] The fifth layer is 1 MFB, which is respectively connected with the first RAB in the fourth layer for downsampling and the second RAB in the fourth layer for downsampling.

[0083] Taking U-Net++ as the backbone network of the model, in order to reduce the model parameters, the skip connections in the decoder part of U-Net++ are reduced. The skip connection structure is the arc connection as Figure 5 shown, Figure 5 which is the dense skip connection structure in the original Unet++ network. Because there are too many connections, it may affect the performance of the network. Therefore, it is optimized. The improved architecture of the present invention is as Figure 4 shown by deleting the skip connections across multiple layers. At the same time, the convolutional parts in the encoder and decoder are replaced by residual attention blocks to increase the depth of the model and make the model pay more attention to the features of the cancerous region; a multi-scale fusion module is added to the middle layer to enable the model to obtain a larger receptive field and acquire more context features in the case of the existing single-resolution dataset.

[0084] Aiming at the problem that the feature level of pathological tissue sections is relatively deep, an end-to-end segmentation network is built. A simplified U-Net++ architecture is used to reduce a large number of skip connections. While reducing the training parameters, richer and more complex image features are captured, enhancing the feature extraction ability for pathological tissue sections.

[0085] To address the problem of the varying morphological sizes of cancerous regions to be segmented, a multi-scale feature fusion module is adopted in the central part of the network. Without using pooling methods and while retaining image information, multi-receptive field features are obtained, thereby enhancing the network's feature extraction ability and robustness for multi-resolution pathological sections.

[0086] Regarding the interpretability and segmentation accuracy of the network, a residual attention module is designed to replace the convolutional blocks in the encoder and decoder, thereby deepening the depth and complexity of the network. The feature responses of each layer of the network to the background region are suppressed, and the recognition accuracy for the segmentation target is improved.

[0087] S103: Input the whole-slide image to be detected into the trained segmentation model to obtain a segmentation result in the form of a heat map.

[0088] Make a whole-slide segmentation dataset to be detected. Use a sliding window to extract image patches of size 256×256 pixels in the tissue region of the pathological image, and perform color normalization on the image patches.

[0089] Directly input the image patches into the trained segmentation model based on multi-scale fusion and attention mechanism for the prediction of cancerous regions. In the prediction result, each pixel in the image patch can obtain a cancer probability p, where p ∈ [0, 1].

[0090] According to the cancer probability values p of all pixels, draw a heat map of the cancerous region of the image patch and display the region where the malignant tumor is located in red. This segmentation result can predict the existence and rough location of the cancerous region, but cannot truly outline their boundaries.

[0091] S104: Apply a fully connected conditional random field (CRF) model to refine the segmentation result to obtain an image patch.

[0092] Use the heat map of the cancerous region of the image patch in S103 as the input and input it into the improved fully connected CRF model to restore the detailed local structure in the heat map and further improve the segmentation accuracy.

[0093] In a specific embodiment, the structure of the improved fully connected CRF model includes:

[0094] Use the Gibbs energy distribution function to represent the distribution of a certain label in the image, as follows:

[0095]

[0096] Among them, the unary term energy function θ i (x i ) = -logP(x i ), P(x i) represents the predicted cancer probability value of the i-th pixel in the image patch. If the predicted cancer probability value is larger, the energy is smaller and the prediction is more accurate. The binary energy function allows for efficient inference when using a fully connected graph, and the function is as follows:

[0097] θ ij (x i ,x j )=μ(x i ,x j )(ω1k1,ω2k2)

[0098]

[0099]

[0100] where k1 and k2 are the Gaussian kernels of the binary energy function; σ α , σ β , σ γ are the standard deviation parameters of the Gaussian kernel; ω1 and ω2 are the weights of the Gaussian function; p i , p j and I i , I j represent the position information and color information of pixel points i and j respectively; k1 tends to group pixel points with similar positions and colors into the same label category; k2 can fuse isolated points into the same label as the surrounding pixels, thereby increasing the detailed information of the segmentation result; μ(x i ,x j ) represents the compatibility measure between two labels. If the pixel label categories of x i and x j are not compatible with each other, then the corresponding value of μ(xi,xj) is larger, and the overall value of the energy function will also increase.

[0101] S105, splice the obtained image patches to form the segmentation result of the whole slide.

[0102] Set the threshold to 0.5, perform binary segmentation on the pixels according to the threshold, and convert the image patches in the whole slide that have undergone image post-processing operations into binary images, where cancer pixels are 1 and normal pixels are 0.

[0103] Splice the processed segmentation maps into an image with the same size as the original whole slide to form the segmentation result of the entire whole slide.

[0104] For the problem that the tissue boundary of the segmentation result generated by the segmentation network lacks detailed local structure, use a fully connected conditional random field for image post-processing to restore the detailed part of the image and further improve the segmentation accuracy.

[0105] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for segmenting cancerous regions in breast tissue sections, characterized in that, The method includes: Based on the annotation of the whole cancer slide and the cancerous regions therein, a training set and a test set are made; A segmentation model based on multi-scale fusion and attention mechanism is designed, and the training set is used to train the segmentation model to obtain a trained segmentation model; The whole slide image to be detected is input into the trained segmentation model to obtain a segmentation result in the form of a heat map; The fully connected conditional random field (CRF) model is applied to refine the segmentation result to obtain image patches; The obtained image patches are stitched together to form the segmentation result of the whole slide; The step of designing the segmentation model based on multi-scale fusion and attention mechanism includes: Taking the U-Net++ network model as the backbone network of the segmentation model, and optimizing the decoder part of the original U-Net++ network model to reduce a large number of skip connections, so that each node in the decoder only saves one skip connection and is directly connected to the preset side nodes; Using residual attention blocks to replace the convolutional blocks of the original U-Net++ network model; For the given input feature map F, the residual attention block is input into the spatial attention module and the channel attention module respectively after passing through two convolutional blocks, the features output by the two modules are fused, and the input feature map F and the double attention feature after fusion are used as the output through identity mapping; Adding a multi-scale feature fusion module to the intermediate layer between the encoder and the decoder; the multi-scale feature fusion module is composed of three dilated convolutions with different dilation rates and with skip connections; the features of three scales and the input features after upsampling operation are fused in the channel dimension by using the skip connections; the fused features are converted into a feature map with a fixed size by using a 1×1 convolution.

2. The method for segmenting cancerous regions of breast tissue sections according to claim 1, wherein The step of making the training set and the test set based on the annotation of the whole cancer slide and the cancerous regions therein includes: Based on the pre-provided data set, the non-cancerous regions, cancerous regions and the corresponding annotations of each whole cancer slide therein are obtained, and each whole cancer slide is converted into a corresponding cancer region mask map. In the mask map, the pixel value of the cancerous region is 255, and the pixel value of the non-cancerous region is 0; Using the Ostu segmentation method, the whole cancer slide is converted into a grayscale image, and the grayscale when the tissue region and the background variance are the largest is found according to the pixel distribution. The extracted tissue region is converted into a tissue region mask file corresponding to the slide according to the obtained grayscale; In the non-cancerous regions and cancerous regions of the whole cancer slide respectively, image patches with a size of 256×256 pixels are randomly sampled, and segmentation operations are performed on the corresponding positions of the mask image of the whole cancer slide to extract mask patches of the same size; the proportion of the tissue region contained in the extracted image patch in the whole image patch should be not less than 50%, and at the same time, the normal image patch does not contain cancer pixels, and the tumor image patch should contain at least one cancer pixel; The color of the extracted image patches is normalized by the Reinhard color transfer method, and the histogram equalization technique is used to enhance the image patches to more prominently show the pathological cell-level features of the breast tissue. Finally, the data set is divided into a training set and a test set according to a preset ratio.

3. The method for segmenting cancerous regions of breast tissue sections according to claim 1, wherein Training the segmentation model with the training set to obtain a trained segmentation model includes: Performing data augmentation on the image patches and mask images in the training set using flipping, rotation, and cropping methods, using the Adam optimization algorithm with a relatively fast convergence rate, and adopting the cross-entropy loss function as the loss function in the initial stage during training. When the model converges to a certain extent, the Dice loss function is used to adjust the model parameters.

4. The method for segmenting cancerous regions of breast tissue sections according to any one of claims 1-3, characterized in that, The step of inputting the whole-slide image to be detected into the trained segmentation model to obtain a segmentation result in the form of a heat map includes: Making a segmentation dataset for the whole-slide image to be detected, using a sliding window to extract image patches to be detected with a size of 256×256 pixels in the tissue region of the pathological image, and performing color normalization on the image patches to be detected; Inputting the image patches to be detected into the trained segmentation model to predict the cancerous regions; in the prediction results, each pixel in the image patches to be detected can obtain a cancer probability; According to the cancer probability values of all pixels, drawing a heat map of the cancerous regions of the image patches to be detected, and displaying the region where the malignant tumor is located in a set color region. This segmentation result is used to predict the existence and location of the cancerous regions.

5. The method for segmenting cancerous regions of breast tissue sections according to claim 1, wherein The step of refining the segmentation result using the fully connected conditional random field (CRF) model to obtain image patches includes: The fully connected conditional random field (CRF) model includes: Representing the label distribution situation E(x|I) in the image using the Gibbs energy distribution function: Among them, the unary term energy function θ i (x i ) = -logP(x i ), where P(x i ) represents the predicted cancer probability value of the i-th pixel in the image block. If the predicted cancer probability value is larger, the energy is smaller and the prediction is more accurate. The binary term energy function θ ij (x i , x j ) is specifically expressed as: θ ij (x i ,x j ) = μ(x i ,x j )(ω1k1, ω2k2) Among them, k1 and k2 are the Gaussian kernels of the binary term energy function; σ α , σ β , σ γ are the standard deviation parameters of the Gaussian kernels; ω1 and ω2 are the weights of the Gaussian functions; p i , p j and I i , I j respectively represent the position information and color information of pixel points i and j, and μ(x i , x j ) represents the compatibility measure between two labels.

6. The method for segmenting cancerous regions of breast tissue sections according to claim 1, wherein The step of stitching the obtained image patches to form a segmentation result of the whole slide includes: Converting the image patches in the whole slide that have undergone image post-processing operations into binary images through a set threshold, where cancerous pixels are 1 and normal pixels are 0; Stitching the processed segmentation maps into an image with the same size as the original whole slide to form a segmentation result of the entire whole slide.

7. A device for segmenting cancerous regions in breast tissue sections, characterized in that, Including: A production module for making a training set and a test set based on the cancerous whole slides and the annotations of the cancerous regions therein; A training module for designing a segmentation model based on multi-scale fusion and attention mechanism, and training the segmentation model with the training set to obtain a trained segmentation model; A segmentation module for inputting the whole-slide image to be detected into the trained segmentation model to obtain a segmentation result in the form of a heat map; including using the U-Net++ network model as the backbone network of the segmentation model, and optimizing the decoder part of the original U-Net++ network model to reduce a large number of skip structures, so that each node in the decoder only saves one skip structure and is directly connected to the preset side nodes; Replace the convolutional blocks of the original U-Net++ network model with residual attention blocks; for a given input feature map F, the residual attention block passes through two convolutional blocks and is respectively input into the spatial attention module and the channel attention module, fuses the features output by the two modules, and the input feature map F and the dual-attention features after fusion are used as the output through identity mapping; add a multi-scale feature fusion module in the intermediate layer between the encoder and the decoder; the multi-scale feature fusion module consists of three dilated convolutions with different dilation rates with skip connections; use the skip connections to fuse the features of three scales and the input features after upsampling operation in the channel dimension; use 1×1 convolution to convert the fused features into a feature map of a fixed size; A refinement module for refining the segmentation result using a fully connected conditional random field (CRF) model to obtain image patches; A stitching module for stitching the obtained image patches to form a full-slice segmentation result.

Citation Information

Patent Citations

  • Method for segmenting choroidal atrophy in fundus medical image

    CN112634234A