Method and device for pathological image segmentation and classification based on semi-supervised learning

Through the U-shaped network structure and weak supervision segmentation method based on Swin-Transformer, the problems of large image size, difficult labeling, complex multi-scale and unbalanced categories in pathological image analysis are solved, and efficient pathological image segmentation and classification are achieved.

CN114037720BActive Publication Date: 2025-07-04BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111211187.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-07-04
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

In pathological image analysis, there are problems such as huge image size, difficult labeling data, complex multi-scale information, and unbalanced categories, resulting in poor segmentation and classification performance.

Method used

Using a U-shaped network structure based on Swin-Transformer Blocks, a weakly supervised segmentation method combined with point annotation and geometric constraints is used to extract multi-scale information, dense connections and edge prior knowledge to achieve simultaneous segmentation and classification.

Benefits of technology

Under weak annotation conditions, the segmentation and classification performance of pathological images is improved, gradient vanishing and overfitting are reduced, and category imbalance problem is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114037720B_ABST
    Figure CN114037720B_ABST
Patent Text Reader

Abstract

Method and device for pathological image segmentation and classification based on semi-supervised learning, effectively solving the problem of image scale change, changing the network structure to enable it to perform segmentation and classification simultaneously, solving the problem of class imbalance of samples in whole-slide images, and improving the performance of classification and segmentation. The method includes: (1) Using a U-shaped network structure based on Swin-Transformer Blocks to extract multi-scale information from image data; (2) In upsampling, adopting dense connections; (3) Using point annotations of cell images for weakly supervised segmentation, adopting negative boundary sparse supervision of annotation points and geometric constraints, and combining the Voronoi partitioning strategy of spatial expansion from points to regions to perform rough segmentation; in the fine segmentation stage, further using the edge prior knowledge in the unmodified image to adjust the kernel contour through a contour-sensitive constraint function; (4) In the network, modifying the last linear mapping layer to enable it to input the results of segmentation and classification simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method for pathological image segmentation and classification based on semi-supervised learning, and an apparatus for cell pathological image segmentation and classification based on semi-supervised learning. Background Art

[0002] In pathological image analysis, images are mainly stored as whole slide images (WSIs). Traditional pathological sections are scanned using a digital scanner to obtain high-resolution digital images, which are then seamlessly stitched and integrated by a computer to produce visual digital images. Compared with traditional glass slides, it has the advantages of easy storage, non-fading, non-loss, and easy retrieval. When segmenting cells in pathological images, annotating images is a time-consuming and laborious task that requires a large amount of manpower to perform pixel-level annotation for each cell. Usually, a complete WSI section has tens of millions of cell images. Currently, there is less research on point-supervised segmentation kernels. Due to the complex segmentation background and multi-scale information in pathological images, the segmentation and classification performance of images will be affected.

[0003] The difficulties in pathological image analysis are as follows:

[0004] 1. The image size is huge, it is difficult to label data, and the image contains many different scales of information.

[0005] 2. When performing segmentation, a gold standard for segmentation is required instead of a class label, and obtaining the gold standard for segmentation is time-consuming and laborious.

[0006] 3. The classes of the images have a class imbalance problem between positive and negative samples. Summary of the Invention

[0007] To overcome the defects of the prior art, the technical problem to be solved by the present invention is a method for pathological image segmentation and classification based on semi-supervision, which can effectively solve the image segmentation problem under weak annotation, and change the network structure so that it can perform segmentation and classification simultaneously, and can solve the problem of difficult image annotation.

[0008] The technical solution of the present invention is: This method for pathological image segmentation and classification based on semi-supervised learning includes the following steps:

[0009] (1) Use a U-shaped network structure based on Swin-Transformer Blocks to adaptively extract multi-scale information from image data;

[0010] (2) Use bilinear interpolation dense connection during the upsampling process to reduce the loss of fineness of the decoder while alleviating the problems of gradient disappearance and overfitting;

[0011] (3) Weakly supervised segmentation is performed using point annotations of cell images. In the rough segmentation stage, negative boundary sparse supervision of annotation points and geometric constraints is adopted, combined with the Voronoi partitioning strategy for spatial expansion from points to regions. In the fine segmentation stage, through the contour-sensitive constraint function, the edge prior knowledge in the unmodified image is further utilized to adjust the nuclear contour.

[0012] (4) In the network, the results of segmentation and classification are output simultaneously by modifying the last linear mapping layer.

[0013] Through the U-shaped network structure based on Swin-Transformer, the present invention can effectively extract multi-scale information from images. During the upsampling process, a dense connection is adopted, which can reduce the loss of fineness of the decoder and also alleviate the problems of gradient disappearance and overfitting. And the structure of the network is changed so that it can perform classification and segmentation simultaneously. The easier-to-obtain cell point annotations are used for segmentation in two stages. In the first stage, a negative boundary sparse supervision of annotation points and geometric constraints, combined with the Voronoi partitioning strategy for spatial expansion from points to regions, is adopted to perform rough segmentation. In the second stage, a contour-sensitive constraint function is proposed to further utilize the edge prior knowledge in the unmodified image to adjust the nuclear contour.

[0014] There is also provided an apparatus for semi-supervised learning-based pathological image segmentation and classification, which includes:

[0015] An extraction module, based on the U-shaped network structure of Swin-Transformer Blocks, enables the network to adaptively extract multi-scale information from images;

[0016] An upsampling module, adopting a dense connection structure based on bilinear interpolation, reduces the loss of fineness of the decoder and alleviates the problems of gradient disappearance and overfitting;

[0017] A weakly supervised module, using weak annotations of cell images, adopts a strategy based on annotation points and geometric constraints, boundary constraints, combined with the Voronoi partitioning strategy from points to regions for rough segmentation, and in the fine segmentation stage, utilizes the edge prior information of the image to adjust the nuclear contour through the contour-sensitive constraint function;

[0018] A result output module, which is configured to change the last fully connected layer of the network so that it outputs both the segmentation result and the classification result. Description of the Drawings

[0019] Figure 1 The flowchart of the first semi-supervised learning-based pathological image segmentation and classification method according to the present invention is shown.

[0020] Figure 2Shows the structure diagram of the Swin-Transformer Block adopted by the present invention.

[0021] Figure 3 Respectively represent hierarchical connection (a), residual connection (b), and dense connection (c).

[0022] Figure 4 Shows a schematic diagram of the bilinear interpolation algorithm.

[0023] Figure 5 Shows the overall network structure diagram of the deep learning adopted by the present invention.

[0024] Figure 6 Shows the original image and the point annotation image (a) Voronoi region constraint image (b) point interval image (c).

[0025] Figure 7 Shows the flowchart of the first-stage rough segmentation.

[0026] Figure 8 Shows the original image in the second stage (a) the rough segmentation image in the first stage (b) the Kirsch operator edge extraction image (c) the morphological edge extraction image (d) the sparse region image (e). Detailed implementation manner

[0027] As Figure 1 shown, this method for pathological image segmentation and classification based on semi-supervised learning includes the following steps:

[0028] (1) Use a U-shaped network structure based on Swin-Transformer Blocks to adaptively extract multi-scale information from image data;

[0029] (2) Use bilinear interpolation dense connection in the upsampling process to reduce the loss of fineness of the decoder while alleviating the problems of gradient disappearance and overfitting;

[0030] (3) Use the point annotation of cell images for weakly supervised segmentation. In the rough segmentation stage, adopt negative boundary sparse supervision of annotation points and geometric constraints, and combine the Voronoi partitioning strategy of spatial expansion from points to regions; in the fine segmentation stage, further use the edge prior knowledge in the unmodified image to adjust the kernel contour through a contour-sensitive constraint function;

[0031] (4) In the network, modify the last linear mapping layer to output the segmentation and classification results simultaneously.

[0032] Based on the U-shaped network structure of Swin-Transformer, the present invention can effectively extract multi-scale information from images. During the upsampling process, a dense connection is adopted, which can reduce the loss of fineness of the decoder and also alleviate the problems of gradient disappearance and overfitting. Moreover, the network structure is changed to enable it to perform classification and segmentation simultaneously; two-stage segmentation is carried out using more easily obtained cell point annotations. In the first stage, a negative boundary sparse supervision of annotation points and geometric constraints is adopted, combined with the Voronoi partitioning strategy of spatial expansion from points to regions to perform rough segmentation. In the second stage, a contour-sensitive constraint function is proposed to further utilize the edge prior knowledge in the unmodified image to adjust the kernel contour.

[0033] Preferably, in the step (1), the Swin-Transformer Bolcks perform adaptive feature extraction based on the multi-head attention module with moving windows, residual connections, and multi-layer perceptrons. The calculation method when the Swin-Transformer Bolcks of the i-th layer perform feature extraction is formula (1)

[0034]

[0035] Where: and z l respectively represent the outputs of the (S)W-MSA and multi-layer perceptron of the i-th layer;

[0036] The self-attention mechanism is calculated by formula (2):

[0037]

[0038] Where: represent the query, key, and value matrices, M 2 and d respectively represent the number of patches under a window and the dimensions of the query and key, and B represents the values in the confusion matrix in it.

[0039] Preferably, in the step (2), during the upsampling process, a dense connection structure is adopted to reduce the loss of fineness of the decoder and alleviate the problems of gradient disappearance and overfitting at the same time. In the dense connection, the upsampling process is formula (3):

[0040]

[0041] Where: f n represents the interpolation method for upsampling, and this method is the bilinear interpolation method; in the bilinear interpolation, the known function Q 11 =(x1,y1), Q 12 =(x1,y2), Q21 =(x2, y1), Q 22 =(x2, y2) for the values of four points, the formula for bilinear interpolation of a pixel point (x, y) is (4):

[0042]

[0043] Preferably, in the step (3), the cell image is first subjected to point annotation, and two distance maps are generated respectively focusing on positive pixels and negative pixels, including the distance map of point annotation and the edge map

[0044] The distance map of point annotation is used to focus on positive pixels with high confidence. Assuming that the annotation points of each nucleus are close to the center of the nucleus, and then the point annotation is expanded to a reliable nuclear supervision area through a distance filter Each element in is calculated by formula (5) as follows:

[0045]

[0046] where: m and n are the marked points of the distance annotation map of point annotation, and α is a scaling parameter for controlling the distribution ratio;

[0047] The Voronoi diagram is used to focus on negative pixels with high confidence, denoted as The partition edges are obtained through the Voronoi diagram, and these edges can be further expanded by the rapidly decreasing response of the distance filter (formula (5)); used to describe negative pixels with high confidence.

[0048] Preferably, in the weak supervised learning of the step (3):

[0049] First, the polar loss function is used to update the parameters of Swin-Transformer-Unet, and is used to represent the output segmentation map. The plolar loss function is formula (6):

[0050]

[0051] where H(x) performs self-supervised learning on the output segmentation map by correcting it into a binary image mask;

[0052] At the same time, two sparse loss functions are set to update the parameters in the Transformer network, respectively focusing on some positive and negative pixels:

[0053]

[0054]

[0055] Wherein: the (·) dot operation represents pixel-by-pixel multiplication, the ReLu operation extracts a reliable weight mask for sparse loss calculation, and let L point only focus on positive pixels with high confidence, and L voronoi only focus on negative pixels with high confidence.

[0056] Preferably, in the step (3), in the first-stage rough segmentation stage, the initial segmentation map is used as an extension of the point annotation in the initial state; the extended point distance map is used to iterate the segmentation model, and these maps are updated by the latest trained model, and the point distance map is updated according to Equation (5), where the annotation map P is replaced by the rough segmentation result of the previous round; this operation is repeated several times to obtain the rough segmentation result denoted as R coarse .

[0057] Preferably, in the step (3), in the first-stage fine segmentation stage, the local region map supervision method is used; by extracting the apparent contour of the input image as additional supervision, the edge map is first refined and the result is denoted as E r :

[0058] E r =(dilation(R coarse , k)-erosion(R coarse , k))&E Kirsch (9)

[0059] Wherein, dilation and erosion are respectively the dilation and erosion morphological operations of the image on k pixels, and E Kirsch represents the image after the Kirsch operator extracts the edges of the input image.

[0060] Preferably, in the second stage of the step (3), in the sparse supervised learning, in order to achieve supplementary boundary supervision, an additional contour-sensitive loss L contour is used to fine-tune the nuclear contour. Similarly, the local region map is used for supervision

[0061]

[0062] Wherein, E r represents the refined edge map, and Kirsch represents the Kirsch operator.

[0063] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the methods of the above embodiments, and the storage medium can be: ROM / RAM, magnetic disk, optical disk, memory card, etc. Therefore, corresponding to the method of the present invention, the present invention also simultaneously includes an apparatus for pathological image segmentation and classification based on semi-supervised learning, and this apparatus is usually represented in the form of functional modules corresponding to the steps of the method. The apparatus includes:

[0064] An extraction module, based on the U-shaped network structure of Swin-Transformer Blocks, enables the network to adaptively extract multi-scale information from images;

[0065] An upsampling module, adopting a dense connection structure based on bilinear interpolation, reduces the loss of decoder fineness, alleviates the problems of gradient disappearance and overfitting;

[0066] A weak supervision module, using the weak annotations of cell images, adopts a strategy based on annotation points, geometric constraints, boundary constraints, and combines the Voronoi partitioning strategy from points to regions for rough segmentation. In the fine segmentation stage, using the prior information of the image edge, through a contour-sensitive constraint function, the nuclear contour is adjusted;

[0067] A result output module, which is configured to change the last fully connected layer of the network so that it outputs both the segmentation result and the classification result.

[0068] As described above, the above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for pathological image segmentation and classification based on semi-supervised learning, characterized in that: It includes the following steps: (1) Use a U-shaped network structure based on Swin-Transformer Blocks to adaptively extract multi-scale information from image data; (2) Use bilinear interpolation dense connection during the upsampling process to reduce the loss of fineness in the decoder while alleviating the problems of gradient disappearance and overfitting; (3) Use the point annotations of cell images for weakly supervised segmentation. In the rough segmentation stage, use negative boundary sparse supervision of annotation points and geometric constraints, combined with the Voronoi partitioning strategy of spatial expansion from points to regions; in the fine segmentation stage, further use the edge prior knowledge in the unmodified image to adjust the nuclear contour through a contour-sensitive constraint function; (4) In the network, modify the last linear mapping layer to output the segmentation and classification results simultaneously; In the step (3), point annotation is first performed on the cell image, and two distance maps are respectively generated to focus on positive pixels and negative pixels, including the distance map with point annotation and the edge map Distance map with point annotations To focus on positive pixels with high confidence, assuming that the annotated points of each kernel are close to the center of the kernel, and then expanding the point annotations to a reliable kernel supervision area through a distance filter, Each element in is calculated by equation (5) as follows: where: m and t are the coordinates of the marked points of the distance annotation map of the point annotation, and ɑ is the scaling parameter that controls the distribution ratio; Use the Voronoi diagram to focus on high-confidence negative pixels, denoted as Obtain the partition edge of each cell in the image through the Voronoi diagram, and further expand its distance from the edge through the calculation method of the distance filter (Equation 5); Used to describe high-confidence negative pixels.

2. The method for pathological image segmentation and classification based on semi-supervised learning according to claim 1, wherein: In step (1), the Swin-Transformer Bolcks performs adaptive feature extraction based on the multi-head attention module with moving windows, residual connection, and multi-layer perceptron. The calculation method when the l-th layer of Swin-Transformer Bolcks performs feature extraction is formula (1) Among them: WMSA represents the window self-attention mechanism, and SWMSA represents the sliding window self-attention mechanism. and z l respectively represent the outputs after the cross-window attention layer and the fully connected layer; In the WMSA and SWMSA network layers, each window calculates self-attention through formula (2), and the formula is as follows: Wherein: represents the query, key, and value matrices, and B represents the value in the confusion matrix.

3. The method for pathological image segmentation and classification based on semi-supervised learning according to claim 2, wherein: In the weakly supervised learning of step (3): First, use the polar loss function to update the parameters of Swin-Transformer-Unet. Let S represent the output segmentation map, and the polar loss function is formula (6): where f(I) represents the image output after passing through the network, and H(x) represents the calculation process of the self-supervised learning module; At the same time, set two sparse loss functions to update the parameters in the Transformer network, respectively focusing on some positive and negative pixels: Among them: the (·) dot operation represents pixel-by-pixel multiplication, and the ReLu operation extracts a reliable weight mask for sparse loss calculation. Through the settings of these two loss functions, L point only focuses on positive pixels with high confidence, and L voronoi only focuses on negative pixels with high confidence.

4. The method for pathological image segmentation and classification based on semi-supervised learning according to claim 3, characterized in that: In the step (3), in the first-stage rough segmentation stage, the initial segmentation map is used as the expansion of the point annotation in the initial state; the expanded point distance map is used to iterate the segmentation model, and these maps are updated by the latest trained model. The point distance map D is updated according to Equation (5), where the annotation map P is replaced by the rough segmentation result of the previous round; this operation is repeated several times to obtain the rough segmentation result denoted as R coarse .

5. The method for pathological image segmentation and classification based on semi-supervised learning according to claim 4, characterized in that: In the step (3), in the first-stage rough segmentation stage, a local region map supervision method is used; by extracting the apparent contour of the input image as additional supervision, the edge map is first refined and the result is denoted as E r : E r = (dilation(R coarse , k) - erosion(R coarse , k)) & E Edge (9) Among them, dilation and erosion are respectively the dilation and erosion morphological operations performed on the image with k pixels, and E Edge represents the operator for image edge extraction.

6. The method for pathological image segmentation and classification based on semi-supervised learning according to claim 5, wherein: After the rough segmentation stage of the first phase in step (3), in sparse supervised learning, in order to achieve supplementary boundary supervision, an additional contour-sensitive loss L contour is used to fine-tune the kernel contour. Similarly, a local region map is used for supervision Among them, E r represents the edge map after refined processing, and Edge represents the operator for edge extraction.

7. An apparatus for pathological image segmentation and classification based on semi-supervised learning, which is used to execute the method according to claim 1, characterized in that: It includes: An extraction module, a U-shaped network structure based on Swin-Transformer Blocks, enables the network to adaptively extract multi-scale information from images; An upsampling module, adopts a dense connection structure based on bilinear interpolation, reduces the loss of fineness in the decoder, and alleviates the problems of gradient disappearance and overfitting; A weakly supervised module, uses the weakly supervised annotations of cell images, combines annotation point constraints, geometric constraints, and boundary constraints, and combines the Voronoi partitioning strategy from points to regions for rough segmentation. In the fine segmentation stage, use the edge prior information of the image to adjust the nuclear contour through a contour-sensitive constraint function; A result output module, which is configured to change the last fully connected layer of the network so that it outputs both the segmentation result and the classification result.

Citation Information

Patent Citations

  • Interactive annotation refinement method for digital pathological image

    CN111986150A

  • Pathological image segmentation method based on domain adversarial self-supervised learning

    CN113379764A