Pathological tissue semantic segmentation method and device based on pixel-by-pixel contrast learning

Through the pathological tissue semantic segmentation method of image enhancement processing and pixel-by-pixel comparison learning module, the robustness and accuracy of the pathological image segmentation model under color differences and tumor heterogeneity are solved, and the segmentation performance of the model in identifying tissue heterogeneity is improved.

CN120355915APending Publication Date: 2025-07-22THE SECOND AFFILIATED HOSPITAL OF CHONGQING MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510392857.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing pathological image segmentation models are difficult to generalize when facing the challenges of color differences in pathological images and tumor heterogeneity, resulting in a decrease in model robustness and accuracy, especially in areas with blurred boundaries or complex structures.

Method used

The semantic segmentation method of pathological tissue based on pixel-by-pixel comparison learning is adopted. Through image enhancement processing and pixel-by-pixel comparison learning module, the model's learning ability and resolution ability of pathological image features are enhanced, and the robustness of color changes and the segmentation accuracy of heterogeneous tissues is improved.

Benefits of technology

The performance of the model in identifying tissue heterogeneity is improved, the ability to judge pixels of different categories is enhanced, and the segmentation performance and segmentation accuracy of local areas is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355915A_ABST
    Figure CN120355915A_ABST
Patent Text Reader

Abstract

The invention discloses a pathological tissue semantic segmentation method and device based on pixel-by-pixel contrast learning, and the method comprises the steps: carrying out the data preprocessing of a training set through employing a shape-controllable region cutting and mixing data enhancement scheme, and guaranteeing that an input image used for model training at least comprises two semantic categories; extracting multi-scale features by using an encoder, and then splicing and fusing the features by using a decoder; the fused features are learned and constrained by pixel-by-pixel contrast; and further, the features can be mapped to a category space, a pixel prediction result is obtained, and a segmentation result mask is generated. According to the method, segmentation does not depend on a specific dyeing mode, so that the robustness to color change is improved; or the model can more meticulously capture the internal structure difference and boundary information of the tissue, so that the segmentation precision of the heterogeneous tissue is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a pathological tissue semantic segmentation method and device based on per-pixel contrast learning. Background Art

[0002] In modern pathological research, pathological image analysis has become a key tool for cancer diagnosis, treatment, and prognosis assessment. Pathological images are obtained by taking tissue sections through a microscope, and these sections usually undergo different staining techniques, such as hematoxylin-eosin (H&E) staining, to enhance the visibility of tissue structures. Doctors rely on these images to analyze cell morphology, tissue structure, and detect potential tumor regions.

[0003] Due to the complex and time-consuming nature of image analysis tasks, in recent years, deep learning techniques, especially convolutional neural networks (CNNs), have been widely applied to automated pathological image analysis, particularly in tissue segmentation tasks. Currently, for semantic segmentation of pathological images, the commonly used technical solutions are mainly based on CNN and Transformer architectures, especially architectures such as the fully convolutional network (FCN), DeepLab series, U-Net, and SegFormer. These frameworks are basically encoder-decoder architectures to achieve pixel-level prediction. Among them, the encoder extracts features, and the decoder reconstructs based on the features extracted by the encoder. In the process of using artificial intelligence for pathological image analysis, tissue semantic segmentation is one of the most basic steps in pathological image processing. Color differences caused by staining or scanning in pathological images, as well as tumor heterogeneity, may affect the performance of tissue segmentation models.

[0004] Most semantic segmentation models such as FCN, DeepLab series models, and SegFormer are designed for natural images, and the U-Net model is designed for medical image processing. However, the impact of color differences and tumor heterogeneity in pathological images is not further considered in these model designs. The challenges of color differences and tumor heterogeneity in pathological image analysis may affect model performance and generalization. Among them, color differences refer to the color differences presented in pathological images due to staining or scanning. This difference makes it difficult for segmentation methods based on traditional deep learning models to generalize, that is, a model that performs well during training may perform poorly on new images. Tumor heterogeneity is because tumor tissue itself is heterogeneous, and the lesion areas may show significant structural differences in the same image or images of different patients, which makes it difficult for models based on standard segmentation algorithms to accurately identify and segment all lesion areas, especially tumor areas with blurred boundaries or complex morphologies. These challenges greatly limit the performance of existing deep learning segmentation models in clinical practical applications, and the robustness and accuracy of the models often decrease significantly. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a method for semantic segmentation of pathological images. The provided method can make the segmentation independent of specific staining patterns, thereby improving the robustness to color changes; or enable the model to more carefully capture the structural differences and boundary information inside the tissue, thereby improving the segmentation accuracy of heterogeneous tissues, especially in regions with blurred boundaries and complex structures.

[0006] In order to achieve any of the above invention purposes, the technical solutions of the present invention are as follows: On the one hand, the present invention provides a method for semantic segmentation of pathological tissues based on per-pixel contrast learning. The method includes: Obtain a first set of pathological images, where the first set of pathological images includes a plurality of first pathological images, and each of the first pathological images has a class label; After enhancing the first set of pathological images according to the class label, obtain a second set of pathological images, where the second set of pathological images includes a plurality of second pathological images; Obtain a first set of label images, where the first set of label images includes a plurality of first label images, the first label images and the first pathological images are in one-to-one correspondence, and each of the first label images has the same class label as the corresponding first pathological image; After performing the same enhancement processing on the first set of label images according to the class label, obtain a second set of label images, where the second set of label images includes a plurality of second label images; Input the training set into a semantic segmentation model for model training. The semantic segmentation model includes an encoder, a decoder, and a per-pixel contrast learning module. The training set includes the second set of pathological images and the second set of label images; Obtain a third set of pathological images, input the third set of pathological images into the trained semantic segmentation model, and output a segmentation result. The third set of pathological images includes a plurality of third pathological images.

[0007] Since the image enhancement process is added to the method, it is ensured that the images input for model training contain at least two semantic categories, strengthening the learning ability and discrimination ability of the method for features in pathological images.

[0008] In some embodiments, the class labels of all the first pathological images in the first set of pathological images include at least 2; The step of obtaining the second set of pathological images after enhancing the first set of pathological images according to the class label includes: Sequentially select the first pathological images with the class label of the first category in the first set of pathological images as target pathological images; Randomly select a first pathological image with a class label other than the first category from the first pathological image set as a reference pathological image; Intercept a first preset graphic at a first preset position of the reference pathological image; Overlay the intercepted graphic on the first preset position of the target pathological image to obtain an enhanced first pathological image; The enhanced first pathological image and all first pathological images with class labels other than the first category are used as second pathological images, and all second pathological images constitute the second pathological image set.

[0009] The second label image set obtained by performing the same enhancement processing on the first label image set according to the class label includes: Sequentially select a first label image with a class label of the first category in the first label image set as a target label image; Randomly select a first label image with a class label other than the first category from the first label image set as a reference label image; Intercept a second preset graphic at a second preset position of the reference label image, where the second preset position is the same as the first preset position, and the second preset graphic is the same as the second preset graphic; Overlay the intercepted graphic on the second preset position of the target label image to obtain an enhanced first label image; The enhanced first label image and all first label images with class labels other than the first category are used as second label images, and all second label images constitute the second label image set.

[0010] In some embodiments, the first preset graphic is a circle, and the first preset position is the center of the circle. The intercepting the first preset graphic at the first preset position of the reference pathological image includes: Randomly set the center of the circle; Randomly set the radius of the circle, where the radius is less than one-fourth of the height of the reference pathological image; Intercept the preset graphic from the reference pathological image according to the randomly set center of the circle and radius.

[0011] In some embodiments, the inputting the training set into the semantic segmentation model for model training includes: Use an encoder to extract features from the training set; Use a decoder to fuse the extracted features; Use a per-pixel contrast learning module to perform per-pixel contrast learning on the fused features.

[0012] In some embodiments, the encoder includes 4 parts, each part including an overlapping block embedding module, a flattening module, and a Transformer module; the use of the encoder to extract features from the training set includes: Performing an overlapping block embedding operation using the overlapping block embedding module to divide the input image into multiple first image block tensors; the overlapping block embedding operation includes a convolution operation, and the convolution kernel size of the convolution is greater than the stride; Flattening the multiple first image block tensors using the flattening module; Inputting the flattened first image block tensors into the Transformer module; Retaining the high-resolution coarse features generated by the previous part for the decoder to generate a segmentation result, and at the same time using them as the input for the next part to continue obtaining low-resolution features of different scales.

[0013] In some embodiments, the attention calculation formula of the Transformer module is , where Q is the query vector, K is the key vector, V is the key-value vector, and Q, K, and V have the same size of N and C, , N is the sequence length, P is the size of the first image block, C is the number of channels, H is the height of the images in the training set, W is the width of the images in the training set, and dk is the dimension of K; Before inputting the flattened image block tensors into the Transformer module, the following processing is performed: Reducing the sequence length through the Reshape and Linear layers, and reducing N to , where is the reduction ratio.

[0014] In some embodiments, the output of the Transformer module is executed according to the following formula where, is the feature output after the self-attention operation, The operation of is to create a linear layer to convert the input features into the number of hidden layer features, and the outermost MLP is to convert the hidden layer features into the number of output features. GELU is an activation function,

[0015] In some embodiments, the method further includes: The decoder obtains the features extracted by the four parts of the encoder, and performs interpolation operations using bilinear interpolation to adjust the sizes of the features generated by the second, third, and fourth parts of the encoder to be the same as the size of the features generated by the first part; Concatenate the features generated by the first part, and the adjusted features generated by the second, third, and fourth parts; Input the concatenated features into the classification head, which are converted into segmentation masks for semantic segmentation after convolution, batch normalization, and convolution operations in the classification head. The cross-loss function is calculated on the output of the classification head. The cross-loss function uses the following formula , where H and W respectively represent the height and width of the images in the training set, represents the true label, represents the probability that the output of the classification head is , and C is the number of classes; Perform per-pixel contrastive learning to complete model training. Among them, the following formula is used to calculate the contrastive loss where, and respectively represent the class identifications of pixel i and pixel j, z represents the feature map obtained after passing through the feature projection module, , and all represent pixel features, N is the number of pixel points in the feature map, represents the temperature hyperparameter; The total loss function is where takes a value of 0.1.

[0016] In some embodiments, the obtaining of the third pathological image set and inputting the third pathological image set into the trained semantic segmentation model to output a segmentation result includes: Obtain multiple whole-slide images, crop each whole-slide image to obtain the third pathological image set of each whole-slide image, and input the third pathological images in each of the third pathological image sets into the trained semantic segmentation model to output the segmentation result of each third pathological image; The size of the third pathological image is the same as the size of the first pathological image; On the other hand, the present invention provides a pathological image semantic segmentation device based on per-pixel contrastive learning. The device includes an enhancement module and a semantic segmentation model; The enhancement module is used for: Obtain a first set of pathological images, where the first set of pathological images includes multiple first pathological images, and each of the first pathological images has a class label; Perform enhancement processing on the first set of pathological images according to the class label to obtain a second set of pathological images, where the second set of pathological images includes multiple second pathological images; Obtain a first set of label images, where the first set of label images includes multiple first label images, the first label images and the first pathological images are in one-to-one correspondence, and each of the first label images has the same class label as the corresponding first pathological image; Perform the same enhancement processing on the first set of label images according to the class label to obtain a second set of label images, where the second set of label images includes multiple second label images; The semantic segmentation model includes an encoder, a decoder, and a per-pixel contrastive learning module. The semantic segmentation model is trained through an input training set, and the training set includes the second set of pathological images and the second set of label images; the semantic segmentation model is used to: Obtain a third set of pathological images, input the third set of pathological images into the trained semantic segmentation model, and output a segmentation result, where the third set of pathological images includes multiple third pathological images.

[0017] In some embodiments, the obtaining the third set of pathological images, inputting the third set of pathological images into the trained semantic segmentation model, and outputting a segmentation result includes: Obtain multiple whole-slide images, crop each whole-slide image to obtain a third set of pathological images for each whole-slide image, input the third pathological images in each third set of pathological images into the trained semantic segmentation model, and output a segmentation result for each third pathological image; the size of the third pathological image is the same as the size of the first pathological image; On the other hand, the present invention provides a terminal device, including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the pathological image semantic segmentation method as described above.

[0018] On the other hand, the present invention provides a computer-readable storage medium, in which a processor-executable program is stored, and the processor-executable program is used to implement the pathological image semantic segmentation method as described above when executed by a processor.

[0019] Compared with the prior art, the technical effects of the present invention are as follows in at least one of the following items: Introduce an image enhancement step in the semantic segmentation of pathological images to alleviate the problem that the model may not be able to accurately distinguish images containing multiple categories when facing tissue heterogeneity, and improve the performance in identifying tissue heterogeneity.

[0020] Adopt a per-pixel contrastive learning method to encourage the model to learn the correlations between pixels, deepen the model's judgment of different category pixels, and improve the segmentation performance of the model's local area. Brief Description of the Drawings

[0021] Figure 1 It is a flowchart of the semantic segmentation method of the present invention.

[0022] Figure 2 It is a diagram showing the scanned images of six scanners and the corresponding RGB channel distributions.

[0023] Figure 3 It is a diagram showing inter-tumor heterogeneity.

[0024] Figure 4 It is a diagram showing intra-tumor heterogeneity.

[0025] Figure 5 It is a module diagram of the semantic segmentation device of the present invention.

[0026] Figure 6 It is another module diagram of the semantic segmentation device of the present invention. Detailed Embodiments

[0027] The following non-limiting embodiments can enable those of ordinary skill in the art to more comprehensively understand the present invention, but do not limit the present invention in any way. The following content is only an exemplary illustration of the scope claimed in the present application. Those skilled in the art can make various changes and modifications to the invention of the present application according to the disclosed content, and it should also fall within the scope claimed in the present application.

[0028] The present invention will be described below by way of specific embodiments in some embodiments.

[0029] Histological grading is an index routinely evaluated in the pathological diagnosis of rectal cancer (CRC). It is evaluated by a pathologist by observing the proportion of different differentiations of tumors in the tissue, and can be used to evaluate the tumor progression degree of patients and predict the prognosis of patients. At present, the visual-based evaluation method brings a large workload to pathologists, and there are large observer differences in this method. Developing a computer-aided diagnosis system can effectively optimize the CRC histological grading process, reduce the burden on pathologists, and improve the speed and accuracy of evaluation.

[0030] It should be noted that the embodiments of the present invention are not only applicable to rectal cancer (CRC) and tumor grade assessment, but can be applied to different pathological diagnoses of any other tissues. For the convenience of description, in some specific embodiments, the pathological diagnosis of rectal cancer (CRC) is used as an example in certain cases.

[0031] During diagnosis, since most of the sliced tissues are in a transparent state, it is necessary to stain the tissues for observation. There are many types of stains to choose from. Among them, hematoxylin (H) and eosin (E) are commonly used stains. The pathological images obtained by scanning after staining with this stain are called H&E images. However, the digitally presented H&E images may show obvious color differences, which will lead to a decline in model performance. The reasons for this difference mainly come from two stages. One is the preparation stage of the glass slide, and the other is the scanning stage. First, analyze from the process of preparing the glass slide: tissue acquisition (thickness, sectioning process, section storage time) and tissue staining (different dye batches, staining time, staining operation) are all possible reasons for color differences; secondly, analyze from the scanning stage: differences between different brands or models of scanners from different manufacturers (different light source systems, optical systems, image sensors, data processing units, algorithms and file formats for image processing, etc.) may lead to color differences in the scanned images and different scales of representation per pixel. Regarding the problem of image color differences, most researchers tend to perform staining normalization on the data or perform color enhancement on the training set. The former can make the colors in the image more consistent and reduce color deviations caused by staining, scanning equipment, etc. However, staining normalization may lead to the loss of detailed textures; the latter aims to let the model learn data with different color distributions to improve robustness by adjusting parameters such as contrast and brightness. This method can improve the generalization of the model to a certain extent, but the effect is limited. Figure 2 shows the images obtained by scanning the same H&E-stained glass slide with six pathological scanners. In order to more objectively show the color differences in the images brought by different scanners, in Figure 2 shows the R, G, B channel distributions of each image. It can be seen that there are differences in the distributions of different scanners in the R, G, B channels, and the pathological images scanned by scanner E and scanner F have obvious color differences from the pathological images scanned by the first four scanners.

[0032] Tumor heterogeneity is divided into inter-tumor heterogeneity and intra-tumor heterogeneity. Inter-tumor heterogeneity refers to the differences in tumors between different patients with the same histological grade, that is, the tumor cells in the same grade of tumors may show different morphologies, sizes and arrangements, as shown in Figure 3 . Intra-tumor heterogeneity refers to the differences in tumors within a patient, as shown in Figure 4, which can be subdivided into spatial heterogeneity and temporal heterogeneity. Spatial heterogeneity refers to the differences observable within a single tumor, while temporal heterogeneity means that as time goes by, the tumor develops dynamically and changes occur within the tumor. Therefore, in some pathological images, changes from low-grade to high-grade tumors can be observed, and it is more difficult to determine the category of tumors in intermediate forms. In addition, some high-grade tumors are scattered in the tumor stroma with unclear boundaries, which also increases the difficulty of tissue segmentation.

[0033] Essentially, both of the above two challenges test the learning ability and discrimination ability of the segmentation method or the segmentation model for features. In different embodiments of the present invention, an enhancement module and a per-pixel contrast learning module are introduced separately or simultaneously to strengthen the above capabilities of the segmentation method or the segmentation model. The enhancement module effectively increases the diversity of training data by fusing image regions of different categories. The per-pixel contrast learning module uses per-pixel contrast loss constraints to strengthen the model's learning of the features of the same-class pixels, reduce the impact of color differences. At the same time, this module improves the model's ability to distinguish different types of tissues by pulling the pixels of the same class closer and pushing the pixels of different classes farther away in the feature space.

[0034] To automatically and accurately quantify histological grades on WSI, an embodiment of a semantic segmentation method for pathological images is proposed in the present invention, as Figure 1 shown, including the following steps: S1: Obtain a first pathological image set, the first pathological image set includes a plurality of first pathological images, and each of the first pathological images has a category identifier; Perform enhancement processing on the first pathological image set according to the category identifier to obtain a second pathological image set, the second pathological image set includes a plurality of second pathological images; Obtain a first label image set, the first label image set includes a plurality of first label images, the first label images and the first pathological images are in one-to-one correspondence, and each of the first label images has the same category identifier as the corresponding first pathological image; Perform the same enhancement processing on the first label image set according to the category identifier to obtain a second label image set, the second label image set includes a plurality of second label images.

[0035] S2: Input the training set into the semantic segmentation model for model training, the semantic segmentation model includes an encoder, a decoder and a per-pixel contrast learning module, and the training set includes the second pathological image set and the second label image set.

[0036] S3: Take a third pathological image set, input the third pathological image set into the trained semantic segmentation model, and output a segmentation result, the third pathological image set includes a plurality of third pathological images.

[0037] Among the patches cropped from whole slide digital pathology images (WSIs) of CRC patients, most patches are of a single category. If only single-category patches are used to train the model, it will limit the model's ability to learn the commonalities and differences between multiple categories. Especially when facing tissue heterogeneity, it may cause the model to be unable to accurately distinguish images containing multiple categories. To address this challenge, this embodiment introduces a data augmentation step to improve the model's performance in identifying tissue heterogeneity by combining features of different categories.

[0038] In this embodiment, the first pathological image can be a patch cropped from the collected WSIs. For example, the size of the patch can be 1024×1024, which is the first pathological image set in the present invention. After data augmentation processing and cropping, a second pathological image is obtained. The third pathological image can be a patch cropped from the patient's WSI, and the size can be 1024×1024.

[0039] In this embodiment, the first label image corresponds to the first pathological image one by one. The first label image is the label image of the corresponding first pathological image.

[0040] In some embodiments, the class identifiers of all the first pathological images in the first pathological image set include at least 2; The obtaining of the second pathological image set by performing enhancement processing on the first pathological image set according to the class identifier includes: Sequentially select the first pathological image with the class identifier of the first category in the first pathological image set as the target pathological image; Randomly select a first pathological image with a class identifier other than the first category from the first pathological image set as the reference pathological image; Intercept the first preset graphic at the first preset position of the reference pathological image; Cover the intercepted graphic at the first preset position of the target pathological image to obtain the enhanced first pathological image; The enhanced first pathological image and all the first pathological images with class identifiers other than the first category are used as the second pathological images, and all the second pathological images constitute the second pathological image set.

[0041] The obtaining of the second label image set by performing the same enhancement processing on the first label image set according to the class identifier includes: Sequentially select the first label image with the class identifier of the first category in the first label image set as the target label image; Randomly select a first label image with a class identifier other than the first category from the first label image set as the reference label image; Intercept a second preset graphic at a second preset position of the reference label image, where the second preset position is the same as the first preset position, and the second preset graphic is the same as the second preset graphic; Overlay the intercepted graphic on the second preset position of the target label image to obtain an enhanced first label image; The enhanced first label image and all first label images with category identifications other than the first category are used as second label images, and all second label images constitute a second label image set.

[0042] In this embodiment, the first pathological images have the same size, with category identifications, including at least a first category and other labels. Exemplarily, the first category represents normal tissue, such as normal glands. Other labels are other labels outside the abnormal glands, such as background, low-grade tumors, and high-grade tumors. In this embodiment, a partial screenshot (i.e., the preset graphic at the preset position) on the pathological image of a non-first category is overlaid on the pathological image of the first category to obtain an enhanced pathological image with the first category, and then the enhanced pathological image and the pathological images with category identifications other than the first category are together used to form a second pathological image set training set. Specifically, the preset graphic at the preset position in this embodiment must be within the selected reference pathological image and will not exceed the range of the reference pathological image. The same enhancement process is also performed on the first label object set. Exemplarily, the same operations are performed on the first label image corresponding to the first pathological image. For example, when selecting a reference label image for the target label image from the first label image set, the selected reference label image is the label image corresponding to the reference pathological image selected for the target pathological image corresponding to the target label image. That is, the target label image corresponds one-to-one with the target pathological image, and the reference label image of the target label image also corresponds to the reference pathological image of the target pathological image, that is, the reference label image of the target label image is the label image of the pathological image corresponding to the target label image.

[0043] In some embodiments, the first preset graphic is a circle, and the first preset position is the center of the circle. The intercepting of the first preset graphic at the first preset position of the reference pathological image includes: Randomly set the center of the circle; Randomly set the radius of the circle, where the radius is less than one-fourth of the height of the reference pathological image; Intercept the preset graphic from the reference pathological image according to the randomly set center of the circle and radius.

[0044] Specifically, taking 4 category identifiers, namely background, normal gland, low-grade tumor, and high-grade tumor as examples. The specific operation of data augmentation is as follows: If the target image is a normal gland, randomly select one image from the other three types of images, cut out the circle defined by the randomly selected center and randomly set radius, and paste it at the same position of the target image. At the same time, perform the same operation on the corresponding label image of the target image. Taking an image with a size of 1024*1024 as an example, the following settings can be made. The range of the horizontal and vertical coordinates of the center is (100, 512), and the range of the radius is expressed as: where h represents the height of the image input, and ⌊⌋ is the floor symbol.

[0045] In this embodiment, reasonably selecting the size and shape of the cropped image can conveniently and appropriately achieve the data augmentation effect, which is beneficial to improving the performance of the model in identifying tissue heterogeneity.

[0046] In some embodiments, the inputting the training set into the semantic segmentation model for model training includes: Using an encoder to extract features from the images in the training set; Using a decoder to fuse the extracted features; Using a per-pixel contrast learning module to perform per-pixel contrast learning on the fused features.

[0047] In this embodiment, the encoder is used to extract features from the images in the training set, and the decoder is used to fuse the extracted features and perform per-pixel contrast learning.

[0048] After obtaining the fused features through the decoder, the model will further calculate the segmentation loss and then train the model through forward propagation. However, the segmentation loss does not consider the inherent context relationship of the image nor the relationship between different categories. In addition, reducing the sequence length when calculating self-attention may also lead to information loss. These two reasons may both result in poor segmentation performance of the model in the local focus area. This embodiment proposes per-pixel contrast learning to encourage the model to learn the association between pixels, deepen the model's judgment of different category pixels, and improve the segmentation performance of the model in the local area.

[0049] In some embodiments, the encoder includes 4 parts, and each part includes an overlapping block embedding module, a flattening module, and a Transformer module; the using the encoder to extract features from the training set includes: Performing an overlapping block embedding operation using the overlapping block embedding module to divide the input image into multiple first image block tensors; the overlapping block embedding operation includes a convolution operation, and the convolution kernel size of the convolution is greater than the stride; Flatten multiple of the first image patch tensors using a flattening module; Input the flattened first image patch tensors into a Transformer module; Retain the high-resolution coarse features generated in the previous part for the decoder to generate a segmentation result, and at the same time use them as the input for the next part to continue obtaining low-resolution features of different scales.

[0050] Suppose a pathological image data is selected from the training set For model training (for example, both H and W are 1024). First, to generate a sequence suitable for the Transformer architecture, the image needs to be subdivided into image patches. Smaller image patches are more conducive to capturing finer local features.

[0051] The high-resolution coarse features generated in the first part (the first stage) will be retained for the decoder to generate the final semantic segmentation result. In addition, they will also be used as the input for the next part to continue obtaining more detailed low-resolution features of different scales.

[0052] Specifically, in some embodiments, when inputting into the Transformer module of the next part, first use an overlapping patch embedding module for dimensionality reduction operation to obtain feature layers of different scales. However, different from the operation of the overlapping patch embedding module (overlapping patch embedding 1) in the first part, the overlapping embedding module here uses overlapping patch embedding 2, and its convolution parameters are: kernel size = 3, stride = 2, padding = 1. The convolution parameters of the third part (stage three) and the fourth part (stage four) are the same as those of the second part (the second stage). The operations of other parts of the encoder (other stages) are the same as those of the first stage.

[0053] In the ViT model, positional encoding is also introduced when inputting into the Transformer module. This is because after dividing the image into image patches, there is still a positional relationship between each image patch, so positional information needs to be added. The model used in this embodiment cancels the use of positional encoding and uses convolution as an alternative. Convolution operations have local perception ability. When processing an image, when the convolution kernel slides to different positions of the feature map, it will operate on the pixels in the local area, thus retaining the feature information of different positions in the input image in the output feature map. Therefore, in the feature map obtained after convolution, the relationship between adjacent pixels can reflect their relative positional relationship in the input image. In the semantic segmentation task, the positional relationship retained through convolution operations is sufficient to support the model to perform this task.

[0054] In some embodiments, the attention calculation formula of the Transformer module is , where Q, K, and V have the same size of N and C, , where N is the sequence length, P is the size of the first image patch, C is the number of channels, H is the height of the images in the training set, W is the width of the images in the training set, and d k is the dimension of K; Before inputting the flattened image patch tensor into the Transformer module, the following processing is performed: Reduce the sequence length through the Reshape and Linear layers, and reduce N to , where is the reduction ratio.

[0055] In this embodiment, the sequence restoration process in PVT (Pyramid Vision Transformer) is used. The sequence length is reduced by adding Reshape and Linear layers, and is reduced to to adapt to self-attention calculation.

[0056] In some embodiments, the output of the Transformer module is executed according to the following formula where, is the feature output after the self-attention operation, The operation of is to create a linear layer to convert the input features into the number of hidden layer features, and the outermost MLP is to convert the hidden layer features into the number of output features. GELU is an activation function,

[0057] In some embodiments, the method further includes: The decoder obtains the features extracted by the four parts of the encoder, and performs an interpolation operation using bilinear interpolation to adjust the sizes of the features generated by the second, third, and fourth parts of the encoder to be the same as the size of the features generated by the first part; Concatenate the features generated by the first part and the adjusted features generated by the second, third, and fourth parts; Input the concatenated features into the classification head, and after convolution, batch normalization, and convolution operations in the classification head, they are converted into a segmentation mask for semantic segmentation. The cross-loss function is calculated on the output of the classification head, and the cross-loss function uses the following formula , where H and W represent the height and width of the features, Represents the true label, indicating that the classification head output is with a probability of, where C is the number of classes; Perform per-pixel contrastive learning to complete model training. Among them, the following formula is used to calculate the contrastive loss Among them, and respectively represent the class identifiers of pixel i and pixel j, z represents the feature map obtained after passing through the feature projection module, 、 and both represent pixel features, N is the number of pixel points in the feature map, represents the temperature hyperparameter; The total loss function is , where takes a value of 0.1.

[0058] Exemplarily, the number of class identifiers can be 3, which are respectively used to identify normal glands, high-grade tumors, and low-grade tumors.

[0059] In this embodiment, the decoder uses an MLP layer to splice and fuse features of different scales. The spliced and fused features are projected into a low-dimensional feature space through the PROJ module, and then per-pixel contrastive learning calculation is performed.

[0060] In some embodiments, the segmentation result indicates the class identifier of each pixel of the third pathological image.

[0061] Optionally, the obtaining of the third pathological image set, inputting the third pathological image set into the trained semantic segmentation model, and outputting a segmentation result includes: Obtain a plurality of whole-slide images, crop each whole-slide image to obtain the third pathological image set of each whole-slide image, input the third pathological images in each of the third pathological image sets into the trained semantic segmentation model, and output the segmentation result of each third pathological image; the size of the third pathological image is the same as the size of the first pathological image.

[0062] It can be understood that the tissue images obtained during the examination often include multiple whole-slide images. After each whole-slide image is cropped, multiple patches of a preset size can be obtained, which are the third pathological images in the present invention. For example, patches of 1024×1024 pixels can be obtained to facilitate the determination of histological grade.

[0063] Exemplarily, multiple whole-slide images of organisms are preprocessed and cropped into multiple third pathological image sets. Each third pathological image set corresponds to a whole-slide image and contains multiple third pathological images (which can be set as Patches of 1024×1024 pixels). These Patches are input into a trained semantic segmentation model to predict the category of each Patch.

[0064] After obtaining the fused features through the decoder, the model will further calculate the segmentation loss and then train the model through forward propagation. However, the segmentation loss does not consider the inherent context relationship of the image, nor does it consider the relationship between different categories. In addition, reducing the sequence length when calculating self-attention may also lead to information loss. These two reasons may both result in poor segmentation performance of the model in the local focus area. In this embodiment, per-pixel contrastive learning is adopted to encourage the model to learn the association between pixels, deepen the model's judgment of pixels of different categories, and improve the segmentation performance of the local area of the model.

[0065] The core of calculating the contrastive loss is to narrow the distance between features of the same class (positive sample pairs) and expand the distance between features of different classes (positive and negative samples). In this embodiment, pixels of the same class as the anchor point (query) are regarded as positive samples, and pixels of different classes are regarded as negative samples. Therefore, calculating the distance between two features becomes the key. The exponential cosine similarity used in SimCLR is adopted in this embodiment, and the formula for calculating the contrastive loss is: where, and respectively represent the class identifiers of pixel i and pixel j, z represents the feature map obtained after passing through the feature projection module, 、 and all represent pixel features, N is the number of pixel points in the feature map, represents the temperature hyperparameter.

[0066] In this formula, only the calculations that satisfy the condition are summed up. The purpose of doing this is to regard the pixels of all sample pairs belonging to the same class as positive samples.

[0067] Data augmentation steps: In this study, data augmentation was introduced to improve the model's performance in identifying tissue heterogeneity by combining features of different classes. The training data included a total of 4 classes including the background, namely the background, normal glands, low-grade tumors, and high-grade tumors. The specific operation of data augmentation was as follows: If the target image was a normal gland, a random image was selected from the other three types of images, a circle with a random radius and center was cut out, and pasted at the same position on the target image. At the same time, the same operation was performed on the label corresponding to the target image. The range of the horizontal and vertical coordinates of the center was (100, 512), and the range of the radius was expressed as: where h represents the height of the image input, and ⌊⌋ is the floor function.

[0068] The specific architecture of the model for CRC tissue segmentation was constructed based on the SegFormre model. The encoder was mainly used to generate features of different scales, and the decoder used MLP layers to splice and fuse features of different scales. The spliced and fused features were projected into a low-dimensional feature space through the PROJ module, and then pixel-wise contrastive learning was calculated.

[0069] In the encoder operation part, it included the following content: Suppose a pathological image data was selected from the training set for model training (in this embodiment, both H and W were 1024). First, in order to generate a sequence suitable for the Transformer architecture, the image needed to be subdivided into image patches. Different from dividing the image into 16×16-sized image patches in ViT, in this embodiment, the image was divided into 4×4 image patches. The smaller image patches were more conducive to capturing finer local features. After performing the above operation, the tensor size was 256×256, and then the vector was flattened to obtain a sequence length of 65536. The sequence length to be processed in ViT was 197 (plus the classification head), and this computational amount was acceptable for subsequent self-attention calculations. However, a sequence length of 65536 was a computational complexity that was not very acceptable for the Transformer model. Therefore, the sequence length needed to be reduced before performing self-attention calculations, and the calculation formula was: , where Q, K, and V had the same size of N and C, , N was the sequence length, P was the size of the first image patch, C was the number of channels, H was the height of the input image, W was the width of the input image, and dk was the dimension of K.

[0070] Here, the sequence reduction process in PVT was introduced. The sequence length was reduced by adding Reshape and Linear layers. Through the following formula, the dimension of was reduced from .

[0071] Among them is the reduction ratio.

[0072] When the ViT model inputs the Transformer module, positional encoding is also introduced. This is because after the image is divided into image patches, there is still a positional association between each image patch, so positional information needs to be added. The model used in this embodiment cancels the use of positional encoding and uses convolution as a substitute. Convolution operations have local perception capabilities. When processing images, when the convolution kernel slides to different positions of the feature map, it will operate on the pixels in the local area, thereby retaining the feature information of different positions in the input image in the output feature map. Therefore, in the feature map obtained after convolution, the relationship between adjacent pixels can reflect their relative positional relationship in the input image. In the semantic segmentation task, the positional relationship retained through convolution operations is sufficient to support the model to perform this task. The convolution parameters are (convolution kernel = 3, stride = 1, padding = 1). The specific formula is: Among them, is the feature output after the self-attention operation, The operation of is to create a linear layer to convert the input features into the number of hidden layer features. The outermost MLP is to convert the hidden layer features into the number of output features. GELU is an activation function, is the output of this part.

[0073] The high-resolution coarse features generated in the first stage will be retained for the decoder to generate the final semantic segmentation result. In addition, they will also be used as the input of the next module to continue to obtain lower-resolution and more detailed features of different scales. When input into the next Transformer module, first, an overlapping patch embedding operation will be performed to reduce the dimension to obtain feature layers of different scales. However, different from the operation of overlapping patch embedding 1, overlapping patch embedding 2 is used here, and its convolution parameters are: convolution kernel = 3, stride = 2, padding = 1. The convolution parameters of this part in the third and fourth stages are the same as those in the second stage. The operations in other stages of the encoder are similar to those in the first stage. The number of layers of the Transformer is {3, 6, 40, 3} respectively. The size of the output feature vector in the first stage is (256, 256, 64), where 64 is the number of channels. The size of the output feature vector in the second stage is (128, 128, 128). The size of the output feature vector in the third stage is (64, 64, 320). The size of the output feature vector in the third stage is (32, 32, 512).

[0074] In the decoder operation part, it includes the following content: All the features extracted in the four stages of the encoder part are fed into the decoder part for operation. First, the dimensions of the features are unified. In this embodiment, bilinear interpolation is used for the interpolation operation, that is, the pixel values in the input feature map are weighted and averaged between the surrounding pixel points to generate new pixel values. The feature maps generated in the second, third, and fourth stages are all adjusted to the same size as the feature map generated in the first stage in this way. Here, only the size of the feature map changes, and the number of channels remains unchanged. Then the four feature maps are concatenated. The size of the concatenated feature map is 256×256, and the number of channels is 1024. The concatenated features are input into the classification head. In the classification head, the input features are converted into a segmentation mask for semantic segmentation after convolution (convolution kernel = 3, stride = 1, padding = 1), batch normalization, Dropout (p = 0.1), and convolution (convolution kernel = 1, stride = 1, padding = 0). The size of the segmentation mask is 256×256×4, and 4 is the number of classes. In this process, a feature projection module is used to reduce the dimension of the concatenated features. The module first linearly transforms the input features with a 1×1 convolution without changing the spatial dimension of the features, and then enhances the non-linear representation and stability of the features through batch normalization and activation function operations. Then, another 1×1 convolution is used to project the processed features into a low-dimensional feature space, and the dimension of the projected features is set to 256.

[0075] After obtaining the fused features through the decoder, the per-pixel contrast learning part needs to be executed.

[0076] After obtaining the fused features through the decoder, the model will further calculate the segmentation loss and then train the model through forward propagation. However, the segmentation loss does not consider the inherent context relationship of the image, nor does it consider the relationship between different classes. In addition, reducing the sequence length when calculating self-attention may also lead to information loss. Both of these reasons may cause the model to have poor segmentation performance in the local focus area. In this embodiment, per-pixel contrast learning is adopted to encourage the model to learn the association between pixels, deepen the model's judgment of pixels of different classes, and improve the segmentation performance of the local area of the model. The contrast loss is calculated per-pixel in the reduced-dimensional feature space. The core of calculating the contrast loss is to reduce the distance between features of the same class (positive sample pairs) and expand the distance between features of different classes (positive and negative samples). In this part, pixels of the same class as the anchor (query) are regarded as positive samples, and pixels of different classes are regarded as negative samples. Therefore, calculating the distance between two features becomes the key. The exponential cosine similarity used in SimCLR is used in this embodiment, and the formula for calculating the contrast loss is: where, and respectively represent the class labels of pixel i and pixel j, z represents the feature map obtained after passing through the feature projection module, 、 and both represent pixel features, where N is the number of pixel points in the feature map, represents the temperature hyperparameter.

[0077] In this formula, only the calculations that satisfy condition are summed up. The purpose of doing this is to regard the pixels of all sample pairs belonging to the same class as positive samples.

[0078] The loss function in this embodiment is set as follows: Apply the cross-entropy loss function to the output of the classification head and apply the contrastive loss function to the output of the feature projection module.

[0079] The cross-entropy loss function is a commonly used loss function in classification tasks, which is used to measure the difference between two probability distributions. Usually, this function is minimized to make the probability distribution output by the model as close as possible to the true label, so as to effectively train a model with better performance. The calculation method of the cross-entropy function in this embodiment is: where H and W represent the height and width of the feature, represents the true label, represents the probability that the output of the classification head is and C represents the number of classes; The total loss function is expressed as: where takes the value of 0.1.

[0080] After obtaining the segmentation result through the above steps, enter the histological grading automatic quantification method.

[0081] This embodiment adopts a per-pixel contrastive learning model based on Transformer (PCL-Trans) to cope with the challenges of image color differences and tumor heterogeneity during tissue segmentation. Finally, the PCL-Trans model can accurately distinguish low-grade tumors and high-grade tumors in CRC. The comprehensive results on the test sets of six different pathological scanners and an external test set show that the segmentation model proposed in this embodiment has the optimal performance.

[0082] The present invention also provides a pathological image semantic segmentation device based on per-pixel contrastive learning, as Figure 5 shown. The device includes an enhancement module and a semantic segmentation model.

[0083] The enhancement module is used for: Obtain a first set of pathological images, where the first set of pathological images includes a plurality of first pathological images, and each of the first pathological images has a class identifier; After performing enhancement processing on the first pathological image set according to the category identifier, a second pathological image set is obtained, and the second pathological image set includes multiple second pathological images; Obtain a first label image set, where the first label image set includes multiple first label images, the first label images and the first pathological images correspond one by one, and each first label image has the same category identifier as the corresponding first pathological image; After performing the same enhancement processing on the first label image set according to the category identifier, a second label image set is obtained, and the second label image set includes multiple second label images; The semantic segmentation model includes an encoder, a decoder, and a per-pixel contrastive learning module. The semantic segmentation model is trained through an input training set, and the training set includes the second pathological image set and the second label image set; the semantic segmentation model is used to: Obtain a third pathological image set, input the third pathological image set into the trained semantic segmentation model, and output a segmentation result. The third pathological image set includes multiple third pathological images.

[0084] In some embodiments, as Figure 6 shown, the pathological image semantic segmentation device further includes a histological grading determination module; The obtaining of the third pathological image set, inputting the third pathological image set into the trained semantic segmentation model, and outputting a segmentation result includes: Obtain multiple whole slide images, crop each whole slide image to obtain a third pathological image set for each whole slide image, input the third pathological images in each third pathological image set into the trained semantic segmentation model, and output a segmentation result for each third pathological image; the size of the third pathological image is the same as the size of the first pathological image.

[0085] The execution steps and beneficial effects of each module in the pathological image semantic segmentation device provided by the present invention are the same as those in the foregoing embodiments of the pathological image semantic segmentation method, and will not be elaborated here.

[0086] The present invention also provides an embodiment of a terminal device, and the terminal device includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the pathological image semantic segmentation method as described above.

[0087] The present invention also provides an embodiment of a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the pathological image semantic segmentation method as described above when executed by the processor.

[0088] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements can be made without departing from the principle of the present invention, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A pathological image semantic segmentation method based on per-pixel contrast learning, characterized in that, The method includes: Obtaining a first pathological image set, the first pathological image set including a plurality of first pathological images, each of the first pathological images having a class identifier; Performing enhancement processing on the first pathological image set according to the class identifier to obtain a second pathological image set, the second pathological image set including a plurality of second pathological images; Obtaining a first label image set, the first label image set including a plurality of first label images, the first label images corresponding one-to-one with the first pathological images, each of the first label images having the same class identifier as the corresponding first pathological image; Performing the same enhancement processing on the first label image set according to the class identifier to obtain a second label image set, the second label image set including a plurality of second label images; Inputting the training set into a semantic segmentation model for model training, the semantic segmentation model including an encoder, a decoder, and a per-pixel contrast learning module, the training set including the second pathological image set and the second label image set; Obtaining a third pathological image set, inputting the third pathological image set into the trained semantic segmentation model, and outputting a segmentation result, the third pathological image set including a plurality of third pathological images.

2. The pathological image semantic segmentation method according to claim 1, wherein The class identifiers of all the first pathological images in the first pathological image set include at least two; The obtaining of the second pathological image set by performing enhancement processing on the first pathological image set according to the class identifier includes: Sequentially selecting the first pathological images with the class identifier of the first category in the first pathological image set as target pathological images; Randomly selecting the first pathological images with class identifiers other than the first category from the first pathological image set as reference pathological images; Cropping a first preset graphic at a first preset position of the reference pathological image; Covering the cropped graphic at the first preset position of the target pathological image to obtain an enhanced first pathological image; The enhanced first pathological images and all the first pathological images with class identifiers other than the first category are used as second pathological images, and all the second pathological images constitute the second pathological image set; The obtaining of the second label image set by performing the same enhancement processing on the first label image set according to the class identifier includes: Sequentially selecting the first label images with the class identifier of the first category in the first label image set as target label images; Randomly selecting the first label images with class identifiers other than the first category from the first label image set as reference label images; Cropping a second preset graphic at a second preset position of the reference label image, the second preset position being the same as the first preset position, and the second preset graphic being the same as the second preset graphic; Covering the cropped graphic at the second preset position of the target label image to obtain an enhanced first label image; The enhanced first label images and all the first label images with class identifiers other than the first category are used as second label images, and all the second label images constitute the second label image set.

3. The pathological image semantic segmentation method according to claim 2, wherein, The first preset figure is a circle, the first preset position is the center of the circle, and the first preset figure for intercepting the first preset position of the reference pathological image includes: Randomly set the center of the circle; Randomly set the radius of the circle, where the radius is less than one-fourth of the height of the reference pathological image; Intercept the preset figure from the reference pathological image according to the randomly set center of the circle and radius.

4. The pathological image semantic segmentation method according to claim 3, wherein The inputting the training set into the semantic segmentation model for model training includes: Using an encoder to extract features from the images in the training set; Using a decoder to fuse the extracted features; Using a per-pixel contrast learning module to perform per-pixel contrast learning on the fused features.

5. The pathological image semantic segmentation method according to claim 4, wherein, The encoder includes 4 parts, and each part includes an overlapping block embedding module, a flattening module, and a Transformer module; The using the encoder to extract features from the training set includes: Performing overlapping block embedding operations using the overlapping block embedding module to divide the input image into multiple first image block tensors; the overlapping block embedding operation includes a convolution operation, and the convolution kernel size of the convolution is greater than the stride; Using the flattening module to flatten the multiple first image block tensors; Inputting the flattened first image block tensors into the Transformer module; Retaining the high-resolution coarse features generated by the previous part for the decoder to generate a segmentation result, and at the same time using them as the input for the next part to continue obtaining low-resolution features of different scales.

6. The pathological image semantic segmentation method according to claim 5, characterized in that The attention calculation formula of the Transformer module is , Among them, Q is the query vector, K is the key vector, and V is the key-value vector. Q, K, and V have the same size of N and C. , where N is the sequence length, P is the size of the first image patch, C is the number of channels, H is the height of the images in the training set, W is the width of the images in the training set, and d k is the dimension of K; Before inputting the flattened image block tensors into the Transformer module, the following processing is performed: Reduce the sequence length through Reshape and Linear layers, and reduce N to , wherein is the reduction ratio.

7. The pathological image semantic segmentation method according to claim 6, wherein The output of the Transformer module is executed according to the following formula Among them, is the feature output after the self-attention operation, The operation of is to create a linear layer to convert the input features into the number of hidden layer features, and the outermost MLP is to convert the hidden layer features into the number of output features. GELU is an activation function, is the output of this part.

8. The pathological image semantic segmentation method according to any one of claims 6 to 7, characterized in that The method further includes: The decoder obtains the features extracted by the 4 parts of the encoder, and performs interpolation operations using bilinear interpolation to adjust the sizes of the features generated by the second part, the third part, and the fourth part of the encoder to be the same as the size of the features generated by the first part; Concatenate the features generated by the first part, and the adjusted features generated by the second part, the third part, and the fourth part; Input the concatenated features into a classification head, and after convolution, batch normalization, and convolution operations in the classification head, they are converted into a segmentation mask for semantic segmentation. The cross-loss function is calculated on the output of the classification head, and the cross-loss function uses the following formula , where H and W respectively represent the height and width of the images in the training set, represents the true class label, represents that the output of the classification head is the probability, and C is the number of classes; Perform per-pixel contrast learning to complete model training, where the contrast loss is calculated using the following formula Among them, and represent the class identifiers of pixel i and pixel j respectively, z represents the feature map obtained after the feature projection module, , and all represent pixel features, N is the number of pixel points in the feature map, represents the temperature hyperparameter; The total loss function is , where takes the value of 0.

1.

9. The pathological image semantic segmentation method according to claim 8, wherein The obtaining the third pathological image set, inputting the third pathological image set into the trained semantic segmentation model, and outputting a segmentation result includes: obtaining a plurality of whole-slide images, cropping each whole-slide image to obtain the third pathological image set of each whole-slide image, inputting the third pathological images in each third pathological image set into the trained semantic segmentation model, and outputting the segmentation result of each third pathological image. The size of the third pathological image is the same as the size of the first pathological image.

10. A pathological image semantic segmentation device based on pixel-by-pixel contrast learning, characterized in that, The device includes an enhancement module and a semantic segmentation model; The enhancement module is configured to: Obtain a first pathological image set, where the first pathological image set includes multiple first pathological images, and each of the first pathological images has a class identifier; Perform enhancement processing on the first pathological image set according to the class identifier to obtain a second pathological image set, where the second pathological image set includes multiple second pathological images; Obtain a first label image set, where the first label image set includes multiple first label images, the first label images and the first pathological images are in one-to-one correspondence, and each of the first label images has the same class identifier as the corresponding first pathological image; Perform the same enhancement processing on the first label image set according to the class identifier to obtain a second label image set, where the second label image set includes multiple second label images; The semantic segmentation model includes an encoder, a decoder, and a per-pixel contrast learning module. The semantic segmentation model is trained by inputting a training set, and the training set includes the second pathological image set and the second label image set; the semantic segmentation model is configured to: Obtain a third pathological image set, input the third pathological image set into the trained semantic segmentation model, and output a segmentation result, where the third pathological image set includes multiple third pathological images.