Multi-component image segmentation method and system in pathological image based on category separation

Through the global and category-specific feature extraction and fusion of the category separation network model, combined with the optimization of multiple loss function, the problem of difficult long-distance dependencies in pathological images is solved, and a more efficient multi-component segmentation effect is achieved.

CN120340027AActive Publication Date: 2025-07-18GENERAL HOSPITAL OF NUCLEAR IND

Patent Information

Application Number
CN202510837448.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-18
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The existing convolutional neural networks are difficult to effectively capture long-distance dependencies in pathological image segmentation, and the large local differences in pathological images and imbalance in tissue components distribution, resulting in poor results in the existing methods in multi-component segmentation.

Method used

The pathological image processing method based on category separation is adopted, and the features of the pathological image are extracted separately through the global feature extractor and the category-specific feature extractor, the feature fusion module is used for feature fusion, and the category image reconstruction decoder is reconstructed, combining alignment loss, category cross loss, segmentation loss and reconstruction loss to improve segmentation performance.

Benefits of technology

It significantly improves the accuracy and accuracy of multi-component segmentation of pathological images, especially in the tissue component segmentation task in the tumor microenvironment, and improves segmentation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340027A_ABST
    Figure CN120340027A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-component image segmentation method and system in a pathological image based on category separation, and relates to the technical field of image processing of computational pathology, and the method comprises the steps: obtaining a pathological image, inputting the pathological image into a pre-established category separation network model for processing, comprising a global feature extractor, a category specific feature extractor, a category fusion module, a segmentation module and a category image reconstruction decoder. The global feature extractor and the category specific feature extractor extract global features and category specific features of the pathological image respectively, and the category fusion module fuses the global features and the category specific features of the pathological image to obtain fused global features and category specific features; the segmentation module carries out segmentation based on the global features of the pathological image to obtain a multi-component segmentation result, and the category image reconstruction decoder carries out category image reconstruction based on the fused category specific features and the multi-component segmentation result to obtain a reconstructed image, thereby further improving the segmentation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing in computational pathology, specifically a method and system for multi-component segmentation of pathological images based on class separation. Background Art

[0002] In the field of pathology, identifying the morphological characteristics of pathological images, the expression of molecular targets, tumor-infiltrating lymphocytes, etc. under a microscope is the basis for disease diagnosis, treatment, and prognosis assessment. Currently, with the rapid development of artificial intelligence (AI), more and more research efforts are dedicated to the establishment and clinical application of AI pathology models, among which accurate whole-slide image (WSIs) segmentation of tissues is a crucial step. However, obtaining manual annotations of images is a very cumbersome task, usually requiring a large amount of time from pathologists with domain expertise.

[0003] U-Net is an efficient image segmentation model. It is a convolutional neural network (CNN) based on a decoder architecture and a skip connection mechanism, which minimizes the loss of spatial information and is particularly valuable in medical image applications. To enhance the ability of U-Net to extract features at different stages, various improvements have emerged based on U-Net. They aggregate multi-stage features through skip connections to generate high-resolution segmentation maps and perform well in medical image segmentation. However, due to the inherent limitations of convolutional operations, CNNs cannot learn long-range relationships between pixels. Some methods attempt to address this problem by adding self-attention in the encoder, transmitting significant features in the encoder to the decoder and suppressing irrelevant information to accurately reconstruct the segmentation map. Although these methods alleviate this problem, CNN-based methods are still insufficient to capture long-range dependencies. Additionally, pathological images have characteristics such as large local differences and uneven distribution of tissue components, and there is a particular need to seek better methods for segmenting tissue components in the tumor microenvironment. Summary of the Invention

[0004] To address the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a method and system for multi-component segmentation of pathological images based on class separation.

[0005] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: A method for multi-component segmentation of pathological images based on class separation, the method comprising the following steps: Obtain a pathological image, and input the pathological image into a pre-established class separation network model for processing, wherein the class separation network model includes a global feature extractor, a class-specific feature extractor, a class fusion module, a segmentation module, and a class image reconstruction decoder; The global feature extractor and the class-specific feature extractor extract the global features and class-specific features of the pathological image respectively, and the class fusion module fuses the global features and class-specific features of the pathological image to obtain the fused global features and class-specific features; The segmentation module performs segmentation based on the global features of the pathological image to obtain a multi-component segmentation result, and the class image reconstruction decoder performs class image reconstruction based on the fused class-specific features and the multi-component segmentation result to obtain a reconstructed image.

[0006] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the pre-established class separation network model is obtained by jointly optimizing through an alignment loss, a class cross loss, a segmentation loss function, a contrast loss function, and a reconstruction loss.

[0007] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the pathological image is obtained through a constructed pathological image training set, and includes pathological image I i and its corresponding multi-component class label denotes the number of pathological images in the training set, and the global features of the pathological image I are extracted through a constructed global feature extractor E based on the U-Net encoder g extract the global features of the pathological image I i through a constructed global feature extractor E based on the U-Net encoder and the class-specific features of the pathological image I are extracted through M constructed class-specific feature extractors based on the U-Net encoder extract the class-specific features of the pathological image I i through M constructed class-specific feature extractors based on the U-Net encoder where M represents the number of component segmentation classes, denotes element-wise multiplication.

[0008] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of the class fusion module fusing the global features of the pathological image: Send into the class fusion module CF to obtain the fused , and use the KL divergence to align the global feature with the fused global feature , and the alignment loss is defined as: where, denotes calculating the KL divergence between the global feature and the fused global feature ; The class fusion module is composed of an intra-class transformation module, an inter-class transformation module, a concat, a linear layer, and a layer normalization layer, and are respectively sent into the intra-class transformation module to obtain , , As Input1, it is jointly fed into the inter-class transformation module with as Input2 to obtain the fused class-specific features , are successively fed into the concat, linear layer, and layer normalization layer together and then used as Input2. As Input1, it is jointly fed into the inter-class transformation module to obtain the fused global features .

[0009] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the input of the intra-class transformation module is flattened to , and then the dimensions are permuted to , where C represents the number of channels, H represents the height, and W represents the width. After being fed into the linear layer Linear, the positional encoding of the learnable parameter PE is added. The process is expressed as where represents the weight of the linear layer. After being fed into the layer normalization LN layer, it is respectively fed into three linear layers Linear to calculate three components , , . Each component evenly divides the last dimension into R heads, that is , , , . The R heads respectively calculate the attention as where represents the size of the last dimension of represents the r-th head of the intra-class transformation IaT module component , represents the r-th head of the intra-class transformation IaT module component , T represents the transpose. represents the r-th head of the intra-class transformation IaT module component , A1 represents the attention of the first head, A2 represents the attention of the second head of the intra-class transformation IaT module, and A r represents the attention of the r-th head of the intra-class transformation IaT module. Denote the attention of the $R$-th head of the intra-class transformation $I_aT$ module. $CL$ represents concatenating first and then sending to a linear layer. Send it to a layer normalization layer first, and then add the feature after passing through the FFN module composed of a linear layer and a GELU activation layer , and then perform the inverse transformation of dimension swapping and flattening to obtain .

[0010] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the input of the inter-class transformation module Flattened into , and then swap the dimensions to , add the position encoding of the learnable parameter after sending it to a linear layer, and then send it to a layer normalization layer and a linear layer. The process is expressed as: where , represent the weights of the linear layer. Divide the last dimension of into $R$ heads equally to obtain ; the input Flattened into , and then swap the dimensions to , add the position encoding of the learnable parameter $PE$ after sending it to a linear layer. The process is expressed as: where represents the weight of the linear layer, After sending it to a layer normalization layer, send it to a linear layer respectively to obtain two components , , and divide the last dimension of each component into $R$ heads equally to obtain , , , and calculate the attention of $R$ heads respectively as: where represents the $r$-th head of the component of the inter-class transformation $I_tT$ module, represents the $r$-th head of the component of the inter-class transformation $I_tT$ module, represents the $r$-th head of the component of the inter-class transformation $I_tT$ module, represents the attention of the 1st head of the inter-class transformation $I_tT$ module, represents the attention of the 2nd head of the inter-class transformation $I_tT$ module, represents the attention of the $r$-th head of the inter-class transformation $I_tT$ module, Represents the attention of the $R$-th head of the inter-class transformation ItT module, which is successively sent into the layer normalization layer, and the features after passing through the FFN module composed of a linear layer and a GELU activation layer, plus , and then undergoes the inverse transformation of dimension swapping and flattening to obtain .

[0011] Combined with the first aspect, in some implementations of the first aspect, the method further includes: within the class-specific feature extractor, the pathological image $I$ i and other classes are sent into the extracted class cross features , where , and the class cross loss is calculated: where calculate the L1 norm of.

[0012] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the process by which the segmentation module performs segmentation based on the global features of the pathological image to obtain a multi-component segmentation result: The global features are sent into the segmentation module based on the U-Net decoder , and the multi-component segmentation result $P$ i of the pathological image $I$ i is output, and the segmentation loss is calculated: where represents the cross-entropy loss, represents the Dice similarity loss, is the multi-component class label the label value of the $m$-th class of the $t$-th pixel, is the segmentation result the predicted probability of the $m$-th class of the $t$-th pixel, and $T$ is the number of pixels.

[0013] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the process by which the class image reconstruction decoder performs class image reconstruction based on the fused class-specific features and the multi-component segmentation result: The fused class-specific features are sent into the class image reconstruction decoder for class image reconstruction, and the reconstructed image is output, and the reconstruction loss is calculated: Among them, calculate the L2 norm of.

[0014] In a second aspect, to achieve the above object, the present invention discloses a multi-component segmentation image system for pathological images based on class separation, including: An image processing unit, configured to obtain a pathological image and input the pathological image into a pre-established class separation network model for processing, where the class separation network model includes a global feature extractor, a class-specific feature extractor, a class fusion module, a segmentation module, and a class image reconstruction decoder; A feature fusion unit, configured to respectively extract the global feature and the class-specific feature of the pathological image by the global feature extractor and the class-specific feature extractor, and the class fusion module fuses the global feature and the class-specific feature of the pathological image to obtain the fused global feature and class-specific feature; A component segmentation unit, configured to perform segmentation based on the global feature of the pathological image by the segmentation module to obtain a multi-component segmentation result, and the class image reconstruction decoder performs class image reconstruction based on the fused class-specific feature and the multi-component segmentation result to obtain a reconstructed image.

[0015] Advantages of the present invention: The present invention can further improve the segmentation performance and is applicable to the multi-component segmentation of pathological images. The present invention achieves better results in the multi-class segmentation task of pathological images. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts; Figure 1 is a schematic flowchart of the method of the present invention; Figure 2 is a schematic diagram of the intra-class transformation IaT module of the present invention; Figure 3 is a schematic diagram of the inter-class transformation ItT module of the present invention; Figure 4 is a schematic diagram of the class fusion module CF of the present invention; Figure 5 is a schematic diagram of the class separation network of the present invention; Figure 6 is a schematic diagram of the segmentation comparison between the present invention and the existing method; Figure 7 is a schematic diagram of the system structure of the present invention. Specific Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0018] Embodiment 1: As Figure 1 shown, a method for multi-component segmentation of pathological images based on class separation, the method includes the following steps: S101: Obtain a pathological image, and input the pathological image into a pre-established class separation network model for processing. Among them, the class separation network model includes a global feature extractor, a class-specific feature extractor, a class fusion module, a segmentation module, and a class image reconstruction decoder; The pre-established class separation network model is obtained by jointly optimizing through alignment loss, class cross loss, segmentation loss function, contrast loss function, and reconstruction loss; Use the total loss to optimize the entire network, and the total loss is expressed as: The pathological image is obtained through a constructed pathological image training set, and includes pathological image I i and its corresponding multi-component class labels represents the number of pathological images in the training set. The global feature extractor E based on the U-Net encoder is constructed g to extract the global features of pathological image I i , and through M class-specific feature extractors based on the U-Net encoder constructed , the class-specific features of pathological image I are extracted. M represents the number of component segmentation classes i , and represents pixel multiplication; represents pixel multiplication; S102: The global feature extractor and the class-specific feature extractor respectively extract the global features and class-specific features of the pathological image, and the class fusion module fuses the global features and class-specific features of the pathological image to obtain the fused global features and class-specific features; The class fusion module sends to the class fusion (abbreviated as CF) module to obtain the fused ​, and the Kullback-Leibler Divergence is used for the global feature and the fused global feature to be aligned. The alignment loss is defined as: where represents calculating the KL divergence between the global feature and the fused global feature .

[0019] The class fusion module consists of an Intra-class Transformer (IaT) module, an Inter-class Transformer (ItT) module, concat (channel concatenation), a Linear layer, and a Layer Normalization (LN) layer. are respectively fed into the intra-class transformation module to obtain , as Input1, and jointly fed into the inter-class transformation module with as Input2 to obtain the fused class-specific feature , are jointly fed into the concat, Linear layer, and Layer Normalization layer in sequence and then used as Input2, as Input1 and jointly fed into the inter-class transformation module to obtain the fused global feature .

[0020] The input of the intra-class transformation module where C represents the number of channels, H represents the height, and W represents the width, is flattened to , and then the dimensions are permuted to , and after being fed into the Linear layer, the positional encoding of the learnable parameter PE is added. This process can be expressed as where represents the weight of the Linear layer, after being fed into the Layer Normalization LN layer, they are respectively fed into three Linear layers to calculate three components , , , and each component evenly divides the last dimension into R heads, that is , , , , and the attention is calculated for the R heads respectively as Among them, represents the size of the last dimension, and CL means that after concatenation operation, it is sent into the linear layer. After being sent into the layer normalization layer, the features after passing through the FFN module composed of 2 linear layers and the GELU activation layer are added with , and then the inverse transformation of dimension swapping and flattening is performed to obtain .

[0021] The input of the inter-class transformation module is flattened into , and then the dimension is swapped to , after being sent into the linear layer, the position encoding of the learnable parameter is added, and then it is sent into the layer normalization layer and the linear layer. The process is expressed as: Among them, , represents the weight of the linear layer. The last dimension is evenly divided into R heads to obtain ; the input is flattened into , and then the dimension is swapped to , after being sent into the linear layer, the position encoding of the learnable parameter PE is added. The process is expressed as: Among them, represents the weight of the linear layer, after being sent into the layer normalization layer, it is then sent into the linear layer respectively to obtain two components , , each component evenly divides the last dimension into R heads to obtain , , , and the attention of the R heads is calculated respectively as: After being sent into the layer normalization layer, the features after passing through the FFN module composed of 2 linear layers and the GELU activation layer are added with , and then the inverse transformation of dimension swapping and flattening is performed to obtain .

[0022] As a preferred scheme of the multi-component segmentation image method for pathological images based on class separation described in the present invention, the pathological image I i and other classes are sent into the class cross features extracted by , among which, , calculate the cross-category loss : Among them, Calculate The L1 norm of.

[0023] S103: The segmentation module performs segmentation based on the global features of the pathological image to obtain a multi-component segmentation result, and the category image reconstruction decoder performs category image reconstruction based on the fused category-specific features and the multi-component segmentation result to obtain a reconstructed image.

[0024] The process of the segmentation module performing segmentation based on the global features of the pathological image to obtain a multi-component segmentation result: Send the global feature Into the segmentation module based on the U-Net decoder , output the multi-component segmentation result P i Of the pathological image I i , and calculate the segmentation loss : Among them, Represents the cross-entropy loss, Represents the Dice similarity loss, Is the multi-component category label The label value of the m-th category of the t-th pixel, Is the segmentation result The predicted probability of the m-th category of the t-th pixel, and T is the number of pixels.

[0025] The process of the category image reconstruction decoder performing category image reconstruction based on the fused category-specific features and the multi-component segmentation result: Send the fused category-specific features Into the category image reconstruction decoder Perform category image reconstruction and output the reconstructed image , and calculate the reconstruction loss : Among them, Calculate The L2 norm of.

[0026] Specifically, the solution of the present invention will be further elaborated below through embodiments: In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0027] First, the data used in the present invention is collected from the international public dataset (Semantic Segmentation of Breast Cancer, BCSS). It contains whole-slide pathological stained images from 151 different patients, and all images are stained with HE. These whole-slide pathological images are annotated by an experienced pathologist, including the regions of all 5 types of components in the images: Tumor, Stroma, Lymphocytic infiltrate, Necrosis, and Others. 100 images are used as the training set, and 51 images are used as the test set. The following steps are used to process the whole-slide pathological images: Using a 512×512 sliding window, the whole-slide pathological images are cropped into non-overlapping local slice images; among the cropped images, the training set contains 32,000 local slice images, and the test set contains 15,000 local slice images.

[0028] The present invention uses the pytorch deep learning framework, torch version 1.13.0. All training and validation processes are completed on an NVIDIA GeForce RTX 3090 graphics card with 24G of video memory. During the training process, the neural network reads data using the mini-batch method, and the batch size is set to 2. The optimizer selects the Stochastic Gradient Descent (SGD) method, the initial learning rate is set to 0.01, the optimizer momentum is set to 0.9, and the optimizer regularization coefficient is set to 0.0001. The learning rate adjustment strategy selects the Poly method.

[0029] The experimental results of the present invention are as follows: To quantitatively evaluate the performance of the method proposed in the present invention, two different evaluation metrics are selected to evaluate the performance of the neural network in the multi-component segmentation task of breast cancer: Dice similarity coefficient (DSC) and Intersection over Union (IoU). The calculation formula of DSC is: Among them, is the number of correctly segmented pixels (true positives), is the number of incorrectly segmented pixels (false positives), is the number of missed segmented pixels (false negatives). mDSC calculates the average DSC of all classes. The calculation formula of the Intersection over Union metric is: The mIoU calculates the average IoU for all classes. The present invention compares the segmentation results with existing methods. As shown in Table 1, Table 1 compares the segmentation results of the present invention with other methods.

[0030] As Figure 2 shown, it is a schematic diagram of the Intra-class Transformation (IaT) module of the present invention. The input of the Intra-class Transformation (IaT) module (where C represents the number of channels, H represents the height, and W represents the width) is flattened (Flatten) into , and then the dimensions are permuted to . After being fed into a linear layer (Linear), the position encoding of the learnable parameter PE is added. This process can be expressed as where, represents the weight of the linear layer, After being fed into the layer normalization LN layer, it is respectively fed into three linear layers Linear to calculate three components , , . Each component evenly divides the last dimension into R heads, that is , , , . The r-th head of the Intra-class Transformation (IaT) module calculates the attention as where, represents the size of the last dimension of represents the r-th head of the component of the Intra-class Transformation (IaT) module, represents the r-th head of the component of the Intra-class Transformation (IaT) module. T represents transpose, represents the r-th head of the component of the Intra-class Transformation (IaT) module. A1 represents the attention of the first head, A2 represents the attention of the second head of the Intra-class Transformation (IaT) module, and A r represents the attention of the r-th head of the Intra-class Transformation (IaT) module, represents the attention of the R-th head of the Intra-class Transformation (IaT) module. CL represents concatenation operation followed by feeding into a linear layer, After being fed into the layer normalization layer and then the FFN module composed of a linear layer and a GELU activation layer, the features are added with , and then the inverse transformation of dimension permutation and flattening is performed to obtain .

[0031] As Figure 3As shown, it is a schematic diagram of the inter-class transformation ItT module of the present invention. The input of the inter-class transformation ItT module is flattened into , and then the dimensions are swapped to . After being fed into the linear layer, the position encoding with learnable parameters is added. Then it is fed into the layer normalization layer and the linear layer. This process can be expressed as where , represent the weights of the linear layer. Flatten the last dimension of into R heads, obtaining ; The input is flattened into , and then the dimensions are swapped to . After being fed into the linear layer, the position encoding of the learnable parameter PE is added. The process is expressed as: where represents the weight of the linear layer. After being fed into the layer normalization layer, it is then fed into the linear layer respectively to obtain two components , . Each component flattens the last dimension into R heads, obtaining , , . The r-th head of the inter-class transformation ItT module calculates the attention respectively as: where represents the r-th head of the component of the inter-class transformation ItT module, represents the r-th head of the component of the inter-class transformation ItT module, represents the r-th head of the component of the inter-class transformation ItT module, represents the attention of the 1st head of the inter-class transformation ItT module, represents the attention of the 2nd head of the inter-class transformation ItT module, represents the attention of the r-th head of the inter-class transformation ItT module, represents the attention of the R-th head of the inter-class transformation ItT module, After being fed into the layer normalization layer, the features after passing through the FFN module composed of the linear layer and the GELU activation layer are added with , and then the inverse transformation of dimension swapping and flattening is performed to obtain , They are successively sent to the layer normalization layer, and the features after passing through the FFN module composed of 2 linear layers and the GELU activation layer, plus , and then the inverse transformation of dimension exchange and flattening is performed to obtain .

[0032] As Figure 4 shown, it is a schematic diagram of the category fusion module CF of the present invention. The category fusion module CF is composed of an intra-class transformation IaT module, an inter-class transformation ItT module, concat (channel splicing), a linear layer (Linear), and a layer normalization (LN) layer. They are respectively sent to the intra-class transformation module to obtain , . As Input1, it is jointly sent to the inter-class transformation module with as Input2 to obtain the fused category-specific features . They are successively sent to concat, the linear layer, and the layer normalization layer together and then used as Input2, as Input1 and jointly sent to the inter-class transformation module to obtain the fused global features ; As Figure 5 shown, it is a schematic diagram of the category separation network of the present invention; the category separation network is composed of a global feature extractor, a category-specific feature extractor, a category fusion module, a segmentation module, and a category image reconstruction decoder. The global feature extractor , which is used to extract the global features i of the image I ; M category-specific feature extractors ( , M represents the number of component segmentation categories), which are used to extract the category-specific features i of the image I ; the category fusion module CF, which is used to fuse the global features and the category-specific features ; the category cross features i extracted by sending the image I and other categories are sent to for category separation; the segmentation module D g , which is used to obtain the multi-component segmentation result P by using the global features i as the input; the fused category-specific features are sent to the category image reconstruction decoder for category image reconstruction, and the reconstructed image is output.

[0033] As shown Figure 6 in the figure, it is the segmentation result diagram of different methods. The first row is the original pathological section image. The second row is the corresponding gold standard. Next are the segmentation result diagrams of DRD-UNet, TestFit, DETisSeg, PCSformer, HisynSeg, and the method of the present invention. The last row is the segmentation label, which is the area of 5 types of components in the image: Tumor (tumor), Stroma (stroma), Lymphocyticinfiltrate (lymphocytic infiltration), Necrosis (necrosis), and Others (others).

[0034] Table 1 Comparison table of experimental data Referring to Table 1, it can be seen that compared with other existing methods, in terms of the average segmentation accuracy of five categories, the method proposed in the present invention is better than other existing methods in both the Dice similarity coefficient and the intersection over union. Compared with the optimal HisynSeg method among other methods, the five-category mIoU of the method proposed in the present invention is increased by 1.39%, and the five-category mDSC reaches 79.83%. Referring Figure 6 to it, there are still many cases of incorrect segmentation in other methods. For example, in the image in the first column of the second row segmented by DRD-UNet, others are incorrectly segmented as stroma; in the image in the second column of the third row segmented by TestFit, necrosis is incorrectly segmented as tumor; in the image in the first column of the fourth row segmented by DETisSeg, some others are incorrectly segmented as stroma; in the image in the first column of the fifth row segmented by PCSformer, some others are incorrectly segmented as stroma; in the image in the first column of the sixth row segmented by HisynSeg, some others are incorrectly segmented as stroma; the segmentation result of the method of the present invention is very close to the segmentation label in the last row.

[0035] To prove the beneficial effects of each part of the present invention on image segmentation, ablation experiments were further carried out below. The basic network is a global feature extractor and a segmentation module to form a UNet, and is used to constrain multi-class segmentation; the basic network + class-specific feature extractor + class image reconstruction decoder means adding a class-specific feature extractor and a class image reconstruction decoder to the basic network, and using to constrain multi-class segmentation and to constrain class image reconstruction; the basic network + class-specific feature extractor + class image reconstruction decoder It shows that a class - specific feature extractor, a class - image reconstruction decoder, and a class fusion module are added to the basic network, and is used to constrain multi - class segmentation, is used to constrain class - image reconstruction, and is used to constrain class separation; basic network + class - specific feature extractor + class - image reconstruction decoder + inter - class transformation module It shows that a class - specific feature extractor, a class - image reconstruction decoder, and an inter - class transformation module are added to the basic network, and is used to constrain multi - class segmentation, is used to constrain class - image reconstruction, and is used to constrain class separation; basic network + class - specific feature extractor + class - image reconstruction decoder + class fusion module It shows that a class - specific feature extractor, a class - image reconstruction decoder, and a class fusion module are added to the basic network, and is used to constrain multi - class segmentation, is used to constrain class - image reconstruction, and is used to constrain class separation; finally, it is the method of the present invention. It can be seen from Table 2 that each newly added module or loss function improves the evaluation metrics mIoU and mDSC.

[0036] Table 2 Ablation Experiment Example 2: Second, as Figure 7 shown, to achieve the above - mentioned purpose, the present invention discloses a multi - component segmentation image system for pathological images based on class separation, including: An image processing unit 11, configured to obtain a pathological image and input the pathological image into a pre - established class - separation network model for processing, wherein the class - separation network model includes a global feature extractor, a class - specific feature extractor, a class fusion module, a segmentation module, and a class - image reconstruction decoder; A feature fusion unit 12, configured to respectively extract the global feature and the class - specific feature of the pathological image by the global feature extractor and the class - specific feature extractor, and the class fusion module fuses the global feature and the class - specific feature of the pathological image to obtain the fused global feature and class - specific feature; A component segmentation unit 13, configured to perform segmentation on the basis of the global feature of the pathological image by the segmentation module to obtain a multi - component segmentation result, and the class - image reconstruction decoder performs class - image reconstruction based on the fused class - specific feature and the multi - component segmentation result to obtain a reconstructed image.

[0037] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0038] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0039] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0040] The foregoing has shown and described the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure claimed.

Claims

1. A method for multi-component segmentation of pathological images based on class separation, characterized in that, The method includes the following steps: Obtain a pathological image, and input the pathological image into a pre-established class separation network model for processing, where the class separation network model includes a global feature extractor, a class-specific feature extractor, a class fusion module, a segmentation module, and a class image reconstruction decoder; The global feature extractor and the class-specific feature extractor respectively extract the global feature and the class-specific feature of the pathological image, and the class fusion module fuses the global feature and the class-specific feature of the pathological image to obtain the fused global feature and class-specific feature; The segmentation module performs segmentation based on the global feature of the pathological image to obtain a multi-component segmentation result, and the class image reconstruction decoder performs class image reconstruction based on the fused class-specific feature and the multi-component segmentation result to obtain a reconstructed image.

2. The method for segmenting multi-component images in pathological images based on class separation according to claim 1, wherein The pre-established class separation network model is obtained by jointly optimizing through an alignment loss, a class cross loss, a segmentation loss function, a contrast loss function, and a reconstruction loss.

3. The method for segmenting multi-component images in pathological images based on class separation according to claim 1, characterized in that, The pathological image is obtained through the constructed pathological image training set and includes pathological image I i and its corresponding multi-class labels denotes the number of pathological images in the training set. The global feature extractor E based on the U-Net encoder is constructed g to extract the global features of pathological image I i ; the M class-specific feature extractors based on the U-Net encoder are constructed to extract the class-specific features of pathological image I ; M represents the number of component segmentation categories i ; denotes pixel multiplication ​​ 4. The method for segmenting multi-component images in pathological images based on class separation according to claim 1, characterized in that, The process of the class fusion module fusing the global feature of the pathological image: Send into the category fusion module CF to obtain the fused , and use the KL divergence to globally feature align with the fused global feature , and the alignment loss is defined as: Among them, represents the calculation of the global feature and the fused global feature the KL divergence between them; The category fusion module consists of an intra-class transformation module, an inter-class transformation module, concat, a linear layer, and a layer normalization layer. They are respectively fed into the intra-class transformation module to obtain and , As Input1, it is jointly sent into the inter-class transformation module with Input2 to obtain the fused class-specific features , After being successively fed into the concatenation layer, the linear layer, and the layer normalization layer, it is used as Input2. As Input1, it is jointly fed into the inter-class transformation module to obtain the fused global features. .

5. The method for segmenting multi-component images in pathological images based on class separation according to claim 4, characterized in that, The input of the intra-class transformation module Flattened into , and then the dimensions are permuted to , where C represents the channels, H represents the height, and W represents the width. After being fed into the linear layer Linear, the positional encoding of the learnable parameter PE is added. The process is expressed as Among them, represents the weight of the linear layer, After being sent into the layer normalization LN layer, it is respectively sent into three linear layers Linear to calculate three components , , Each component evenly divides the last dimension into R heads, that is , , , , and the R heads respectively calculate the attention as Among them, represents the size of the last dimension of the in-class transformation IaT module component of the r-th head, the in-class transformation IaT module component of the r-th head, T represents transpose, the in-class transformation IaT module component of the r-th head, A1 represents the attention of the first head, A2 represents the attention of the second head of the in-class transformation IaT module, A r represents the attention of the r-th head of the in-class transformation IaT module, represents the attention of the R-th head of the in-class transformation IaT module, CL represents concatenation operation followed by feeding into a linear layer, fed into a layer normalization layer, the features after passing through an FFN module composed of a linear layer and a GELU activation layer, plus and then performing the inverse transformation of dimension swapping and flattening to obtain .

6. The method for segmenting multi-component images in pathological images based on class separation according to claim 5, wherein, The input of the inter-class transformation module is flattened into , and then the dimensions are swapped to . After being fed into the linear layer, the positional encoding with learnable parameters is added, and then it is fed into the layer normalization layer and the linear layer. The process is expressed as: Among them, , represent the weights of the linear layer. The last dimension of is evenly divided into R heads to obtain ; The input is flattened into , and then the dimensions are swapped to . After being fed into the linear layer, the position encoding of the learnable parameter PE is added. The process is expressed as: Among them, represents the weight of the linear layer, After being sent into the layer normalization layer, they are respectively sent into the linear layer to obtain two components , , and each component evenly divides the last dimension into R heads to obtain , , , the attention of R heads is calculated respectively as: Among them, represents the r-th head of the inter-class transformation ItT module component ; represents the r-th head of the inter-class transformation ItT module component ; represents the r-th head of the inter-class transformation ItT module component ; represents the attention of the first head of the inter-class transformation ItT module, represents the attention of the second head of the inter-class transformation ItT module, represents the attention of the r-th head of the inter-class transformation ItT module, represents the attention of the R-th head of the inter-class transformation ItT module, which are successively sent into the layer normalization layer, and the features after passing through the FFN module composed of a linear layer and a GELU activation layer, plus , and then an inverse transformation of dimension swapping and flattening is performed to obtain .

7. The method for segmenting multi-component images in pathological images based on category separation according to claim 1, characterized in that Within the class-specific feature extractor, the pathological image I i and other classes are sent into the extracted class-crossing features , where , calculate the class-crossing loss : Among them, Calculate the L1 norm of.

8. The method for segmenting multi-component images in pathological images based on class separation according to claim 1, wherein, The process of the segmentation module performing segmentation based on the global feature of the pathological image to obtain a multi-component segmentation result: Globally send the features to the segmentation module based on the U-Net decoder to output the multi-component segmentation result P i of the pathological image I i and calculate the segmentation loss : Among them, represents the cross-entropy loss, represents the Dice similarity loss, is the multi-class category label of the label value of the m-th class of the t-th pixel in is the segmentation result of the predicted probability of the m-th class of the t-th pixel, and T is the number of pixels.

9. The method for segmenting multi-component images in pathological images based on category separation according to claim 1, wherein The process of the class image reconstruction decoder performing class image reconstruction based on the fused class-specific feature and the multi-component segmentation result: Send the fused class-specific features to the class image reconstruction decoder for class image reconstruction to output the reconstructed image and calculate the reconstruction loss : Among them, Calculate the L2 norm of.

10. A multi-component segmentation image system for pathological images based on class separation, which adopts the method for multi-component segmentation of pathological images based on class separation according to any one of claims 1 to 9, characterized in that, Includes: An image processing unit for obtaining a pathological image and inputting the pathological image into a pre-established class separation network model for processing, where the class separation network model includes a global feature extractor, a class-specific feature extractor, a class fusion module, a segmentation module, and a class image reconstruction decoder; A feature fusion unit for the global feature extractor and the class-specific feature extractor to respectively extract the global feature and the class-specific feature of the pathological image, and the class fusion module to fuse the global feature and the class-specific feature of the pathological image to obtain the fused global feature and class-specific feature; A component segmentation unit for the segmentation module to perform segmentation based on the global feature of the pathological image to obtain a multi-component segmentation result, and the class image reconstruction decoder to perform class image reconstruction based on the fused class-specific feature and the multi-component segmentation result to obtain a reconstructed image.

Citation Information

Patent Citations

  • Small sample image classification method fusing global and adaptive local information

    CN119580011A

  • Cancer patient prognosis prediction method based on graph alignment and multi-modal alignment

    CN119833116A

  • Rapid pathological image analysis method and apparatus based on magnification-aligned transformer

    WO2025065803A1

Cited By

  • Breast ultrasound image segmentation method and system of cross-graph fusion network

    CN120689357A

  • A method and system for breast ultrasound image segmentation using cross-graph fusion networks

    CN120689357B