Pathological image segmentation-based pseudo label generation method, system and device, and medium

By introducing graph attention network and global classification maximum pooling technology in pathological image segmentation, the problems of poor quality of pseudo-label generation and insufficient weak supervision signal transmission in the existing technology are solved, which significantly improves the accuracy and stability of pathological image segmentation and meets clinical diagnosis needs.

CN120125815APending Publication Date: 2025-06-10SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510136195.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When processing pathological images, the existing weakly supervised semantic segmentation methods face poor quality of pseudo-label generation, insufficient recognition ability on a few categories, large impact on model initialization parameters, easy error accumulation, and insufficient weakly supervised signal transmission, resulting in low accuracy and stability of pathological image segmentation.

Method used

The graph attention network and global classification maximum pooling technology are introduced to improve the accuracy and stability of pathological image segmentation by establishing contextual relationships between pixels, enhancing the transmission of weakly supervised signals, and optimizing the pseudo-label generation process.

Benefits of technology

Generate finer-grained and accurate pseudo-labels through graph attention networks, combining global classification maximum pooling and global average pooling technologies to strengthen the transmission of weakly supervised signals, improve the performance of pathological image segmentation, and meet clinical diagnostic needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125815A_ABST
    Figure CN120125815A_ABST
Patent Text Reader

Abstract

The invention discloses a pseudo label generation method, system and device based on pathological image segmentation and a medium, and the method comprises the steps: segmenting a pathological full-slice image into a plurality of square small-block images, and carrying out the multi-label labeling of the square small-block images; inputting the marked square small block images into a pre-trained segmentation model, and extracting features of the marked square small block images to obtain a feature map; based on the feature map, information spreading is carried out by using an attention mechanism, and feature representation of the nodes is updated; according to the feature representation of the updated node and the spatial information of the node, generating probability distribution of each node belonging to different categories, and obtaining a pseudo label; refining the pseudo labels; and optimizing the refined pseudo tag to obtain a final pseudo tag. The image attention network and the global classification maximum pooling technology are introduced, and the accuracy and stability of pathological image segmentation are improved by establishing the context relation between pixels, enhancing transmission of weak supervision signals and optimizing the pseudo label generation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image pseudo-labeling, and in particular to a pseudo-label generation method, system, device and medium based on pathological image segmentation. Background Art

[0002] In recent years, computer artificial intelligence-assisted diagnosis tools derived from the development of deep learning have gradually been implemented, and some related technical requirements have also emerged. Especially in the field of medical image analysis, the automatic analysis of pathological images has become a research hotspot. Traditional pathological image analysis relies on a large amount of manual annotation, and these annotation processes are not only time-consuming and laborious, but also easily affected by subjective factors, resulting in inconsistent annotation quality. In order to reduce the workload of manual annotation and improve the accuracy of pathological image segmentation, weakly supervised learning methods have emerged. This weakly supervised learning method trains the model by using a small number of coarse-grained annotations (such as image-level labels), greatly reducing the annotation cost. However, when dealing with pathological images, the existing weakly supervised semantic segmentation methods still face the following technical challenges: poor pseudo-label generation quality, insufficient recognition ability for minority classes, large influence of model initialization parameters, easy error accumulation, and insufficient weakly supervised signal transmission, resulting in low accuracy and stability of pathological image segmentation. Summary of the Invention

[0003] The main purpose of this application is to overcome the disadvantages and deficiencies of the prior art, and provide a pseudo-label generation method, system, device and medium based on pathological image segmentation. By introducing a graph attention network and global classification max pooling technology, the accuracy and stability of pathological image segmentation are improved by establishing the context relationship between pixels, enhancing the transmission of weakly supervised signals, and optimizing the pseudo-label generation process.

[0004] To achieve the above purpose, this application adopts the following technical solutions:

[0005] In the first aspect, this application provides a pseudo-label generation method based on pathological image segmentation, including the following steps:

[0006] Cut the digital pathological whole-slide image into several square small-block images of a preset size, and perform multi-label annotation on the several square small-block images to obtain the annotated square small-block images;

[0007] Input the annotated square small-block images into a pre-trained segmentation model, and the feature extraction module in the pre-trained segmentation model extracts the features of the annotated square small-block images to obtain a feature map;

[0008] Input the feature map into the graph attention network module in the pre-trained segmentation model. The graph attention network module uses the attention mechanism for information propagation to update the feature representation of the nodes;

[0009] Generate the probability distribution of each node belonging to different categories based on the updated feature representation of the nodes and the spatial information of the nodes to obtain pseudo-labels;

[0010] The refinement module in the pre-trained segmentation model refines the pseudo-labels to obtain refined pseudo-labels;

[0011] The optimization module in the pre-trained segmentation model optimizes the refined pseudo-labels to obtain the final pseudo-labels.

[0012] As a preferred technical solution, the generating the probability distribution of each node belonging to different categories based on the updated feature representation of the nodes and the spatial information of the nodes to obtain pseudo-labels includes:

[0013] The graph attention network module divides the feature map into multiple regions, where each region serves as a node;

[0014] Connect the nodes through a graph structure to obtain a complete connected graph;

[0015] Based on the complete connected graph, perform information propagation through the attention mechanism to gradually update the feature representation of the nodes;

[0016] Obtain the coordinates of each node in the digital pathology whole slide image and normalize the coordinates;

[0017] Encode the normalized coordinates using sine and cosine functions to generate the spatial information of each node;

[0018] Based on the updated feature representation of the nodes and the spatial information of the nodes, obtain the probability distribution of each node belonging to different categories;

[0019] Normalize the probability distribution to obtain pseudo-labels.

[0020] As a preferred technical solution, the refining the pseudo-labels includes:

[0021] Use a fixed convolution kernel to extract the RGB values and position encoding information of neighboring pixels from the feature map;

[0022] Calculate the comprehensive similarity between pixels based on the RGB values and position encoding information of the neighboring pixels;

[0023] Use the similarity between pixels to perform weighted summation on the feature representation in the pseudo-labels to obtain refined pseudo-labels.

[0024] As a preferred technical solution, the pseudo-label is optimized, including:

[0025] Obtain the refined pseudo-label and a preset confidence threshold;

[0026] Traverse each pixel position in the refined pseudo-label in sequence;

[0027] Calculate the maximum probability value of each pixel position in all categories;

[0028] Compare the maximum probability value with the preset confidence threshold;

[0029] If the maximum probability value is greater than the preset confidence threshold, then use the category of the maximum probability value as the output of the pixel;

[0030] If the maximum probability value is less than the preset confidence threshold, then ignore the pixels of the category of the maximum probability value;

[0031] Based on the output of the pixel and the pixels that ignore the category of the maximum probability value, obtain the optimized pseudo-label.

[0032] As a preferred technical solution, a siamese network structure is used to train the segmentation model to obtain a trained segmentation model;

[0033] Among them, the siamese network structure includes two branches with shared parameters, and the results output by the two branches will be used to guide each other;

[0034] Both of the two branches include the feature extraction module, the graph attention network module, the refinement module and the optimization module.

[0035] As a preferred technical solution, the siamese network structure optimizes the segmentation model through the pseudo-label loss, the graph attention network weak supervision loss and the segmentation model weak supervision loss. Specifically:

[0036] Sum the pseudo-label loss, the pseudo-label loss function and the segmentation model weak supervision loss function to obtain the total loss;

[0037] Use the total loss to optimize the segmentation model.

[0038] As a preferred technical solution, it also includes aggregating the features in the feature map by using global classification max pooling and global average pooling, including:

[0039] Perform global classification max pooling on the feature map to obtain the feature representation of the significant region;

[0040] Perform global average pooling on the feature map to obtain a global mean feature representation;

[0041] Aggregate the feature representation of the significant region and the global mean feature representation to obtain an aggregated feature.

[0042] In a second aspect, the present application provides a pseudo-label generation system based on pathological image segmentation, which is applied to the pseudo-label generation method based on pathological image segmentation, and includes an annotation module, a feature extraction module, a feature update module, a pseudo-label generation module, a refinement module, and an optimization module;

[0043] The annotation module is used to divide the digital pathology whole slide image into several square small block images of a preset size, and perform multi-label annotation on the several square small block images to obtain the annotated square small block images;

[0044] The feature extraction module is used to input the annotated square small block images into a pre-trained segmentation model, and the feature extraction module in the pre-trained segmentation model extracts the features of the annotated square small block images to obtain a feature map;

[0045] The feature update module is used to input the feature map into the graph attention network module in the pre-trained segmentation model, and the graph attention network module uses the attention mechanism to perform information propagation and update the feature representation of the nodes;

[0046] The pseudo-label generation module is used to generate the probability distribution of each node belonging to different categories according to the feature representation of the updated nodes and the spatial information of the nodes to obtain pseudo-labels;

[0047] The refinement module is used to refine the pseudo-labels by the refinement module in the pre-trained segmentation model to obtain refined pseudo-labels;

[0048] The optimization module is used to optimize the refined pseudo-labels by the optimization module in the pre-trained segmentation model to obtain the final pseudo-labels.

[0049] In a third aspect, the present application provides an electronic device, and the electronic device includes:

[0050] At least one processor; and a memory communicatively connected to the at least one processor;

[0051] Wherein, the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the pseudo-label generation method based on pathological image segmentation.

[0052] In summary, compared with the prior art, the technical solutions provided by this application bring at least the following beneficial effects:

[0053] This application proposes a pseudo-label generation method based on pathological image segmentation, which divides a digital pathological whole-slide image into several square small-block images of a preset size, and performs multi-label annotation on the several square small-block images; inputs the annotated square small-block images into a pre-trained segmentation model, and the feature extraction module in the pre-trained segmentation model extracts the features of the annotated square small-block images to obtain a feature map; inputs the feature map into the graph attention network module in the pre-trained segmentation model, and the graph attention network module uses the attention mechanism to perform information propagation and update the feature representation of the nodes; generates the probability distribution of each node belonging to different categories according to the updated feature representation of the nodes and the spatial information of the nodes to obtain pseudo-labels; the refinement module in the pre-trained segmentation model refines the pseudo-labels to obtain refined pseudo-labels; the optimization module in the pre-trained segmentation model optimizes the refined pseudo-labels to obtain the final pseudo-labels. This application introduces the graph attention network and global classification max pooling technology, and improves the accuracy and stability of pathological image segmentation by establishing the context relationship between pixels, enhancing the transmission of weak supervision signals, and optimizing the pseudo-label generation process, so as to better meet the clinical diagnosis requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a flowchart of a pseudo-label generation method based on pathological image segmentation provided by an embodiment of this application;

[0056] Figure 2 It is a schematic flowchart of the training of the segmentation model provided by an embodiment of this application;

[0057] Figure 3 It is a block diagram of a pseudo-label generation system based on pathological image segmentation provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of this application.

[0059] Reference to "embodiment" in this application means that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in this application may be combined with other embodiments.

[0060] In recent years, computer artificial intelligence-assisted diagnosis tools derived from the development of deep learning have been gradually implemented, and some related technical requirements have also emerged. Especially in the field of medical image analysis, the automated analysis of pathological images has become a research hotspot. Traditional pathological image analysis relies on a large number of manual annotations, and these annotation processes are not only time-consuming and laborious, but also easily affected by subjective factors, resulting in inconsistent annotation quality. In order to reduce the workload of manual annotation and improve the accuracy of pathological image segmentation, weakly supervised learning methods have emerged. This method trains the model by using a small number of coarse-grained annotations (such as image-level labels), greatly reducing the annotation cost. However, when existing weakly supervised semantic segmentation methods are applied to pathological images, they still face the following technical challenges: 1. Poor quality of pseudo-labels: Although it is simple to generate pseudo-labels using class activation maps in existing methods, the generated pseudo-labels often have low resolution and blurred boundaries, making it difficult to accurately segment pathological tissue regions; 2. Low efficiency of supervised signal transmission: Commonly used global average pooling or global max pooling techniques often perform poorly in transmitting weakly supervised signals. Especially when the details in the image are complex, the supervised signals cannot be effectively transmitted to all regions, resulting in the model being unable to capture sufficient local information and affecting the segmentation accuracy; 3. Insufficient recognition ability for minority classes: Due to the problem of sample imbalance in pathological images, existing methods tend to focus on processing dominant classes and have weak recognition ability for non-dominant classes (such as lymphocytes and necrotic tissues around tumor tissues), resulting in poor performance when dealing with minority classes; 4. Great influence of model initialization parameters and easy error accumulation: In weakly supervised tasks, the segmentation performance is significantly affected by the model initialization parameters. If the initialization is improper, it may lead to large fluctuations during the model training process, affecting the convergence speed and even the final performance. In addition, errors usually accumulate, and the training process is sensitive, prone to overfitting or unstable training; The above problems lead to low accuracy and stability of pathological image segmentation.

[0061] To solve the above problems, this application proposes a pseudo-label generation method based on pathological image segmentation, which introduces a graph attention network and global classification max pooling technology, aiming to improve the accuracy and stability of pathological image segmentation by establishing the context relationship between pixels, enhancing the transmission of weakly supervised signals, and optimizing the pseudo-label generation process, so as to better meet the clinical diagnosis needs.

[0062] Please refer to Figure 1 , in an embodiment of this application, a pseudo-label generation method based on pathological image segmentation is provided, including the following steps:

[0063] S1. Cut the digital pathological whole slide image into several square small images of a preset size, and perform multi-label annotation on the several square small images to obtain the annotated square small images.

[0064] The digital pathological whole-slide image in this application refers to the digitization of the entire glass slide of a pathological tissue section through a scanner to generate a high-resolution digital image.

[0065] Furthermore, the preset size is 224×224 pixels, that is, the size of the square small-block image is 224×224 pixels. This size can not only retain sufficient tissue details but also facilitate subsequent processing.

[0066] Each square small-block image may contain multiple tissue types or lesion areas. Therefore, using multi-label annotation can effectively describe the complex pathological information in the square small-block image. For example, a square small-block image may be simultaneously labeled as "tumor", "necrosis", and "lymphocyte infiltration". Finally, a sample will correspond to a label, such as: 224×224→[1,0,0,1]. In addition, during the annotation process, the pathologist does not need to perform pixel-level or contour-drawing annotation on specific regions in the image, but rather performs multi-label annotation on the entire square small-block image.

[0067] S2. Input the annotated square small-block image into a pre-trained segmentation model. The feature extraction module in the pre-trained segmentation model extracts the features of the annotated square small-block image to obtain a feature map.

[0068] Furthermore, the feature extraction module in this application uses the DeeplabV3+ backbone network for feature extraction. Among them, DeeplabV3+ is a deep learning model commonly used in image semantic segmentation tasks. Its core feature is the introduction of atrous convolution and atrous spatial pyramid pooling (ASPP) to enhance the multi-scale feature extraction ability.

[0069] Specifically, DeeplabV3+ extracts rich features through depthwise separable convolutional layers, atrous convolutional layers, and atrous spatial pyramid pooling. Among them, the depthwise separable convolutional layer realizes a larger receptive field by introducing an extended convolutional kernel, which helps to capture multi-scale information. The atrous convolutional layer realizes a larger receptive field by introducing an extended convolutional kernel, which helps to capture multi-scale information. The atrous spatial pyramid pooling fuses the feature map through multi-scale atrous convolutional pooling, thereby enhancing the model's perception ability of different-scale structures.

[0070] S3. Input the feature map into the graph attention network module in the pre-trained segmentation model. The graph attention network module uses the attention mechanism for information propagation to update the feature representation of the nodes.

[0071] S4. Generate the probability distribution of each node belonging to different categories based on the feature representation of the updated node and the spatial information of the node, so as to obtain pseudo-labels.

[0072] The Graph Attention Network (GAT) of the present application is a model that introduces the self-attention mechanism into the graph neural network and is used to process graph-structured data. In order to overcome the boundary blur and noise problems in generating pseudo-labels by traditional Class Activation Map (CAM) methods, the present application uses the Graph Attention Network (GAT) to generate seed pseudo-labels. The graph attention network module can effectively solve the problems of "missing and mislabeling" and "incorrect labeling" in pseudo-labels, and improve the quality of pseudo-labels through stronger expressive power and non-linear feature transformation.

[0073] A pseudo-label refers to a non-manually annotated label generated by a model prediction or an algorithm, which is used to replace the missing accurate label in the weak supervision learning process.

[0074] Specifically, generating the probability distribution of each node belonging to different categories based on the feature representation of the updated node and the spatial information of the node to obtain pseudo-labels includes:

[0075] S41. In the graph attention network module, first split the feature map extracted from the DeeplabV3+ backbone network into multiple regions, and each region is regarded as a node in the graph attention network;

[0076] S42. Connect the nodes through a graph structure to form a fully connected graph;

[0077] In this way, information can be transmitted between different regions regardless of their positions, so as to realize the integration of global context information.

[0078] S43. Based on the complete connected graph, perform information propagation through the attention mechanism to layer-by-layer update the feature representation of the node;

[0079] Furthermore, the feature representation of each node in the graph attention network module (i.e., the feature of each small square image) will be updated layer by layer. Information propagation is performed through a non-linear attention mechanism, and the update process of each layer can be described by the following formula:

[0080] Input feature transformation:

[0081]

[0082] Among them, represents the feature vector of the i-th node in the (l - 1)-th layer; W in represents the weight matrix; LeakyReLU is the activation function;

[0083] Attention calculation mechanism:

[0084]

[0085] where a ijh represents the correlation between node i and node j under different attention heads h; f softmax (x i ) is the softmax function for normalization; a ij is the correlation between node i and node j under different attention heads h;

[0086] Node feature update:

[0087]

[0088] where x id is the original feature of node i; pos id represents the position encoding information of node i.

[0089] S44. Obtain the coordinates of each node (i.e., the square patch image) in the original digital pathology whole slide image, and normalize the coordinates;

[0090] S45. Encode the normalized coordinates using sine and cosine functions to generate the spatial information of each node;

[0091] To enable the segmentation model to have spatial perception ability, position encoding is performed for each node. This position encoding not only provides the spatial position of the node but also combines the structural information in the image, enhancing the understanding of local and global information by the graph attention network.

[0092] S46. Based on the updated feature representation of the node and the spatial information of the node, obtain the probability distribution of each node belonging to different categories;

[0093] S47. Normalize the probability distribution to obtain the pseudo-labels.

[0094] Through the multi-layer graph attention mechanism, the final feature representation of each node can capture the global context information and perform spatial perception in combination with the position encoding; this information is converted into the probability distribution of each node belonging to different categories in the last layer of the graph attention network; specifically, after the final node features are normalized by Softmax, they can be regarded as the probabilities of the node belonging to each category; through this process, the graph attention network module generates more fine-grained and accurate seed pseudo-labels. These pseudo-labels can be used as the supervision signals for subsequent training to guide the weakly supervised semantic segmentation model to continue learning.

[0095] The graph attention network module in this application not only promotes information exchange between different layers, but also enhances the expressive ability of features by introducing an attention mechanism.

[0096] S5. The refinement module in the pre-trained segmentation model refines the pseudo-labels to obtain refined pseudo-labels.

[0097] S6. The optimization module in the pre-trained segmentation model optimizes the refined pseudo-labels to obtain the final pseudo-labels.

[0098] To solve the problem of poor quality of generated pseudo-labels, this application refines and optimizes the generated pseudo-labels to improve the accuracy and detail of the pseudo-labels.

[0099] In this application, the refinement module optimizes the alignment of the low-level image appearance by combining the RGB information and position information of local pixels, realizes the boundary refinement and noise suppression of the pseudo-labels, and improves the quality of the generated pseudo-labels; the optimization module masks the low-confidence regions to avoid the interference of these regions on the training of the segmentation model, thereby further improving the performance of the segmentation model.

[0100] Specifically, refining the pseudo-labels includes:

[0101] S51. Use a fixed convolution kernel to extract the RGB values and position encoding information of neighboring pixels from the feature image;

[0102] Among them, the RGB convolution kernel focuses on pixel color information, while the position convolution kernel focuses on the spatial position of pixels, thereby enhancing local context information; its formula is:

[0103]

[0104] Among them, I ij and P ij are the RGB color and position information of pixel i respectively; and are the standard deviations calculated according to the pixel neighborhood; the weights w 1 and w 2 are learnable coefficients used to adjust the contributions of RGB and position features to the similarity calculation.

[0105] S52. Calculate the comprehensive similarity between pixels according to the RGB values and position encoding information of the neighboring pixels;

[0106] S53. Use the similarity between pixels to perform weighted summation on the feature representation in the pseudo-labels to obtain refined pseudo-labels.

[0107] Furthermore, the similarity between comprehensive pixels is calculated through the SoftMax function, and the feature representation in the pseudo-label output by the graph attention module is weighted using this similarity to obtain a refined pseudo-label. The formula is as follows:

[0108]

[0109] Among them, κ ij,kl is the similarity between comprehensive pixels; is the feature representation of the refined pseudo-label; represents the neighborhood set of pixel i; is the set of convolution expansion factors; t is the refinement iteration number. Through multiple iterations, the refinement module can gradually improve the accuracy of the pseudo-label.

[0110] To further optimize the quality of the pseudo-label, the optimization module performs masking processing on low-confidence regions to avoid the influence of low-confidence regions on model training. The optimization module sets a threshold, and pixels below this threshold are regarded as "invalid", that is, the gradients of these pixels are not calculated. The optimization module is based on the maximum probability value of each pixel category and the set confidence threshold δ fg for comparison. If the probability value is lower than the threshold, the region is ignored and does not participate in subsequent gradient calculations.

[0111] Specifically, optimizing the pseudo-label includes:

[0112] S61. Obtain the refined pseudo-label and the preset confidence threshold;

[0113] S62. Traverse each pixel position in the refined pseudo-label in turn;

[0114] S63. Calculate the maximum probability value of each pixel position in all categories;

[0115] S64. Compare the maximum probability value with the preset confidence threshold;

[0116] S65. If the maximum probability value is greater than the preset confidence threshold, use the category of the maximum probability value as the output of the pixel;

[0117] S66. If the maximum probability value is less than the preset confidence threshold, ignore the pixels of the category with the maximum probability value;

[0118] S67. Based on the output of the pixel and the pixels that ignore the category with the maximum probability value, obtain the optimized pseudo-label.

[0119] The specific formula of the optimization module is:

[0120]

[0121] Among them, 255 represents an ignored label (i.e., the gradient of this pixel is not calculated).

[0122] In the embodiments of the present application, please refer to Figure 2 , and a siamese network structure is adopted to train the segmentation model to obtain a trained segmentation model;

[0123] Among them, the siamese network structure includes two branches with shared parameters, and the results output by the two branches will be used to guide each other;

[0124] Both of the two branches include a feature extraction module, a graph attention network module, a refinement module, and an optimization module.

[0125] The siamese network structure is a special neural network structure composed of two or more branch networks with shared parameters, and is used to learn the similarity between input data.

[0126] Specifically, in the siamese network structure, DeeplabV3+ in the feature extraction module also extracts rich features through depthwise separable convolutional layers, atrous convolutional layers, and atrous spatial pyramid pooling; among them, the depthwise separable convolutional layer realizes a larger receptive field by introducing an extended convolutional kernel, which helps to capture multi-scale information; the atrous convolutional layer realizes a larger receptive field by introducing an extended convolutional kernel, which helps to capture multi-scale information; the atrous spatial pyramid pooling fuses the feature maps through multi-scale atrous convolutional pooling, thereby enhancing the model's perception ability of different scale structures. In the siamese network structure, its specific extraction process includes:

[0127]

[0128] Among them, X A and X B represent the inputs of the two branches in the siamese network structure, that is, the input is a square small image with a size of 224×224, and θ A and θ B are the parameters of the DeeplabV3+ models in the two branches respectively; and are the feature outputs (feature maps) of the two branches after being processed by the DeeplabV3+ model; and are the feature representations obtained by the atrous spatial pyramid pooling (ASPP) modules in the two branches.

[0129] Then, the convolutional layers of DeeplabV3+ are used to extract high-dimensional features layer by layer, capturing local texture and global context information; the Atrous Spatial Pyramid Pooling (ASPP) module aggregates information from different receptive fields through parallel operations of atrous convolutions at different scales, enhancing the recognition effect for complex pathological regions; finally, a feature map after extraction by DeeplabV3+ is obtained, which will have multi-scale information and boundary details.

[0130] During the training process of the segmentation model using the siamese network structure, the processes of the graph attention network module, the refinement module, and the optimization module are the same as above, that is, the graph attention network module generates two pseudo-labels for the feature maps of the two branches respectively, and the refinement modules in the two branches refine the pseudo-labels of their respective branches respectively; the optimization modules in the two branches optimize the pseudo-labels refined by their respective branches respectively.

[0131] The siamese network structure is adopted in this application to solve the problems of the influence of the initial parameters of the segmentation model on the training process and error accumulation. The siamese network structure effectively alleviates the sensitivity of the initial parameters and improves the stability of the segmentation model training by using two network branches with shared parameters for differential training; in the application of the siamese network, the two branches process the input data in parallel, and these two branches share the same weights. By using the siamese structure, the results output by each branch will be used to guide each other to ensure the consistency of the network throughout the training process. This design significantly reduces the dependence of the segmentation model on initialization and avoids performance degradation caused by incorrect initialization. For example, when one branch processes a certain square patch image of a pathological image, the other branch may process a different square patch image, and through contrast learning, the segmentation model can be optimized from global and local information. This method enhances the model's ability to learn features, especially in non-dominant classes (such as small lesion areas) in the image, and shows more robustness.

[0132] In addition, to address the deficiencies of existing methods in aggregating weakly supervised signals, when training the segmentation model using the siamese network structure in this application, a feature aggregation method combining global classification max pooling (GCMP) and global average pooling (GAP) for feature maps is proposed to enhance the segmentation model's perception ability of complex structures in pathological images.

[0133] Among them, global classification max pooling (GCMP) is an improved feature aggregation method that enhances the transmission effect of weakly supervised signals by performing max pooling on the features of each category and weighted aggregation. Global average pooling (GAP) is a commonly used feature aggregation method that reduces the dimension by calculating the average value of all pixels in the feature map.

[0134] Aggregate the features in the feature map using global classification max pooling and global average pooling, including:

[0135] Perform global classification max pooling on the feature map to obtain the feature representation of the significant region;

[0136] Perform global average pooling on the feature map to obtain the global mean feature representation;

[0137] Aggregate the feature representation of the significant region and the global mean feature representation to obtain the aggregated feature.

[0138] Specifically, global classification max pooling (GCMP) aims to improve the traditional global max pooling (GMP) method. GCMP enhances the recognition ability for minority classes (such as tumors, necrosis regions, etc.) by performing max pooling on the features of each class and weighted averaging according to the class weights. The basic idea of GCMP is: in the semantic segmentation task, the classification label of each pixel is crucial for the classification result of the entire image. Through the GCMP method, on the basis of maintaining global information, the most important regional features can be extracted by max pooling. Specifically, the GCMP aggregation method first performs max pooling on the probability features of each pixel point, and then weights the result according to the weight of this class, thus focusing on the features of the target class; assuming i is all pixel positions in the image, is the predicted probability of a certain class, and the aggregated output result The formula can be expressed as:

[0139]

[0140] where, represents the total number of pixels, and 1e-6 is a very small constant used to avoid division by zero. The role of this formula is to aggregate global information through max pooling and balance the importance of each class through class weights.

[0141] At the same time, global average pooling (GAP) is added. The basic idea of GAP: in the GAP method, all pixel features in the image are averaged to obtain a mean representation of the entire image. This method retains global information and has a low computational cost. The output of its GAP formula can be expressed as:

[0142]

[0143] where HW is the number of pixels, is the classification probability value of each pixel, and by calculating the average value of all pixels, it provides a simple and efficient way to aggregate global information.

[0144] Finally, to balance the prediction effects of GCMP and GAP outputs, the combination of GCMP and GAP is introduced, and a weighting coefficient α is introduced to perform weighted averaging on the prediction results of GCMP and GAP outputs (aggregate the features in the feature map). The formula is as follows:

[0145]

[0146] Among them, α is an adjustable coefficient, usually set to 0.5, which means that the contributions of GCMP and GAP to the final output are of equal weight. By adjusting α, the influence of GCMP and GAP can be adjusted according to the actual task requirements.

[0147] This application adopts an aggregation method that combines global classification maximum pooling (GCMP) and global average pooling to strengthen the effective transmission of weak supervision signals and improve the overall segmentation performance.

[0148] The siamese network structure of this application also optimizes the segmentation model through pseudo-label loss, graph attention network weak supervision loss, and segmentation model weak supervision loss. Specifically:

[0149] Sum the pseudo-label loss, pseudo-label loss function, and weak supervision loss function to obtain the total loss;

[0150] Use the total loss to optimize the segmentation model.

[0151] Specifically, first introduce an order loss function in the siamese network structure to compare the differences between each pair of samples (square patch images); the purpose of the order loss function is to measure the similarity between two inputs and optimize the siamese network parameters through this similarity, so that the outputs of the two networks can better match. Suppose there is a pair of square image patches X A and X B of pathological images, and their pseudo-labels are Y A and Y B respectively. By calculating the distance between this pair of square patch images to judge their similarity, if they belong to the same category, it is expected that the feature representations of this pair of square patch images are similar; if they belong to different categories, there should be a large difference in the feature representations of this pair of square patch images. The formula of the order loss function can be defined as:

[0152]

[0153] where Y ic represents the true label of the c-th category in the i-th sample (usually 0 or 1, indicating whether it belongs to this category), represents the predicted probability of the c-th category in the i-th sample by the model; N is the number of samples, and C is the number of categories.

[0154] (1) The pseudo-label loss function defines that the effectiveness of the pseudo-label losses of the two twins at different backbone steps of the twin network is asymmetric. The pseudo-label loss L pseudo The calculation formula is as follows:

[0155]

[0156] where and represent the pseudo-label predictions of the two different branches A and B of the twin network; and are the true labels of branches A and B of the twin network respectively; λ step is a coefficient used to control the weighted average in each training step. λ step = step mod2, so that different weighting methods will be alternately used every two training steps; θ A and θ B are the parameters of backbones A and B of the twin network respectively.

[0157] (2) The weakly supervised loss of the graph attention network further improves the quality of the pseudo-labels and enhances the segmentation model's perception ability of local and global features by combining the graph attention network with the pseudo-label generation process. In this step, the segmentation model generates pseudo-labels through the graph attention network and uses the graph attention mechanism to enhance the learning of different regions and categories; the formula for this process is as follows:

[0158]

[0159] where represents the output after being processed by the graph attention network (GAT), usually through global average pooling to obtain the global feature representation of each category; and are the category probabilities generated by the GAT network respectively; finally, through the standard cross-entropy loss function of the two branches in the twin network structure is used to calculate the difference between the pseudo-labels generated by the graph attention network module and the true labels, and the weakly supervised loss L GAT of the graph attention network can be obtained. The formula is:

[0160]

[0161] where L is the standard cross-entropy loss function; Y is the true label; is the pseudo-label output by the graph attention module, c represents the segmentation category; θ is the parameter of the graph attention module.

[0162] (3) Weak supervision loss of the segmentation model: In actual semantic segmentation tasks, it is usually necessary to improve the segmentation accuracy of the model for images through multi-level feature fusion. Feature aggregation is performed by combining global classification max pooling (GCMP) and global average pooling (GAP); in this way, both global information can be retained and the focus on important regions can be enhanced, thereby effectively improving the performance of the weak supervision segmentation task; then, using the standard cross-entropy loss function, the pseudo-label obtained from the aggregated feature representation is compared with the true label, and for each network branch, the difference between the pseudo-label obtained from the aggregated feature representation and the true label is calculated; for the differences obtained from different branches, they are all used as the performance evaluation indicators for this branch in the segmentation task; finally, the cross-entropy losses calculated by the two branches are added together to obtain the weak supervision loss of the segmentation model, and the weak supervision loss L of the segmentation model is calculated. DeeplabV3+ The specific formula is as follows:

[0163]

[0164] where f GCMP is global classification max pooling, which is used to focus on the most important regions; f GAP is global average pooling, which is used to retain global information. Y is the true label; The result after aggregating all pixels i, A and B respectively represent the outputs of the two backbone branches in the Siamese network. θ A is the parameter of the Siamese network backbone A, and θ B is the parameter of the Siamese network backbone B.

[0165] Finally, the pseudo-label loss, the pseudo-label loss function, and the weak supervision loss function of the segmentation model are summed to obtain the total loss, and the formula is:

[0166] L total = L GAT + L DeeplabV3+ + L pseudo .

[0167] The cross-entropy loss function in this application is the loss function in the classification task, which is used to measure the difference between the probability distribution predicted by the model and the actual label distribution. By minimizing the cross-entropy loss, the segmentation model can gradually improve the accuracy of classification and segmentation. After optimizing the pseudo-label in this application, the cross-entropy loss function is used to calculate the final segmentation error and perform backpropagation optimization.

[0168] In summary, for a pseudo-label generation method based on pathological image segmentation in this application, by introducing a graph attention network to establish the context relationship between pixels, the recognition effect for non-dominant categories is significantly improved; an aggregation method combining global classification maximum pooling (GCMP) and global average pooling is adopted to strengthen the effective transmission of weak supervision signals and improve the overall segmentation performance; through multi-scale feature fusion and pseudo-label optimization strategies, high-precision pseudo-labels are generated to improve the effect of the segmentation model, thereby reducing the manual annotation workload; thus, the accuracy and stability of pathological image segmentation are improved.

[0169] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously.

[0170] Based on the same idea as the pseudo-label generation method based on pathological image segmentation in the above embodiment, this application also provides a pseudo-label generation system based on pathological image segmentation, and this system can be used to execute the above-mentioned pseudo-label generation method based on pathological image segmentation. For the sake of convenience of description, in the structural schematic diagram of the pseudo-label generation system embodiment based on pathological image segmentation, only the parts related to the embodiment of this application are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the system, and it can include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.

[0171] Please refer to Figure 3 , in another embodiment of this application, a pseudo-label generation system based on pathological image segmentation is provided. This system includes an annotation module 101, a feature extraction module 102, an updated feature module 103, a pseudo-label generation module 104, a refinement module 105, and an optimization module 106;

[0172] The annotation module 101 is used to cut a digital pathological whole-slide image into several square small-block images of a preset size, and perform multi-label annotation on the several square small-block images to obtain the annotated square small-block images;

[0173] The feature extraction module 102 is used to input the annotated square small-block images into a pre-trained segmentation model, and the feature extraction module in the pre-trained segmentation model extracts the features of the annotated square small-block images to obtain a feature map;

[0174] The updated feature module 103 is used to input the feature map into the graph attention network module in the pre-trained segmentation model, and the graph attention network module uses the attention mechanism for information propagation to update the feature representation of the nodes;

[0175] The pseudo-label generation module 104 is configured to generate a probability distribution of each node belonging to different categories based on the feature representation of the updated node and the spatial information of the node, so as to obtain pseudo-labels.

[0176] The refinement module 105 is configured to refine the pseudo-labels by the refinement module in the pre-trained segmentation model to obtain refined pseudo-labels.

[0177] The optimization module 106 is configured to optimize the refined pseudo-labels by the optimization module in the pre-trained segmentation model to obtain the final pseudo-labels.

[0178] It should be noted that the pseudo-label generation system based on pathological image segmentation in this application corresponds one-to-one with the pseudo-label generation method based on pathological image segmentation in this application. The technical features and beneficial effects described in the embodiments of the above-mentioned pseudo-label generation method based on pathological image segmentation are applicable to the embodiments of the pseudo-label generation system based on pathological image segmentation. For specific content, reference can be made to the description in the method embodiments of this application, which will not be elaborated here. This is hereby declared.

[0179] In addition, in the implementation manner of the pseudo-label generation system based on pathological image segmentation in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the pseudo-label generation system based on pathological image segmentation is divided into different program modules to complete all or part of the functions described above.

[0180] In another embodiment, an electronic device for implementing a pseudo-label generation method based on pathological image segmentation is provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the pseudo-label generation method based on pathological image segmentation in any embodiment of this application is implemented.

[0181] Exemplarily, in this embodiment, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory and executed by the processor to complete this application. The one or more module elements can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the device.

[0182] The device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include, but is not limited to, a processor and a memory.

[0183] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the device, and connects various parts of the entire device using various interfaces and lines.

[0184] The memory may be used to store the computer program and / or modules. The processor realizes various functions of the device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0185] Correspondingly, the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the pseudo-label generation method based on pathological image segmentation described in any one of the above embodiments.

[0186] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0187] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0188] The above embodiments are preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present application shall be equivalent replacement methods and are all included in the protection scope of the present application.

Claims

1. A pseudo-label generation method based on pathological image segmentation, characterized in that: The steps include: The digital pathology full-slice image is divided into a number of square small-block images of preset sizes, and the square small-block images are annotated with multiple labels to obtain annotated square small-block images; Inputting the labeled small square image into a pre-trained segmentation model, wherein a feature extraction module in the pre-trained segmentation model extracts features of the labeled small square image to obtain a feature map; Inputting the feature graph into the graph attention network module in the pre-trained segmentation model, wherein the graph attention network module uses the attention mechanism to propagate information and update the feature representation of the node; According to the feature representation of the updated node and the spatial information of the node, the probability distribution of each node belonging to different categories is generated to obtain a pseudo label; The refinement module in the pre-trained segmentation model refines the pseudo-label to obtain a refined pseudo-label; The optimization module in the pre-trained segmentation model optimizes the refined pseudo-labels to obtain final pseudo-labels.

2. The pseudo-label generation method based on pathological image segmentation according to claim 1, characterized in that: The method generates a probability distribution of each node belonging to different categories based on the feature representation of the updated node and the spatial information of the node to obtain a pseudo label, including: The graph attention network module divides the feature graph into multiple regions, wherein each region serves as a node; Connecting the nodes through a graph structure to obtain a complete connection graph; Based on the complete connection graph, information is propagated through the attention mechanism to update the feature representation of the nodes layer by layer; Obtaining the coordinates of each node in the digital pathology full-slice image, and normalizing the coordinates; The normalized coordinates are encoded using sine and cosine functions to generate the spatial information of each node; Based on the feature representation of the updated node and the spatial information of the node, a probability distribution of each node belonging to different categories is obtained; The probability distribution is normalized to obtain a pseudo label.

3. The pseudo-label generation method based on pathological image segmentation according to claim 1, characterized in that: The pseudo-label is refined, including: Use a fixed convolution kernel to extract the RGB values ​​and position encoding information of neighboring pixels from the feature map; Calculating the similarity between the integrated pixels according to the RGB values ​​and position coding information of the adjacent pixels; The feature representations in the pseudo labels are weightedly summed using the similarity between the pixels to obtain a refined pseudo label.

4. The pseudo-label generation method based on pathological image segmentation according to claim 3 is characterized in that: Optimizing the pseudo-label includes: Obtaining the refined pseudo-label and a preset confidence threshold; Traversing each pixel position in the refined pseudo-label in turn; Calculate the maximum probability value of each pixel position in all categories; Comparing the maximum probability value with the preset confidence threshold; If the maximum probability value is greater than the preset confidence threshold, the category of the maximum probability value is used as the output of the pixel; If the maximum probability value is less than the preset confidence threshold, the pixels of the category with the maximum probability value are ignored; Based on the output of the pixel and ignoring the pixel of the category with the maximum probability value, an optimized pseudo label is obtained.

5. The pseudo-label generation method based on pathological image segmentation according to claim 1, characterized in that: The segmentation model is trained using a twin network structure to obtain a trained segmentation model; The twin network structure includes two branches that share parameters, and the output results of the two branches are used to guide each other; The two branches both include the feature extraction module, the graph attention network module, the refinement module and the optimization module.

6. The pseudo-label generation method based on pathological image segmentation according to claim 5, characterized in that: The twin network structure optimizes the segmentation model through pseudo-label loss, graph attention network weak supervision loss and segmentation model weak supervision loss. Specifically: The pseudo label loss, the pseudo label loss function and the segmentation model weak supervision loss function are summed to obtain a total loss; The segmentation model is optimized using the total loss.

7. The pseudo-label generation method based on pathological image segmentation according to claim 5, characterized in that: It also includes the use of global classification maximum pooling and global average pooling to aggregate features in the feature map, including: Perform global classification and maximum pooling on the feature map to obtain the feature representation of the salient area; Perform global average pooling on the feature map to obtain the global mean feature representation; The feature representation of the salient region and the global mean feature representation are aggregated to obtain an aggregated feature.

8. A pseudo-label generation system based on pathological image segmentation, characterized in that: A pseudo-label generation method based on pathological image segmentation applied to any one of claims 1 to 7, comprising a labeling module, a feature extraction module, a feature update module, a pseudo-label generation module, a refinement module, and an optimization module; The labeling module is used to divide the digital pathology full-slice image into a plurality of square small-block images of preset sizes, and to perform multi-label labeling on the plurality of square small-block images to obtain labeled square small-block images; The feature extraction module is used to input the labeled small square image into a pre-trained segmentation model, and the feature extraction module in the pre-trained segmentation model extracts the features of the labeled small square image to obtain a feature map; The feature update module is used to input the feature map into the graph attention network module in the pre-trained segmentation model, and the graph attention network module uses the attention mechanism to propagate information and update the feature representation of the node; The pseudo label generation module is used to generate a probability distribution of each node belonging to different categories according to the feature representation of the updated node and the spatial information of the node to obtain a pseudo label; The refinement module is used for the refinement module in the pre-trained segmentation model to refine the pseudo-label to obtain a refined pseudo-label; The optimization module is used in the pre-trained segmentation model to optimize the refined pseudo-labels to obtain final pseudo-labels.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the pseudo-label generation method based on pathological image segmentation as described in any one of claims 1-7.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the pseudo-label generation method based on pathological image segmentation described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Semi-supervised medical image segmentation method and device of collaborative training network based on integrated attention correction, and medium

    CN121259317A