Weak supervision tissue pathology image segmentation method based on fuzzy enhancement Transform

By using the method of fuzzy enhancement Transformer and fuzzy logic to extract fuzzy features in histopathological image segmentation, the problem of dependence on high-quality labeling data and blurred tissue category boundaries in the prior art is solved, and high-precision and robust histopathological image segmentation is achieved.

CN119941760AActive Publication Date: 2025-05-06NANTONG UNIV

Patent Information

Application Number
CN202510086101.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Existing histopathological image segmentation methods rely on high-quality data annotations and are difficult to effectively deal with the problem of blurred boundaries of tissue categories, resulting in insufficient segmentation accuracy and robustness.

Method used

Weakly supervised tissue pathological image segmentation method based on fuzzy enhancement Transformer is used to perform pixel-level segmentation through image-level labels, fuzzy features are extracted in combination with fuzzy logic, and pooling operations are guided through fuzzy enhancement attention maps to generate the final segmentation result.

Benefits of technology

It effectively reduces the dependence on high-quality labeled data, improves the accuracy and robustness of segmentation, can accurately identify tumor areas and other key tissue types, and supports the construction of a trusted and low-cost medical-assisted decision-making system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941760A_ABST
    Figure CN119941760A_ABST
Patent Text Reader

Abstract

The invention provides a weak supervision tissue pathology image segmentation method based on fuzzy enhancement Transform, and belongs to the technical field of medical image intelligent processing. The technical problem that medical image data tags are difficult to put into practical application due to high acquisition cost is solved, and the technical scheme is as follows: firstly, reading images from a histopathological data set, and performing data preprocessing; inputting the image into a Transform, and obtaining an attention matrix and a final layer output; then, a fuzzy augmented attention module is used for generating a fuzzy augmented attention map, and the fuzzy augmented attention map is used for guiding pooling operation to train a Transform model; and finally, extracting a fuzzy enhanced attention map of the model, and generating a pathological image segmentation result through post-processing. The method has the beneficial effects that accurate tissue pathology image segmentation is realized by using the image-level labels which are easy to obtain, and practical application of deep learning in the field of medical diagnosis is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image intelligent processing, and in particular to a weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer. Background Art

[0002] The segmentation of tissue pathology images plays an important role in medical image analysis, especially in the diagnosis of tumors, inflammation and other pathological changes. By accurately segmenting different tissue regions, doctors can better identify the type, range and invasiveness of tumors, providing important basis for clinical treatment and prognosis evaluation. Especially in the study of tumor microenvironment, segmenting different tissue types such as tumor epithelium, tumor infiltrating lymphocytes and tumor-related stroma is of great significance for revealing the biological behavior of tumors, predicting treatment response and evaluating patient prognosis. With the continuous development of deep learning technology, image segmentation methods based on convolutional neural networks (CNN) and Transformer have become mainstream. Deep learning models can automatically learn complex feature representations from data, greatly improving the efficiency and accuracy of medical image segmentation, thereby assisting doctors in clinical decision-making.

[0003] For example, Van et al. proposed a multi-branch segmentation framework HookNet based on convolutional neural networks in "HookNet: Multi-resolution convolutional neural networks for semantic segmentation in histopathology whole-slide images". The framework generates fine-grained segmentation maps of histopathology images by combining contextual information and high-resolution details. However, the collection of medical image annotation data is usually expensive and time-consuming, especially for high-resolution histopathology images such as whole-slide images (WSIs). The resolution of gigapixels makes manual annotation extremely difficult and labor-intensive. In addition, the high heterogeneity of tumor regions makes their morphology and tissue structure diverse, which brings greater challenges to the image segmentation task. There is an urgent need for a new segmentation method that can use simple and easy-to-obtain image-level labels to perform pixel-level histopathology image segmentation, thereby reducing the dependence on a large amount of high-quality annotated data. At the same time, by combining theoretical methods such as fuzzy logic, it can effectively deal with the problem of blurred boundaries of tissue categories in histopathology images and improve the accuracy and robustness of weakly supervised segmentation, which is of certain significance for building a reliable and low-cost medical decision-making support system. Summary of the invention

[0004] The purpose of the present invention is to provide a weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer, which solves the problems of existing tissue pathology image segmentation methods such as dependence on high-quality data annotation and blurred tissue category boundaries, and effectively improves the accuracy of weakly supervised segmentation.

[0005] In order to achieve the above-mentioned invention object, the present invention is implemented by the following measures: a weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer comprises the following steps:

[0006] S1: Read images from the histopathology dataset Perform data preprocessing and then cut the image into 196 non-overlapping image blocks of size 16×16;

[0007] S2: Constructing preprocessed histopathology images into input tokens Send it to the Transformer encoder layer, get the attention matrix of the Transformer encoder layer and extract the affinity matrix Final layer image block token output and class token output Where D is the dimension of the token;

[0008] S3: Output the image block token The fuzzy enhanced attention module constructed by the reshaped input extracts the fuzzy features in the image block token output based on fuzzy rules. The fuzzy features are concatenated with the image block token output and fused through convolution operation to generate a fuzzy enhanced attention map.

[0009] S4: The generated fuzzy enhanced attention map is processed as a weight-guided pooling operation to generate image block token prediction scores, and the category token output is reduced in dimension to generate category token prediction scores. The two scores are respectively compared with the true image labels to calculate the multi-label soft-margin loss to train the Transformer model;

[0010] S5: Extract the fuzzy enhanced attention map from the trained model, use the ReLU activation function and maximum normalization operation, then use the affinity matrix for post-processing, and generate the segmentation result after the argmax operation, where argmax is the index operation for obtaining the maximum value position.

[0011] Furthermore, the specific steps of step S2 are as follows:

[0012] Step S2.1: Input tensor T input The input Transformer encoder layer undergoes L rounds of iterations, during which the query matrix is ​​generated through projection transformation Key Matrix Sum Matrix Extracting the attention matrix And the final layer token output

[0013]

[0014] T output =AV (2)

[0015] where d k is the dimension of the key matrix, softmax(·) is the activation function, and the attention matrix of L layers is aggregated to obtain the aggregated attention matrix

[0016] Step S2.2: Aggregate attention matrix A all The average attention matrix is ​​obtained by averaging over the layer dimensions

[0017]

[0018] in is the i-th aggregate attention matrix;

[0019] Step S2.3: From the average attention matrix A mean Extracting affinity matrix Output T from the final layer token output Extract the final layer image block token output and class token output

[0020] Furthermore, the specific steps of step S3 are as follows:

[0021] Step S3.1: Output the extracted image block tokens Reshape to It is represented by D pieces of size The feature graph is then used to characterize the correlation between the elements in the feature graph and the fuzzy concept, effectively capturing the uncertainty and fuzziness in the features. K Gaussian membership functions are assigned to each feature graph. The calculation formula of the Gaussian membership function is as follows:

[0022]

[0023] where c d,k and σ d,k The table is the center and width of the kth Gaussian membership function on the dth feature map, d = 1, ..., D, k = 1, ..., K, exp(·) is the exponential function, T d represents the d-th feature map, μ d,k (Td ) represents the fuzzy membership matrix obtained by assigning the k-th Gaussian membership function on the d-th feature map. The final fuzzified result is calculated using formula (4):

[0024] Step S3.2: Sampling the fuzzy Gaussian kernel is used to synthesize fuzzy features with semantic information. The sampling formula is as follows:

[0025]

[0026] Where T d ' ,k Represents the feature maps obtained by sampling, and the pixel values ​​on these feature maps obey the mean value c d,k , the standard deviation is Gaussian distribution of

[0027] Step S3.3: Different membership values ​​represent the membership of different positions in the image to a specific fuzzy set. Three sets of fuzzy rules are established to extract fuzzy weights, which are multiplied with the sampled feature maps to extract fuzzy features. The first set of fuzzy logic is used to extract highly matched semantic fuzzy features, and the maximum membership value of each feature point on the feature map is selected as the output weight. To obtain highly matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0028]

[0029] Where τ is the scaling temperature parameter that controls the smoothness, the second set of fuzzy logic is used to extract the semantic fuzzy features with low matching, and the minimum membership value of each feature point on the feature map is selected as the output weight To obtain low matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0030]

[0031] The third group of fuzzy logic is used to extract common semantic fuzzy features, and the medium membership value of each feature point on the feature map is selected as the output weight To obtain general semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0032]

[0033] in Fuzzy result The result of taking the mean in the second dimension;

[0034] Step S3.4: The three extracted fuzzy features are concatenated with the original features and fused using a convolution operation:

[0035]

[0036] Concat(·) is a concatenation operation, Conv(·) is a 3×3 convolution operation, and the concatenated feature channels are adjusted to C, where C is the number of categories in the dataset. Finally, a fuzzy enhanced attention map is generated.

[0037] Furthermore, the specific steps of step S4 are as follows:

[0038] Step S4.1: The fuzzy enhanced attention map contains position information, which is used to guide the pooling operation. First, softmax normalization is performed on the channel dimension to generate the pooling weight α c,x,y :

[0039]

[0040] in represents the feature point on the cth channel in the fuzzy enhanced attention map, c = 1,...,C,

[0041] Step S4.2: Take α c,x,y is the weight pair A fuzzy Perform weighted pooling to enhance the weight of the correct category at the corresponding position, thereby guiding the attention map to expand within a reasonable range and generating image block token prediction scores

[0042]

[0043] Step S4.3: Output the category token obtained in step S2.3 Generate class token prediction scores through dimensionality reduction through linear layers

[0044] y class =Linear(T class ) (15)

[0045] Where Linear(·) represents a linear layer;

[0046] Step S4.4: Patch Token Prediction Scores Prediction score with class token Respectively with the true label

[0047] Compute the multi-label soft margin loss:

[0048]

[0049] Where σ(·) represents the sigmoid activation function, and finally L patch and L class The sum is used as the final loss function, and the loss function is minimized through back propagation to optimize the model parameters.

[0050] Furthermore, the specific steps of step S5 are as follows:

[0051] Step S5.1: After the model training is completed, the image is put into the model again and the blur enhancement attention map is extracted And the fuzzy enhanced attention map corresponding to each category is activated by the ReLU function and the maximum normalization operation:

[0052]

[0053] in Fuzzy enhanced attention map corresponding to category c, max(·) represents the maximum value operation;

[0054] Step S5.2: Use the affinity matrix extracted in step S2.3 Perform post-processing of the fuzzy enhanced attention map:

[0055]

[0056] in Represents the matrix multiplication operation, R N×C (·)and Represents reshaping the matrix into N×C and shape;

[0057] Step S5.3: Finally, A final Upsample to the input image size and apply the argmax operation to generate the final segmentation result pred:

[0058]

[0059] Where Upsample(·) represents the upsampling operation, where It means taking the maximum value index in the category dimension and visually displaying the segmentation results.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] 1. The present invention proposes a weakly supervised learning method based on image-level labels, which can effectively alleviate the high cost and high labor intensity problems in the process of medical image annotation. Unlike traditional methods that require a large amount of pixel-level annotation data, the present invention can achieve high-precision pixel-level segmentation through image-level labels, greatly reducing the demand for data annotation, and reducing annotation costs and time consumption.

[0062] 2. This invention proposes a fuzzy enhanced class-aware attention map generation method, which sets multiple Gaussian membership functions and fuzzy rules on the output feature map to extract fuzzy high-matching semantic features, fuzzy low-matching semantic features and fuzzy general semantic features. The fuzzy enhanced attention map generated by integrating these fuzzy features can provide accurate object location information. At the same time, the fuzzy enhanced attention map is used to guide the pooling operation, which can reasonably expand the attention area.

[0063] 3. The present invention combines the fuzzy enhanced attention map in the trained model with the self-attention map in the sub-attention module of the Transformer model to generate the final attention map result as the tissue pathology image segmentation result. The segmentation result can accurately identify key tissue types such as tumor areas, tumor-infiltrating lymphocytes, and tumor-related stroma, providing strong support for the segmentation of tissue pathology images. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention but do not constitute a limitation of the present invention.

[0065] Figure 1 It is the overall framework diagram of the weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer of the present invention;

[0066] Figure 2 It is a flow chart of the weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer of the present invention;

[0067] Figure 3 This is a flow chart of the fuzzy enhanced attention module of the present invention;

[0068] Figure 4 This is an example of the weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer of the present invention. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0070] Example 1

[0071] See also Figures 1 to 4 The technical solution provided in this embodiment is a weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer. Taking a tissue pathology image dataset as an example, from loading data to obtaining segmentation results, the following steps are included:

[0072] S1: Read images from the histopathology dataset Perform data preprocessing and then cut the image into 196 non-overlapping image blocks of size 16×16;

[0073]

[0074] S2: Constructing preprocessed histopathology images into input tokens Send it to the Transformer encoder layer, get the attention matrix of the Transformer encoder layer and extract the affinity matrix Final layer image block token output and class token output Where D is the dimension of the token;

[0075] S3: Output the image block token The fuzzy enhanced attention module constructed by the reshaped input extracts the fuzzy features in the image block token output based on fuzzy rules. The fuzzy features are concatenated with the image block token output and fused through convolution operation to generate a fuzzy enhanced attention map.

[0076] S4: The generated fuzzy enhanced attention map is processed as a weight-guided pooling operation to generate image block token prediction scores, and the category token output is reduced in dimension to generate category token prediction scores. The two scores are respectively compared with the true image labels to calculate the multi-label soft-margin loss to train the Transformer model;

[0077] S5: Extract the fuzzy enhanced attention map from the trained model, use the ReLU activation function and maximum normalization operation, then use the affinity matrix for post-processing, and generate the segmentation result after the argmax operation, where argmax is the index operation for obtaining the maximum value position.

[0078] Furthermore, the specific steps of step S2 are as follows:

[0079] Step S2.1: Input tensor T input The input Transformer encoder layer undergoes L rounds of iterations, during which the query matrix is ​​generated through projection transformation Key Matrix Sum Matrix Extracting the attention matrix And the final layer token output

[0080]

[0081] T output =AV (3)

[0082] where d k is the dimension of the key matrix, softmax(·) is the activation function, and the attention matrix of L layers is aggregated to obtain the aggregated attention matrix

[0083] Step S2.2: Aggregate attention matrix A all The average attention matrix is ​​obtained by averaging over the layer dimensions

[0084]

[0085] in is the i-th aggregate attention matrix;

[0086] Step S2.3: From the average attention matrix A mean Extracting affinity matrix Output T from the final layer token output Extract the final layer image block token output and class token output A affinity 、T patch and T class As shown in formula (5), formula (6) and formula (7):

[0087]

[0088] T class =(0.36540.0239...-1.3109) 1×384 (7)

[0089] Furthermore, the specific steps of step S3 are as follows:

[0090] Step S3.1: Output the extracted image block tokens Reshape to It is represented by D pieces of size The feature graph is then used to characterize the correlation between the elements in the feature graph and the fuzzy concept, effectively capturing the uncertainty and fuzziness in the features. K Gaussian membership functions are assigned to each feature graph. The calculation formula of the Gaussian membership function is as follows:

[0091]

[0092] where c d,k and σ d,k The table is the center and width of the kth Gaussian membership function on the dth feature map, d = 1, ..., D, k = 1, ..., K, exp(·) is the exponential function, T d represents the d-th feature map, μ d,k (T d ) represents the fuzzy membership matrix obtained by assigning the k-th Gaussian membership function on the d-th feature map. The final fuzzified result is calculated using formula (8): The membership matrix μ 0,0 As shown in formula (9):

[0093]

[0094] Step S3.2: Sampling the fuzzy Gaussian kernel is used to synthesize fuzzy features with semantic information. The sampling formula is as follows:

[0095]

[0096] Where T d ' ,k Represents the feature maps obtained by sampling, and the pixel values ​​on these feature maps obey the mean value c d,k , the standard deviation is Gaussian distribution, where the sampling feature map T0' ,0 As shown in formula (11):

[0097]

[0098] Step S3.3: Different membership values ​​represent the membership of different positions in the image to a specific fuzzy set. Three sets of fuzzy rules are established to extract fuzzy weights, which are multiplied with the sampled feature maps to extract fuzzy features. The first set of fuzzy logic is used to extract highly matched semantic fuzzy features, and the maximum membership value of each feature point on the feature map is selected as the output weight. To obtain highly matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0099]

[0100]

[0101] Where τ is the scaling temperature parameter that controls the smoothness, and the semantic fuzzy features with high matching As shown in formula (14):

[0102]

[0103] The second set of fuzzy logic is used to extract low-matching semantic fuzzy features, and the minimum membership value of each feature point on the feature map is selected as the output weight To obtain low matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0104]

[0105] Semantically fuzzy features with low matching As shown in formula (17):

[0106]

[0107] The third group of fuzzy logic is used to extract common semantic fuzzy features, and the medium membership value of each feature point on the feature map is selected as the output weight To obtain general semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0108]

[0109] in Fuzzy result The result of taking the mean value in the second dimension is a semantically ambiguous feature with low matching. As shown in formula (20):

[0110]

[0111] Step S3.4: The three extracted fuzzy features are concatenated with the original features and fused using a convolution operation:

[0112]

[0113] Concat(·) is a concatenation operation, Conv(·) is a 3×3 convolution operation, and the concatenated feature channels are adjusted to C, where C is the number of categories in the dataset. Finally, a fuzzy enhanced attention map is generated. As shown in formula (22):

[0114]

[0115] Furthermore, the specific steps of step S4 are as follows:

[0116] Step S4.1: The fuzzy enhanced attention map contains position information, which is used to guide the pooling operation. First, softmax normalization is performed on the channel dimension to generate the pooling weight α c,x,y :

[0117]

[0118] in represents the feature point on the cth channel in the fuzzy enhanced attention map, c = 1,...,C, α c,x,y As shown in formula (24):

[0119]

[0120] Step S4.2: Take α c,x,y is the weight pair A fuzzy Perform weighted pooling to enhance the weight of the correct category at the corresponding position, thereby guiding the attention map to expand within a reasonable range and generating image block token prediction scores

[0121]

[0122] y patch As shown in formula (26):

[0123] y patch =(7.7413-5.4955-8.54437.2036) 1×4 (26)

[0124] Step S4.3: Output the category token obtained in step S2.3 Generate class token prediction scores through dimensionality reduction through linear layers

[0125] y class =Linear(T class ) (27)

[0126] Where Linear(·) represents the linear layer, y class As shown in formula (28):

[0127] y class =(8.2471 -4.5638 -8.5215 6.5667) 1×4 (28)

[0128] Step S4.4: Patch Token Prediction Scores Prediction score with class token Respectively with the true label Compute the multi-label soft margin loss:

[0129]

[0130] Where σ(·) represents the sigmoid activation function, and finally L patch and L class The sum is used as the final loss function, and the loss function is minimized through back propagation to optimize the model parameters.

[0131] Furthermore, the specific steps of step S5 are as follows:

[0132] Step S5.1: After the model training is completed, the image is put into the model again and the blur enhancement attention map is extracted And the fuzzy enhanced attention map corresponding to each category is activated by the ReLU function and the maximum normalization operation:

[0133]

[0134] in Fuzzy enhanced attention map corresponding to category c, max(·) represents the maximum value operation, As shown in formula (32):

[0135]

[0136] Step S5.2: Use the affinity matrix extracted in step S2.3 Perform post-processing of the fuzzy enhanced attention map:

[0137]

[0138] in Represents the matrix multiplication operation, R N×C (·)and Represents reshaping the matrix into N×C and The shape of A final As shown in formula (34):

[0139]

[0140] Step S5.3: Finally, A final Upsample to the input image size and apply the argmax operation to generate the final segmentation result pred:

[0141]

[0142] Where Upsample(·) represents the upsampling operation, It means taking the maximum value index in the category dimension and visualizing the segmentation result. The segmentation result pred is as shown in formula (36):

[0143]

[0144] All images in the dataset are put into the model to obtain the segmentation results, and the real segmentation labels are used to evaluate the performance of the segmentation results. The accuracy is 84.51% and the average intersection-over-union ratio is 75.65%, which is significantly better than the traditional weakly supervised segmentation method.

[0145]

[0146] Example 2

[0147] See also Figures 1 to 4 The technical solution provided in this embodiment is a weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer. Taking a tissue pathology image dataset as an example, from loading data to obtaining segmentation results, the following steps are included:

[0148] S1: Read images from the histopathology dataset Perform data preprocessing and then cut the image into 196 non-overlapping image blocks of size 16×16;

[0149]

[0150] S2: Constructing preprocessed histopathology images into input tokens Send it to the Transformer encoder layer, get the attention matrix of the Transformer encoder layer and extract the affinity matrix Final layer image block token output and class token output Where D is the dimension of the token;

[0151] S3: Output the image block token The fuzzy enhanced attention module constructed by the reshaped input extracts the fuzzy features in the image block token output based on fuzzy rules. The fuzzy features are concatenated with the image block token output and fused through convolution operation to generate a fuzzy enhanced attention map.

[0152] S4: The generated fuzzy enhanced attention map is processed as a weight-guided pooling operation to generate image block token prediction scores, and the category token output is reduced in dimension to generate category token prediction scores. The two scores are respectively compared with the true image labels to calculate the multi-label soft-margin loss to train the Transformer model;

[0153] S5: Extract the fuzzy enhanced attention map from the trained model, use the ReLU activation function and maximum normalization operation, then use the affinity matrix for post-processing, and generate the segmentation result after the argmax operation, where argmax is the index operation for obtaining the maximum value position.

[0154] Furthermore, the specific steps of step S2 are as follows:

[0155] Step S2.1: Input tensor T input The input Transformer encoder layer undergoes L rounds of iterations, during which the query matrix is ​​generated through projection transformation Key Matrix Sum Matrix Extracting the attention matrix And the final layer token output

[0156]

[0157] T output =AV (3)

[0158] where d k is the dimension of the key matrix, softmax(·) is the activation function, and the attention matrix of L layers is aggregated to obtain the aggregated attention matrix

[0159] Step S2.2: Aggregate attention matrix A all The average attention matrix is ​​obtained by averaging over the layer dimensions

[0160]

[0161] in is the i-th aggregate attention matrix;

[0162] Step S2.3: From the average attention matrix A mean Extracting affinity matrix Output T from the final layer token output Extract the final layer image block token output and class token output A affinity 、T patch and T class As shown in formula (5), formula (6) and formula (7):

[0163]

[0164] T class =(-10.1635-0.3858...18.8144) 1×384 (7)

[0165] Furthermore, the specific steps of step S3 are as follows:

[0166] Step S3.1: Output the extracted image block tokens Reshape to It is represented by D pieces of size The feature graph is then used to characterize the correlation between the elements in the feature graph and the fuzzy concept, effectively capturing the uncertainty and fuzziness in the features. K Gaussian membership functions are assigned to each feature graph. The calculation formula of the Gaussian membership function is as follows:

[0167]

[0168] where c d,k and σ d,k The table is the center and width of the kth Gaussian membership function on the dth feature map, d = 1, ..., D, k = 1, ..., K, exp(·) is the exponential function, T d represents the d-th feature map, μ d,k (T d ) represents the fuzzy membership matrix obtained by assigning the k-th Gaussian membership function on the d-th feature map. The final fuzzified result is calculated using formula (8): The membership matrix μ 0,0 As shown in formula (9):

[0169]

[0170] Step S3.2: Sampling the fuzzy Gaussian kernel is used to synthesize fuzzy features with semantic information. The sampling formula is as follows:

[0171]

[0172] Where T d ' ,k Represents the feature maps obtained by sampling, and the pixel values ​​on these feature maps obey the mean value c d,k , the standard deviation is Gaussian distribution, where the sampling feature map T0' ,0 As shown in formula (11):

[0173]

[0174] Step S3.3: Different membership values ​​represent the membership of different positions in the image to a specific fuzzy set. Three sets of fuzzy rules are established to extract fuzzy weights, which are multiplied with the sampled feature maps to extract fuzzy features. The first set of fuzzy logic is used to extract highly matched semantic fuzzy features, and the maximum membership value of each feature point on the feature map is selected as the output weight. To obtain highly matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0175]

[0176] Where τ is the scaling temperature parameter that controls the smoothness, and the semantic fuzzy features with high matching As shown in formula (14):

[0177]

[0178] The second set of fuzzy logic is used to extract low-matching semantic fuzzy features, and the minimum membership value of each feature point on the feature map is selected as the output weight To obtain low matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0179]

[0180] Semantically fuzzy features with low matching As shown in formula (17):

[0181]

[0182] The third group of fuzzy logic is used to extract common semantic fuzzy features, and the medium membership value of each feature point on the feature map is selected as the output weight To obtain general semantic fuzzy features Implement fuzzy logic using a scaled softmax operation:

[0183]

[0184] in Fuzzy result The result of taking the mean value in the second dimension is a semantically ambiguous feature with low matching. As shown in formula (20):

[0185]

[0186] Step S3.4: The three extracted fuzzy features are concatenated with the original features and fused using a convolution operation:

[0187]

[0188] Concat(·) is a concatenation operation, Conv(·) is a 3×3 convolution operation, and the concatenated feature channels are adjusted to C, where C is the number of categories in the dataset. Finally, a fuzzy enhanced attention map is generated. As shown in formula (22):

[0189]

[0190] Furthermore, the specific steps of step S4 are as follows:

[0191] Step S4.1: The fuzzy enhanced attention map contains position information, which is used to guide the pooling operation. First, softmax normalization is performed on the channel dimension to generate the pooling weight α c,x,y :

[0192]

[0193] in represents the feature point on the cth channel in the fuzzy enhanced attention map, c = 1,...,C, α c,x,y As shown in formula (24):

[0194]

[0195] Step S4.2: Take α c,x,y is the weight pair A fuzzy Perform weighted pooling to enhance the weight of the correct category at the corresponding position, thereby guiding the attention map to expand within a reasonable range and generating image block token prediction scores

[0196]

[0197] y patch As shown in formula (26):

[0198] y patch =(7.7469-5.9008-8.12457.8962) 1×4 (26)

[0199] Step S4.3: Output the category token obtained in step S2.3 Generate class token prediction scores through dimensionality reduction through linear layers

[0200] y class =Linear(T class ) (27)

[0201] Where Linear(·) represents the linear layer, y class As shown in formula (28):

[0202] y class =(6.8711-7.5792-8.94048.5667) 1×4 (28)

[0203] Step S4.4: Patch Token Prediction Scores Prediction score with class token Respectively with the true label Compute the multi-label soft margin loss:

[0204]

[0205] Where σ(·) represents the sigmoid activation function, and finally L patch and L class The sum is used as the final loss function, and the loss function is minimized through back propagation to optimize the model parameters.

[0206] Furthermore, the specific steps of step S5 are as follows:

[0207] Step S5.1: After the model training is completed, the image is put into the model again and the blur enhancement attention map is extracted And the fuzzy enhanced attention map corresponding to each category is activated by the ReLU function and the maximum normalization operation:

[0208]

[0209] in Fuzzy enhanced attention map corresponding to category c, max(·) represents the maximum value operation, As shown in formula (32):

[0210]

[0211] Step S5.2: Use the affinity matrix extracted in step S2.3 Perform post-processing of the fuzzy enhanced attention map:

[0212]

[0213] in Represents the matrix multiplication operation, R N×C (·)and Represents reshaping the matrix into N×C and The shape of A final As shown in formula (34):

[0214]

[0215] Step S5.3: Finally, A final Upsample to the input image size and apply the argmax operation to generate the final segmentation result pred:

[0216]

[0217] Where Upsample(·) represents the upsampling operation, It means taking the maximum value index in the category dimension and visualizing the segmentation result. The segmentation result pred is as shown in formula (36):

[0218]

[0219] All images in the dataset are put into the model to obtain the segmentation results, and the real segmentation labels are used to evaluate the performance of the segmentation results. The accuracy is 83.68% and the average intersection-over-union ratio is 67.68%, which is significantly better than the traditional weakly supervised segmentation method.

[0220]

[0221] The above description is only a preferred embodiment of this embodiment and is not intended to limit this embodiment. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this embodiment shall be included in the protection scope of this embodiment.

Claims

1. A weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer, characterized in that: The following steps are involved: S1: Read images from the histopathology dataset Where H and W are the height and width of the image, respectively. Data preprocessing is performed, and then the image is cut into N non-overlapping image blocks of size P×P, where N is the number of image blocks and P is the side length of the image block; S2: Constructing preprocessed histopathology images into input tokens Send it to the Transformer encoder layer, get the attention matrix of the Transformer encoder layer and extract the affinity matrix Final layer image block token output and class token output Where D is the dimension of the token; S3: Output the image block token The fuzzy enhanced attention module constructed by the reshaped input extracts the fuzzy features in the image block token output based on fuzzy rules. The fuzzy features are concatenated with the image block token output and fused through convolution operation to generate a fuzzy enhanced attention map. S4: The generated fuzzy enhanced attention map is processed as a weight-guided pooling operation to generate image block token prediction scores. The category token output is reduced in dimension to generate category token prediction scores. The two scores are respectively compared with the true image labels to calculate the multi-label soft edge loss to train the Transformer model. S5: Extract the fuzzy enhanced attention map from the trained model, use the ReLU activation function and maximum normalization operation, and then use the affinity matrix for post-processing. After upsampling, generate the segmentation result through the argmax operation, where argmax is the index operation for obtaining the maximum value position.

2. The weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer according to claim 1 is characterized in that: The specific steps of step S2 are as follows: Step S2.1: Input tensor T input The input Transformer encoder layer undergoes L rounds of iterations, during which the query matrix is ​​generated through projection transformation Key Matrix Sum Matrix Extracting the attention matrix And the final layer token output T output =OFF (2) where d k is the dimension of the key matrix, softmax(·) is the activation function, and the attention matrix of L layers is aggregated to obtain the aggregated attention matrix Step S2.2: Aggregate attention matrix A all The average attention matrix is ​​obtained by averaging over the layer dimensions Among them A a i ll is the i-th aggregate attention matrix; Step S2.3: From the average attention matrix A mean Extracting affinity matrix Output T from the final layer token output Extract the final layer image block token output and class token output 3. The weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer according to claim 1 is characterized in that: The specific steps of step S3 are as follows: Step S3.1: Output the extracted image block tokens Reshape to It is represented by D pieces of size The feature graph is then used to characterize the correlation between the elements in the feature graph and the fuzzy concept, effectively capturing the uncertainty and fuzziness in the features. K Gaussian membership functions are assigned to each feature graph. The calculation formula of the Gaussian membership function is as follows: where c d,k and σ d,k The table is the center and width of the kth Gaussian membership function on the dth feature map, d = 1, ..., D, k = 1, ..., K, exp(·) is the exponential function, T d represents the d-th feature map, μ d,k (T d ) represents the fuzzy membership matrix obtained by assigning the k-th Gaussian membership function on the d-th feature map. The final fuzzified result is calculated using formula (4): Step S3.2: Sampling the fuzzy Gaussian kernel is used to synthesize fuzzy features with semantic information. The sampling formula is as follows: Where T d ' ,k Represents the feature maps obtained by sampling, and the pixel values ​​on these feature maps obey the mean value c d,k , the standard deviation is Gaussian distribution of Step S3.3: Different membership values ​​represent the membership of different positions in the image to a specific fuzzy set. Three sets of fuzzy rules are established to extract fuzzy weights, which are multiplied with the sampled feature maps to extract fuzzy features. The first set of fuzzy logic is used to extract highly matched semantic fuzzy features, and the maximum membership value of each feature point on the feature map is selected as the output weight. To obtain highly matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation: where μ d,k represents the fuzzy membership matrix obtained by assigning the kth Gaussian membership function on the dth feature map, τ is the scaling temperature parameter that controls the smoothness, and the second set of fuzzy logic is used to extract low-matching semantic fuzzy features. The minimum membership value of each feature point on the feature map is selected as the output weight To obtain low matching semantic fuzzy features Implement fuzzy logic using a scaled softmax operation: The third group of fuzzy logic is used to extract common semantic fuzzy features, and the medium membership value of each feature point on the feature map is selected as the output weight To obtain general semantic fuzzy features Implement fuzzy logic using a scaled softmax operation: in Fuzzy result The result of taking the mean in the second dimension; Step S3.4: The three extracted fuzzy features are concatenated with the original features and fused using a convolution operation: Concat(·) is a concatenation operation, Conv(·) is a 3×3 convolution operation, and the concatenated feature channels are adjusted to C, where C is the number of categories in the dataset. Finally, a fuzzy enhanced attention map is generated.

4. The weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer according to claim 1 is characterized in that: The specific steps of step S4 are as follows: Step S4.1: The fuzzy enhanced attention map contains position information, which is used to guide the pooling operation. First, softmax normalization is performed on the channel dimension to generate the pooling weight α c,x,y : in represents the feature point on the cth channel in the fuzzy enhanced attention map, c = 1,...,C, Step S4.2: Take α c,x,y is the weight pair A fuzzy Perform weighted pooling to enhance the weight of the correct category at the corresponding position, thereby guiding the attention map to expand within a reasonable range and generating image block token prediction scores Step S4.3: Output the category token obtained in step S2.3 Generate class token prediction scores through dimensionality reduction through linear layers y class =Linear(T class ) (15) Where Linear(·) represents a linear layer; Step S4.4: Patch Token Prediction Scores Prediction score with class token Respectively with the true label Compute the multi-label soft margin loss: Where σ(·) represents the sigmoid activation function, and finally L patch and L class The sum is used as the final loss function, and the loss function is minimized through back propagation to optimize the model parameters.

5. The weakly supervised tissue pathology image segmentation method based on fuzzy enhanced Transformer according to claim 1 is characterized in that: The specific steps of step S5 are as follows: Step S5.1: After the model training is completed, the image is put into the model again and the blur enhancement attention map is extracted And the fuzzy enhanced attention map corresponding to each category is activated by the ReLU function and the maximum normalization operation: in Fuzzy enhanced attention map corresponding to category c, max(·) represents the maximum value operation; Step S5.2: Use the affinity matrix extracted in step S2.3 Perform post-processing of the fuzzy enhanced attention map: in Represents the matrix multiplication operation, R N×C (·)and Represents reshaping the matrix into N×C and shape; Step S5.3: Finally, A final Upsample to the input image size and apply the argmax operation to generate the final segmentation result pred: Where Upsample(·) represents the upsampling operation, where It means taking the maximum value index in the category dimension and visually displaying the segmentation results.

Citation Information

Patent Citations

  • Medical image depth segmentation method based on fuzzy logic

    CN116188435A

  • Methods and systems for training a model to diagnose abnormalittes in tissue samples

    US20230420133A1

Cited By

  • Multi-wavelength AI control device, method and system for minimally invasive spine surgery

    CN120240944A

  • Hyperspectral image and laser radar data joint classification method based on fuzzy enhanced Transform

    CN121259593A

  • Weakly supervised histopathological image segmentation method based on fuzzy-enhanced transformer

    WO2026152735A1