Medical Image Segmentation Method and Device Based on Text Driving and Affinity Learning
By introducing text-driven and affinity learning methods in medical image segmentation, the problem of pseudo-label defects in medical image segmentation under weak supervision is solved, and more efficient and accurate pathological tissue segmentation is achieved.
Patent Information
- Application Number
- CN202510187449.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In medical image segmentation, it is difficult to accurately locate the complete target area based on weak supervision methods, especially in histopathological images. Due to homogeneity, overlapping phenomena and low color contrast, pseudo-label defects are caused, affecting segmentation performance.
Using text-driven and affinity learning-based medical image segmentation method, we use text supervision and affinity prediction values to improve the quality of pseudo-labels by constructing models including image vision encoder, tag encoder, expertise encoder, knowledge attention module and affinity prediction module.
This method effectively reduces the dependence on pixel-level fine annotation, improves the accuracy and efficiency of medical image segmentation through text supervision and affinity learning, and can more precisely locate and segment pathological tissues.
Smart Images

Figure CN119672347B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of weakly supervised learning, and particularly relates to a medical image segmentation method and device based on text-driven and affinity learning. Background Art
[0002] In the field of computer-aided diagnosis (CAD), the automatic segmentation technology of histopathological images plays a crucial role. It can help doctors quickly lock in abnormal tissue areas, quantitatively analyze the tumor microenvironment, thereby providing important bases for cancer grading and prognosis, and then significantly improving the diagnostic efficiency of clinicians. However, the fully supervised method highly depends on a large amount of high-quality labeled data. Especially in cases where professional medical knowledge is required, it is a time-consuming and laborious task for doctors to manually perform pixel-level fine annotation on histopathological images. Therefore, it has become an urgent need to develop an automated tool based on a weakly supervised method to segment and track regions of interest (ROIs) in histopathological images. It should be noted, however, that factors such as the morphological homogeneity, obvious overlapping phenomenon, and low color contrast in histopathological images result in differences between the same categories and similarities between different categories, which pose great challenges to weakly supervised segmentation techniques. In addition, the organizational structure in histopathological images may present random arrangement and dispersion characteristics, making it extremely difficult to identify complete tissues and regions of interest.
[0003] Currently, the mainstream method for weakly supervised segmentation is to use class activation mapping (CAM) to locate the attention region and then generate pseudo-labels to train the segmentation network. However, this method has a significant problem: the CAM generated using image-level labels can only highlight the most discriminative regions and cannot locate the complete object, resulting in defective pseudo-labels. Therefore, numerous studies have tried different methods to improve the quality of CAM, thereby improving the performance of weakly supervised semantic segmentation. Nevertheless, these improved CAM variants still face challenges in capturing complete tissues. The main problem is that in medical images, the symptoms and manifestations of subtypes are difficult to comprehensively describe with abstract semantic categories, and the supervision information relying solely on image-level labels is often insufficient to accurately locate the complete target region.
[0004] Recently, the Contrastive Language-Image Pretrained model (CLIP) has achieved remarkable success. By pre-training on 400 million image-text pairs from the Internet to predict whether image and text segments match, it demonstrates powerful capabilities in zero-shot classification. This dataset-independent model has the flexibility of direct transfer. More notably, the strong text-to-image generation ability shown by CLIP, such as that demonstrated by DALL-E2, reveals a profound connection between corresponding elements in text and images. Given these advantages, some researchers have attempted to introduce CLIP into weakly supervised natural image segmentation. Xie et al. proposed a novel Cross-Language Image Matching (CLIMS) framework for weakly supervised semantic segmentation by introducing natural language supervision, successfully activating more complete object regions and suppressing interference related to open background regions. Lin et al. reexamined the text input in the weakly supervised semantic segmentation setting and carefully designed two text-driven strategies: clarity-based prompt selection and synonym fusion. However, the application of CLIP in the field of weakly supervised histopathological image segmentation is very scarce. On the other hand, the multi-head self-attention (MHSA) layer in Transformer can capture semantic affinity between patches and can thus be used to improve rough pseudo-labels. However, the affinity captured in MHSA is still inaccurate, and directly applying MHSA to predict affinity to modify labels does not work well in practice. Summary of the Invention
[0005] The purpose of this application is to propose a medical image segmentation method and device based on text-driven and affinity learning for the above-mentioned technical problems.
[0006] In the first aspect, the present invention provides a medical image segmentation method based on text-driven and affinity learning, including the following steps:
[0007] Construct and train a text-driven medical image segmentation model to obtain a trained medical image segmentation model. The trained medical image segmentation model includes an image visual encoder, a label encoder, an expertise encoder, a knowledge attention module, and an affinity prediction module. The image visual encoder includes Transformer blocks in four consecutive stages. A fully connected layer is connected behind both the expertise encoder and the label encoder;
[0008] Obtain the medical image to be segmented, its corresponding expertise and text labels, and input them into the trained medical image segmentation model. The medical image to be segmented passes through the image visual encoder to obtain the hierarchical feature maps output by the Transformer blocks in four stages. The expertise passes through the expertise encoder to obtain the knowledge features. The hierarchical feature map of the 4th stage and the knowledge features after passing through the fully connected layer pass through the knowledge attention module to obtain the first image feature; the text labels pass through the label encoder and the fully connected layer to obtain the label features; add the hierarchical feature maps of the 3rd stage and the 4th stage to obtain the second image feature, calculate the similarities between the first image feature and the second image feature and the label feature respectively to obtain the first similarity feature map and the second similarity feature map, add the first similarity feature map and the second similarity feature map, and then perform the argmax operation to obtain the initial pseudo-label; add the hierarchical feature maps of the 1st stage and the 4th stage to obtain the third image feature, the third image feature passes through the affinity prediction module to obtain the affinity prediction value, and use the affinity prediction value to refine the initial pseudo-label through random walk to obtain the pseudo-label corresponding to the medical image to be segmented.
[0009] Preferably, the Transformer block adopts the structure of the Transformer block in the SegFormer network model, the label encoder adopts the MedCLIP model, and the expertise encoder adopts the ClinicalBert model.
[0010] Preferably, the knowledge attention module includes a concatenation layer, two multi-head self-attention layers and a segmentation layer connected in sequence. The hierarchical feature map of the 4th stage and the knowledge features after passing through the fully connected layer are input into the concatenation layer and concatenated in the token dimension to obtain the mixed features. The mixed features pass through the two multi-head self-attention layers to obtain the token sequence, and the token sequence is input into the segmentation layer to be split and reconstructed to obtain the first image feature.
[0011] Preferably, calculating the similarities between the first image feature and the second image feature and the label feature respectively to obtain the first similarity feature map and the second similarity feature map, adding the first similarity feature map and the second similarity feature map, and then performing the argmax operation to obtain the initial pseudo-label specifically includes:
[0012] Calculate the similarity between the first image feature and the label feature to obtain the first similarity feature map, as shown in the following formula:
[0013] ;
[0014] where represents the first image feature, represents the label feature, represents the first similarity feature map;
[0015] Calculate the similarity between the second image feature and the label feature to obtain a second similarity feature map, as shown in the following formula:
[0016] ;
[0017] Among them, represents the second image feature, represents the second similarity feature map;
[0018] Add the first similarity feature map and the second similarity feature map to obtain a fused feature, as shown in the following formula:
[0019] ;
[0020] Among them, represents the fused feature;
[0021] Perform an argmax operation on the fused feature to obtain an initial pseudo-label, as shown in the following formula:
[0022] ;
[0023] Among them, represents the argmax function, represents the initial pseudo-label.
[0024] Preferably, the third image feature passes through an affinity prediction module to obtain an affinity prediction value, specifically including:
[0025] The affinity prediction module includes a multi-head self-attention layer and a multi-layer perceptron connected in sequence. The third image feature is input into the multi-head self-attention layer in the affinity prediction module to obtain an affinity matrix, as shown in the following formula:
[0026] ;
[0027] Among them, represents the third image feature, represents the multi-head self-attention layer, represents the affinity matrix;
[0028] ;
[0029] Among them, represents the multi-layer perceptron, represents the affinity matrix of the transposed matrix, represents the affinity prediction value.
[0030] Preferably, during the training of the medical image segmentation model, the parameters of the label encoder and the expertise encoder are frozen, and the total loss function used by the medical image segmentation model during the training process is expressed as:
[0031] ;
[0032] wherein, represents the classification loss function, represents the weight coefficient, represents the affinity loss function;
[0033] ;
[0034] wherein, represents the first similarity loss function, represents the second similarity loss function;
[0035] The expression of the first similarity loss function is as follows:
[0036] ;
[0037] wherein, represents the class label of the nth sample, represents the first type of prediction vector obtained after the first similarity feature map passes through the global average pooling layer, represents the sigmoid activation function, and N represents the total number of samples in the training data;
[0038] The expression of the second similarity loss function is as follows:
[0039] ;
[0040] wherein, represents the class label of the nth sample, represents the second type of prediction vector obtained after the second similarity feature map passes through the global average pooling layer;
[0041] The expression of the affinity loss function is as follows:
[0042] ;
[0043] wherein, represents the affinity prediction value of pixel and pixel , and respectively represent the positive sample set and the negative sample set, and respectively represent and The number of samples in the medium; in the initial pseudo-label In, if the pixel and the pixel have the same semantics, they are marked as positive samples and put into the positive sample set, otherwise they are marked as negative samples and put into the negative sample set.
[0044] In a second aspect, the present invention provides a medical image segmentation device based on text-driven and affinity learning, including:
[0045] A model construction module configured to construct and train a text-driven medical image segmentation model to obtain a trained medical image segmentation model. The trained medical image segmentation model includes an image visual encoder, a label encoder, an expertise encoder, a knowledge attention module, and an affinity prediction module. The image visual encoder includes Transformer blocks in four consecutive stages. A fully connected layer is connected behind both the expertise encoder and the label encoder;
[0046] A prediction module configured to obtain a medical image to be segmented, its corresponding expertise, and text labels and input them into the trained medical image segmentation model. The medical image to be segmented passes through the image visual encoder to obtain hierarchical feature maps output by the Transformer blocks in four stages respectively. The expertise passes through the expertise encoder to obtain knowledge features. The hierarchical feature map of the 4th stage and the knowledge features after passing through the fully connected layer pass through the knowledge attention module to obtain the first image feature; the text labels pass through the label encoder and the fully connected layer to obtain label features; the hierarchical feature map of the 3rd stage and the hierarchical feature map of the 4th stage are added to obtain the second image feature. The similarities between the first image feature and the second image feature and the label feature are calculated respectively to obtain the first similarity feature map and the second similarity feature map. The first similarity feature map and the second similarity feature map are added, and then the argmax operation is performed to obtain the initial pseudo-label; the hierarchical feature map of the 1st stage and the hierarchical feature map of the 4th stage are added to obtain the third image feature. The third image feature passes through the affinity prediction module to obtain an affinity prediction value. The initial pseudo-label is refined by random walk using the affinity prediction value to obtain the pseudo-label corresponding to the medical image to be segmented.
[0047] In a third aspect, the present invention provides an electronic device including one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method described in any implementation manner of the first aspect.
[0048] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.
[0049] In a fifth aspect, the present invention provides a computer program product, including a computer program which, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] (1) The medical image segmentation method based on text-driven and affinity learning proposed by the present invention introduces text supervision (including text labels and expertise) and affinity learning to address the limitations of weakly supervised semantic segmentation on histopathological images. Since single image-level labels cannot provide sufficient supervision information, text supervision can provide additional guidance for the model. Therefore, a label encoder and an expertise encoder are used to encode text labels and expertise. And the expertise is used to guide the model to focus on relevant regions of the target tissue through a knowledge attention module.
[0052] (2) The medical image segmentation method based on text-driven and affinity learning proposed by the present invention introduces an affinity prediction module to learn reliable affinity prediction values, and uses the learned affinity prediction values to correct the initial pseudo-labels through random walk. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0054] Figure 1 It is a schematic flowchart of the medical image segmentation method based on text-driven and affinity learning according to the embodiments of the present application;
[0055] Figure 2 It is a schematic diagram of the training process of the medical image segmentation model of the medical image segmentation method based on text-driven and affinity learning according to the embodiments of the present application;
[0056] Figure 3 It is a schematic diagram of the medical image segmentation device based on text-driven and affinity learning according to the embodiments of the present application;
[0057] Figure 4 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0059] Figure 1 A medical image segmentation method based on text driving and affinity learning provided by an embodiment of the present application is shown, including the following steps:
[0060] S1. Construct and train a medical image segmentation model based on text driving to obtain a trained medical image segmentation model. The trained medical image segmentation model includes an image visual encoder, a label encoder, a professional knowledge encoder, a knowledge attention module, and an affinity prediction module. The image visual encoder includes Transformer blocks in four consecutive stages, and a fully connected layer is connected behind both the professional knowledge encoder and the label encoder.
[0061] In a specific embodiment, the Transformer block adopts the structure of the Transformer block in the SegFormer network model, the label encoder adopts the MedCLIP model, and the professional knowledge encoder adopts the ClinicalBert model.
[0062] In a specific embodiment, the knowledge attention module includes a concatenation layer, two multi-head self-attention layers, and a segmentation layer connected in sequence. The hierarchical feature map of the 4th stage and the knowledge feature after passing through the fully connected layer are input into the concatenation layer and connected in the token dimension to obtain a mixed feature. The mixed feature passes through two multi-head self-attention layers to obtain a token sequence, and the token sequence is input into the segmentation layer to be split and reconstructed to obtain the first image feature.
[0063] Specifically, refer to Figure 2, the text-driven medical image segmentation model (CLIPAL) proposed in the embodiments of this application aims to improve the accuracy of weakly supervised histopathological image segmentation. CLIPAL uses two types of text information: firstly, the text labels of the image types in the segmentation task, which provide basic category guidance for the image feature points; secondly, external descriptions of the performance of each type are introduced, and these descriptions are rich in the expertise provided by experts on subtype tissue morphology, color, and relationships with other tissues. In CLIPAL, the pre-trained MedCLIP model is used to match each text label with the feature points in the image space, and the corresponding first image features are extracted. This matching process is based on the similarity between features. The higher the similarity, the greater the possibility that the position belongs to the corresponding semantic category. Discriminative information is explored from the expertise to help accurately identify and locate complete pathological tissues by jointly modeling the long-range correlation between the image and the text. In addition, the inherent affinity in the Transformer is further utilized to improve the initial pseudo-labels. Through random walk propagation, the affinity prediction values are used to correct the initial pseudo-labels, which can spread the object regions and suppress the wrongly activated regions.
[0064] Furthermore, the medical image segmentation model mainly consists of five parts: three encoders (an image visual encoder, a label encoder, and an expertise encoder), a knowledge attention module, and an affinity prediction module. The image visual encoder is a hierarchical encoder, which consists of four stages. The Transformer block in each stage adopts the structure of the Transformer block in the SegFormer network model, and this structure includes an efficient self-attention layer and an overlapping block merger. The Transformer block in each stage generates multi-level multi-scale features of the given input image, and these features provide high-resolution coarse features and low-resolution fine-grained features. Specifically, for an input image with a size of , the input medical image is encoded into image features, and the hierarchical feature map of the i-th stage with a resolution of is obtained , where .
[0065] The label encoder encodes the text labels in the dataset into label features, denoted as , where represents the number of classes in the dataset, and represents the dimension of the label features. Since the label features will be used to calculate the similarity with the image features and the image features, it is very important to select an appropriate image-text pre-trained language model. In the embodiments of this application, the MedCLIP model is used as the label encoder, which is a model fine-tuned based on CLIP on the ROCO dataset.
[0066] The professional knowledge encoder is responsible for embedding the professional knowledge describing the subtype into the knowledge features, denoted as . The knowledge features guide the image features to focus on the regions related to the target tissue. To encode the professional knowledge of the subtype into more general semantic features, the embodiments of the present application use the ClinicalBert model as the professional knowledge encoder. The ClinicalBert model is a language model fine-tuned based on BioBert on the MIMIC-II dataset.
[0067] To improve the training efficiency, the weights of the label encoder and the professional knowledge encoder are frozen, but a fully connected layer (FC) is added after them to adjust the features extracted from the label encoder and the professional knowledge encoder.
[0068] To enhance the model's understanding of the color, morphology, and relationships between different tissues, pathology experts are asked to provide the professional knowledge of different subtype manifestations and encode it into knowledge features through the professional knowledge encoder. The knowledge attention module uses these external knowledge features to guide the image features to focus on the relevant regions of the target tissue. The knowledge attention module consists of a splicing layer, two multi-head self-attention modules, and a segmentation layer. The hierarchical feature map of the fourth stage and the knowledge features after passing through the fully connected layer are concatenated in the token dimension to obtain the mixed features . The mixed features are input into the multi-head self-attention module in the knowledge attention module. Finally, the output tokens are split through the segmentation layer, and the part corresponding to the hierarchical feature map of the fourth stage is taken out as the first image feature, while the hierarchical feature map of the third stage and the hierarchical feature map of the fourth stage are added together to form the second image feature. To save computing resources, the knowledge attention module is added only after the last layer of the image visual encoder.
[0069] In a specific embodiment, the similarities between the first image feature and the second image feature and the label feature are calculated respectively to obtain the first similarity feature map and the second similarity feature map. The first similarity feature map and the second similarity feature map are added together, and then the argmax operation is performed to obtain the initial pseudo-label, which specifically includes:
[0070] Calculate the similarity between the first image feature and the label feature to obtain the first similarity feature map, as shown in the following formula:
[0071] ;
[0072] where represents the first image feature, represents the label feature, Represents the first similarity feature map;
[0073] Calculate the similarity between the second image feature and the label feature to obtain the second similarity feature map, as shown in the following formula:
[0074] ;
[0075] Wherein, Represents the second image feature, Represents the second similarity feature map;
[0076] Add the first similarity feature map and the second similarity feature map to obtain a fused feature, as shown in the following formula:
[0077] ;
[0078] Wherein, Represents the fused feature;
[0079] Perform an argmax operation on the fused feature to obtain an initial pseudo-label, as shown in the following formula:
[0080] ;
[0081] Wherein, Represents the argmax function, Represents the initial pseudo-label.
[0082] Specifically, the embodiments of the present application use the inner product to calculate the similarity between the image feature and the label feature. The image feature includes the first image feature and the second image feature. Therefore, it is necessary to calculate the similarity between the first image feature and the second image feature and the label feature respectively. In order to make full use of the information of the network at different stages, CLIPAL adds and fuses the first similarity feature map and the second similarity feature map. Finally, an argmax operation is performed on the fused feature to obtain the initial pseudo-label .
[0083] In a specific embodiment, the third image feature passes through an affinity prediction module to obtain an affinity prediction value, specifically including:
[0084] The affinity prediction module includes a multi-head self-attention layer and a multi-layer perceptron connected in sequence. The third image feature is input into the multi-head self-attention layer in the affinity prediction module to obtain an affinity matrix, as shown in the following formula:
[0085] ;
[0086] Wherein, Represents the third image feature, Represents the multi-head self-attention layer, Represents the affinity matrix;
[0087] ;
[0088] Among them, represents a multi-layer perceptron, represents the affinity matrix of the transposed matrix, represents the affinity prediction value.
[0089] Specifically, the multi-head self-attention layer (MHSA) can capture the semantic affinity between patches, which can be used to improve the rough initial pseudo-labels. However, the affinity captured in MHSA is still inaccurate, and directly applying MHSA as the affinity to correct the initial pseudo-labels has poor effects in practice. Therefore, an affinity prediction module is introduced. Assuming that the output feature of MHSA is denoted as . In the affinity prediction module, the output features of the multi-head self-attention layer are directly linearly combined through a multi-layer perceptron to generate the affinity prediction value. Essentially, the self-attention mechanism is a directed graph model, and the affinity matrix should be symmetric because nodes sharing the same semantics should be equal. To perform such a transformation, simply add and its transpose.
[0090] In a specific embodiment, during the training process of the medical image segmentation model, the parameters of the label encoder and the expertise encoder are frozen. The total loss function used in the training process of the medical image segmentation model is expressed as:
[0091] ;
[0092] Among them, represents the classification loss function, represents the weight coefficient, represents the affinity loss function;
[0093] ;
[0094] Among them, represents the first similarity loss function, represents the second similarity loss function;
[0095] The expression of the first similarity loss function is as follows:
[0096] ;
[0097] Among them, represents the class label of the nth sample, represents the first type of prediction vector obtained after the first similarity feature map passes through the global average pooling layer, represents the sigmoid activation function, and N represents the total number of samples in the training data;
[0098] The expression of the second similarity loss function is as follows:
[0099] ;
[0100] Where, represents the class label of the nth sample, represents the second type of prediction vector obtained after the second similarity feature map passes through the global average pooling layer;
[0101] The expression of the affinity loss function is as follows:
[0102] ;
[0103] Where, represents the pixel and the pixel 's affinity prediction value, and represent the positive sample set and the negative sample set respectively, and represent respectively and 's number of samples; In the initial pseudo-label , if the pixel and the pixel have the same semantics, they are marked as positive samples and put into the positive sample set, otherwise they are marked as negative samples and put into the negative sample set.
[0104] Specifically, in order to obtain a more favorable affinity prediction value , a key step is to derive a reliable pseudo-affinity label as supervision. The pseudo-affinity label is derived from . Specifically, in , if the pixel and the pixel have the same semantics, set their pseudo-affinity label to positive, otherwise their pseudo-affinity label is negative. Therefore, the pseudo-affinity labels of the pixel pairs can be classified into the positive sample set and the negative sample set, and different methods are used to calculate the affinity prediction values for the positive sample set and the negative sample set respectively to achieve affinity learning. In addition, the embodiments of this application only consider the pixel and the pixel In the case of the same local window, the affinity of distant pixel pairs is ignored.
[0105] The affinity loss function is used to force the network to learn highly confident semantic affinity relationships from the MHSA. On the other hand, since the affinity prediction value is a linear combination of the MHSA, the affinity loss function is also beneficial to the learning of self-attention, which helps the network to further discover the overall object region.
[0106] S2. Obtain the medical image to be segmented, its corresponding expertise and text labels, and input them into the trained medical image segmentation model. The medical image to be segmented passes through the image visual encoder to obtain the hierarchical feature maps output by the Transformer blocks in four stages. The expertise passes through the expertise encoder to obtain the knowledge features. The hierarchical feature map of the fourth stage and the knowledge features after passing through the fully connected layer pass through the knowledge attention module to obtain the first image feature; the text labels pass through the label encoder and the fully connected layer to obtain the label features; the hierarchical feature maps of the third stage and the fourth stage are added to obtain the second image feature. The similarities between the first image feature and the second image feature and the label feature are calculated respectively to obtain the first similarity feature map and the second similarity feature map. The first similarity feature map and the second similarity feature map are added, and then the argmax operation is performed to obtain the initial pseudo-label; the hierarchical feature maps of the first stage and the fourth stage are added to obtain the third image feature. The third image feature passes through the affinity prediction module to obtain the affinity prediction value. The initial pseudo-label is refined by random walk using the affinity prediction value to obtain the pseudo-label corresponding to the medical image to be segmented.
[0107] Specifically, after deploying the trained medical image segmentation model, the medical image to be segmented, its corresponding expertise and text labels can be input into the trained medical image segmentation model. By introducing text information (including text labels and expertise), additional guidance is provided for the model to improve the CAM activation effect. At the same time, the affinity prediction value is calculated through the affinity prediction module, and this affinity prediction value is used to refine the initial pseudo-label to generate a more refined pseudo-label.
[0108] Considering that the fully supervised method often has a very high dependence on a large amount of high-quality labeled data, CLIPAL shows significant advantages. It effectively reduces the dependence on pixel-level fine annotation and can achieve accurate image segmentation only by using image-level annotation and text supervision, providing new possibilities for the automatic analysis and interpretation of pathological tissue images.
[0109] For further reference Figure 3, as an implementation of the methods shown in the above figures, an embodiment of a medical image segmentation device based on text-driven and affinity learning is provided in this application. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0110] An embodiment of a medical image segmentation device based on text-driven and affinity learning is provided in this application embodiment, including:
[0111] A model construction module 1, configured to construct and train a text-driven medical image segmentation model to obtain a trained medical image segmentation model. The trained medical image segmentation model includes an image visual encoder, a label encoder, an expertise encoder, a knowledge attention module, and an affinity prediction module. The image visual encoder includes Transformer blocks in four consecutive stages. A fully connected layer is connected behind both the expertise encoder and the label encoder;
[0112] A prediction module 2, configured to obtain a medical image to be segmented, its corresponding expertise, and text labels and input them into the trained medical image segmentation model. The medical image to be segmented passes through the image visual encoder to obtain hierarchical feature maps output by the Transformer blocks in four stages respectively. The expertise passes through the expertise encoder to obtain knowledge features. The hierarchical feature map of the 4th stage and the knowledge features after passing through the fully connected layer pass through the knowledge attention module to obtain the first image feature; the text labels pass through the label encoder and the fully connected layer to obtain label features; the hierarchical feature maps of the 3rd stage and the 4th stage are added to obtain the second image feature. The similarities between the first image feature and the second image feature and the label feature are calculated respectively to obtain the first similarity feature map and the second similarity feature map. The first similarity feature map and the second similarity feature map are added, and then the argmax operation is performed to obtain the initial pseudo-label; the hierarchical feature maps of the 1st stage and the 4th stage are added to obtain the third image feature. The third image feature passes through the affinity prediction module to obtain an affinity prediction value. The initial pseudo-label is refined by random walk using the affinity prediction value to obtain the pseudo-label corresponding to the medical image to be segmented.
[0113] Figure 4 It is a schematic hardware structure diagram of the electronic device provided in the embodiment of the present invention. As Figure 4 shown, the electronic device in this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer execution instructions; the processor 401 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiment.
[0114] Optionally, the memory 402 can be either independent or integrated with the processor 401.
[0115] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401.
[0116] An embodiment of the present invention further provides a computer storage medium, in which computer-executable instructions are stored. When the processor 401 executes the computer-executable instructions, the above method is implemented.
[0117] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.
[0118] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or modules, and can be in electrical, mechanical or other forms.
[0119] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0120] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The units formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0121] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or the processor 401 to execute some steps of the methods in various embodiments of the present application.
[0122] It should be understood that the above-mentioned processor 401 can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor 401 can also be any conventional processor 401, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being completed by the hardware processor 401, or completed by a combination of hardware and software modules in the processor 401.
[0123] The memory 402 may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0124] The bus 403 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 403 in the drawings of this application is not limited to only one bus 403 or one type of bus 403.
[0125] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0126] An exemplary storage medium is coupled to the processor 401, enabling the processor 401 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 401. The processor 401 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a master device.
[0127] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A medical image segmentation method based on text-driven and affinity learning, characterized in that: The following steps are involved: Constructing and training a text-driven medical image segmentation model to obtain a trained medical image segmentation model, wherein the trained medical image segmentation model includes an image vision encoder, a label encoder, an expertise encoder, a knowledge attention module, and an affinity prediction module, wherein the image vision encoder includes four stages of Transformer blocks connected in sequence, and each of the expertise encoder and the label encoder is connected to a fully connected layer; A medical image to be segmented and its corresponding professional knowledge and text labels are obtained and input into the trained medical image segmentation model, the medical image to be segmented passes through the image visual encoder to obtain hierarchical feature maps output by the Transformer blocks of four stages respectively, the professional knowledge passes through the professional knowledge encoder to obtain knowledge features, the hierarchical feature map of the fourth stage and the knowledge features after passing through the fully connected layer pass through the knowledge attention module to obtain the first image features; the text label passes through the label encoder and the fully connected layer to obtain the label feature; The hierarchical feature map of the third stage and the hierarchical feature map of the fourth stage are added to obtain a second image feature, and the similarities between the first image feature and the second image feature and the label feature are respectively calculated to obtain a first similarity feature map and a second similarity feature map, and the first similarity feature map and the second similarity feature map are added, and then an argmax operation is performed to obtain an initial pseudo-label; the hierarchical feature map of the first stage and the hierarchical feature map of the fourth stage are added to obtain a third image feature, and the third image feature is passed through the affinity prediction module to obtain an affinity prediction value, and the initial pseudo-label is refined by random walk using the affinity prediction value to obtain a pseudo-label corresponding to the medical image to be segmented.
2. The medical image segmentation method based on text-driven and affinity learning according to claim 1, characterized in that: The Transformer block adopts the structure of the Transformer block in the SegFormer network model, the label encoder adopts the MedCLIP model, and the expertise encoder adopts the ClinicalBert model.
3. The medical image segmentation method based on text-driven and affinity learning according to claim 1, characterized in that: The knowledge attention module includes a splicing layer, two multi-head self-attention layers and a segmentation layer connected in sequence. The hierarchical feature map of the fourth stage and the knowledge features after passing through the fully connected layer are input into the splicing layer, connected in the token dimension to obtain mixed features. The mixed features pass through two multi-head self-attention layers to obtain a token sequence. The token sequence is input into the segmentation layer, split and reconstructed to obtain the first image feature.
4. The medical image segmentation method based on text-driven and affinity learning according to claim 1, characterized in that: The similarities between the first image feature and the second image feature and the label feature are calculated respectively to obtain a first similarity feature map and a second similarity feature map, the first similarity feature map and the second similarity feature map are added, and then an argmax operation is performed to obtain an initial pseudo label, specifically including: The similarity between the first image feature and the label feature is calculated to obtain a first similarity feature map, as shown in the following formula: ; in, represents the first image feature, Represents label features, represents a first similarity feature map; The similarity between the second image feature and the label feature is calculated to obtain a second similarity feature map, as shown in the following formula: ; in, represents the second image feature, represents the second similarity feature map; The first similarity feature map and the second similarity feature map are added to obtain a fusion feature, as shown in the following formula: ; in, Indicates fusion features; Perform argmax operation on the fused features to obtain the initial pseudo label, as shown in the following formula: ; in, represents the argmax function, represents the initial pseudo-label.
5. The medical image segmentation method based on text-driven and affinity learning according to claim 1, characterized in that: The third image feature is passed through the affinity prediction module to obtain an affinity prediction value, which specifically includes: The affinity prediction module includes a multi-head self-attention layer and a multi-layer perceptron connected in sequence. The third image feature is input into the multi-head self-attention layer in the affinity prediction module to obtain an affinity matrix, as shown in the following formula: ; in, represents the third image feature, represents a multi-head self-attention layer, represents the affinity matrix; ; in, represents a multi-layer perceptron, Represents the affinity matrix The transposed matrix of represents the affinity prediction value.
6. The medical image segmentation method based on text-driven and affinity learning according to claim 1, characterized in that: During the training of the medical image segmentation model, the parameters of the label encoder and the expertise encoder are frozen, and the total loss function used in the training of the medical image segmentation model is It is expressed as: ; in, represents the classification loss function, represents the weight coefficient, represents the affinity loss function; ; in, represents the first similarity loss function, represents the second similarity loss function; The expression of the first similarity loss function is as follows: ; in, represents the class label of the nth sample, Represents the first type of prediction vector obtained after the first similar feature map passes through the global average pooling layer, represents the sigmoid activation function, and N represents the total number of samples in the training data; The expression of the second similarity loss function is as follows: ; in, represents the class label of the nth sample, Represents the second type of prediction vector obtained after the second similar feature map passes through the global average pooling layer; The expression of the affinity loss function is as follows: ; in, Represents pixels and pixels The predicted affinity value of and Represent the positive sample set and the negative sample set respectively, and Respectively and The number of samples in the initial pseudo label If the pixel and pixels If they have the same semantics, they are marked as positive samples and put into the positive sample set; otherwise, they are marked as negative samples and put into the negative sample set.
7. A medical image segmentation device based on text-driven and affinity learning, characterized in that: include: A model building module is configured to build and train a text-driven medical image segmentation model to obtain a trained medical image segmentation model, wherein the trained medical image segmentation model includes an image vision encoder, a label encoder, an expertise encoder, a knowledge attention module, and an affinity prediction module, wherein the image vision encoder includes four stages of Transformer blocks connected in sequence, and each of the expertise encoder and the label encoder is connected to a fully connected layer; The prediction module is configured to obtain the medical image to be segmented and its corresponding professional knowledge and text labels and input them into the trained medical image segmentation model, the medical image to be segmented passes through the image visual encoder to obtain the hierarchical feature maps output by the Transformer blocks of the four stages respectively, the professional knowledge passes through the professional knowledge encoder to obtain the knowledge features, the hierarchical feature map of the fourth stage and the knowledge features after passing through the fully connected layer pass through the knowledge attention module to obtain the first image features; the text label passes through the label encoder and the fully connected layer to obtain the label feature; The hierarchical feature map of the third stage and the hierarchical feature map of the fourth stage are added to obtain a second image feature, and the similarities between the first image feature and the second image feature and the label feature are respectively calculated to obtain a first similarity feature map and a second similarity feature map, and the first similarity feature map and the second similarity feature map are added, and then an argmax operation is performed to obtain an initial pseudo-label; the hierarchical feature map of the first stage and the hierarchical feature map of the fourth stage are added to obtain a third image feature, and the third image feature is passed through the affinity prediction module to obtain an affinity prediction value, and the initial pseudo-label is refined by random walk using the affinity prediction value to obtain a pseudo-label corresponding to the medical image to be segmented.
8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Dynamic gesture recognition method, system and equipment and medium
CN116524593A
Emotion enhancement continuous training method combining knowledge distillation and comparative learning
CN117115505A