Adaptive medical image segmentation method based on adversarial attention mechanism and deep discrimination
Patent Information
- Application Number
- CN202310784227.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-29
AI Technical Summary
[0007]本发明主要针对现有模型在医学图像细胞分割上存在的局限性,提出基于对抗注意机制和深度判别的自适应医学图像分割方法
[0039]1)本发明方法利用域自适应方法,在医学图像标注稀缺且标注成本高的情况下,充分利用了标注样本并挖掘不同域间的共同特征。
Smart Images

Figure CN116740096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image semantic segmentation technology, specifically to an adaptive medical image segmentation method based on adversarial attention mechanism and depth discrimination. Background Technology
[0002] Currently, one of the main research areas in the field of intelligent medical image-assisted surgery is tumor cell segmentation. Precise cell segmentation achieved with the help of artificial intelligence technology can greatly assist in preoperative diagnosis, treatment planning, and postoperative recovery, while also reducing the burden of manual annotation, allowing valuable medical resources to be used more effectively.
[0003] Domain adaptation is widely used to solve the distribution transfer problem between a source domain where labels are available and a target domain where labels are unavailable. It learns the similarity between target domain features and source domain features, reduces the differences, and aligns the two types of features, thereby constructing a domain-agnostic common feature space. Inspired by generative adversarial networks, domain adaptation networks also focus on aligning feature distributions through different levels of adversarial learning.
[0004] Attention mechanisms were originally developed for natural language modeling. They model the relevance of target features to a specific task and selectively adjust the weights of features based on the attention map. This feature selection approach is now widely used in computer vision tasks and has achieved excellent results.
[0005] Deep supervision methods were initially proposed to train deeper networks for classification tasks. Later, in deep supervision segmentation research, multi-scale information from complementary information sources was introduced by adding a large amount of information to the intermediate layers of the network, thus enabling it to segment targets with better robustness.
[0006] While current cell segmentation models have achieved good segmentation results, this is based on the assumption that the features of the images used for segmentation testing are independent and identically distributed with those of the training set. In real life, medical images exhibit diverse cell morphologies and distributions, which means that segmentation models that perform exceptionally well in other fields often fail to achieve satisfactory results in the medical field. Summary of the Invention
[0007] This invention addresses the limitations of existing models in medical image cell segmentation by proposing an adaptive medical image segmentation method based on adversarial attention mechanisms and depth discrimination. Domain adaptation methods effectively improve the model's segmentation ability for target domain data by aligning the data of the target domain with the data of the source domain. However, traditional domain adaptation methods mostly treat all spatial features equally in a class-agnostic manner, neglecting the importance of classes in the image and discarding potentially discriminative information that may be meaningful for downstream tasks. Furthermore, due to the irregular shape, large number, and crowded nature of cells in medical images, as well as low image contrast, both cue-based and non-cue-based segmentation methods suffer from deficiencies in segmentation performance.
[0008] To solve the above-mentioned problems, the present invention adopts the following solution:
[0009] One aspect of the present invention provides an adaptive medical image segmentation method based on adversarial attention mechanisms and depth discrimination, comprising the following steps:
[0010] 1) Divide medical image data into labeled data and unlabeled data, and perform data augmentation processing on each medical image data.
[0011] 2) The enhanced labeled and unlabeled data from step 1) are respectively input into the dual adversarial mechanism of the adaptive network as source and target domain data. This mechanism performs attention calculation mapping from both spatial and class perspectives, and adaptively aligns the features of the source and target domains.
[0012] 3) After continuous iterative training, when the difference between the features of the target domain and the features of the source domain is less than a certain threshold, the adversarial attention mechanism converges and outputs the preliminary segmentation results of the target domain.
[0013] 4) Enhance the target domain data as input data for the deep discriminative supervision mechanism, and output two results: boundary detection and internal segmentation. The boundary detection results are further improved by learning the instance boundaries and refining the instance-level internal segmentation results at the pixel level.
[0014] 5) The improved boundary detection results from step 4) are fused with the segmentation results from step 3) to obtain a fine-grained segmentation mask for the target domain data.
[0015] Furthermore, in step 1), the data augmentation process employs contrast variation and normalization.
[0016] Furthermore, the calculation function for the spatial attention map is:
[0017]
[0018] f spatial(·) is the spatial attention computation function; φ(·,·) is used to measure the cosine distance between two semantic features at all spatial locations; This represents an encoder for data in a certain domain, where It is for the source domain. It refers to the target domain; x s ′ With x t ′ These are the enhanced source and target domain data, respectively.
[0019] The computation function for attention maps is designed as follows:
[0020]
[0021] f class (·) is the category attention calculation function; Softmax(·) is the normalization function.
[0022] Furthermore, the loss function in step 3) is defined as:
[0023]
[0024] This represents the loss function used to calculate the difference between the source domain encoder and decoder, where CE(·,·) represents the cross-entropy loss function. This represents the decoder that decodes source domain data after encoding. s ′ This indicates the annotation corresponding to the source domain data.
[0025] Furthermore, in step 4), the network's loss function is designed as follows:
[0026] L H =L F +L GD ,
[0027] Among them, L H Let L be the total loss function. F and L GD These represent focus loss and generalized dice loss, respectively.
[0028] Furthermore, in step 4), the boundary detection results are further improved by learning the instance boundaries and refining the instance-level internal segmentation results at the pixel level. Specifically:
[0029] Determine the entire instance boundary P of the learning objective B Simultaneously, the intermediate feature mapping F of the internal segmentation results is utilized. In This allows for further deduction of the location of the contact boundary;
[0030] The enhanced boundary detection result is obtained by fusing and adding the pixels of adjacent instance boundaries with the original boundary detection result:
[0031]
[0032] in, For adjacent instance boundaries, P B ′ This is the enhanced instance boundary.
[0033] Furthermore, the improved boundary detection results in step 4) and the segmentation results in step 3) are fused to obtain a fine semantic segmentation mask for the target domain data.
[0034] P Seg +P B ′ =P S ′ eg
[0035] P Seg P is the preliminary segmentation result output in step 3). S ′ eg This is the enhanced segmentation result.
[0036] Another aspect of the present invention provides an adaptive medical image segmentation device based on adversarial attention mechanism and depth discrimination, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned adaptive medical image segmentation method based on adversarial attention mechanism and depth discrimination.
[0037] Another aspect of the present invention provides a computer-readable storage medium storing a computer program for executing the above-described adaptive medical image segmentation method based on adversarial attention mechanism and depth discrimination.
[0038] Compared with existing technologies, the beneficial effects of the technical solution of this invention are:
[0039] 1) The method of this invention utilizes a domain adaptation approach, which makes full use of labeled samples and mines common features between different domains when medical image annotation is scarce and costly.
[0040] 2) This invention combines a dual adversarial attention mechanism and a deep discriminative supervision mechanism to align the feature differences between data from different domains starting from space and class, and applies pixel-level target internal information to instance-level target boundary detection, thereby improving the generalization and accuracy of the medical image semantic segmentation model. Attached Figure Description
[0041] Figure 1 This is a diagram of the architecture of the present invention;
[0042] Figure 2 A flowchart for countering attention mechanisms;
[0043] Figure 3 A flowchart for a deep discrimination monitoring mechanism;
[0044] Figure 4 This is a flowchart of the method of the present invention;
[0045] Figure 5 This is the segmentation effect of the present invention. Detailed Implementation
[0046] The present invention will be further described below with reference to the accompanying drawings.
[0047] The overall architecture of the network of this invention is as follows: Figure 1 As shown, by integrating the dual adversarial attention mechanism with the deep discriminative supervision mechanism into the adaptive network, the network's ability to extract data features between different domains of medical images is increased, and its ability to finely segment dense cells is improved.
[0048] like Figure 4 As shown, the specific implementation steps of one embodiment of this application are as follows:
[0049] 1) Based on whether the medical images have been labeled with tumor regions by medical experts, the data is divided into labeled and unlabeled data. After the division, to make the training data more valuable and to eliminate the noise inherent in the data, data augmentation processing such as contrast adjustment and normalization is performed:
[0050] x s ′ =Aug(x) s )
[0051] x t ′ =Aug(x) t )
[0052] Here x s With x t The source and target domain data before enhancement are shown below, x s ′ With x t ′ These are the enhanced source and target domain data, respectively; Aug(·) represents the data augmentation operation.
[0053] 2) Because the segmentation targets in medical images have different features, the cell size, shape and staining degree in each image are different. The domain adaptation method can effectively improve the model's ability to segment data in the target domain by aligning the data in the target domain with the data in the source domain.
[0054] However, traditional domain adaptation methods mostly treat all spatial features equally in a class-agnostic manner, ignoring the importance of classes in the image. Therefore, a dual adversarial attention mechanism is first proposed, the process of which is as follows: Figure 2 As shown.
[0055] The enhanced labeled and unlabeled data from step 1) are respectively input into the dual adversarial mechanism of the adaptive network as source and target domain data. This mechanism performs attention calculation mapping from both spatial and class perspectives, and adaptively aligns the features of the source and target domains.
[0056] In a preferred embodiment, the computation function for the spatial attention map is designed as follows:
[0057]
[0058] f spatial (·) is the spatial attention computation function; φ(·,·) is defined to measure the cosine distance between two semantic features at all spatial locations; This represents an encoder for data in a certain domain, where It is for the source domain. It is targeted at the target domain.
[0059] In a preferred embodiment, the computation function for the class attention map is designed as follows:
[0060]
[0061] f class (·) is the category attention calculation function; Softmax(·) is the normalization function, which is often used in deep learning.
[0062] 3) After continuous iterative training, when the difference between the features of the target domain and the features of the source domain is less than a certain threshold, the adversarial attention mechanism converges and outputs the preliminary segmentation results of the target domain.
[0063] In a preferred embodiment, the loss function is defined as follows:
[0064]
[0065] This represents the loss function used to calculate the difference between the source domain encoder and decoder, where CE(·,·) represents the cross-entropy loss function. y represents the decoder that encodes and then decodes the source domain data. s ′ This indicates the annotation corresponding to the source domain data.
[0066] 4) Simultaneously, due to the irregular shape, large number, and crowded nature of cells in medical imaging, the initial segmentation results based on adversarial attention mechanisms may result in incorrect or incomplete segmentation at cell overlap areas, indicating room for improvement in the field of medical imaging where accurate results are required. Therefore, a depth-discriminatory supervision mechanism is introduced into the adaptive network. By applying instance-level target internal information to pixel-level target boundary detection, the adaptive network's ability to segment adjacent cell boundaries is improved. The process of this mechanism is as follows: Figure 3 As shown.
[0067] Specifically, the enhanced target domain data x from step 1) is taken in this manner. t ′ The network outputs two results: boundary detection and interior segmentation, serving as input data for the deep discriminative supervision mechanism. Specifically, for these two detection components, in a preferred example, the network's loss function is designed as follows:
[0068] L H =L F +L GD ,
[0069] Among them, L H Let L be the total loss function. F and L GD These represent focus loss and generalized dice loss, respectively.
[0070] The boundary detection results are further enhanced by refining the instance-level internal segmentation results at the pixel level. Specifically, the depth discrimination algorithm first relearns the entire instance boundary of the target, and then uses the feature mapping of the internal segmentation results to further infer the position of the contact boundary. Finally, the pixels of the adjacent instance boundaries are fused and added to the original boundary detection results to obtain the enhanced boundary detection results.
[0071] 5) The improved boundary detection results from step 4) are fused with the segmentation results from step 3) to obtain a fine-grained semantic segmentation mask for the target domain data. From the results... Figure 5 It can be seen from ( Figure 5 From left to right, the data consists of source domain data, target domain data, and the network's segmentation effect on the target domain data. The network's adaptive ability and segmentation ability for a large number of crowded cell targets are both good.
[0072] Another embodiment of this application discloses an adaptive medical image segmentation device based on adversarial attention mechanisms and depth discrimination; the device includes:
[0073] The annotation unit is used to divide medical image data into labeled data and unlabeled data, and to perform data augmentation processing on each medical image data.
[0074] Adaptive Alignment Unit: This unit is used to input the enhanced labeled and unlabeled data as source and target domain data into the dual adversarial mechanism of the adaptive network. This mechanism performs attention calculation mapping from both spatial and class perspectives, and adaptively aligns the features of the source and target domains.
[0075] Preliminary segmentation result unit: After continuous iterative training, when the difference between the features of the target domain and the features of the source domain is less than a certain threshold, the adversarial attention mechanism converges and outputs the preliminary segmentation result of the target domain.
[0076] The target enhancement unit is used to take the enhanced target domain data as input data for the deep discriminative supervision mechanism, and output two results: boundary detection and internal segmentation. It further enhances the boundary detection results by learning the instance boundaries and refining the instance-level internal segmentation results at the pixel level.
[0077] The target fusion unit is used to fuse the enhanced boundary detection results with the segmentation results to obtain a fine-grained segmentation mask for the target domain data.
[0078] Embodiments of the device of the present invention can be applied to network devices. The device embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of the device loading corresponding computer program instructions from non-volatile memory into memory and executing them. The computer program is used to execute an adaptive medical image segmentation method based on adversarial attention mechanisms and depth discrimination. From a hardware perspective, in addition to the processor, network interface, memory, and non-volatile memory, the device may typically include other hardware for hardware-level expansion. On the other hand, this application also provides a computer-readable storage medium storing a computer program that executes an adaptive medical image segmentation method based on adversarial attention mechanisms and depth discrimination.
[0079] For the apparatus embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The apparatus embodiments described above are merely illustrative, and those skilled in the art can understand and implement them without creative effort.
[0080] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only.
[0081] It should also be noted that the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0082] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An adaptive medical image segmentation method based on adversarial attention mechanism and depth discrimination, characterized in that... The method includes the following steps: 1) Divide medical image data into labeled data and unlabeled data, and perform data augmentation processing on each medical image data; 2) The enhanced labeled and unlabeled data from step 1) are respectively input into the dual adversarial mechanism of the adaptive network as source and target domain data. This mechanism performs attention calculation mapping from both spatial and class perspectives, and adaptively aligns the features of the source and target domains. 3) After continuous iterative training, when the difference between the features of the target domain and the features of the source domain is less than a certain threshold, the adversarial attention mechanism converges and outputs the preliminary segmentation result of the target domain. 4) Enhance the target domain data as input data for the deep discriminative supervision mechanism, and output two results: boundary detection and internal segmentation. The boundary detection results are further improved by learning the instance boundaries and refining the instance-level internal segmentation results at the pixel level. 5) The improved boundary detection results from step 4) are fused with the segmentation results from step 3) to obtain a fine-grained segmentation mask for the target domain data.
2. The method according to claim 1, characterized in that, In step 1), the data augmentation process employs contrast variation and normalization.
3. The method according to claim 1, characterized in that, In step 2), the calculation function for the spatial attention map is: f spatial (·) is the spatial attention computation function; φ(·,·) is used to measure the cosine distance between two semantic features at all spatial locations; This represents an encoder for data in a certain domain, where It is for the source domain. It refers to the target domain; x s ′ With x t ′ These are the enhanced source and target domain data, respectively. The computation function for attention maps is designed as follows: f class (·) is the category attention calculation function; Softmax(·) is the normalization function.
4. The method according to claim 1, characterized in that, The loss function in step 3) is defined as: This represents the loss function used to calculate the difference between the source domain encoder and decoder, where CE(·,·) represents the cross-entropy loss function. This represents the decoder that decodes source domain data after encoding. s ′ This indicates the annotation corresponding to the source domain data.
5. The method according to claim 1, characterized in that, In step 4), the network loss function is designed as follows: L H L F +L GD , Among them, L H Let L be the total loss function. F and L GD These represent focus loss and generalized dice loss, respectively.
6. The method according to claim 1, characterized in that, In step 4), the boundary detection results are further improved by learning the instance boundaries and refining the instance-level internal segmentation results at the pixel level; specifically, the method is as follows: Determine the entire instance boundary P of the learning objective B Simultaneously, the intermediate feature mapping F of the internal segmentation results is utilized. In This allows for further deduction of the location of the contact boundary; The enhanced boundary detection result is obtained by fusing and adding the pixels of adjacent instance boundaries with those of the original boundary detection result: in, For adjacent instance boundaries, P B ′ This refers to the enhanced instance boundaries.
7. The method according to claim 6, characterized in that, The improved boundary detection results in step 4) and the segmentation results in step 3) are fused to obtain a fine semantic segmentation mask for the target domain data. P Seg +P B ′ =P S ′ eg P Seg P is the preliminary segmentation result output in step 3). S ′ eg This is the enhanced segmentation result.
8. An adaptive medical image segmentation device based on adversarial attention mechanism and depth discrimination, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the adaptive medical image segmentation method based on adversarial attention mechanism and depth discrimination as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing an adaptive medical image segmentation method based on adversarial attention mechanism and depth discrimination as described in any one of claims 1-7.
Citation Information
Patent Citations
Feature adaptive alignment unsupervised domain adaptive remote sensing image semantic segmentation method
CN113378906A
Image semantic segmentation method based on double category level adversarial network
CN114612658A