Medical image segmentation method based on intensity-space double mask auto-encoder
By using intensity-space double mask autoencoder in medical image segmentation and combining SSIM loss and contrast loss, the problems of blurred lesions, unclear boundaries and multi-scale features in medical images are solved, and higher segmentation accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510488694.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-13
AI Technical Summary
Existing medical image segmentation technology is difficult to effectively deal with focal features, unclear boundaries and multi-scale features.
The medical image segmentation method based on the intensity-space double mask autoencoder is adopted. Through the dual strategies of intensity mask and spatial mask, grayscale distribution characteristics and local spatial structure information are captured, and the feature distinction is optimized by combining SSIM loss and comparison loss.
It significantly improves the segmentation accuracy of blurred boundaries and multi-scale lesion areas in medical images, enhances adaptability to noise and complex backgrounds, and provides efficient feature extraction capabilities.
Smart Images

Figure CN120147345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and more precisely, it relates to a medical image segmentation method based on an intensity-spatial dual masked autoencoder. Background Art
[0002] Medical images usually have complex structures and features. Self-supervised pre-training models can learn general image features, such as textures, shapes, and edges, from a large amount of unlabeled medical image data. These general features are crucial for subsequent medical image segmentation tasks because they can provide richer information for the model. However, medical image processing still faces many challenges, such as blurred lesion features, unclear boundaries, and multi-scale feature problems, which have long troubled researchers. To effectively overcome these challenges, researchers have continuously explored various methods and technologies. Among them, masking strategies, as an important pre-training method, have gradually received attention.
[0003] Among many self-supervised learning methods, Masked Image Modeling (MIM) is becoming increasingly popular. This method randomly occludes parts of an image and allows the model to predict the occluded parts based on the visible parts, thereby learning the correlations and dependencies between different regions in the image and enabling the model to more comprehensively understand the semantic information of the image.
[0004] In the MIM task, the model not only needs to learn low-level features of the image (such as colors, textures, etc.), but also needs to learn high-level semantic information (such as object categories, scene understanding, etc.). However, during the training process, the model may over-focus on pixel-level reconstruction and neglect the learning of high-level semantic information. This situation may lead to poor performance of the model in downstream tasks (such as medical image segmentation) because these tasks rely more on the model's understanding of the high-level semantics of the image. Summary of the Invention
[0005] The purpose of the present invention is to address the deficiencies of the prior art and propose a medical image segmentation method based on an intensity-spatial dual masked autoencoder.
[0006] In the first aspect, a medical image segmentation method based on an intensity-spatial dual masked autoencoder is provided, including: S1. Obtain the original image and preprocess the original image; S2. Perform intensity masking operation and spatial masking operation respectively according to the preprocessed original image to generate an intensity masked image and a spatial masked image; S3. Input the intensity masked image and the spatial masked image into a dual-branch encoder respectively for feature extraction; S4. Upsample and reconstruct the feature maps output by the encoder through a decoder, and respectively output an intensity mask reconstructed image and a spatial mask reconstructed image; S5. Calculate the loss functions between the intensity mask reconstructed image and the original image, and between the spatial mask reconstructed image and the original image respectively; S6. Train the image segmentation model according to the loss function, and then apply the trained image segmentation model for image segmentation.
[0007] Preferably, S1 includes: S101. Adjust the size, window width and window level of the original image, and extract the lung window image and the mediastinal window image; S102. Perform Sobel operator processing on the lung window image and the mediastinal window image to generate an edge image; S103. Combine the lung window image, the mediastinal window image and the edge image into an RGB three-channel image.
[0008] Preferably, in S2, the intensity mask operation includes: Normalize the preprocessed original image; Convert the normalized image into a grayscale image; Divide the grayscale image into K grayscale intervals, and randomly select some intervals according to a preset mask ratio; Mask the pixels falling into the selected grayscale intervals, set their RGB values to zero, and generate an intensity mask image.
[0009] Preferably, in S2, the spatial mask operation includes: Divide the preprocessed original image into image blocks of size P, and randomly select some image blocks according to a preset mask ratio; Mask all the pixels in the selected image blocks, set their RGB values to zero, and generate a spatial mask image.
[0010] Preferably, in S3, the encoder adopts an architecture based on SegFormer.
[0011] Preferably, in S5, the loss function includes a structural similarity loss and a contrast loss.
[0012] In a second aspect, a medical image segmentation system based on an intensity-spatial dual mask autoencoder is provided for performing any of the methods in the first aspect, including: An acquisition module, configured to acquire an original image and preprocess the original image; A mask module, configured to perform an intensity mask operation and a spatial mask operation respectively according to the preprocessed original image, and generate an intensity mask image and a spatial mask image; An encoding module, configured to input the intensity mask image and the spatial mask image into a dual-branch encoder respectively for feature extraction; A decoding module, configured to upsample and reconstruct the feature maps output by the encoder through a decoder, and output an intensity mask reconstructed image and a spatial mask reconstructed image respectively; A calculation module, configured to calculate the loss functions between the intensity mask reconstructed image and the original image, and between the spatial mask reconstructed image and the original image respectively; A training module, configured to train an image segmentation model according to the loss function, and then apply the trained image segmentation model to perform image segmentation.
[0013] In a third aspect, a computer storage medium is provided. A computer program is stored in the computer storage medium; when the computer program runs on a computer, the computer is enabled to execute the method according to any one of the first aspect.
[0014] In a fourth aspect, an electronic device is provided, including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the method according to any one of the first aspect.
[0015] The beneficial effects of the present invention are as follows: 1. Through the dual strategies of intensity mask and spatial mask, the model of the present invention can simultaneously capture the gray distribution characteristics and local spatial structure information, and significantly improve the segmentation accuracy of fuzzy boundaries and multi-scale lesion regions in medical images.
[0016] 2. By combining the SSIM loss and the contrast loss, the model of the present invention optimizes the feature distinctiveness while reconstructing the image, and enhances the adaptability to noise and complex backgrounds.
[0017] 3. Based on the encoder architecture of SegFormer, the present invention realizes efficient feature extraction in 2D tasks, and provides an extensible technical path for high-dimensional medical image segmentation through the proposed 3D optimization direction. Description of the Drawings
[0018] Figure 1 It is a flowchart of the medical image segmentation method based on the intensity-spatial dual mask autoencoder provided by the present application; Figure 2 It is a schematic diagram of the effects of three channels and RGB images synthesized by using the method provided by the present application; Figure 3 It is a schematic diagram of the network structure of the intensity-spatial dual mask autoencoder provided by the present application; Figure 4 It is a schematic diagram of the visualization of the segmentation result provided by the present application; Figure 5 Online evaluation schematic diagram of pneumonia and mediastinal tumors provided by this application ; Figure 6 It is a schematic diagram of the model loss function curve on the TotalSegmentator dataset. Detailed implementation manners
[0019] The present invention will be further described below in conjunction with embodiments. The description of the following embodiments is only for helping to understand the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0020] Embodiment 1: Contrastive Learning (CL) is a self-supervised learning method for learning data representations, and its core idea is to compare the similarities and differences between different samples. Currently, there is the Tissue Contrastive Semi-Masked Autoencoder (TCS-MAE). The method adopts a density-based masking strategy. According to the HU spectral feature ranges occupied by different tissues, HU values are randomly masked on the basis of uniform intervals, and the masked tissue regions are reconstructed during the training process, so that the model can more effectively capture the feature representations of different tissues.
[0021] On this basis, this application discovers that the spatial masking strategy has its unique advantages. The spatial mask divides the image into multiple small blocks for masking, enabling the model to pay more attention to the local structure and detail information of the image. For example, in lung CT images, features such as small nodules and textures in the lungs are more easily learned and understood at the spatial level, while randomly masking HU values may, to a certain extent, damage the integrity of this local structure information. Therefore, this application proposes a new chest CT pre-training method called Intensity-Spatial Dual-Masked Autoencoder (ISD-MAE). Based on TCS-MAE, this application introduces a second branch - the Masked Autoencoder (MAE), and performs contrastive learning reconstruction through the two branches to optimize the performance of medical image classification and segmentation tasks.
[0022] As Figure 1 shown, the medical image segmentation method based on the intensity-spatial dual-masked autoencoder provided by this application includes: S1. Obtain the original image and preprocess the original image.
[0023] In S1, the original image is a 2D image or a 3D image. Exemplarily, when the original image is a 2D lung CT image, S1 includes: S101. Adjust the size, window width, and window level of the original image, and extract the lung window image and mediastinal window image.
[0024] In S101, adjust the size of the original image so that its dimension changes from (3, H, W) to (3, 256, 256). Then, this application performs window width and window level adjustment. Among them, the window width controls the displayed gray level range and affects the image contrast. A larger window width will stretch the displayed gray level range, making the image details more clearly visible. The window level sets the central gray level value for display, and adjusting the window level can change the brightness of the image.
[0025] S102. Perform Sobel operator processing on the lung window image and the mediastinal window image to generate an edge image.
[0026] In S102, the edge image is sharpened by the Sobel operator to highlight the outline of the tissue structure, helping the model better understand the internal structure and boundaries of the image.
[0027] S103. Combine the lung window image, the mediastinal window image, and the edge image into an RGB three-channel image.
[0028] Generally, a 2D lung CT image only contains single-channel information. Therefore, this application combines the original lung window image, the mediastinal image, and the edge image into an RGB three-channel image as the input of the 2D image.
[0029] Specifically, the composition of the input RGB image is as follows: Among them: Here, represents the original image, where the lung window parameters are set to , and setting these parameters can make the tissue structure, tumors, inflammation, and other lesion areas in the lungs more clearly visible.
[0030] is obtained by performing Sobel operator processing on the lung window image and the mediastinal window image . The Sobel operator includes a horizontal direction convolution kernel and a vertical direction convolution kernel : Among them, represents or , and represents the Sobel operator as follows: Then, the edge images of the lungs and mediastinum are obtained by the following calculations: in, Representation and or Finally, the application combines the obtained edge images of the lungs and mediastinum to obtain the final edge image: By taking the maximum value of the two images processed by the Sobel operator, not only can the noise interference be reduced, but also more information and contrast can be provided to the network. Figure 2 The figure shows the effect of synthesizing three channels using the above method. These three channels have better contrast and richer edge information than the original image.
[0031] In addition, when the original image is a 3D image, similar to the preprocessing method of 2D images, the 3D chest CT image also needs to be resized, changing the dimensions from (1, H, W, D) to (1,256, 256, 256). This application uses an interpolation method to adjust the interlayer thickness in the H, W, and D directions to 1 mm to reduce image differences caused by different layer thicknesses. Images of uniform thickness are usually smoother and more continuous, which not only improves the visual quality, but also makes the segmentation results clearer and more accurate. Subsequently, this application slices the 3D image layer by layer in the D dimension, and constructs RGB channels for each slice according to the aforementioned method, and then stacks them layer by layer. Finally, the preprocessed image is expanded from (1, 256, 256, 256) to (3, 256, 256, 256).
[0032] S2. Perform intensity masking operation and spatial masking operation on the preprocessed original image to generate an intensity masked image and a spatial masked image.
[0033] In S2, it is assumed that the input chest CT image is represented as , intensity mask operations, including: Normalize the preprocessed original image.
[0034] Specifically, the chest CT images Normalize to the range [0,1]: After that, the normalized image is converted into a grayscale image.
[0035] Specifically, the normalized image is converted into a grayscale image using the Luminosity Method : : where , and respectively represent the slices of the red, green, and blue channels of the image .
[0036] In addition, the interval is equally divided into sub - intervals to form a set . Here, is a hyper - parameter preset according to the mask scale. Then, according to the intensity mask ratio parameter , randomly selects intervals from the set .
[0037] For each pixel in the grayscale image , if it falls within any of the intervals in the set , the RGB value at the corresponding position in the original image is set to zero: Finally, an intensity mask image is formed.
[0038] In S2, the spatial mask operation includes: First, set the hyper - parameters and , which represent the size and mask ratio of the spatial mask respectively. The image is divided into grids of size , and then a certain number of image patches are selected according to the ratio , and the pixel values in these image patches are set to 0. Finally, we obtain an image with a spatial mask, denoted as .
[0039] S3. Input the intensity mask image and the spatial mask image into a dual - branch encoder for feature extraction.
[0040] Figure 3 shows the ISD-MAE framework of this application. Overall, the entire network follows the U-Net structure, but there are some differences. First, this application replaces the encoder part with SegFormer. SegFormer is an efficient semantic segmentation framework that combines Transformer and a lightweight MLP decoder. It does not rely on Positional Encoding, combines local attention and global attention, and can efficiently output multi-scale features, making it superior to the traditional U-Net encoder in feature extraction.
[0041] Before entering the network, the original image is first processed through the steps in S1 and S2 to generate an Intensity Mask Image and a Spatial Mask Image respectively. Subsequently, the image is randomly horizontally flipped, vertically flipped, rotated at any angle, and scaled with a certain probability to enhance the robustness of the model and reduce overfitting.
[0042] Then, they are respectively input into the encoder. After 5 rounds of downsampling, the size of the feature map gradually decreases while the number of channels increases. The feature map of the final layer has 512 channels, and the image size is 8 ∗ 8 (for 3D images, the size is 8 ∗ 8 ∗ 8).
[0043] S4. The feature map output by the encoder is upsampled and reconstructed by the decoder to output an Intensity Mask Reconstruction Image and a Spatial Mask Reconstruction Image respectively.
[0044] The feature map of the last layer of the encoder is input into the decoder and upsampled using nearest neighbor interpolation. The upsampled feature map is concatenated with the feature map of the corresponding layer of the encoder. After 5 rounds of such upsampling operations, the size of the feature map gradually recovers, and the number of channels decreases to 128, forming the decoder output with a size of (128, 256, 256). Finally, after passing through the reconstruction head module, it is restored to the original image size (3, 256, 256) to generate the reconstruction image.
[0045] S5. Calculate the loss functions between the Intensity Mask Reconstruction Image and the original image, and between the Spatial Mask Reconstruction Image and the original image respectively.
[0046] S6. Train the image segmentation model according to the loss function, and then apply the trained image segmentation model for image segmentation.
[0047] Embodiment 2: Based on Embodiment 1, Embodiment 2 of this application provides a more specific medical image segmentation method based on an intensity-spatial dual-mask autoencoder, including: S1. Obtain the original image and preprocess the original image.
[0048] S2. According to the preprocessed original image, perform intensity masking operation and spatial masking operation respectively to generate an intensity mask image and a spatial mask image.
[0049] S3. Input the intensity mask image and the spatial mask image into a dual-branch encoder respectively for feature extraction.
[0050] S4. Upsample and reconstruct the feature maps output by the encoder through a decoder, and output an intensity mask reconstruction image and a spatial mask reconstruction image respectively.
[0051] S5. Calculate the loss functions between the intensity mask reconstruction image and the original image, and between the spatial mask reconstruction image and the original image respectively.
[0052] In S5, the loss function includes a structural similarity (SSIM) loss and a contrast loss.
[0053] The first part of the loss function is the similarity loss between the decoder output and the real image, and this loss function adopts the structural similarity index measure Loss function. For two images and , The calculation formula is as follows: represents and , which are the output results of the intensity mask image and the spatial mask image after passing through the model respectively. and are the means of the images and respectively, and are their standard deviations, is and 's covariance. and are constants used to stabilize the calculation. The loss function of the model is the mean of the loss for each original image and its intensity mask image and spatial mask image : The second part of the loss function is the contrastive loss. The output feature map of the encoder passes through a projection head module, which performs average pooling on the feature map, flattens it, and passes it through a linear layer to obtain a 128-dimensional feature embedding. represents the embedding of the intensity mask image, Represents the embedding of the spatial mask image. Then calculate the contrast loss: pass Norm pair vector and Normalize and calculate the similarity score matrix of these two vectors: is a constant used to scale the embedding vector similarity score matrix, usually set to .
[0054] Define the number of categories as , create category labels , and calculate the cross entropy loss twice. The final contrast loss is the average of the two cross entropy losses.
[0055] Loss Function yes loss and contrast loss sum: S6. Train the image segmentation model according to the loss function, and then use the trained image segmentation model to perform image segmentation.
[0056] To verify the method of this application, this application first performed pre-training tasks on the TotalSegmentator lung CT scan dataset, and then performed downstream segmentation and classification tasks on eight 2D datasets and two 3D datasets, as shown in Table 1. Among these datasets, GRAM, dataset A, and dataset B are private datasets. These scans come from lung CT scans from multiple hospitals. The GRAM dataset is a 2D image obtained by slicing the 3D lung CT dataset, containing a total of 4,163 slices with a thickness of 1 mm. Dataset A contains 1,238 images, and dataset B is obtained by slicing dataset A.
[0057] Table 1 Datasets used in the experiment Dataset Name Shape Number of Samples Training / Testing Ratio Region Epoch Task Total Segmentator 2D 1227 / 8805 / None 50 Pre-training Dataset C 2D 2729 4:1 Pneumonia 30 Segmentation Task06 Lung 2D 463 4:1 Lung Cancer 100 Segmentation Dataset B 2D 1655 4:1 Pneumonia 40 Segmentation Lung nodule seg 2D 1657 4:1 Lung Nodule 50 Segmentation Lung CT nodule 2D 920 4:1 Lung Nodule 40 Segmentation LIDC IDRI 2D 1560 4:1 Lung Nodule 40 Segmentation Dataset D 2D 13980 4:1 Pneumonia 40 Classification GRAM 2D 4163 4:1 Fungal Infection 20 Classification Dataset E 3D 20 4:1 Pneumonia 30 Segmentation Dataset A 3D 1238 4:1 Pneumonia 20 Segmentation This application uses the Dice similarity coefficient ( ) and the Hausdorff distance ( ) to evaluate the performance of the pre-trained model of this application in downstream segmentation tasks, and uses to evaluate classification tasks.
[0058] The Dice coefficient is a metric for evaluating the accuracy of image segmentation and is commonly used in the field of medical image segmentation. It measures the similarity of segmentation by calculating the ratio of the overlapping part of two segmented regions to their total volume. The formula for the Dice coefficient is as follows: where represents the region of the model segmentation result, represents the true segmentation region.
[0059] The Hausdorff distance is a measure that describes the distance between two sets of points and is commonly used in fields such as image processing and computer vision to quantify the difference between two shapes or contours. The Hausdorff distance is defined as: where represents the distance metric between two points, usually using the Euclidean distance.
[0060] In addition, in Table 2, this application shows a comparison of the downstream segmentation results of the ISD-MAE of this application with multiple state-of-the-art self-supervised learning models on a series of chest CT datasets. The data in Table 2 represent the segmentation Dice scores and distances of different models on various datasets. The compared models include AE, DAE, MAE, CMAE, CAE, and MaskFeat.
[0061] Table 2 Results of downstream segmentation tasks using different self-supervised methods In each cell of Table 2, the upper number represents the Dice similarity coefficient (%), and the lower number represents the Hausdorff distance (HD). The best score in each row is shown in bold.
[0062] As can be seen from the segmentation results in Table 2, the Dice scores of ISD-MAE are generally higher than those of other methods on all pneumonia segmentation datasets (Datasets A, B, C, E). Especially on Dataset A, its Dice score reaches as high as 90.10 ± 0.54%, significantly exceeding other methods, while the Hausdorff value is relatively low. On the lung nodule segmentation datasets (Lungnodule seg, Lung CT nodule, LIDC IDRI), ISD-MAE also performs well. Especially on the Lungnodule seg and Lung CT nodule datasets, the Dice scores reach 93.37 ± 1.55% and 92.11 ± 2.23% respectively, also higher than other methods.
[0063] From the perspective of standard deviation, the standard deviation of ISD-MAE on multiple datasets is relatively small, indicating that its performance is relatively stable and less affected by data fluctuations, which is crucial for tasks that require high-precision prediction such as medical image segmentation. Although the performance of ISD-MAE on 3D images (such as Dataset E) is not as good as that on 2D images, it still has competitiveness. This also reflects the potential and advantages of ISD-MAE in processing complex and multi-dimensional medical image data, as well as its strong adaptability. In addition, on the mediastinal tumor dataset (Task06Lung), ISD-MAE also performs excellently.
[0064] It is worth noting that compared with its performance in the pneumonia segmentation task, the ISD-MAE pre-trained model of this application performs even better on the downstream mediastinal tumor segmentation dataset, showing a significant performance improvement. The mediastinal region contains multiple important organs and tissues, which are usually adjacent to each other and have blurred boundaries, resulting in irregular shapes and blurred edges. ISD-MAE focuses on the local details of tumors through masks in terms of intensity and space, such as texture and edge features at the spatial level. For mediastinal tumors with irregular shapes, capturing these local features is particularly important. At the tissue level, considering the overall relationship between the tumor and the surrounding tissues and obtaining context information helps to more accurately determine the location and boundary of the tumor. While the manifestation of pneumonia in chest CT images is relatively intuitive and may not require such a fine multi-level mask strategy to achieve effective segmentation, so the advantages of ISD-MAE cannot be fully exerted in such scenarios.
[0065] In addition to the segmentation results, ISD-MAE also performs excellently in downstream classification tasks, with ROC-AUC scores reaching 0.94 and 0.91 on Dataset A and GRAM dataset respectively. This highlights the robustness and effectiveness of ISD-MAE in processing complex medical image data, as well as its good generalization ability across different datasets. Moreover, the high ROC-AUC scores indicate that the model has strong predictive power and great application potential in classification tasks in medical image analysis.
[0066] Figure 4 shows the segmentation results of this application. In Figure 4, the green areas in the "Ground Truth" column represent the target structures with professional medical annotations, which have clear anatomical characteristics: in the Task06_Lung dataset, this area accurately outlines the lung targets related to mediastinal tumors, with clear boundaries and regular shapes, strictly defining the actual location and scope of the tumors in the lungs; in Dataset B, it presents the distribution characteristics of pneumonia in the lungs, mostly showing multifocal patterns with relatively clear boundaries, reflecting the spatial distribution of the actual infection areas; for the Lung_CT_nodule dataset, the true label areas appear as multiple small areas with irregular shapes, accurately annotating the actual location and shape of the lung nodules. These true label areas provide an intuitive benchmark for evaluating the segmentation effects of various methods. By comparing with the segmentation results of other methods (such as CMAE, DAE, etc.), the accuracy and deviation of each method in capturing characteristics such as target boundaries, shapes, and positions can be clearly judged.
[0067] From the segmentation results in Figure 4, it can be observed that compared with the other four methods (CMAE, DAE, CAE, AE), ISD-MAE can capture the boundaries and details of the lesion areas more accurately. Especially, it performs well in the pneumonia (Dataset B) segmentation task and can identify tumors with different shapes and sizes in the Lung CT nodule dataset. This indicates that the ISD-MAE of this application has learned more general image features, demonstrating stronger generalization ability and having obvious advantages in delineating the boundaries of lesion areas.
[0068] In addition, Figure 5 shows the online DSC evaluation results of ISD-MAE, CMAE, and AE in four segmentation tasks. These tasks include pneumonia (datasets C, E) and mediastinal tumors (Lung nodule seg and Lung CT nodule), where dataset E is a 3D dataset. It can be observed from the figure that in datasets C and the Lung nodule seg task, the Dice coefficient of ISD-MAE is always higher than that of CMAE and AE, reaching a DSC of over 80% almost at the initial stage. For the Lung CT nodule dataset, ISD-MAE has a lower score in the first two training rounds but then quickly surpasses CMAE and AE. ISD-MAE may require more time to adapt to the characteristics of the dataset at the initial stage, so its performance is relatively weak in the first few training rounds. However, once the model starts to converge, its performance quickly improves and exceeds other methods.
[0069] Furthermore, during the image pre-training process, the present application plotted the loss function curves of the ISD-MAE, CMAE, and AE methods (as Figure 6 shown) to evaluate their performance during training. At the beginning of training, the loss function value of ISD-MAE rapidly drops from a relatively high level, with a significant decline. In contrast, the initial loss function value of CMAE is slightly lower than that of ISD-MAE, and the decline rate is relatively slow; while the initial loss function value of AE is higher, and the decline rate is also slow.
[0070] As training progresses, the loss function of ISD-MAE continues to decline steadily and reaches a relatively low level (about 0.1) at around 300 steps. The loss function of CMAE also declines steadily, but the value at 300 steps is still higher than that of ISD-MAE (about 0.18). AE's decline accelerates during this stage, but the overall loss function value is still higher than that of ISD-MAE and CMAE. In the later stage of training, the loss functions of the three methods tend to stabilize. The final stable value of CMAE is slightly higher than that of ISD-MAE and also higher than that of AE.
[0071] The loss function curves indicate that ISD-MAE converges faster during the pre-training process and has a lower final loss value, demonstrating better training performance compared to CMAE and AE.
[0072] Furthermore, to evaluate the contributions of the intensity mask and the spatial mask to the model performance, a series of experiments were designed in this application. For example, while keeping other components unchanged, we trained the model using only the intensity mask or the spatial mask respectively, and observed the changes in metrics such as the Dice coefficient and the Hausdorff distance in the downstream segmentation task.
[0073] Table 3 shows the performance in the downstream segmentation task after pre-training using different mask strategies. "Only intensity mask" means using only the intensity mask strategy, while "only spatial mask" means using only the spatial mask strategy. As can be seen from the table, for the 2D task of pneumonia segmentation, the intensity mask strategy performs better in terms of the Dice coefficient, but the spatial mask strategy performs better in terms of the Hausdorff distance ( ), indicating more precise boundary segmentation.
[0074] For the 3D tasks (datasets E, A), although the difference in the Dice coefficient is not significant, the spatial mask strategy still performs better in terms of HD. For the pulmonary nodule datasets, the advantage of the spatial mask strategy is more obvious. On the three pulmonary nodule datasets, both the Dice coefficient and the Hausdorff distance under the spatial mask strategy are higher.
[0075] This indicates that the intensity mask strategy performs better in terms of the Dice coefficient, which may mean better overall segmentation performance and the ability to capture more target regions. While the spatial mask strategy performs better in terms of the Hausdorff distance ( ), indicating more accurate boundary segmentation results. Therefore, considering combining the two mask strategies can utilize their respective advantages to make the model more robust and adaptable to different datasets and tasks.
[0076] Table 3 Influence of Different Mask Strategies on Downstream Segmentation Tasks Intensity Mask Only Spatial Mask Only Dataset C 85.02 ± 0.67 22.71 ± 0.19 85.65 ± 0.45 21.78 ± 0.22 Task06 Lung 87.42 ± 0.897.87 ± 0.11 85.78 ± 0.547.41 ± 0.09 Dataset B 79.89 ± 1.9623.09 ± 0.18 79.12 ± 1.8423.01 ± 0.19 Lung nodule seg 89.45 ± 0.66 4.63 ± 0.13 90.33 ± 0.78 4.24 ± 0.11 Lung CT nodule 84.35 ± 0.46 14.12 ± 0.37 85.33 ± 0.64 13.85 ± 0.28 LIDC IDRI 86.55 ± 0.84 8.02 ± 0.21 87.05 ± 0.75 8.00 ± 0.17 Dataset E 83.68 ± 2.8414.79 ± 0.77 82.2 ± 3.4214.62 ± 0.64 Dataset A 67.47 ± 0.6021.26 ± 0.72 67.08 ± 0.64 21.39 ± 0.68 In Table 3, "only intensity mask" means using only the intensity mask strategy, while "only spatial mask" means using only the spatial mask strategy. In each cell of the table, the upper number represents the Dice similarity coefficient (%), and the lower number represents the Hausdorff distance (HD). The best scores in each row are shown in bold.
[0077] Although ISD-MAE performs excellently on traditional 2D images, it lacks significant advantages on 3D datasets. The multi-layer mask strategy of ISD-MAE may not be flexible or effective enough in 3D space to fully capture the complex features in 3D images. To make improvements in this regard, this application works in the following three main areas: Loss function: Consider using the multi-scale SSIM loss function and combine it with weighted aggregation to more comprehensively evaluate structural similarity. Alternatively, integrate the DICE loss to simultaneously consider structural similarity and regional overlap in the image.
[0078] Enhanced 3D convolutional block: Explore the use of methods such as deformable convolution and depthwise separable convolution, which are more robust than traditional convolution and more computationally efficient when dealing with images with scale and rotation variations.
[0079] Dataset: Process 3D images from multiple perspectives (such as axial, coronal, and sagittal planes) and enhance the segmentation performance by fusing information from these perspectives. This can be integrated through simple feature concatenation, weighted averaging, or attention mechanisms.
[0080] It should be noted that the same or similar parts in this embodiment and Embodiment 1 can be referred to each other and will not be elaborated in this application.
[0081] Embodiment 3: Based on Embodiments 1 and 2, Embodiment 3 of this application provides a medical image segmentation system based on an intensity-spatial dual masked autoencoder, including: An acquisition module for acquiring the original image and preprocessing the original image; A masking module for performing intensity masking operations and spatial masking operations respectively on the preprocessed original image to generate an intensity masked image and a spatial masked image; An encoding module for respectively inputting the intensity masked image and the spatial masked image into a dual-branch encoder for feature extraction; A decoding module for upsampling and reconstructing the feature maps output by the encoder through a decoder, and respectively outputting an intensity masked reconstructed image and a spatial masked reconstructed image; A calculation module for respectively calculating the loss functions between the intensity masked reconstructed image and the original image and between the spatial masked reconstructed image and the original image; A training module for training the image segmentation model according to the loss function and then applying the trained image segmentation model for image segmentation.
[0082] It should be noted that the system provided in this embodiment is the system corresponding to the methods provided in Embodiments 1 and 2. Therefore, the same or similar parts in this embodiment and Embodiments 1 and 2 can be referred to each other and will not be elaborated in this application.
[0083] In summary, this application proposes an improved self-supervised learning method, namely the Intensity-Spatial Dual Mask Autoencoder (ISD-MAE). This method introduces a Mask Autoencoder (MAE) branch into the model to perform intensity masking and spatial masking operations on images, so as to achieve multi-scale feature learning and the segmentation task of chest CT images. By combining the dual-branch structure and contrastive learning, the ability of the model to learn tissue features and boundary details is enhanced. In addition, in terms of data preprocessing, this application combines the original lung images, mediastinal images, and edge images into an RGB three-channel image to highlight the edge details of the lesion area. Experimental results show that ISD-MAE is significantly better than other methods in the segmentation tasks of 2D pneumonia and mediastinal tumor datasets, but there is still room for improvement in its performance on 3D datasets.
Claims
1. A medical image segmentation method based on intensity-space dual mask autoencoder, characterized in that: include: S1, obtaining an original image, and preprocessing the original image; S2, performing intensity mask operation and spatial mask operation respectively according to the preprocessed original image to generate an intensity mask image and a spatial mask image; S3, inputting the intensity mask image and the spatial mask image into a dual-branch encoder for feature extraction respectively; S4, upsampling and reconstructing the feature map output by the encoder through the decoder, and outputting an intensity mask reconstructed image and a spatial mask reconstructed image respectively; S5, respectively calculating the loss function between the intensity mask reconstructed image and the original image and between the spatial mask reconstructed image and the original image; S6. Training the image segmentation model according to the loss function, and then applying the trained image segmentation model to perform image segmentation.
2. The medical image segmentation method based on intensity-space dual mask autoencoder according to claim 1, characterized in that: S1 includes: S101, adjusting the size, window width and window position of the original image, and extracting the lung window image and the mediastinum window image; S102, performing Sobel operator processing on the lung window image and the mediastinum window image to generate an edge image; S103, combining the lung window image, the mediastinum window image and the edge image into an RGB three-channel image.
3. The medical image segmentation method based on intensity-space dual mask autoencoder according to claim 2, characterized in that: In S2, the intensity mask operation includes: Normalize the preprocessed original image; Convert the normalized image to a grayscale image; Divide the grayscale image into K grayscale intervals, and randomly select some intervals according to the preset mask ratio; The pixels falling into the selected grayscale interval are masked and their RGB values are set to zero to generate an intensity mask image.
4. The medical image segmentation method based on intensity-space dual mask autoencoder according to claim 3, characterized in that: In S2, the spatial mask operation includes: The preprocessed original image is divided into image blocks of size P, and some image blocks are randomly selected according to the preset mask ratio; All pixels within the selected image block are masked and their RGB values are set to zero to generate a spatial mask image.
5. The medical image segmentation method based on intensity-space dual mask autoencoder according to claim 4, characterized in that: In S3, the encoder adopts a SegFormer-based architecture.
6. The medical image segmentation method based on intensity-space dual mask autoencoder according to claim 5, characterized in that: In S5, the loss function includes structural similarity loss and contrast loss.
7. A medical image segmentation system based on intensity-spatial dual mask autoencoder, characterized in that: Used to perform the method according to any one of claims 1 to 6, comprising: An acquisition module, used for acquiring an original image and preprocessing the original image; The mask module is used to perform intensity mask operation and spatial mask operation respectively according to the preprocessed original image to generate an intensity mask image and a spatial mask image; An encoding module, used for inputting the intensity mask image and the spatial mask image into a dual-branch encoder for feature extraction; A decoding module, used for upsampling and reconstructing the feature map output by the encoder through a decoder, and outputting an intensity mask reconstructed image and a spatial mask reconstructed image respectively; A calculation module, used for respectively calculating the loss function between the intensity mask reconstructed image and the original image and between the spatial mask reconstructed image and the original image; The training module is used to train the image segmentation model according to the loss function, and then use the trained image segmentation model to perform image segmentation.
8. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is executed on a computer, the computer executes any one of the methods described in claims 1 to 6.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 6.
Citation Information
Cited By
Medical image segmentation method based on deep learning
CN121639726A
A medical image segmentation method based on deep learning
CN121639726B