Apparatus and method for segmentation of medical image
The Diffusion Transformer Segmentation model addresses the challenges of noisy and varying anatomical structures in medical images by using anatomical learning techniques, enhancing segmentation accuracy and improving diagnostic precision.
Patent Information
- Application Number
- JP2025078923
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-09
- Publication Date
- 2025-11-27
AI Technical Summary
Medical images often contain noise and artifacts that degrade image quality, and anatomical structures vary in shape, size, and texture, complicating accurate segmentation, especially in areas with complex or ambiguous boundaries.
A Diffusion Transformer Segmentation (DTS) model is employed for medical image segmentation, utilizing anatomical learning techniques like adjacent label smoothing and inverse boundary attention to enhance segmentation accuracy.
The DTS model improves the accuracy of organ segmentation in medical images with complex boundaries, capturing spatial relationships and enhancing object boundaries, enabling more precise diagnosis and treatment planning.
Smart Images

Figure 2025173487000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to medical image segmentation technology, and more particularly to an anatomical-based medical image segmentation device and method specialized for medical image segmentation. [Background technology]
[0002] Medical images acquired using equipment such as computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound often contain noise generated during image acquisition and processing. Furthermore, artifacts such as motion artifacts, metal artifacts, and aliasing artifacts degrade image quality, further challenging accurate segmentation. Anatomical structures in the human body vary in shape, size, and texture, resulting in different image shapes for the same anatomical structure. Variations in imaging protocols, such as differences in parameters and imaging artifacts, can cause inconsistencies in image shapes, further complicating the segmentation task. Pathological phenomena, such as tumors, lesions, and anomalies, can further blur organ boundaries, further challenging segmentation.
[0003] The background art of the present invention is disclosed in Patent Document 1 below. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Korean Patent Publication No. 10-2023-0165284 Summary of the Invention [Problem to be solved by the invention]
[0005] The present invention provides a medical image segmentation device and method. It is expected that the accuracy of organ segmentation in areas with complex or ambiguous boundaries in medical images will be significantly improved using a DTS (Diffusion Transformer Segmentation) model. Furthermore, the present invention aims to overcome the inherent problems of existing segmentation models through anatomical learning techniques such as adjacent label smoothing and inverse boundary attention, thereby providing a more accurate segmentation method.
[0006] The technical problems that the present invention aims to achieve are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0007] According to one aspect of the present invention, an apparatus and method for medical image segmentation is provided.
[0008] A medical image segmentation device according to one embodiment of the present invention may include an image input unit that inputs a medical image, a processing unit that embeds the input image into two encoders, a prediction unit that inputs the embedded image into a decoder to predict a global feature map, and a segmentation unit that divides the predicted region into regions of accurate organ positions. [Effects of the Invention]
[0009] According to an embodiment of the present invention, a Diffusion Transformer Segmentation (DTS) model can be used to significantly improve the accuracy of organ segmentation in medical images containing regions with complex or ambiguous boundaries. The DTS model captures spatial relationships within anatomical structures and enhances object boundaries with adjacent structures or background, enabling more accurate diagnosis and treatment planning in medical imaging applications.
[0010] Furthermore, the present invention provides models for various modalities, such as CT, MRI, and lesion images, thereby increasing efficiency and promoting future research and development of medical image software in medical image training, ultimately contributing to the advancement of medical image analysis.
[0011] The effects of the present invention are not limited to the effects described above, but include all effects that can be inferred from the configuration of the invention described in the description of the present invention or the claims. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a diagram illustrating a medical image segmentation device according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating the structure of a medical image segmentation device according to an embodiment of the present invention; [Figure 3] 1 is a diagram illustrating the structure of a medical image segmentation device according to an embodiment of the present invention. [Figure 4] 1 is a diagram illustrating the structure of a medical image segmentation device according to an embodiment of the present invention. [Figure 5] 1 is a diagram illustrating the structure of a medical image segmentation device according to an embodiment of the present invention. [Figure 6] 1 is a diagram illustrating the structure of a medical image segmentation device according to an embodiment of the present invention. [Figure 7] FIG. 2 is a diagram for explaining an algorithm of a medical image segmentation device according to an embodiment of the present invention. [Figure 8] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. [Figure 9] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. [Figure 10] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. [Figure 11] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. [Figure 12] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. [Figure 13] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. [Figure 14] 10A and 10B show experimental results of a medical image segmentation apparatus according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] Since the present invention can be modified in various ways and can have various embodiments, specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit the present invention to the specific embodiments, and includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention. In describing the present invention, if a detailed description of related publicly known technology is deemed to unnecessarily obscure the gist of the present invention, such a detailed description will be omitted. Furthermore, the term "a" or "an" used in the specification and claims generally means "one or more" unless otherwise specified.
[0014] Throughout this specification, when a part is said to be "connected (connected, contacted, or coupled)" to another part, this includes not only "directly connected" but also "indirectly connected" through an intervening member. Furthermore, when a part is said to "include" a certain component, this does not exclude other components, and means that the part may further include other components, unless otherwise specified.
[0015] The terms used in this specification are merely used to describe specific embodiments and are not intended to limit the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and do not preclude the presence or additional possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0016] The present invention will now be described with reference to the accompanying drawings. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. In the drawings, parts that are not relevant to the description will be omitted in order to clearly explain the present invention, and like parts will be designated by like reference numerals throughout the specification.
[0017] FIG. 1 is a diagram for explaining a medical image segmentation apparatus according to an embodiment of the present invention.
[0018] Referring to FIG. 1, the medical image segmentation apparatus includes an image input unit 110, a processing unit 130, a prediction unit 150, and a segmentation unit 170.
[0019] The image input unit 110 inputs a medical image to the medical image segmentation device. The medical image may include CT, MRI, lesion image data, and labels.
[0020] The processing unit 130 embeds the input image into two encoders. The processing unit 130 calculates and divides the input image and pre-labeled image into patch units, and then embeds them. The medical image is embedded in a first feature encoder, and the image and label-calculated image are embedded and added to the encoder of the present invention, with an emphasis on image representation learning. Specific details are described in detail in Figure 3.
[0021] The processing unit 130 can effectively encode anatomical information of the human body from images using self-supervised learning (SSL). The present invention includes three proxy operations for learning comprehensive semantic representations in masked images without using labels. Self-supervised learning (SSL) performs contrastive learning, which encodes masked images to improve the ability to distinguish between different samples with hidden feature representations; masked location prediction, which predicts the location of samples; and partial reconstruction prediction, which reconstructs masked patch regions of each subvolume to learn feature representations.
[0022] In contrastive training, positive examples are derived from identical inputs to represent semantic similarities. In particular, latent representations derived from identical inputs are considered positive examples. The unique image representations within the mini-configuration are utilized to generate negative examples for contrastive training. These negative examples highlight the differences between representations, allowing the model to learn and distinguish between diverse inputs.
[0023]
number
[0024] TIFF2025173487000003.tif25166
[0025] TIFF2025173487000004.tif25166
[0026]
number
[0027] TIFF2025173487000006.tif13166
[0028] The masked image modeling method with partial reconstruction prediction reconstructs all pixel values in the masked region by the decoder of the image to learn the feature representation. Considering the complex characteristics of medical images, a multidimensional decoder is necessary for exhaustive image reconstruction. TIFF2025173487000007.tif12166
[0029]
number
[0030] TIFF2025173487000009.tif13166
[0031] The present invention minimizes a total objective loss function that combines partial reconstruction prediction, masked position prediction, and contrastive learning loss as shown in Equation 4 below.
[0032]
number
[0033] TIFF2025173487000011.tif14166
[0034] The prediction unit 150 inputs the embedded image to the decoder and predicts the global feature map. The prediction unit 150 primarily predicts the global feature map using the decoder. The process of generating the global feature map will be described in detail in FIG. 4.
[0035] The segmentation unit 170 divides the predicted region into regions of the correct organ location. At this time, the segmentation unit 170 uses a Reverse Boundary Attention (RBA) module to provide attention to incorrectly predicted regions. The RBA module is described in detail in FIG. 5. The segmentation unit 170 applies a k-neighbor label smoothing algorithm to medical data of body parts, such as the abdomen and brain, which have structural positions within a compact space. k-neighbor label smoothing smooths the k neighboring labels for a given class or organ to utilize the relative position of the organ. In complex multi-class (k>2) situations like this, the present invention is advantageous if there is a spatial relationship between them. The spatial relationship refers to the relative position of the organs anatomically. The equation for k-neighbor label smoothing (k-NLS) is given by Equation 5 below.
[0036]
number
[0037] In Equation 5, the distance is calculated for each channel and is calculated as the distance between an arbitrary point and the center of the i-th class.
[0038]
number
[0039] TIFF2025173487000014.tif18166 is the set of distances between the classes. The scale factor, represented by α, determines the degree of smoothing applied to the predicted probabilities. The pseudo-code applied to the present invention is detailed in Figure 7.
[0040] 2 to 6 are diagrams illustrating the structure of a medical image segmentation device according to an embodiment of the present invention.
[0041] The inverse process, parameterized by θ, trains a neural network to invert the noise to recover the original data.
[0042]
number
[0043]
number
[0044] TIFF2025173487000018.tif25166
[0045] Referring to Figure 3, the image segmentation device inputs an original image 210. Then, the original input image and the correct mask image labeled by the medical staff are concatenated, and patch partition 220 is performed to divide the image into patches, after which tokens with sequences are embedded.
[0046] Referring to Figure 4, the present invention uses self-attention in a diffusion encoder of a Swin transformer to learn global dependencies between patches. Here, the present invention adds weights learned from existing CT image feature representations in a pre-training conditional encoder to better understand the features of the input image. Then, the present invention generates a global feature map 230 using a diffusion decoder.
[0047] Referring to FIG. 5, RBA can easily detect the boundary of an incorrectly predicted object by focusing on a non-object area in the recognition of the image boundary. TIFF2025173487000019.tif12166
[0048] Referring to Figure 6, the Reverse-Boundary Attention (RBA) method improves the prediction of a segmentation model by progressively capturing and designating regions that may have been initially ambiguous. Thus, the present invention removes previously estimated predicted regions with a higher-level output function where the existing estimate is upsampled at a deeper level, and then sequentially searches for detailed information including the region and its boundaries, ultimately refining the segmentation model. TIFF2025173487000020.tif16166
[0049]
number
[0050]
number
[0051] TIFF2025173487000023.tif22166
[0052]
number
[0053] TIFF2025173487000025.tif22166TIFF2025173487000026.tif21166
[0054] TIFF2025173487000027.tif21166
[0055]
number
[0056] TIFF2025173487000029.tif12166 and SW-MSA outputs, where LN and MLP indicate layer normalization and multilayer perceptron.
[0057] In addition, the present invention calculates self-attention by including the relative position bias as shown in Equation 13.
[0058]
number
[0059] TIFF2025173487000031.tif40166
[0060] The TIFF2025173487000032.tif21166 representation is propagated through normalization to a residual block consisting of a 3x3x3 convolutional layer. The features processed at each step are then upsampled using a deconvolutional layer and concatenated with the features processed at the previous step. The segmentation task combines the output of the Swine Transformer encoder with the features processed in the input volume. The concatenated information is propagated through the residual block and a final 1x1x1 convolutional layer, where an appropriate activation function (softmax) is applied to calculate the segmentation probability. TIFF2025173487000033.tif13166
[0061]
number
[0062] In Equation 14, DTS represents a new diffusion transformer segmentation model, which serves as an alternative to the traditional denoising U-Net.
[0063] Figures 8 to 14 show experimental results of a medical image segmentation device according to an embodiment of the present invention. Quantitative results indicate performance on CT, MRI, and skin lesion image datasets. In cases such as the BTCV dataset, the proposed model, which is similar to a diffusion segmentation model but for smaller organs, performs well and outperforms previous studies on MRI and skin lesion images.
[0064] Referring to Figure 8, Figure 8 shows the results of the BTCV challenge for multi-organ segmentation. The top and bottom parts of Figure 8 show the non-diffusion and diffusion-based segmentation models, respectively.
[0065] Referring to Figure 9, Figure 9 shows quantitative results for the BraTS dataset. Here, the loss function combines the DICE loss [Sudre et al., 2017], BCE loss, and MSE loss. For BTCV, training utilizes randomly cropped images with a resolution of 96x96 and a constellation size of 4 per GPU. However, for BraTS, the random crop size is set to 128x128, and the constellation size is configured as 2 per GPU. We introduce random flips, rotations, intensity scaling, and shifts for data augmentation, set the number of diffusion steps to 1000, and use a sliding window overlap ratio of 0.8 for the final prediction.
[0066] Figure 10 shows the performance test results for both non-diffusion and diffusion-based segmentation models on the ISIC dataset. Here, the average accuracy of Dice and HD95 scores achieved 92.12 and 2.18, respectively, demonstrating high performance. In this dataset, there is no structural relationship between labels, and there is only a single label, so K-nearest neighbor label smoothing cannot be applied.
[0067] Referring to Figure 11, we conduct comprehensive ablation experiments on the BTCV dataset to evaluate the efficiency of self-supervised learning. Figure 11 shows the results using a specific setting to calculate the experimental losses. The experiments include three loss functions: LRec (partial reconstruction prediction), LLoc (mask location prediction), and LCL (contrastive learning). In particular, LRec is learned on a pixel basis, LLoc is learned on a region basis, and LCL is learned at an extended sample level with an emphasis on contrastive learning. As can be seen from the experimental results, LRec plays an important role in understanding meaningful representation learning in medical images.
[0068] Referring to FIG. 12, the present invention continues to explore model improvements, investigating ablation studies on the BTCV dataset, focusing in particular on the effect of the scale factor α on k-nearest neighbor label smoothing performance. Generally, label smoothing prevents model overfitting and improves generalization performance. In Equation 6, the scale factor α determines the degree of smoothing applied to the predicted probability. By increasing the α value, the present invention features a wider and softer probability distribution, with a Dice result of 3.29 (%), outperforming existing baseline models.
[0069] Referring to Figure 13, the Scratch model uses a hybrid model that combines the existing dominant CNN-based denoising U-Net with a Swin Transformer encoder, which has proven effective in various fields such as natural language and image processing. The model with the Swin Transformer encoder captures the contextual meaning of organs during image feature extraction, improving segmentation.
[0070] Referring to Figure 14, Figure 14 shows qualitative results for the BTCV dataset. When the rectangular box displayed in the ground truth data in Figure 14 is enlarged, it can be seen that the representation of the corresponding part is smoothly divided, and the segmentation of small organs by learning the feature representation shows performance close to that of the representation of the ground truth data.
[0071] The medical image segmentation method described above may be embodied as computer-readable code on a computer-readable recording medium. The computer-readable recording medium may be, for example, a portable recording medium (CD, DVD, Blu-ray Disc, USB storage device, portable hard disk) or a fixed recording medium (ROM, RAM, computer-internal hard disk). The computer program recorded on the computer-readable recording medium may be transmitted to another computing device via a network such as the Internet, installed in the other computing device, and used in the other computing device.
[0072] Although all components constituting the embodiments of the present invention have been described as being integrally combined or operating in combination, the present invention is not necessarily limited to these embodiments, and all components may be selectively combined and operated in one or more combinations within the scope of the present invention.
[0073] Although acts are shown in a particular order in the figures, it should not be understood that the acts must be performed in the particular order shown, or in a sequential order, or that desirable results will be achieved if all of the acts shown are not performed. Multitasking and parallel processing may be advantageous in certain situations. Furthermore, the separation of various components in the above-described embodiments should not be understood as necessarily requiring such separation; the program components and systems described may generally be incorporated together in a single software product or packaged in multiple software products.
[0074] The present invention has been described above with reference to the preferred embodiments. Those skilled in the art may embody the present invention in modified forms without departing from the spirit and scope of the present invention. Therefore, the disclosed embodiments should be considered from an illustrative rather than a restrictive perspective. The scope of the present invention is defined by the claims, not the foregoing description, and all variations within the scope of equivalents thereto are encompassed by the present invention. [Explanation of symbols]
[0075] 100 Medical image segmentation device 110 Image input unit 130 Processing section 150 Prediction Department 170 Segmentation Department
Claims
1. A medical image segmentation device, comprising: an image input unit for inputting a medical image; a processing unit for embedding the input image into two encoders; a prediction unit that inputs the embedded image to a decoder and predicts a global feature map; a segmentation unit for dividing the predicted feature regions into regions of accurate organ locations; Contains A medical image segmentation device characterized by:
2. The processing unit The input image and the pre-labeled image are divided into patch units by calculation, and then embedding is performed. The medical image segmentation device according to claim 1 .
3. The processing unit Self-supervised learning (SSL) is applied to the input image to encode anatomical information of the human body, thereby performing partial reconstruction prediction for feature representation learning. The medical image segmentation device according to claim 1 .
4. The prediction unit Apply a Diffusion Decoder to generate a global feature map The medical image segmentation device according to claim 1 .
5. The segmentation unit The RBA (Reverse Boundary Attention) module provides attention to incorrectly predicted regions. The medical image segmentation device according to claim 1 .
6. 1. A medical image segmentation method, comprising: inputting a medical image; Embedding the input image into two encoders; inputting the embedded image to a decoder to predict a global feature map; Segmenting the predicted feature regions into regions of precise organ locations; Contains A medical image segmentation method comprising:
7. The step of embedding the input image into two encoders includes: The input image and the pre-labeled image are divided into patch units by calculation, and then embedding is performed. The medical image segmentation method of claim 6.
8. The step of embedding the input image into two encoders includes: Self-supervised learning (SSL) is applied to the input image to encode anatomical information of the human body, thereby performing partial reconstruction prediction for feature representation learning. The medical image segmentation method of claim 6.
9. The step of inputting the embedded image to a decoder and predicting a global feature map includes: Apply a Diffusion Decoder to generate a global feature map The medical image segmentation method of claim 6.
10. The step of dividing the predicted region into regions of precise organ locations comprises: The RBA (Reverse Boundary Attention) module provides attention to incorrectly predicted regions. The medical image segmentation method of claim 6.
11. Implementing the medical image segmentation method according to any one of claims 6 to 10 1. A computer program recorded on a computer-readable recording medium.
Citation Information
Patent Citations
Medical image segmentation device and method and computer readable storage medium
CN115760810A
Dental structured instance segmentation method based on diffusion prior repair
CN117576396A
Medical image segmentation and severity grading using neural network architectures with semi-supervised learning techniques
US10430946B1
A generalist framework for panoptic segmentation of images and videos
WO2024081778A1
Systems and methods for processing electronic medical images for diagnostic or interventional use
KR1020230165284A