Medical image segmentation apparatus and method
The diffusion transformer model addresses the challenges of noisy and anatomically varied medical images by integrating self-supervised learning and inverse boundary attention, enhancing organ segmentation accuracy and aiding precise medical diagnosis.
Patent Information
- Application Number
- FR2025005072
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-13
- Publication Date
- 2025-11-21
AI Technical Summary
Medical images acquired from CT, MRI, and ultrasound often contain noise and artifacts, leading to degraded image quality and complex segmentation tasks due to anatomical variations and pathological phenomena, making accurate organ segmentation difficult.
A diffusion transformer (DTS) segmentation model that integrates self-supervised learning and inverse boundary attention to enhance organ segmentation accuracy by learning anatomical structures and refining boundary predictions.
Improves the accuracy of organ segmentation in complex or ambiguous regions, enabling more precise medical image analysis and diagnosis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Apparatus and method for segmenting medical images. FIELD OF THE INVENTION
[0001] The present invention relates to a medical image segmentation technique and, more particularly, to an anatomy-based medical image segmentation device and a method specialized in medical image segmentation. PRIOR TECHNIQUE
[0002] Medical images acquired from computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound equipment often contain noise generated during image acquisition or processing. In addition, artifacts such as motion artifacts, metallic artifacts, and folding artifacts can degrade image quality and make accurate segmentation more difficult. Because human anatomies vary in shape, size, and texture, even the same anatomical structures can appear in different image shapes. Since inconsistencies in image shape can occur due to changes in imaging protocols, such as parameter differences and imaging artifacts, the segmentation task can become more complex.Furthermore, in the presence of a pathological phenomenon such as a tumor, lesion, or anomaly, the boundaries of organs become more obscure, and additional difficulties may arise during segmentation.
[0003] The technique underlying the present invention is disclosed in Korean patent no. 10-2023-0165284. Summary of the invention
[0004] The present invention provides an apparatus and method for segmenting medical images. The diffusion transformer (DTS) segmentation model of the present invention is expected to significantly improve the accuracy of organ segmentation in regions with complex or ambiguous boundaries within a medical image. Furthermore, the present invention aims to overcome the fundamental problems of existing segmentation models and to provide a more accurate segmentation method through anatomy-based learning, such as label smoothing by nearest neighbors or inverse attention to boundaries.
[0005] The technical problems to be solved by the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by persons competent in the art from the following descriptions.
[0006] To achieve the aforementioned objective, the present invention proposes a device and a method for segmenting medical images.
[0007] A medical image segmentation device according to an embodiment of the present invention may include: an image input unit for capturing a medical image; a processing unit for integrating the input image into two encoders; a prediction unit for capturing the image integrated into a decoder to predict an overall feature map; and a segmentation unit for segmenting the predicted feature region into regions of precise organ locations. Brief description of the drawings
[0008] [Fig.1] Fig.1 is a view explaining a medical image segmentation device according to an embodiment of the present invention.
[0009] [Fig.2] Fig.2 is a view showing the structure of a medical image segmentation device according to an embodiment of the present invention.
[0010] [Fig.3] The [Fig.3] is a view showing the structure of a medical image segmentation device according to an embodiment of the present invention.
[0011] [Fig.4] The [Fig.4] is a view showing the structure of a medical image segmentation device according to an embodiment of the present invention.
[0012] [Fig.5] The [Fig.5] is a view showing the structure of a medical image segmentation device according to an embodiment of the present invention.
[0013] [Fig.6] The [Fig.6] is a view showing the structure of a medical image segmentation device according to an embodiment of the present invention.
[0014] [Fig.7] The [Fig.7] is a view explaining the algorithm of a medical image segmentation device according to an embodiment of the present invention.
[0015] [Fig.8] The [Fig.8] is a view showing the results of the experiment of a medical image segmentation device according to an embodiment of the present invention.
[0016] [Fig.9] The [Fig.9] is a view showing the results of the experiment of a medical image segmentation device according to an embodiment of the present invention.
[0017] [Fig. 10] The [Fig. 10] is a view showing the results of the experiment of a medical image segmentation device according to an embodiment of the present invention.
[0018] [Fig. 11] The [Fig. 11] is a view showing the results of the experiment of a medical image segmentation device according to an embodiment of the present invention.
[0019] [Fig. 12] [Fig. 12] is a view showing the results of the experiment on an apparatus medical image segmentation according to an embodiment of the present invention.
[0020] [Fig. 13] The [Fig. 13] is a view showing the results of the experiment on an apparatus medical image segmentation according to an embodiment of the present invention.
[0021] [Fig. 14] The [Fig. 14] is a view showing the results of the experiment on an apparatus medical image segmentation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0022] The present invention is capable of various modifications and embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed descriptions. However, this is not intended to limit the present invention to the specific embodiments, and it shall be understood that it includes all modifications, equivalents, and substitutes within the spirit and technical scope of the present invention. Where it is established, during the description of the present invention, that a detailed description of a related known technology may unnecessarily obscure the substance of the present invention, the detailed description shall be omitted. Furthermore, singular expressions used in the specification and claims shall generally be interpreted as meaning "one or more," unless otherwise stated.
[0023] Throughout this specification, when a part is described as "connected (coupled, contacted, joined)" to another part, this includes cases where they are "indirectly connected" through the intervention of other members, as well as cases where they are "directly connected." Furthermore, when a part is described as "including" a certain component, this does not mean that other components are excluded, but rather that other components may be provided unless specifically stated otherwise.
[0024] The terms used in this specification are intended only to describe specific embodiments and not to limit the present invention. Singular expressions include plural expressions unless otherwise indicated. In this specification, the terms "include," "have," and others are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and not to exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0025] The present invention is described below with reference to the accompanying drawings. However, the present invention can be implemented in various forms and is not limited to the designs described herein. Furthermore, in order to clearly depict the present invention in the drawings, parts not relevant to the description are omitted, and similar drawing reference numbers are assigned to similar parts throughout the specification.
[0026] Fig. 1 is a view explaining a medical image segmentation device according to an embodiment of the present invention.
[0027] With reference to [Fig.1], the medical image segmentation device comprises an image input unit 110, a processing unit 130, a prediction unit 150 and a segmentation unit 170.
[0028] The image input unit 110 introduces a medical image into the medical image segmentation device. Medical images may include data and labels from computed tomography, MRI, and lesion images.
[0029] [The processing unit 130 performs an integration operation on the input image in two encoders. The processing unit 130 calculates the input image and a pre-labeled image, divides the image into patch units, and performs the integration. The medical image is integrated into a first feature encoder that focuses on learning the image representation, and the image and the labeled image are added to the encoder of the present invention for integration. Specific elements will be described in detail in FIG. 3.
[0030] The processing unit 130 can efficiently encode human anatomical information in an image by self-supervised learning (SSL). The present invention may include three substitution tasks for learning complete semantic representations in a masked image without using labels. The self-supervised learning (SSL) performs contrastive learning to improve the ability to distinguish different samples with hidden feature representations by encoding a masked image, masked location prediction to predict the location of a sample, and partial reconstruction prediction to learn feature representations by reconstructing a masked patch area of each sub-volume.
[0031] Contrastive learning derives positive samples from the same input and expresses semantic similarities. More specifically, latent representations of features from the same input are considered positive samples. Representations of features from a single image within a mini-batch are used to generate negative samples for contrastive learning. These negative samples highlight the differences between the representations of the features to allow the model to learn and distinguish between the different inputs.
[0032] [Math 1] -Qi ~~ “ ? V" St- )
[0033] In equation 1, t is a temperature parameter that controls the regularity of the distribution. 1 is an index that equals 1 if k ≥ i. x denotes a representation of a feature extracted by the encoder. silï^Xj denotes the similarities between the representations of positive samples, and X^ denotes the similarities between representations of negative samples.
[0034] The prediction of masked locations uses a 9-dimensional probability vector to represent a predicted number for the nth subvolume, denoted vn as the masked patch number in [0, 1,..., 8]. When the target v is given, a cross-entropy loss is used to predict the number.
[0035] [Math 2] 1¼
[0036] In equation 2, R denotes the number of sub-volumes and vn is expressed as a one-point vector.
[0037] In partial reconstruction prediction, the masked image modeling process learns the feature representation by reconstructing all the pixel values of a masked region using the image decoder. Given the complex characteristics of medical images, a multidimensional decoder is necessary for complete image reconstruction. The partial reconstruction loss is defined as the distance between the reconstructed region and the masked voxels of a target region.
[0038] [Math 3]
[0039] In equation 3, R is a subset of a subvolume of the target region, | | is the number of related subvolumes, and Yr and y denote respectively a predicted value and an input value.
[0040] The present invention minimizes a total objective loss function that combines the losses of partial reconstruction prediction, masked location prediction, and contrastive learning, as shown in Equation 4.
[0041] [Math 4] total ^Rec + ^2^0,
[0042] In equation 4, and are defined on 0.1 and 0.01 following verification experiments.
[0043] The prediction unit 150 introduces an embedded image into the decoder to predict a global feature map. The prediction unit 150 primarily predicts a global feature map via the decoder. The process of generating a global feature map will be described in detail in [Fig. 4].
[0044] The segmentation unit 170 segments the predicted region into regions corresponding to the precise locations of the organs. At this stage, the segmentation unit 170 pays attention to incorrectly predicted regions using an inverse boundary attention module (RB A). The RB A module will be described in detail in [Fig. 5]. The segmentation unit 170 applies a k-nearest neighbors label smoothing algorithm to medical data of body parts such as the abdomen, brain, and others, having a structural position in a compact space. The k-nearest neighbors label smoothing algorithm uses the relative locations of the organs by smoothing the labels by the k nearest neighbors for a given class or organ. In a complex situation with several classes (k > 2) such as this, the present invention is advantageous when there is a positional relationship between the two.From an anatomical point of view, the position relation means the relative position relation of the organs. The equation for k-nearest neighbors (k-NLS) label smoothing is that represented below in equation 4.
[0045] [Math 4]
[0046] d t = {dxyjpc, y, z eN, x< W, y <H, z<D}
[0047] In equation 4, the distance is calculated for each channel as the distance between an arbitrary point and the center of an i-th class.
[0048] [Math 5]
[0049] yk-NLS = | y _ --2- -1 Y t Ht dt+e |
[0050] Here, Yt is "1" in the case of a target class and "0" in the case of other classes, a is a label smoothing scaling factor, e is le-6, which is constant to avoid division by 0, and yz = { ..., $. | j — £ J is a set of centroids between each pixel and the class. The scaling factor denoted by a determines the degree of smoothing applied to a predicted probability. The pseudocode applied to the present invention is described in detail in [Fig. 7].
[0051] Figs. 2 to 6 are views showing the structure of a medical image segmentation device according to an embodiment of the present invention.
[0052] With reference to [Fig. 2], a diffusion model is configured from a diffusion process and a noise removal process. In the diffusion process, Gaussian noise is progressively added to the segmentation label over a series of steps t. The process does not include a neural network. The inverse process trains a neural network to invert the noise in order to recover the original data. In this case, the inverse process is parameterized by 0.
[0053] [Math 6] 100541 £,,(¾.]¾) = 11^(¾.]¾)
[0055] [Math 7] [°°56] p0(xt, t),
[0057] The distribution p^(Xf) is specified as 0 Process of diffusion, and in equations 6 and 7, 1 denotes a raw image assumed to be an nxn matrix. Then, the inverse process transforms the distribution of the latent variable (Gaussian noise image) into a data distribution £^(Xq) (final segmentation map).
[0058] With reference to [Fig.3], the image segmentation device between an original image 210. Then, after concatenating the original input image and a ground truth mask image labeled by a medical staff member and dividing the image into patch units by performing a patch partition 220, integration is carried out in the form of tokens having a sequence.
[0059] With reference to [Fig. 4], the present invention learns the overall dependence between patches by performing self-attention using a Swin transformer diffusion encoder. Here, the present invention adds weights learned from the existing CT image feature representation through pre-training of a conditional encoder for a better understanding of the input image features. Then, the present invention generates an overall feature map 230 using a diffusion decoder.
[0060] With reference to FIG. 5, the RB A performs inverse attention by focusing on objectless regions through the recognition of parts of the image boundaries in order to easily identify the boundaries of incorrectly predicted objects. The present invention creates an xt-1 image from an image supplemented with Gaussian noise and generates a final predicted image by repeatedly performing a noise removal process.
[0061] With reference to FIG. 6, the inverse boundary attention (RBA) method improves the prediction of the segmentation model by progressively capturing and designating regions that may have been initially ambiguous. Therefore, the present invention removes previously estimated prediction regions from a high-level output function where existing estimated values are sampled in deeper layers, sequentially explores specific information, including the corresponding regions and boundaries, and finally, progressively improves the prediction of the segmentation model. The present invention makes it possible to obtain an inverse attention RA^ 610 by multiplying the high-level outputs {Fi, i = 1, 2, 3, 4} by the inverse attention weights R^ 620.
[0062] [Math 8]
[0063] RA i =F i QR i
[0064] [Math 9]
[0065] R.= e(a(Y(SM)))
[0066] In equations 8 and 9, when U(-), o(-) and 0(-) are upsampling, sigmoid and inverse functions, the inverse function removes the matrix, which corresponds to 1 for all elements. The inverse attention weighting RA^ 610 passes through two convolution layers with normalization to finally obtain an inverse attention at the boundaries Si+1, as shown in equation 10.
[0067] [Math 10] i«'«i
[0069] In the noise suppression process, when the encoder input is a sub-volume the dimension of a 3D token with a resolution of The patch of (H1, W', D') is H'xW'xD'xS. The patch partition layer generates a sequence of 3D tokens of size Æ x AL x 22, projected into a C-dimensional space via an integration layer. For efficient modeling of the interactions between the tokens, the input volume is divided into non-overlapping windows, and local self-attention is calculated in each region. Specifically, in the layer, the 3D tokens are uniformly distributed across windows of size rH' / MlxfW' / MlxrD' / Ml.
[0070] In the next layer 1+1, the split windows are moved into units of [ ^ ] ) voxeL- The output of the Swin transformer encoder block in layers 1 and 1+1 is shown to be that represented in equation 11.
[0071] [Math 11]
[0072] = _ MSA( LN ( zM) ) +
[0073] z 1 =MLP(ln(z 1 ))+z 1
[0074] ÿ +1 =sW-MSA(LN(z^)+z^
[0075] z i+i=mLP(ln(z 1+1 ))+z 1+1
[0076] Here, W-MSA and SW-MSA are windows that divide the regular and multi-head self-attention modules, respectively. and are the outputs of W-MSA and SW-MSA, and LN and MLP represent hierarchical normalization and multilayer perception.
[0077] In addition, the present invention calculates self-attention by including a relative position bias as shown in equation 12.
[0078] [Math 12]
[0079] A n en n on ç qg, V) = Softmax[^\v
[0080] In equation 12, qg VG represent respectively the query, the The key and the value, and d is the size of the query and the key. The present invention uses a layer of patch blending between all stages to reduce resolution twice. During step 1, the number of tokens such that H y JE v maintained 1 J 1 2 A 2 A 2 by the linear integration layer and the transformer block. In steps 2, 3, and 4, the same process is repeated with resolutions of L xjx X -y X^pet Æ x — x J2. respectively. 16 A 16 A 16 1
[0081] A CNN-based decoder is connected to the encoder via a hop connection. At each step of F^io^Qr 1; 2, 3}) called a bottleneck (i=4), the output sequence is captured and the feature size is adjusted to H y — X — • The representation extracted at each step is transferred to block 2' 2J 2' residual configured in 3x3x3 convolutional layers by normalization. Then, the functions processed at each step are upsampled using a deconvolutional layer and connected to the functions processed in the previous steps. The segmentation task combines the processed functions in the input and output volumes of the Swin transformer encoder. The connected information is transferred through the residual block and the final convolutional layer Ixlxl, and a An appropriate activation function (Softmax) is applied to calculate a segmentation probability. At this stage, when conditional diffusion segmentation is applied, the noise xt, the temporal overlay t, and the conditional image I encoded in J form are integrated.
[0082] [Math 13]
[0083] x0 = DTS(concat(x t , I), t,
[0084] In equation 13, DTS denotes a new diffusion transformer segmentation model, which replaces the existing U-Net noise suppression model.
[0085] Figures 8 to 14 are views showing experimental results of a medical image segmentation device according to an embodiment of the present invention. The quantitative results demonstrate performance on CT, MRI, and skin lesion image sets. In the case of the BTCV dataset, the proposed model exhibits high performance on small organs, although similar to that of diffusion segmentation models, and superior performance to that of previous studies on MRI and skin lesion images.
[0086] With reference to [Fig. 8], [Fig. 8] shows a result of the BTCV challenge for multi-organ segmentation. The upper and lower parts of [Fig. 8] show segmentation models without diffusion and with diffusion, respectively.
[0087] With reference to [Fig. 9], [Fig. 9] is a quantitative result of the BraTS dataset. Here, the loss function incorporates DICE loss [Sudre et al., 2017], BCE loss, and MSE loss. In the case of BTCV, the learning uses randomly cropped images with a resolution of 96 x 96 and a batch size of 4 per GPU. However, in the case of BraTS, the random crop size is set to 128 x 128 and the batch size is 2 per GPU. The present invention adopts random flipping, rotation, force scaling, and shifting for data augmentation, and sets the number of diffusion steps to 1,000, and the overlap ratio of the sliding window to the final prediction is 0.8.
[0088] With reference to [Fig. 10], the overall results of the performance tests of the segmentation models with and without diffusion for the CITI dataset are presented. Here, the average accuracy of the Dice and HD95 scores shows high performance, reaching 92.12 and 2.18, respectively. In this dataset, k-nearest neighbor smoothing of the labels cannot be applied because there is only one label with no structural positional relationship between the labels.
[0089] With reference to [Fig. 11], the present invention performs a full ablation experiment on the BTCV dataset to evaluate the effectiveness of self-supervised learning. [Fig. 11] shows the result of using a specific parameter to calculate the loss of the experiment. The experiment includes three loss functions: LRec (partial reconstruction prediction), LLoc (masked location prediction), and LCl (contrastive learning). In particular, LRec is learned on a pixel basis, LLoc is learned on a region basis, and LCl is learned at the augmented sample level, focusing on contrastive learning. As the results of the experiment show, LRec plays an important role in understanding how to learn meaningful representations in medical images.
[0090] With reference to [Fig. 12], the present invention further explores model improvements and investigates ablation studies on the BTCV dataset, focusing particularly on the effect of the scaling factor 'a' on the performance of k-nearest neighbor label smoothing. In general, label smoothing prevents overfitting of the model and improves generalization performance. The scaling factor denoted by 'a' in Equation 5 determines the degree of smoothing applied to a predicted probability. The present invention is characterized by a broader and smoother probability distribution due to the increased value of 'a', and the Dice result is 3.29%, which is higher than existing benchmark models.
[0091] With reference to [Fig. 13], the model created from scratch uses a hybrid model that combines the dominant U-Net noise suppression method based on CNNs and the Swin transformer encoder, whose effectiveness has been proven in various fields such as natural language processing and image processing. The model with the Swin transformer encoder improves segmentation by capturing long-term contextual meaning in feature extraction from images.
[0092] Figure 14 shows the qualitative results for the BTCV dataset. When the square box shown in the ground truth data of Figure 14 is enlarged, it can be seen that the representation of a corresponding part is smoothly segmented, and the segmentation in small organs shows performance close to the representation of the reference data thanks to feature representation learning.
[0093] Fig. 15 shows a computer device implementing a method and a device for generating descriptors according to an embodiment of the present invention.
[0094] The embodiments of the present invention described in FIGS. 1 to 6 can be implemented in the form of a computer device 1500 operating with at least one processor.
[0095] The computer device 1500 may include a processor 1510, a storage 1520, a storage 1530, a communication interface 1540, a system interconnect 1550 and a screen 1560.
[0096] The 1510 processor includes a central processing unit (CPU), a microprocessor unit (MPU), a microcontroller unit (MCU), a graphics processing unit (GPU) and an application processing unit (APU).
[0097] The memory 1520 interacts with the processor 1510 to store data and quickly access the information necessary for the efficient execution of the program. The memory 1520 comprises at least one register, a cache memory, main memory, read-only memory, virtual memory, and non-volatile memory.
[0098] The 1530 storage system is designed to store and manage data permanently. It is used to preserve data even after the computer system is powered off or restarted, and to store operating systems, applications, user files, etc. The 1530 storage system comprises at least one of the following: a hard disk drive (HDD), a solid-state drive (SSD), an optical disk, network storage, and cloud storage.
[0099] The 1540 communication interface provides a path for transmitting and receiving data between various devices inside and outside the computer system. The 1540 communication interface can support at least one communication method from among Universal Serial Bus (USB), Peripheral Component Interconnect Express (PCIe), Serial ATA (SATA), Ethernet, Wi-Fi, Thunderbolt, and High-Definition Multimedia Interface (HDMI).
[0100] The system interconnect 1550 is designed to transmit and receive data and signals between various components of the computer system. The system interconnect 750 can support at least one method among a bus, a point-to-point interconnect, a crossbar switch, and a network-on-chip (NoC).
[0101] The 1560 screen is an output device of the computer system and has the function of providing visual information to users.
[0102] According to the configuration described above, a program according to an embodiment of the present invention is executed on the basis of instructions executed by the processor 1510, and can be stored in memory 1520 or storage 1530.
[0103] The medical image segmentation method described above can be implemented in the form of a computer-readable code on a computer-readable medium. The computer-readable recording medium can be, for example, A portable recording medium (CD, DVD, Blu-ray disc, USB storage device, portable hard drive) or a fixed recording medium (ROM, RAM, hard drive attached to the computer). Computer programs recorded on computer-readable recording media can be transmitted to other computing devices via a network such as the internet and installed on those devices, and can therefore be used on other computing devices.
[0104] As described above, although all the components constituting the embodiments of the present invention have been described to be combined into one or to function in combination, the present invention is not necessarily limited to embodiments. In other words, within the scope of the present invention, all the components can be selectively combined into one or more elements to function.
[0105] Although the operations are illustrated in the drawings in a particular order, it should not be understood that the operations must be performed in the particular order shown in the drawings or in a sequential order, or that all the illustrated operations must be performed to obtain a desired result. In a specific situation, multitasking and parallel processing may be advantageous. Furthermore, it should not be understood that the separation of the various components is necessarily required in the embodiments described above, and it should be understood that the program components and the system described above can generally be integrated together as a single software product or packaged as several software products.
[0106] The present invention has been described above with reference to its embodiments. Those skilled in the art will understand that the present invention can be implemented in modified forms without departing from the essential features of the present invention. Therefore, the disclosed embodiments should be considered from an illustrative rather than a restrictive point of view. The scope of the present invention is not indicated in the above description but in the claims, and any differences in equivalent scope should be interpreted as being included in the present invention.
[0107] According to one embodiment of the present invention, the accuracy of organ segmentation in a medical image comprising regions with complex or ambiguous boundaries can be significantly improved by using a diffusion transformer (DTS) segmentation model. The DTS model can enable more accurate diagnosis and treatment planning in the field of medical imaging applications by capturing spatial relationships within the anatomical structure and highlighting the boundaries of objects between adjacent structures or backgrounds.
[0108] In addition, the present invention can increase efficiency by providing models of various formats such as CT, MRI and lesion images, and contribute to the ultimate progress of medical image analysis by encouraging future research and development of medical imaging software in medical imaging practice.
[0109] It should be understood that the effects of the present invention are not limited to the effects described above and include all the effects which can be deduced from the configuration of the invention described in the description or claims of the present invention. List of reference signs
[0110] 100: Medical image segmentation device
[0111] 110: Image input unit
[0112] 130: Processing Unit
[0113] 150: Prediction Unit
[0114] 170: Segmentation unit
Claims
Demands
1. A medical image segmentation apparatus comprising: an image input unit for capturing a medical image; a processing unit for integrating the input image into two encoders; a prediction unit for capturing the image integrated into a decoder to predict an overall feature map; and a segmentation unit for segmenting the predicted feature region into regions of precise organ locations.
2. Apparatus according to claim 1, wherein the processing unit calculates the input image and a pre-labeled image, divides the images into patch units, and performs the integration.
3. Apparatus according to claim 1, wherein the processing unit performs a partial reconstruction prediction of a part of the feature representation by encoding the anatomical information of a human body by applying self-supervised learning (SSL) to the input image.
4. Device according to claim 1, wherein the prediction unit generates an overall feature map by applying a broadcast decoder.
5. Apparatus according to claim 1, wherein the segmentation unit pays attention to incorrectly predicted regions using an inverse boundary attention (RBA) module.
6. A method for segmenting a medical image comprising the following steps: capturing a medical image; integrating the input image into two encoders; capturing the integrated image into a decoder to predict an overall feature map; and segmenting the predicted region into regions of precise organ locations.
7. A method according to claim 6, wherein the input image integration step in two encoders calculates the input image and a pre-labeled image, divides the images into patch units and performs the integration.
8. A method according to claim 6, wherein the input image integration step in two encoders performs a partial reconstruction prediction of a learning portion of the feature representation by encoding the anatomical information of a human body by applying self-supervised learning (SSL) to the input image.
9. A method according to claim 6, wherein the image capture step integrated into a decoder to predict an overall feature map generates an overall feature map by applying a broadcast decoder.
10. A method according to claim 6, wherein the step of segmenting the predicted region into regions of precise organ locations takes into account incorrectly predicted regions using an inverse boundary attention (RBA) module.
11. Computer program for carrying out one of the medical image segmentation methods of claims 6 to 10 and recorded on a computer-readable recording medium.