Boundary-aware 3d mri tumor segmentation method and tumor diagnosis device
By introducing a boundary-aware loss function and a surface distance field into a convolutional neural network, the problem of low accuracy in spinal cord tumor segmentation in existing technologies is solved, and higher precision tumor segmentation results are achieved.
Patent Information
- Application Number
- CN202411298458.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-09-18
AI Technical Summary
Existing tumor segmentation methods based on convolution, attention, and large models have low accuracy in segmenting spinal cord tumors, making it difficult to accurately segment and identify them. This is especially true because spinal cord tumors are small in size, have different locations and shapes, leading to low accuracy in segmentation using existing methods.
A boundary-aware 3D MRI tumor segmentation method is adopted, which utilizes a convolutional neural network (CNN) model and combines a first model branch and a second model branch. The first branch is used to predict the class label, and the second branch is used to predict the surface distance field of the tumor based on boundary awareness. The model is optimized by a boundary-aware loss function to improve the perception of details of the tumor boundary and surrounding area.
It significantly improves the accuracy of tumor segmentation, especially for spinal cord tumors, by capturing details of the tumor boundary and surrounding area, thereby enhancing the accuracy and consistency of tumor segmentation.
Smart Images

Figure CN119762417B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image segmentation, and particularly relates to a three-dimensional MRI tumor segmentation method based on boundary perception and a tumor diagnosis device. BACKGROUND
[0002] Precise morphological quantification of tumors, including tumor size, location, and type, is expected to enhance monitoring protocols and optimize treatment planning strategies. In the past few years, a series of complex tumor segmentation methods based on convolution, attention, and large models have been proposed. These methods perform well on various medical images, thanks to the availability of large-scale datasets. However, these methods still face great challenges in segmenting certain tumors, such as spinal cord tumors, with low accuracy. Therefore, for these tumors, there is still a lack of automatic models that can accurately segment and identify tumors. SUMMARY
[0003] The embodiments of the present application provide a three-dimensional MRI tumor segmentation method based on boundary perception and a tumor diagnosis device.
[0004] In a first aspect, the embodiments of the present application provide a three-dimensional MRI tumor segmentation method based on boundary perception, which comprises:
[0005] Obtaining three-dimensional MRIs of multiple tumor patients as a tumor dataset, the three-dimensional MRIs including gadolinium-enhanced T1 SC slices;
[0006] Segmenting tumors in the three-dimensional MRIs in the tumor dataset as input of a first segmentation model to obtain a class label of the tumors in the three-dimensional MRIs;
[0007] The first segmentation model is a model based on a convolutional neural network (CNN), and the first segmentation model includes a first model branch and a second model branch. The first model branch is used to predict the class label, and the second model branch is used to predict a surface distance field of the tumors based on boundary perception. The surface distance field is used to describe the nearest distance of each voxel in the three-dimensional MRI to the tumor surface. The training and optimization of the first segmentation model are based on the training and optimization of the first model branch and the second model branch.
[0008] Optionally, the second model branch is trained and optimized based on a first boundary perception loss function. The first boundary perception loss function is used to determine the boundary perception loss of each voxel in the three-dimensional MRI. The boundary perception loss of each voxel is determined based on a surface distance value of each voxel. The surface distance value is used to represent the nearest distance of each voxel to the tumor surface.
[0009] Optionally, the voxels include voxels outside the tumor and voxels inside the tumor. The surface distance value of the voxels outside the tumor is represented as a negative value, and the surface distance value of the voxels inside the tumor is represented as a positive value.
[0010] Optionally, if the absolute value of the surface distance of a voxel outside the tumor is greater than the maximum surface distance of a voxel inside the tumor, the surface distance value is counted as 0.
[0011] Optionally, the boundary perception loss for each voxel is determined based on the normalized value of the surface distance for each voxel.
[0012] Optionally, the training and optimization of the second model branch can be performed according to the tumor category.
[0013] Optionally, the first boundary-aware loss function is expressed as follows:
[0014]
[0015] Among them, l ba The parameters are: h (representing boundary-aware loss), w (representing the horizontal pixel count of the 3D MRI), d (representing the number of slices in the 3D MRI), k (representing the total number of categories in the 3D MRI), and f. (h,w,d,k) The surface distance value of the (h, w, d, k)th voxel predicted by the second model branch is used to represent the value of the voxel. Used to represent the true surface distance value of the (h, w, d, k)th voxel.
[0016] Optionally, the second model branch is trained and optimized based on the average boundary-aware loss of all voxels in the 3D MRI.
[0017] Optionally, the first model branch is trained and optimized based on a first segmentation loss function, which includes a cross-entropy loss function and a dice loss function.
[0018] Secondly, embodiments of this application provide a tumor diagnostic device, comprising:
[0019] The acquisition module is used to acquire three-dimensional MRI images of multiple cancer patients as a tumor dataset. The three-dimensional MRI images include gadolinium-enhanced T1SC slices.
[0020] The segmentation module is used to segment tumors by taking the 3D MRI images from the tumor dataset as input to the first segmentation model, and to obtain the category labels of tumors in the 3D MRI images.
[0021] The first segmentation model is a model based on a convolutional neural network (CNN), and the first segmentation model includes a first model branch and a second model branch.
[0022] Optionally, the second model branch is trained and optimized based on a first boundary-aware loss function, and the first boundary-aware loss function is used to determine a boundary-aware loss of each voxel in the three-dimensional MRI, and the boundary-aware loss of each voxel is determined based on a surface distance value of each voxel, and the surface distance value is used to represent the nearest distance of each voxel to the tumor surface.
[0023] Optionally, the voxels include voxels outside the tumor and voxels inside the tumor, the surface distance value of the voxels outside the tumor is represented as a negative value, and the surface distance value of the voxels inside the tumor is represented as a positive value.
[0024] Optionally, if the absolute value of the surface distance value of the voxels outside the tumor is greater than the maximum surface distance value of the voxels inside the tumor, the surface distance value is counted as 0.
[0025] Optionally, the boundary-aware loss of each voxel is determined based on a normalized value of the surface distance value of each voxel.
[0026] Optionally, the training and optimization of the second model branch are performed according to the category of the tumor.
[0027] Optionally, the first boundary-aware loss function is represented as follows:
[0028]
[0029] wherein l ba is used to represent the boundary-aware loss, h is used to represent the number of transverse pixels of the three-dimensional MRI, w is used to represent the number of longitudinal pixels of the three-dimensional MRI, d is used to represent the number of slices of the three-dimensional MRI, k is used to represent the total number of categories of the three-dimensional MRI, and f (h,w,d,k) is used to represent the surface distance value of the (h, w, d, k)th voxel predicted by the second model branch, is used to represent the real surface distance value of the (h, w, d, k)th voxel.
[0030] Optionally, the second model branch is trained and optimized based on the average value of the boundary-aware loss of all voxels in the three-dimensional MRI.
[0031] Optionally, the first model branch is trained and optimized based on a first segmentation loss function, and the first segmentation loss function comprises a cross-entropy loss function and a dice loss function.
[0032] In a third aspect, an electronic device is provided, which includes a memory, at least one processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the method of any one of the above first aspect when executing the computer program.
[0033] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program, when executed by a processor, implements the method of any one of the above first aspect.
[0034] In a fifth aspect, a computer program product is provided, which, when running on an electronic device, causes the electronic device to execute the method of any one of the above first aspect.
[0035] In the present application, the segmentation of the tumor in the three-dimensional MRI by the first segmentation model can be divided into the prediction of the tumor category by the first model branch and the prediction of the tumor surface distance field based on boundary perception by the second model branch. Based on the second model branch, the three-dimensional surface of the tumor in the three-dimensional MRI can be perceived, so as to capture the details of the tumor boundary and the surrounding area, which helps to improve the accuracy of tumor segmentation. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced.
[0037] Figure 1 is a schematic diagram of MRI of different tumors / organs provided by an embodiment of the present application.
[0038] Figure 2 is a schematic diagram of the flow of the three-dimensional MRI tumor segmentation method based on boundary perception provided by an embodiment of the present application.
[0039] Figure 3 is a structural schematic diagram of the first segmentation model provided by an embodiment of the present application.
[0040] Figure 4 is a structural schematic diagram of the first segmentation model provided by another embodiment of the present application.
[0041] Figure 5 is an example diagram of the tumor surface distance field provided by an embodiment of the present application.
[0042] Figure 6 is a structural schematic diagram of a two-stage training baseline provided by an embodiment of the present application.
[0043] Figures 7A-7D is a schematic diagram of the result of a tumor diagnosis experiment provided by an embodiment of the present application.
[0044] Figure 8 is a schematic diagram of the result of a tumor diagnosis experiment provided by another embodiment of the present application.
[0045] Figure 9 is a schematic diagram of the structure of a tumor diagnosis device provided by an embodiment of the present application.
[0046] Figure 10 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0047] The technical solutions of the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0048] Accurate morphological quantification of tumors, including tumor size, location, and type, is expected to enhance monitoring protocols and optimize treatment planning strategies. In the past few years, a series of complex tumor segmentation methods based on convolution, attention, and large models have been proposed. These methods perform well on various medical images, thanks to the availability of large-scale datasets. However, these methods still face great challenges in segmenting certain tumors, such as spinal cord tumors, with low accuracy. Therefore, for these tumors, there is still a lack of automatic models that can accurately segment and identify tumors.
[0049] Taking spinal cord tumors as an example, spinal cord tumors are an important factor leading to morbidity and mortality in the nervous system. The spinal cord is a key component of the central nervous system and is essential for somatosensory and motor functions. Common types of spinal cord tumors include astrocytoma, ependymoma, hemangioblastoma, and meningioma. Spinal cord tumors are mainly identified through three-dimensional magnetic resonance imaging (MRI) scans. In the clinical field, it is still a daunting challenge to distinguish different types of spinal cord tumors, which can lead to misdiagnosis, failure of tumor growth monitoring, and subsequent delay in treatment intervention.
[0050] The main reason why spinal cord tumors are difficult to accurately segment by the above segmentation methods is that spinal cord tumors often have a small size, different locations, and shapes. However, the above segmentation methods mainly focus on distinguishing shapes with relatively large morphologies, such as brain tumors, left / right ventricles, abdominal organs, etc., while ignoring the challenging problem of spinal cord tumor identification. For example, Figure 1As shown, common spinal cord tumors located in the long cylindrical spinal cord often have small size, different locations and shape changes, while other tumors / organs show relatively uniform and larger size. This means that due to the obvious anatomical differences between spinal cord tumors and other body organs / tumors, the above segmentation method simply applied to spinal cord tumor segmentation will result in low accuracy.
[0051] Based on this, Figure 2 The embodiment of the present application provides a three-dimensional MRI tumor segmentation method based on boundary perception, and the method comprises steps S210-S220.
[0052] In step S210, three-dimensional MRI of multiple tumor patients is acquired as a tumor data set, and the three-dimensional MRI can include gadolinium-enhanced longitudinal relaxation time weighted sagittal image T1SC slices. The type of tumor of the tumor patient is not limited, for example, it can be a spinal cord tumor. The resolution, thickness and number of T1SC slices in the three-dimensional MRI are not limited, and the three-dimensional morphology of the tumor patient and the tumor can be restored based on the T1SC slices. It is worth noting that the T1SC slices in the tumor data set can be preprocessed. For example, the T1SC slices are scaled, scan intensity normalized, etc.
[0053] In step S220, the three-dimensional MRI in the tumor data set is taken as the input of the first segmentation model for tumor segmentation, and the class label of the tumor in the three-dimensional MRI is obtained.
[0054] The first segmentation model can be a model based on a convolutional neural network (CNN). With the popularity of convolutional neural networks, early methods mainly use CNN-based frameworks for medical image segmentation, including U-Net, 3D-UNet, Y-Net, KiU-Net, U-Net++, U-Net 3+, nnUNet, and many other variants. Among them, nnUNet is a general segmentation framework that can automatically configure the network framework and settings to extract features at multiple scales, and has shown particularly strong performance on various medical data sets. As an example, the first segmentation model can select nnUNet as the backbone network.
[0055] The first segmentation model can include a first model branch and a second model branch, the first model branch can be used to predict the class label, and the second model branch can be used to predict the surface distance field of the tumor based on boundary perception. The surface distance field is used to describe the nearest distance of each voxel in the three-dimensional MRI to the tumor surface. For example, the first model branch and the second model branch can be a convolutional neural layer in the first segmentation model, respectively.
[0056] To improve the segmentation performance of two-dimensional natural images, previous studies integrated boundary information into the learning process. However, they mainly focused on regressing edge pixels only, while ignoring the details of the shape surface in a wider context. Some recent studies utilized signed distance for medical image segmentation. However, they either focused on simple two-dimensional cases or required complex morphing or transformation processes. In addition, they also did not consider the boundary changes across multiple classes. It is worth noting that the second model branch in the present application can learn the three-dimensional surface distance field of different types of tumors respectively. Taking spinal cord tumors as an example, the second model branch can learn the three-dimensional surface distance field of four types of spinal cord tumors respectively, thereby helping to focus on the boundary changes across multiple classes.
[0057] The training and optimization of the first segmentation model can be based on the training and optimization of the first model branch and the second model branch. For example, the first model branch and the second model branch can be trained and optimized based on the respective loss functions. Among them, the training and optimization of the first model branch can be performed with reference to the manual tumor annotation of the three-dimensional MRI of the tumor patient by the tumor expert, and the training and optimization of the second model branch can be performed with reference to the measurement of the nearest distance from each voxel in the three-dimensional MRI to the tumor surface.
[0058] As an example, the architecture of the first segmentation model can be as shown in Figure 3 where the backbone network is the first segmentation model, the segmentation branch is the first model branch, and the surface distance field branch is the second model branch. The segmentation branch and the surface distance field branch are trained and optimized based on loss function 1 and loss function 2, respectively.
[0059] In the present application, the segmentation of the tumor in the three-dimensional MRI by the first segmentation model can be divided into the prediction of the tumor class by the first model branch and the prediction of the tumor surface distance field based on boundary perception by the second model branch. Based on the second model branch, the three-dimensional surface of the tumor in the three-dimensional MRI can be perceived, thereby capturing the details of the tumor boundary and the surrounding area, which helps to improve the accuracy of tumor segmentation.
[0060] Optionally, the second model branch can be trained and optimized based on a first boundary-aware loss function, which can be used to determine a boundary-aware loss for each voxel in the three-dimensional MRI, and the boundary-aware loss for each voxel can be determined based on a surface distance value for the voxel, which can be used to represent a closest distance of the voxel to the tumor surface. For example, for each voxel, a difference between a closest distance of the voxel to the tumor surface predicted by the second model branch and a closest distance of the voxel to the tumor surface actually measured can be used to determine a boundary-aware loss for the voxel. That is, the first boundary-aware loss function can be solved based on the difference. Based on the first boundary-aware loss function, the accuracy of the surface distance field of the tumor predicted by the second model branch can be improved, which can help to have a more accurate perception of the details of the tumor and the surrounding region.
[0061] Optionally, the voxels in the three-dimensional MRI can include voxels outside the tumor (which can also be referred to as background voxels) and voxels inside the tumor, and the surface distance values of the voxels outside the tumor can be represented as negative values, and the surface distance values of the voxels inside the tumor can be represented as positive values. That is, the tumor surface (i.e., the tumor boundary) is a demarcation with a surface distance value of 0. Based on this, it is helpful to distinguish the voxels outside and inside the tumor, so as to accurately determine the boundary of the tumor.
[0062] Optionally, if the absolute value of the surface distance value of a voxel outside the tumor is greater than the maximum surface distance value of the voxels inside the tumor, the surface distance value of the voxel is counted as 0. That is, when the surface distance value of a voxel outside the tumor meets the above condition, it is truncated to 0. Generally, when the absolute value of the surface distance value of a voxel outside the tumor is greater than the maximum surface distance value of the voxels inside the tumor, it indicates that the voxel is far away from the tumor boundary and belongs to a background voxel. The background voxel is no longer significant for perceiving the tumor boundary, and therefore can not be included in the tumor surface distance field. Based on this, more attention can be paid to the voxels around the tumor boundary, so as to help to more accurately determine the boundary of the tumor.
[0063] Optionally, the boundary-aware loss for each voxel can be determined based on a normalized value of the surface distance value of the voxel. As an example, the normalization of the surface distance value of each voxel can be based on the maximum surface distance value of all voxels in the three-dimensional MRI. For example, the normalized value of the surface distance value of each voxel = (surface distance value / maximum surface distance value + 1) / 2. Based on the normalization of the surface distance value, the calculation efficiency of the boundary-aware loss can be improved.
[0064] Optionally, the training and optimization of the second model branch can be performed according to the categories of the tumors. Based on the foregoing, the second model branch can learn the three-dimensional surface distance field of different categories of tumors respectively, and thus the training and optimization of the second model branch can also be performed according to the categories of the tumors, thereby helping the second model branch to pay more attention to the boundary changes across multiple categories more accurately.
[0065] Optionally, the first boundary-aware loss function can be expressed as follows:
[0066]
[0067] wherein l ba may be used to represent the boundary-aware loss, h can be used to represent the number of transverse pixels of the three-dimensional MRI, w can be used to represent the number of longitudinal pixels of the three-dimensional MRI, d can be used to represent the number of slices of the three-dimensional MRI, k can be used to represent the total number of categories of the three-dimensional MRI, f (h,w,d,k) may be used to represent the surface distance value of the (h, w, d, k) voxel predicted by the second model branch, may be used to represent the real surface distance value of the (h, w, d, k) voxel.
[0068] It is worth noting that with the predicted tumor surface distance value and the corresponding real surface distance value, a simple method is to select as the loss function to optimize each voxel equally. However, since the spinal cord tumor only occupies a relatively small area, most of the surface distance values are zero, which belong to the truncated background voxels. In order to make the boundary-aware loss really focus on the tumor boundary area, the term is introduced in the present application. Based on this, the model can be driven to pay more attention to the voxels with non-zero distance values and belonging to the boundary area of interest, thereby helping to more accurately determine the boundary of the tumor.
[0069] Optionally, the second model branch can be trained and optimized based on the average value of the boundary-aware loss of all voxels in the three-dimensional MRI. That is, the total boundary-aware loss can take the average value of the boundary-aware loss of all voxels.
[0070] Optionally, the first model branch is trained and optimized based on the first segmentation loss function, and the first segmentation loss function includes a cross-entropy loss function and a dice loss function. For example, the cross-entropy loss function can be determined based on formula (2). For another example, the dice loss function can be determined based on formula (3).
[0071]
[0072]
[0073] wherein, wherein, represents a cross-entropy loss function, represents a Dice Loss function, represents a predicted segmentation result, represents a true segmentation, Pi, k is the output of the flexible max-pooling function of the network, representing the probability that the ith voxel is classified as the kth class; g is a one-hot encoding corresponding to the true segmentation map, and the dimensions of p and g are both N x K; i e N, representing the number of voxels in a given patch or batch; k e K, representing the total number of classes in the segmentation task; according to the true segmentation, g is set to 1 only when the ith voxel does indeed belong to the kth class; otherwise g is set to 0. i,k i,k
[0074] Further, the first segmentation model can be jointly trained with the above three losses (total boundary-aware loss, cross-entropy loss, dice loss). For example, as shown in equation (4), the total loss of the first segmentation model can be simply calculated by equating the weights of the three losses.
[0075] l = l ce + l dice + l ba (4).
[0076] The boundary-aware three-dimensional MRI tumor segmentation method based on the first segmentation model will be described below with reference to the drawings and taking a spinal cord tumor as an example.
[0077] Spinal cord tumor dataset
[0078] Data collection: The spinal cord tumor dataset contains three-dimensional MRI of 653 patients, collected from October 2017 to September 2023, and is specifically used for preoperative assessment of untreated intervention of spinal cord tumors. This dataset contains comprehensive spinal manifestations (covering the neck, chest and waist areas), including four main types of spinal cord tumors: 1) meningioma (247 patients), 2) ependymoma (203 patients), 3) astrocytoma (101 patients) and 4) hemangioblastoma (102 patients), as shown in the first row of Table 1. Figure 1
[0079] The MRI of each patient consists of gadolinium-enhanced T1 SC slices. As shown in Table 1, the original resolution of the sagittal magnetic resonance scans in this dataset varies, with in-plane resolution ranging from 0.34 to 1.06 millimeters and slice thickness ranging from 1.5 to 8 millimeters. The number of slices in the MRI in the sagittal plane varies from a minimum of 9 to a maximum of 36.
[0080] Table 1 Detailed information of magnetic resonance scans in the spinal cord tumor dataset
[0081]
[0082] Tumor Annotation: Each patient's MRI is accompanied by expert-designed manual tumor annotation. Two independent neuroradiologists manually label the actual tumor area on gadolinium-enhanced T1SC. If there is any dispute regarding the delineation of the tumor area, the final decision is made by a senior neuroradiologist to ensure the highest fidelity and accuracy of the data annotation.
[0083] Data partitioning: To enable a more comprehensive and equitable comparison, the spinal cord tumor dataset was evenly divided into five distinct parts to facilitate robust 5-fold cross-evaluation of various methods and future research. Table 2 summarizes the details of these five parts.
[0084] Table 2 Number of subjects in the five sections
[0085]
[0086] Data preprocessing: Following standardized protocols widely used in existing work, sagittal slices were rescaled to 0.47 mm in both the anteroposterior and vertical directions, and to 3.3 mm in the lateral direction. MRI scan intensity values were trilinearly interpolated, and tumor annotations were performed using nearest-neighbor interpolation. The intensity of each volume was normalized by first subtracting the mean intensity value from each voxel and then dividing by the standard deviation of the scan.
[0087] Segmentation model
[0088] like Figure 4 As shown, given an input three-dimensional magnetic resonance scan T∈R H×W×D Where H×W represents the slice resolution and D represents the number of slices. This is input into the backbone network, i.e., the first segmentation model (e.g., nnUNet), to obtain the multi-class segmentation output S∈R for each voxel. H×W×D×K , where K represents the total number of categories. In this example, K is set to 5, representing the 4 types of spinal cord tumors and the background category. Let one convolutional neural layer in nnUNet be the segmentation branch (i.e., the first model branch). In parallel with the segmentation branch, add another branch (i.e., the second model branch) to predict the tumor surface distance field F∈R. H×W×D×K The algorithm employs a single convolutional neural layer. The segmentation branch is supervised using widely adopted cross-entropy and Dice loss, and uses ground truth semantic labels. The surface distance field is supervised using multi-class boundary-aware loss (i.e., the first boundary-aware loss function), and uses ground truth labels.
[0089] Tumor surface distance field
[0090] As Figure 1 shown in the first row, spinal cord tumors exhibit particularly challenging shape variations. This motivates the introduction of tumor surface distance fields to learn the complex and fine details of tumor boundaries and surrounding regions in three-dimensional magnetic resonance scans. The definition of tumor surface distance fields is introduced in the following Figure 5 section with the example of two-dimensional T1 SC slices of spinal cord tumors.
[0091] As Figure 5 shown, given a three-dimensional magnetic resonance scan slice t∈R H× W, which belongs to a patient with an arbitrary type of spinal cord tumor, and its corresponding ground truth tumor mask s∈{0, 1} H×W , where 1 denotes foreground tumor pixels and 0 denotes background. For any particular pixel p i within the tumor mask, one can compute its closest distance (with positive sign) to the tumor mask boundary, denoted as d i . For any particular pixel p j outside the tumor mask, one can also compute its closest distance (with negative sign) to the tumor mask boundary, denoted as d j .
[0092] Since spinal cord tumors usually only occupy a relatively small region, one can choose to truncate the distance values to zero when the absolute distance d j of pixels outside the tumor region is larger than the maximum distance value within the tumor mask. The formal definition is:
[0093]
[0094] where {d0…d i …d I} denotes all distance values within the tumor mask, and is the indicator function. This simple truncation will allow the network to focus on the details around the surface boundary.
[0095] Finally, all untruncated distance values are normalized in the range [0, 1] as shown in equations (6) and (7), while keeping all previously truncated values to be zero.
[0096]
[0097]
[0098] It is worth noting that the definition of tumor surface distance fields on the above two-dimensional slices can be directly extended to three-dimensional MRI. That is, the surface distance value of any voxel in three-dimensional MRI is the closest distance of the voxel to the tumor boundary in three-dimensional space.
[0099] Similarly, given a three-dimensional MRI T and its ground truth tumor mask S, for each voxel in the three-dimensional MRI, its surface distance value can be computed by measuring the nearest distance to the three-dimensional tumor boundary, then truncated and normalized, obtaining the ground truth tumor surface distance values for the entire volume, denoted as It will supervise the newly added network branch. It is worth noting that, considering that the boundary shape of each type of spinal cord tumor tends to be very different, the ground truth surface distance values are defined per class, rather than in a class-agnostic manner.
[0100] Loss function
[0101] Cross-entropy loss function and Dice loss function: the segmentation branch is supervised by the cross-entropy loss function and the Dice loss function, denoted as l ce and l dice , the calculation methods are shown in equations (2) and (3) respectively.
[0102] Boundary-aware loss function: the predicted tumor surface distance field is supervised by the boundary-aware loss function, the calculation method is shown in equation (4).
[0103] Implementation
[0104] All experiments were performed on a single NVIDIA 3009 graphics processing unit (GPU). The model was trained for 1000 epochs with a learning rate of 0.01, and the learning rate was decayed by 0.00001 for each epoch. The backbone network nnUNet contains 7 layers, each containing two convolution operations, when it is trained on the spinal cord tumor dataset based on its automatically searched configuration. The first two layers use a kernel size of 1x1x1, while the subsequent layers use a kernel size of 3x3x3. The stride configuration of the first layer is 1x1x1, the stride configuration of the fourth layer is 2x2x2, and the stride configuration of the remaining layers is 1x2x2. The number of feature channels of all layers is set to 32, 64, 128, 256, 320, 320, and 320.
[0105] In addition to using 5-fold cross-validation to evaluate all models on the spinal cord tumor dataset, all methods can also be evaluated on the public KNIGHT dataset (a kidney dataset), as kidney tumors have similarities with spinal cord tumors. According to the searched configuration, when the backbone network is trained on the public KNIGHT dataset, the network contains six layers, each containing two 3x3x3 convolution operations. The stride configuration of the first layer is 1x1x1, and the stride configuration of the remaining layers is 2x2x2. The number of feature channels of all layers is set to 32, 64, 128, 256, 320, and 320.
[0106] Experiments
[0107] To validate the effectiveness of the method of the present embodiment, it can be compared with other tumor segmentation methods.
[0108] With the success of attention mechanisms and vision transformers in capturing long-range contextual information, many works extended the attention-based methods into CNN-based frameworks to achieve better medical segmentation, including Trans U-Net, UNETR, Swin UNETR, nnFormer, 3D UX-Net, and many other variants. Recently, large segmentation base models have made great progress in natural images, and many subsequent works extended them to the field of medical image segmentation. Benefiting from powerful network architectures and large training datasets, these methods perform well at the cost of a large amount of computational resources and delicate and fast engineering skills.
[0109] With the development of segmentation models, methods specifically designed for spinal cord tumor segmentation have also emerged. In one study, a dual-cascade 3D CNN for chordoma segmentation was proposed. However, the segmented tumors are located within the spinal column region, which shows different intensities and sizes and is juxtaposed with different tissue types. Therefore, this segmentation model designed for the spine is not suitable for spinal cord tumor segmentation in this example due to the inherent differences. In another study, a two-stage cascade architecture consisting of two U-Net models was used to localize the spinal cord and segment the tumor, respectively, but it has deficiencies in determining multiple types of tumors due to its weakness in distinguishing fine shapes.
[0110] Based on this, the present embodiment selects the following two groups of baselines in medical image segmentation for comparison, and all models are trained from scratch.
[0111] End-to-end training baseline: includes established state-of-the-art models nnFormer, 3D UX-Net, Swin UNETR, and nnUNet.
[0112] Two-stage training baseline: for the problem of multi-class 3D segmentation per voxel, a simple procedure is to apply a binary per-voxel segmentation (tumor and background) followed by a per-voxel classification model that takes as input a 3D image with / without the predicted binary tumor mask, as Figure 6The same backbone nnUNet was chosen for binary segmentation in phase 1 for fair comparison. In phase 2, two powerful models were chosen: 3DResNet101 and the encoder part of nnUNet (denoted as UEnc), both followed by a multilayer perceptron (MLP) layer with (1024-512-256-K) neurons. When training in phase 2, the three-dimensional mask with or without estimated class-agnostic tumors was also inputted for comparison, denoted as or
[0113] In summary, there were 4 baselines in the experiment: 1) 2) 3) 4)
[0114] Metrics: Following the Multi-Centre, Multi-Vendor & Multi-Disease Cardiac Image Segmentation Challenge (M&MS Challenge), Dice coefficient and 95th percentile Hausdorff Distance (HD) were used as metrics in this example. For missing predictions, the maximum value of Hausdorff distance was 450 mm.
[0115] Dataset: As mentioned before, in addition to evaluating all models on the spinal cord tumor dataset using 5-fold cross-validation, all methods were also evaluated on the public KNIGHT dataset. This dataset contains 400 three-dimensional computed tomography (CT) scans, split into a training set of 300 scans and a test set of 100 scans. Each voxel is labeled with one of two tumor classes: NoAT (No Adjuvant Therapy) and CanAT (Candidate for Adjuvant Therapy), or the background class. The resolution of all CT scans was rescaled to 2 mm x 2 mm x 2 mm. Intensities were trilinearly interpolated, and tumor annotations were interpolated by nearest neighbor. The intensities of each volume were standardized by first subtracting the mean intensity value from each voxel and then dividing by the standard deviation of the scan.
[0116] Experimental results on spinal cord tumor dataset
[0117] Table 3 compares the quantitative results of all baselines and the methods of this example on the spinal cord tumor dataset, which are the average of 5-fold cross-validation. Detailed quantitative results for each part are provided in Tables 4-8. Qualitative results are shown in Figures 7A-7DThe results are shown in Table 3-Table 8 and Figures 7A-7D It can be seen that:
[0118] - In the first set of baselines, nnUNet outperforms other attention-based methods with better Dice scores and Hausdorff distances.
[0119] - In the second set of baselines, the classifier based on nnUNet encoder achieves higher Dice scores than the model based on ResNet101.
[0120] - Compared with all these powerful baselines, the method of the present application achieves significantly better performance in both Dice scores and Hausdorff distances. It is worth noting that compared with the backbone network nnUNet, the method of the present application achieves 10% higher Dice scores and 53.1mm better Hausdorff distances, which proves the effectiveness of the tumor surface distance field learned by the boundary-aware loss proposed in the present application.
[0121] - It can also be noted that the Dice score of the tumor type astrocytoma (35.2%) is significantly lower than other types. A key reason is that astrocytoma and ependymoma look similar, leading to frequent misclassification between these two tumor types. Therefore, the Dice score of ependymoma is also relatively low (57.8%).
[0122] Table 3 Dice scores (%) and Hausdorff distances (mm) on spinal cord tumor dataset using 5-fold cross-validation
[0123]
[0124] Table 4 Dice scores (%) and Hausdorff distances (mm) on spinal cord tumor dataset Fold1
[0125]
[0126]
[0127] Table 5 Dice scores (%) and Hausdorff distances (mm) on spinal cord tumor dataset Fold2
[0128]
[0129] Table 6 Dice scores (%) and Hausdorff distances (mm) on spinal cord tumor dataset Fold3
[0130]
[0131]
[0132] Table 7 Dice score (%) and Hausdorff distance (mm) on the spinal cord tumor dataset Fold4
[0133]
[0134] Table 8 Dice score (%) and Hausdorff distance (mm) on the spinal cord tumor dataset Fold5
[0135]
[0136] Results on KNIGHT dataset
[0137] Table 9 compares the quantitative results of all methods on the KNIGHT dataset, with qualitative results as shown in Figures 7A-7D Table 10. Based on Table 9 and Figures 7A-7D the following can be obtained:
[0138] - In both groups of baselines, while nnUNet and show excellent performance, the method of the present application outperforms them, achieving a Dice score of 40.7% and a Hausdorff distance of 226.8 mm.
[0139] - It can also be noted that most of the two-stage methods tend to classify all subjects into the “NoAT” category. Therefore, they show high performance in the “NoAT” class, but produce a Dice score close to zero in the “CanAT” class.
[0140] On both datasets, as shown in Figures 7A-7D , the tumor boundaries predicted by the method of the present application show greater consistency with the ground truth. In contrast, other baselines often predict larger tumors, resulting in a large number of false positive voxels. In addition, some baselines tend to predict multiple sub-regions with different tumor types, while the method of the present application consistently generates more accurate masks with fine boundaries and correct tumor types.
[0141] Table 9 Dice score (%) and Hausdorff distance (mm) on the KNIGHT dataset
[0142]
[0143] Ablation experiments
[0144] To verify the effectiveness of the design of the present application, the following groups of ablation studies were conducted on the public KNIGHT dataset.
[0145] - Group 1: To validate the design of the tumor surface distance field, four different settings are chosen in this example: 1) the distance field is neither truncated nor normalized; 2) the distance field is truncated but not normalized; 3) the distance field is truncated and normalized, which is the setting of the method of the present embodiments; 4) the boundary-aware loss l ba in (5) is replaced by the l2 term ; 5) the l1 term in the boundary-aware loss l ba is replaced by the cross-entropy loss l . ce
[0146] - Group 2: For the truncation strategy defined in (5), three settings are validated in this example: the distance d j is truncated at one / two / three times of max{d0…d i …d I}, respectively. Intuitively, the wider the truncation range, the larger the region that needs to be focused on, and thus the less efficient it can be.
[0147] - Group 3: This example further validates the design of the multi-class distance field. To make comparison, the multi-class awareness is simplified to class-agnostic distance field in this example. This means that the surface distance field branch in (5) has an output shape of R H×W×D×1 .
[0148] - Group 4: For the boundary-aware loss defined in (1), the effectiveness of the l term is validated in this example.
[0149] Tables 10-13 compare the ablation experiment results of Groups 1 / 2 / 3 / 4, respectively, Figure 8 showing the qualitative results. Based on Tables 10-13 and Figure 8 , it can be concluded that:
[0150] - In Table 10, the truncation design of the present embodiments improves the Dice score by 5.9%, and normalizing the truncated distance field further improves the Dice score by 9.5%. Nevertheless, the use of l ba based l ce or e2 is only slightly better than that of l ba based l (h,w,d,k) .
[0151] - In Table 11, the wider the region to be truncated, the worse the results obtained. The reason is that the fine details near the tumor boundary can be more important than the pixels far away from the boundary.
[0152] - In Table 12, the multi-class distance field in the present embodiments simplifies the learning of complex boundaries, thus providing a clearer understanding of different tumor types, which improves the Dice score by 4.3% compared to the class-agnostic setting.
[0153] In Table 13, The term-driven model pays more attention to voxels with non-zero distance values, such as voxels of the tumor boundary region of interest, which improves the Dice score by 1%. If the gradient calculation in the training process is stopped , the Dice score decreases by 2%.
[0154] Table 10 Ablation experiment of the first group on the KNIGHT data set
[0155]
[0156] Table 11 Ablation experiment of the second group on the KNIGHT data set
[0157]
[0158]
[0159] Table 12 Ablation experiment of the third group on the KNIGHT data set
[0160]
[0161] Table 13 Ablation experiment of the fourth group on the KNIGHT data set
[0162]
[0163] Based on the above experiments, it can be determined that the method of the embodiments of the present application has high segmentation accuracy, and is obviously better than the existing baseline on the spinal cord tumor data set and the public kidney tumor segmentation data set.
[0164]
[0165] The above describes the method provided by the embodiments of the present application in detail, and the device provided by the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0166] Referring to Figure 9 , a structural schematic diagram of a tumor diagnosis device provided by the embodiments of the present application. As Figure 9 indicated, the tumor diagnosis device 900 can include an acquisition module 910 and a segmentation module 920.
[0167] The acquisition module 910 is configured to acquire three-dimensional MRIs of multiple tumor patients as a tumor data set, and the three-dimensional MRIs include gadolinium-enhanced T1 SC slices.
[0168] The segmentation module 920 is configured to perform tumor segmentation by taking the three-dimensional MRIs in the tumor data set as the input of a first segmentation model, to obtain the category label of the tumor in the three-dimensional MRIs.
[0169] The first segmentation model is a model based on a convolutional neural network (CNN), and the first segmentation model includes a first model branch and a second model branch, the first model branch is used to predict a class label, and the second model branch is used to predict a surface distance field of the tumor based on boundary perception, the surface distance field is used to describe a nearest distance from each voxel in the three-dimensional MRI to a surface of the tumor, and training and optimization of the first segmentation model are based on training and optimization of the first model branch and the second model branch.
[0170] Optionally, the second model branch is trained and optimized based on a first boundary perception loss function, and the first boundary perception loss function is used to determine a boundary perception loss of each voxel in the three-dimensional MRI, and the boundary perception loss of each voxel is determined based on a surface distance value of each voxel, and the surface distance value is used to represent the nearest distance from each voxel to the surface of the tumor.
[0171] Optionally, the voxels include voxels outside the tumor and voxels inside the tumor, the surface distance value of the voxels outside the tumor is represented as a negative value, and the surface distance value of the voxels inside the tumor is represented as a positive value.
[0172] Optionally, if an absolute value of the surface distance value of the voxels outside the tumor is greater than a maximum surface distance value of the voxels inside the tumor, the surface distance value is counted as 0.
[0173] Optionally, the boundary perception loss of each voxel is determined based on a normalized value of the surface distance value of each voxel.
[0174] Optionally, the training and optimization of the second model branch are performed according to the category of the tumor.
[0175] Optionally, the first boundary perception loss function is represented as the following formula:
[0176]
[0177] wherein, l ba is used to represent the boundary perception loss, h is used to represent a transverse pixel number of the three-dimensional MRI, w is used to represent a longitudinal pixel number of the three-dimensional MRI, d is used to represent a slice number of the three-dimensional MRI, k is used to represent a total number of categories of the three-dimensional MRI, and f (h,w,d,k) is used to represent a surface distance value of the (h, w, d, k)th voxel predicted by the second model branch, is used to represent a real surface distance value of the (h, w, d, k)th voxel.
[0178] Optionally, the second model branch is trained and optimized based on an average value of the boundary perception loss of all voxels in the three-dimensional MRI.
[0179] Optionally, the first model branch is trained and optimized based on a first segmentation loss function, and the first segmentation loss function comprises a cross-entropy loss function and a dice loss function.
[0180] Embodiments of the present application further provide an electronic device. As shown in the figure, the electronic device 1000 comprises at least one processor 1001, a memory 1002, and a computer program 1003 stored in the memory 1002 and executable on the at least one processor 1001, wherein the processor 1001 implements the method provided by the present application when executing the computer program 1003. Figure 10
[0181] For example, the computer program 1003 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 1002 and executed by the processor 1001 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, and the instruction segments are used to describe the execution process of the computer program in the electronic device 1000.
[0182] Those skilled in the art can understand that Figure 10 The electronic device 1000 is only an example and does not constitute a limitation on the electronic device, and can include more or fewer components than those shown, or combine certain components, or different components, for example, the electronic device 1000 can also include an input / output device, a network access device, a bus, etc.
[0183] The processor 1001 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0184] The memory 1002 can be an internal storage unit of the electronic device 1000, for example, a hard disk or a memory of the electronic device 1000. The memory 1002 can also be an external storage device of the electronic device 1000, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like, equipped on the electronic device 1000. Further, the memory 1002 can include both an internal storage unit and an external storage device of the electronic device 1000. The memory 1002 is used to store a computer program and other programs and data required by the electronic device 1000. The memory 1002 can also be used to temporarily store data that has been output or will be output.
[0185] The electronic device 1000 provided by the embodiment can execute the method embodiments described above, and the implementation principles and technical effects are similar, which will not be described here.
[0186] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in each of the above method embodiments can be implemented.
[0187] The embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can execute the steps in each of the above method embodiments.
[0188] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application can implement all or part of the processes in the above method embodiments by a computer program to instruct related hardware to complete. The computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each of the above method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the photographing device / electronic device, recording medium, computer memory, ROM (read only memory), RAM (random access memory), CD-ROM (compact disc read-only memory), magnetic tape, floppy disk, and optical data storage device, etc. The computer readable storage medium mentioned in the present application can be a non-volatile storage medium, in other words, a non-transitory storage medium.
[0189] In the above embodiments, the description of each embodiment is focused on, and the parts not described or recorded in a certain embodiment can be referred to the relevant description of other embodiments.
[0190] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software mode depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0191] In the embodiments provided in the present application, it should be understood that the disclosed apparatuses / devices and methods can be implemented in other ways. For example, the above-described apparatus / device embodiments are only schematic. The division of the modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0192] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0193] It should also be understood that the term "and / or" as used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations thereof, and includes these combinations.
[0194] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when" or "upon" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted as meaning "upon determining" or "in response to determining" or "upon detecting [a described condition or event]" or "in response to detecting [a described condition or event]" depending on the context.
[0195] In addition, in the description of the present application and the appended claims, the terms "first", "second", and the like are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0196] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A boundary-aware based three-dimensional magnetic resonance imaging (MRI) tumor segmentation method, characterized in that, The method comprises: obtaining three-dimensional MRI of multiple tumor patients as a tumor dataset, the three-dimensional MRI comprising gadolinium-enhanced longitudinal relaxation time weighted sagittal images T1SC slices; performing tumor segmentation on the three-dimensional MRI in the tumor dataset as input of a first segmentation model to obtain a class label of the tumor in the three-dimensional MRI; wherein the first segmentation model is a model based on a convolutional neural network (CNN), the first segmentation model comprising a first model branch and a second model branch, the first model branch being used for predicting the class label, and the second model branch being used for predicting a surface distance field of the tumor based on boundary perception, the surface distance field being used for describing the nearest distance of each voxel in the three-dimensional MRI to the surface of the tumor, and training and optimization of the first segmentation model being based on training and optimization of the first model branch and the second model branch.
2. The method of claim 1, wherein, The second model branch is trained and optimized based on a first boundary perception loss function, the first boundary perception loss function being used for determining a boundary perception loss of each voxel in the three-dimensional MRI, the boundary perception loss of each voxel being determined based on a surface distance value of the voxel, the surface distance value being used for representing the nearest distance of the voxel to the surface of the tumor.
3. The method of claim 2, wherein, The voxels include voxels outside the tumor and voxels inside the tumor, the surface distance value of the voxels outside the tumor being represented as a negative value, and the surface distance value of the voxels inside the tumor being represented as a positive value.
4. The method of claim 3, wherein, If the absolute value of the surface distance value of the voxels outside the tumor is greater than the maximum surface distance value of the voxels inside the tumor, the surface distance value is counted as 0.
5. The method of claim 4, wherein, The boundary perception loss of each voxel is determined based on a normalized value of the surface distance value of the voxel.
6. The method of claim 5, wherein, The training and optimization of the second model branch are performed according to the class of the tumor.
7. The method of claim 6, wherein, The first boundary perception loss function is represented as follows: wherein the l ba for representing the boundary-aware loss, the h is used for representing the number of transverse pixels of the three-dimensional MRI, the w is used for representing the number of longitudinal pixels of the three-dimensional MRI, the d is used for representing the number of slices of the three-dimensional MRI, the k is used for representing the total number of categories of the three-dimensional MRI, the f (h,w,d,k) for representing the surface distance value of the (h, w, d, k) voxel predicted by the second model branch, the for representing the real surface distance value of the (h, w, d, k) voxel.
8. The method of claim 7, wherein, The second model branch is trained and optimized based on the average value of the boundary perception loss of all voxels in the three-dimensional MRI.
9. The method of claim 8, wherein, The first model branch is trained and optimized based on a first segmentation loss function, the first segmentation loss function comprising a cross-entropy loss function and a dice loss function.
10. A tumor diagnostic apparatus characterized by comprising: The method comprises: an obtaining module, configured to obtain three-dimensional magnetic resonance imaging (MRI) of multiple tumor patients as a tumor dataset, the three-dimensional MRI comprising gadolinium-enhanced longitudinal relaxation time weighted sagittal images T1SC slices; a segmentation module, configured to perform tumor segmentation on the three-dimensional MRI in the tumor dataset as input of a first segmentation model to obtain a class label of the tumor in the three-dimensional MRI; The first segmentation model is a model based on a convolutional neural network (CNN), and the first segmentation model includes a first model branch and a second model branch. The first model branch is configured to predict the category label, and the second model branch is configured to predict a surface distance field of a tumor based on boundary perception. The surface distance field is configured to describe a nearest distance from each voxel in the three-dimensional MRI to a surface of the tumor. Training and optimization of the first segmentation model is based on training and optimization of the first model branch and the second model branch.
Citation Information
Patent Citations
CT image segmentation method and device
CN114782472A
Brain tumor segmentation method based on brain three-dimensional MRI image design
CN115496771A