A Multi-scale U-Net Medical Image Segmentation Method Based on the Joint Spatial Domain

Through the cascade-Cartesian product coordinate spatial segmentation network and multi-scale Atrous convolution, the problem of neglecting coordinate system influence in existing medical image segmentation models is solved, and higher segmentation accuracy and rotation invariance are achieved, which is suitable for the field of medical image segmentation.

CN115760874BActive Publication Date: 2025-07-04CHENGDU TIANHE YICHENG TECH SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211422825.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-07-04
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

When the existing medical image segmentation model is trained in polar coordinate systems and Cartesian coordinate systems, the combined influence of different spatial coordinate systems is ignored, resulting in a reduction in segmentation accuracy and the encoding module fails to effectively utilize multi-scale spatial information.

Method used

A multi-scale U-Net model based on joint spatial domain is adopted, and a cascade-Cartesian product coordinate spatial segmentation network is combined with multi-scale Atrous convolution and residual connection to achieve rotation invariance and translation invariance, and a polar coordinate center point prediction network is used for image conversion and feature fusion.

Benefits of technology

The accuracy of medical image segmentation is improved, especially in the segmentation effect of elliptical shape targets, the rotation invariance and translation invariance of the segmentation model are enhanced, the number of parameters is reduced, and the overfitting problem is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760874B_ABST
    Figure CN115760874B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-scale U-Net medical image segmentation method based on the joint spatial domain, which includes obtaining original medical image data; obtaining the central point coordinates of the original medical image by using a polar coordinate central point prediction network; converting the original medical image into a polar coordinate medical image according to the central point coordinates of the medical image; constructing a multi-scale U-Net network model based on the joint spatial domain and training the model by using the polar coordinate medical image; and generating a medical image segmentation result by using the trained multi-scale U-Net network model based on the joint spatial domain. The present invention adopts a multi-layer dilated convolution encoding module to achieve multi-scale content fusion, and uses the central point and polar coordinates to achieve an attention mechanism and rotational invariance, thereby improving the segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and particularly relates to a multi-scale U-Net medical image segmentation method based on a joint spatial domain. Background Art

[0002] Medical image segmentation is crucial for the diagnosis of related diseases and the formulation of surgical plans. Common application scenarios of medical image segmentation are to identify single structures with an elliptical shape or a similar shape distribution, such as most organs, skin lesions, polyps, cancers, etc. Among them, colorectal cancer (CRC) ranks third in cancer incidence and second in global cancer mortality. The better segmentation of rectal cancer and the staging of rectal cancer are closely related to the segmentation of the elliptical rectum. Since U-Net achieved high precision in image segmentation in 2015, a large number of U-Net variants have been generated, such as DENSEUNET, ResUNet which proposed a U-Net structure with a residual module, and applied the U-Net model to 3D images.

[0003] However, these network models ignore that for the elliptical objects to be segmented, better segmentation performance can be obtained in the polar coordinate system, such as Dense-Unet, ResUNet, Double-Unet, etc. Although some models include the use of the polar coordinate system, they are only trained in one coordinate system, ignoring the joint influence of different spatial coordinate systems on the final segmentation result. As seen in Figure 1 , when there are large differences in the pixels in the internal area of the rectum, the model in the polar coordinates can obtain better performance, such as Figure 1 (b). However, in some cases, when the segmentation model is in the Cartesian coordinate system similar to Figure 1 (a), the segmentation model can also produce good results. The results show that both coordinate systems have a certain influence on the final result. DDNet considers two coordinate systems, but combines them in parallel, which damages the attention effect brought by the polar coordinate transformation. Moreover, most of the encoding modules of these models do not pay attention to multi-scale spatial information, resulting in a reduction in segmentation accuracy. Szegedy et al. only utilize multi-scale spatial information in the last encoding layer and use the Inception module with too many parameters to obtain multi-scale information.

[0004] As is well known, in deep learning, there are two methods to improve the accuracy of the model. The first method is to deepen the depth of the model, that is, the number of layers, but this may bring huge computational overhead. The second method is to increase the width of the network (like the number of convolutional kernels in one layer), but if the width is too large, the model will have a large number of parameters and is prone to overfitting when the amount of training data is insufficient.

[0005] The basic method to solve the above problems is to introduce sparsity. For example, the Inception structure proposed in GoogleNet uses convolutional and pooling layers of different scales to extract features from the output of the previous layer, then combines the results to form the input of the next layer of the network, and uses 1*1 convolution to extract features from the previous layer and reduce the dimensionality of the output of this layer. The Residual block was proposed by Kaiming He et al. In ResNet. Compared with general deep neural networks, the Residual block is defined on two interconnected layers of the network. That is, in the Residual block, the data is not directly input into the non-linear transformation unit, but the elements are added to the original input and then non-linear transformation is performed. The reason for this method is to allow errors to propagate backward all the time. To better accelerate the training speed of the Inception network, the Google team proposed Inception-ResNet. This network combines the advantages of the above two networks and replaces the pooling layer in the traditional Inception structure with a residual connection. Inspired by

[0006] Inception-ResNet and the U-Net structure, the CE-Net proposed a DAC module, adding dense Atrous convolution on the basis of predecessors. The DAC module can capture wider and deeper semantic features through multi-scale dilated convolution and embed them after the last layer of the U-Net encoding module. However, it can be noted that the CE-Net only uses the DAC module once, and the other encoding modules are still traditional residual connection modules, and a large amount of semantic information is also lost after multiple convolutional downsamplings. In addition, due to the large width of the DAC module, if it is directly applied to multi-layer encoding, it will bring a large number of parameters and is prone to overfitting.

[0007] In medical image segmentation, a polar coordinate network has been proposed to improve the accuracy of the segmentation model. In 2018, to enhance the training data, the method of polar coordinate transformation was used to enhance the training data, obtaining different polar coordinate origins and transforming the original image. The original image was converted into different polar coordinate images through these origins, thus increasing the amount of training data. Kim et al. designed a user-guided segmentation method. The expert selects a point as the polar coordinate origin, and then uses a convolutional neural network (CNN) to segment the transformed image. When segmenting the optic disc and the optic cup, a neural network M-Net was designed. The network uses the existing automatic optic disc center localization. Then, based on the detected optic disc center, the fundus image is converted into the polar coordinate system and input into the M-Net network. The M-Net network is a U-Net network that simply integrates the multi-scale idea.

[0008] In DDNet in 2019, it was mentioned that CNN has translational invariance in the Cartesian domain, that is, for any pixel in the image, the feature vector learned through CNN convolution in this coordinate domain is translationally invariant. In the polar coordinate domain, the feature vector learned by CNN is rotationally invariant. Therefore, in order to better segment the optic disc and optic cup, DDNet contains two encoding branches. One branch takes the image in the Cartesian product coordinate system as input, and the other branch takes the same input image after polar coordinate transformation as input. The two neural network branches proceed in parallel, which means that the encoding results of the same layer will perform feature fusion during encoding, merge into a single feature vector, and then be input into the decoding module to form the segmentation result. However, a defect of this network is that the Cartesian coordinate system branch does not obtain the origin of the features, thus losing the attention effect brought by the polar coordinate transformation (because the polar coordinate transformation is based on the origin. The origin of the target area makes the polar coordinate image pay more attention to the area near the target area). Although translational invariance of segmentation can be obtained, it damages the final fusion result of the two branches. In addition, in 2017, a polar coordinate transformation network for image classification was proposed, which consists of a polar coordinate predictor and a neural network that uses a heat map to predict the origin of the target. Then, the centroid of the heat map is used as the origin of the polar coordinate transformation. However, this network is used for image classification. Therefore, in 2021, based on this network, Marin et al. proposed a method called "Polar Image Transformation Training". This network also includes an origin predictor and will convert the image into polar coordinates according to the predicted origin. Then it is used as the input of the U-Net segmentation network. But this network only utilizes the characteristics of polar coordinates and ignores the translational invariance of the Cartesian product.

[0009] The above content is only used to assist in understanding the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0010] In view of the above deficiencies in the prior art, the present invention provides a multi-scale U-Net medical image segmentation method based on the joint spatial domain.

[0011] In order to achieve the above invention object, the technical solution adopted by the present invention is:

[0012] A multi-scale U-Net medical image segmentation method based on the joint spatial domain, comprising the following steps:

[0013] S1. Obtain the original medical image data;

[0014] S2. Use the polar coordinate center point prediction network to obtain the center point coordinates of the original medical image;

[0015] S3. Convert the original medical image into a polar coordinate medical image according to the center point coordinates of the medical image;

[0016] S4. Construct a multi-scale U-Net network model based on the joint spatial domain and use polar coordinate medical images for model training;

[0017] S5. Use the trained multi-scale U-Net network model based on the joint spatial domain to generate medical image segmentation results.

[0018] Optionally, the polar coordinate center point prediction network is specifically an encoder-decoder network based on a stacked hourglass structure, and the output of each stack is fed back as input to the next stack.

[0019] Optionally, step S3 specifically includes the following sub-steps:

[0020] S31. Reset the target center point to the center point of the image data and determine the coordinates of the center point after transformation;

[0021] S32. Perform a linear coordinate transformation on the corresponding coordinate points in the rectangular coordinate system according to the source sample points and the target regular grid in the polar coordinate system;

[0022] S33. Construct a sampling grid of the grid generator according to the linear coordinate transformation parameters, sample the original medical image by the sampler using the sampling grid, and splice it with the original medical image to obtain the final polar coordinate medical image.

[0023] Optionally, the formula for resetting the target center point to the center point of the image data in step S31 is:

[0024] d = min(x0, y0, w - x0, h - y0)

[0025] x ∈ (x0 - d, x0 + d)

[0026] y ∈ (y0 - d, y0 + d)

[0027] where d is the minimum distance between the center point of the target to be segmented and the four sides of the image, (x0, y0) is the center point coordinates of the original image, w and h are the width and height of the original image, and (x, y) are the coordinates of any pixel point after transformation.

[0028] Optionally, the formula for performing a linear coordinate transformation on the corresponding coordinate points in the rectangular coordinate system in step S32 is:

[0029]

[0030] where are the source sample point coordinates in the polar coordinate system, are the corresponding coordinate points in the rectangular coordinate system.

[0031] Optionally, the multi-scale U-Net network model based on the joint spatial domain specifically includes:

[0032] A cascaded first segmentation sub-network and second segmentation sub-network;

[0033] The first segmentation sub-network is used to perform image segmentation on the polar coordinate medical image, re-convert the segmentation result into a rectangular coordinate medical image, and input it into the second segmentation sub-network after feature fusion with the input polar coordinate medical image;

[0034] The second segmentation sub-network is used to perform image segmentation on the input fused medical image, and perform feature fusion on the segmentation result with the segmentation result of the first segmentation sub-network to obtain the final medical image segmentation result.

[0035] Optionally, the first segmentation sub-network specifically includes:

[0036] An encoder and a decoder that make up the U-Net network structure;

[0037] The encoder includes two parallel first dilated convolution channels and second dilated convolution channels, and two parallel first convolution channels and second convolution channels;

[0038] The first dilated convolution channel includes consecutive convolution layers with a first consecutive dilation rate, and is used to perform convolution calculation on the input polar coordinate medical image according to the receptive field determined by the first consecutive dilation rate;

[0039] The second dilated convolution channel includes consecutive convolution layers with a second consecutive dilation rate, and is used to perform convolution calculation on the input polar coordinate medical image according to the receptive field determined by the second consecutive dilation rate, and splice the convolution result with the convolution result of the first dilated convolution channel to obtain a first splicing result;

[0040] The first convolution channel includes a convolution layer, and is used to splice the convolution result with the first splicing result to obtain a second splicing result;

[0041] The second convolution channel includes a convolution layer, and is used to perform feature fusion on the convolution result with the second splicing result to obtain the output result of the encoder.

[0042] Optionally, the calculation formulas for determining the receptive field according to the first consecutive dilation rate and determining the receptive field according to the second consecutive dilation rate are:

[0043] F = 2(rate - 1)*(k - 1) + k

[0044] Where F is the determined receptive field size, rate is the consecutive dilation rate, and k is the convolution kernel size.

[0045] Optionally, the multi-scale U-Net network model based on the joint spatial domain uses a loss function based on deep supervision for model training, expressed as:

[0046]

[0047] where L o ut is the loss function of the model, and L1 loss is the loss value of the first segmentation sub-network, and L2 loss is the loss value of the second segmentation sub-network.

[0048] The present invention has the following beneficial effects:

[0049] (1) The present invention proposes a joint spatial domain network model (Joint U-Net). It is a cascaded polar-Cartesian product coordinate space segmentation network. This network can not only obtain the rotational invariance of convolution, but also has the translational invariance of convolution in the Cartesian coordinate system under the polar coordinate network. The cascaded network structure can avoid some negative impacts on the final segmentation result when the Cartesian coordinate systems are in parallel.

[0050] (2) The present invention designs the AIR module, which consists of multi-scale Atrous convolutions, enabling the convolution process to have a larger receptive field, reducing the number of parameters, and obtaining more semantic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a comparison diagram of the rectal segmentation results of the existing segmentation network in the polar coordinate and Cartesian coordinate systems;

[0052] Figure 2 is a schematic flowchart of a multi-scale U-Net medical image segmentation method based on the joint spatial domain in an embodiment of the present invention;

[0053] Figure 3 is a schematic diagram of the multi-scale U-Net medical image segmentation process based on the joint spatial domain in an embodiment of the present invention;

[0054] Figure 4 is a schematic diagram of the output result of the polar coordinate origin predictor in an embodiment of the present invention;

[0055] Figure 5 is a schematic diagram of the internal structure of the coordinate conversion module Diff-CTP in an embodiment of the present invention;

[0056] Figure 6 is a schematic diagram of the AIR encoding module in an embodiment of the present invention;

[0057] Figure 7 is a schematic diagram of the effect of the dilated convolution in the AIR encoding module in an embodiment of the present invention;

[0058] Figure 8 Schematic diagram of the experimental results of the multi-semantic segmentation network in the embodiments of the present invention;

[0059] Figure 9 Schematic diagram of the segmentation results on the rectal data set in the embodiments of the present invention;

[0060] Figure 10 Schematic diagram of the segmentation results on the skin lesion data set in the embodiments of the present invention. Detailed implementation manners

[0061] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0062] As Figure 2 and Figure 3 shown, the embodiments of the present invention provide a multi-scale U-Net medical image segmentation method based on the joint spatial domain, including the following steps S1 to S5:

[0063] S1. Obtain the original medical image data;

[0064] S2. Use the polar coordinate center point prediction network to obtain the center point coordinates of the original medical image;

[0065] In an optional embodiment of the present invention, the polar coordinate center point prediction network adopted by the present invention uses a series of stacked encoder-decoder networks. The neural network based on the stacked hourglass structure was initially used for human pose estimation and can capture information at various scales. The output of each specific stack is a heat map for predicting the center of the object. The output of each stack is fed back as input to the next stack, thereby allowing continuous improvement of the heat map prediction. The loss function of the network is the output loss of each stack, and then averaged to perform corresponding deep supervision. The final predicted heat map is the output of the last stack in the network. To predict the center point, the present invention uses 8 stacked hourglass neural structures to predict the center point of the target area. To predict the center point of the image target G(x, y), first, the corresponding matrix image M needs to be obtained. The calculation formula of this matrix is:

[0066] M ij = ∑(x, y)·x i ·y j

[0067] where i and j represent the rows and columns in the matrix M.

[0068] The centroid (c x , c y ) of the image is obtained using the following formula:

[0069]

[0070] Each of the eight hourglass stacks in the center point prediction network will generate a heat map. The coordinates of the pixel with the maximum intensity in the heat map output by the last hourglass structure will be used as the center point of the final predicted target area. Figure 4 An example of the corresponding center point heat map is shown in , where (a) is the original image input and (b) is the corresponding center point heat map.

[0071] S3. Convert the original medical image into a polar coordinate medical image according to the center point coordinates of the medical image;

[0072] In an alternative embodiment of the present invention, step S3 specifically includes the following sub-steps:

[0073] S31. Reset the target center point to the center point of the image data and determine the converted center point coordinates;

[0074] S32. Perform a linear coordinate transformation on the corresponding coordinate points in the rectangular coordinate system according to the source sample points and the target regular grid in the polar coordinate system;

[0075] S33. Construct a sampling grid of the grid generator according to the linear coordinate transformation parameters, sample the original medical image using the sampler with the sampling grid, and splice it with the original medical image to obtain the final polar coordinate medical image.

[0076] Specifically, the present invention maps the features on the Cartesian product grid to the features on the polar coordinate grid, and this transformation requires the use of the predicted target center point. To simplify the input parameters, the target center point is reset to the center point of the image data using the following formula:

[0077] d = min(x0, y0, w - x0, h - y0)

[0078] x ∈ (x0 - d, x0 + d)

[0079] y ∈ (y0 - d, y0 + d)

[0080] where d is the minimum distance between the center point of the target to be segmented and the four sides of the image, (x0, y0) are the center point coordinates of the original image, w and h are the width and height of the original image, and (x, y) are the coordinates of any pixel point after conversion. The converted center point coordinates become (w / 2, h / 2).

[0081] Then, according to the source sample points in the polar coordinate system And the target rule grid, for the corresponding coordinate points in the rectangular coordinate system Perform a linear coordinate transformation, expressed as:

[0082]

[0083] Wherein, Is the coordinate of the source sample point in the polar coordinate system, Is the corresponding coordinate point in the rectangular coordinate system.

[0084] Figure 5 Is the specific structure of the Diff-CTP coordinate conversion module. First, input U to obtain transformation parameters according to the above formula to construct a sampling grid, where (θ, R) corresponds to The grid generator obtains the mapping relationship T ((θ,R)) After that, the sampler uses the sampling grid and the input feature map U as inputs at the same time to obtain the final feature map transformation result V.

[0085] S4. Construct a multi-scale U-Net network model based on the joint spatial domain, and use polar coordinate medical images for model training;

[0086] In an optional embodiment of the present invention, the multi-scale U-Net network model constructed by the present invention specifically includes:

[0087] Cascaded first segmentation sub-network and second segmentation sub-network;

[0088] The first segmentation sub-network is used to perform image segmentation on the polar coordinate medical image, and re-convert the segmentation result into a rectangular coordinate medical image, and input it into the second segmentation sub-network after feature fusion with the input polar coordinate medical image;

[0089] The second segmentation sub-network is used to perform image segmentation on the input fused medical image, and perform feature fusion on the segmentation result with the segmentation result of the first segmentation sub-network to obtain the final medical image segmentation result.

[0090] Wherein, the first segmentation sub-network specifically includes:

[0091] An encoder and a decoder constituting a U-Net network structure;

[0092] The encoder includes two parallel first dilated convolution channels and second dilated convolution channels, and two parallel first convolution channels and second convolution channels;

[0093] The first dilated convolution channel includes consecutive convolution layers with a first consecutive dilation rate, and is used to perform convolution calculation on the input polar coordinate medical image according to the receptive field determined by the first consecutive dilation rate;

[0094] The second dilated convolutional channel includes consecutive convolutional layers with a second consecutive dilation rate, which are used to perform convolutional calculations on the input polar coordinate medical image according to the receptive field determined by the second consecutive dilation rate, and splice the convolutional result with the convolutional result of the first dilated convolutional channel to obtain a first splicing result;

[0095] The first convolutional channel includes a convolutional layer, which is used to splice the convolutional result with the first splicing result to obtain a second splicing result;

[0096] The second convolutional channel includes a convolutional layer, which is used to perform feature fusion on the convolutional result and the second splicing result to obtain the output result of the encoder.

[0097] Specifically, the Joint U-Net designed in the present invention includes a total of two segmentation sub-networks. The first segmentation sub-network, Multi-Content P-UNet, trains the input image in polar coordinates. Based on U-Net, it replaces the simple encoding module in U-Net with a multi-scale spatial fusion module to obtain more semantic information. Then, the second segmentation sub-network Cart U-Net is connected in series. This network also uses the U-Net network structure, but trains the image in the Cartesian coordinate system. The final segmentation is the result of the first network, and the second network is only used to make our first network have better performance and has no impact on the final segmentation result. The reason for using the connected network structure is as follows:

[0098] (1) The feature map output by the first segmentation sub-network can be further improved by obtaining the original input image and the corresponding segmentation mask again.

[0099] (2) To obtain rotational invariance in the polar coordinate system and translational invariance of the segmentation result in the Cartesian coordinate system. Because if this process is performed in parallel, such as in DDNet, the Cartesian coordinate branch will not obtain the origin value information of the image, thus losing the attention effect brought by polar coordinate transformation and having a negative impact on the final feature fusion.

[0100] Figure 6Shows the multi-scale spatial fusion module and the corresponding channel changes, where the encoding module combines the ideas of dilated convolution [23, 24, 25, 26], Inception [17, 27, 28], and residual [19, 29]. However, the large convolution kernels in Inception are converted into several consecutive dilated convolutions with smaller convolution kernels. First, for the input feature map Input, it is calculated using two parallel dilated convolution branches. The consecutive dilation rates R of the two branches are (1, 2) and (1, 2, 3) respectively. The corresponding maximum receptive fields can be calculated to be 7 and 15 respectively using the following formula. Where rate is the dilation rate, k is the convolution kernel size, and F is the output receptive field size. The reason for using consecutive convolutions is to alleviate the jagged segmentation results brought by dilated convolution, as Figure 7 shown, where (a) R = 1, (b) R = 2, (c) R = 3.

[0101] F = 2(rate - 1)*(k - 1) + k

[0102] After performing dilated convolution, the convolution result D and the input Input after convolution will be combined in a channel concatenation manner to retain more scale information. Finally, using the idea in the resBlock, the input input and the output result of the foregoing process are pixel-wise added. For the L-th layer convolution, the overall process of the AIR module can be expressed as:

[0103] out L = cat(D L , Conv(input L )) + input L The present invention designs a multi-scale spatial fusion module that can expand the receptive field like the Inception encoding block without a large increase in parameters and can focus on more scale information. Although dilated convolution expands the receptive field, it is still a 3*3 convolution operation, as Figure 7 shown.

[0104] In addition, the present invention also uses two methods to alleviate the gradient loss caused by the overly deep network and enhance the final segmentation result.

[0105] (1) The output of the Multi-Content P-UNet will be reconverted into an image in the Cartesian coordinate system after passing through the Diff-PTC. At this time, after feature fusion with the original input, it is input into the Cart U-Net. After feature fusion of the output of the first sub-network and the output of the second network, the final output is used as the corresponding segmentation result.

[0106] (2) Using the idea of deep supervision, a loss function is designed as shown in the following formula. Where L1 lossis the loss value of the first sub-network, where L2 loss is the loss value of the second sub-network. L o ut is the final loss function:

[0107]

[0108] S5. Use the trained multi-scale U-Net network model based on the joint spatial domain to generate medical image segmentation results.

[0109] Next, the effectiveness of the above-mentioned multi-scale U-Net medical image segmentation method based on the joint spatial domain provided by the present invention will be verified in combination with specific experimental data sets.

[0110] 1. Experimental data sets and related metrics

[0111] The experimental data sets are provided by Zhejiang Sir Run Run Shaw Hospital. The training set contains 2,203 rectal MRI images of 219 patients. The validation set contains 477 MRI images of rectal cancer from 51 patients. And the training set contains 468 MRI images of the rectum from 50 patients, as shown in Table 1.

[0112] Table 1 Information of rectal data set

[0113]

[0114]

[0115] In addition, to verify its universality, experiments were also carried out on the ISIC 2018 skin lesion segmentation data set, which contains 2,694 dermoscopic images of skin lesions. We adjusted the size of each image to 384×512 and divided the training set, validation set, and test set according to the ratio of 8:1:1, and then normalized each image to the range of [-0.5, 0.5].

[0116] To verify the effectiveness of the AIR module, we carried out relevant experiments on PyTorch 1.7.1 and NVIDIA GeForce RTX2090 GPU. The optimizer is Adam with a learning rate of {10}^{-4}, the Batch size is 4, the epoch is 300, and checkpoints are used to store the model after each epoch to obtain the best validation loss. The dice coefficient is used as the loss function, as shown in the following formula:

[0117] Diceloss = 1 - (2|X∩Y| + α) / (|X| + |Y| + α)

[0118] Where X is the label of the input image, Y is the predicted value output by the model. α is the smoothing parameter, which is set to 1 in this experiment.

[0119] In this experiment, we set four metrics to evaluate the model, namely the Dice coefficient, the mean intersection over union mIoU, precision, and recall. Precision and accuracy are both calculated at the pixel level.

[0120] 2. Experiment on the AIR Encoding Module

[0121] For the AIR module, we combined it with U-Net and compared it with DenseBlock+U-Net, ResBlock+U-Net, etc. This experiment was conducted under the polar coordinate grid. The experimental results are shown in Tables 2 and 3.

[0122] Table 2 Comparison Results of the AIR Module and Other Encoding Modules on the Rectum Dataset

[0123]

[0124]

[0125] Table 3 Comparison Results of the AIR Module and Other Encoding Modules on the Skin Dataset

[0126]

[0127] As can be seen from Tables 2 and 3, our experimental results are the highest in DICE, MIOU, and precession. The visualization of the corresponding results is as Figure 8 shown, where (a) U-Net (b) U-Net+resBlock (c) U-Net+DenseBlock (d) U-Net+Inception (e) U-Net+AIR.

[0128] For Experiment 4.2, we conducted a total of five experiments and compared them with the most popular encoding modules DenseBlock, ResBlock, and Inception. The Inception module consists of convolutional kernels of 1*1, 3*3, and 5*5 and max pooling. From Tables 2 and 3, we can analyze that compared with these encoding modules, our AIR module obtained the best results in the Dice coefficient, MIOU, and precession. In addition, the center point of the feature map is also obtained from the label. If the center point is obtained from the center point prediction network we trained, relevant experiments show that the segmentation accuracy generally drops by about 0.01 to 0.02. This conclusion also applies to Experiment 4.3. The visualization of the segmentation results of our Experiment 4.2 is as Figure 8As shown, it can be seen that U-Net has obvious jitter at the segmentation boundary, while ResBlock and DenseBlock are more likely to segment into non-target areas. Our AIR module has better segmentation performance in most cases, but there are also problems with boundary segmentation when the pixel differences in the area to be segmented are huge.

[0129] 3. Experimental analysis of Joint U-Net

[0130] To verify the effectiveness of Joint U-Net, relevant experiments were conducted on PyTorch 1.10.0 and NVIDIA GeForce RTX3090 GPU. The optimizer is Adam, with a learning rate of 0.001 and a batch size of 6. In the experiment, the center points of the dataset were obtained through the corresponding labels.

[0131] Except that our model combines polar coordinates and rectangular coordinates, other experiments were trained and tested in polar coordinates. To verify the effectiveness of the polar-rectangular coordinate concatenation, we conducted a comparative experiment with the polar-polar coordinate concatenation network Double Unet. To ensure a single variable, we replaced the encoding structures in both Double Unet and our network Joint U-Net with ordinary convolutions. Finally, the comparative experiment results with the mainstream segmentation models on the rectal dataset are shown in Table 4, and the comparative experiment results on the skin lesion segmentation dataset are shown in Table 5. The experiment includes three different network architectures: U-Net, U-Net++

[33] +resBlock, and DeepLabV3+

[34] +resBlock.

[0132] Table 4 Comparison results on the rectal dataset

[0133]

[0134] Table 5 Comparison results on the skin lesion dataset

[0135]

[0136]

[0137] Seven experiments were conducted on the rectal dataset and the skin lesion dataset, and the results are shown in Tables 4 and 5. First, experiments on the basic network U-Net were carried out in the Cartesian coordinate system, and it can be found that its effect is the worst. After polar coordinate transformation, the segmentation accuracy of U-Net on the rectal and skin lesion datasets increased by about 0.07 to 0.10. This is sufficient to show that the polar coordinate system can better segment elliptical objects. In addition, the experimental results of Double U-Net and Joint U-Net (OURS) with a common coding structure show that the information obtained only by segmentation in the polar coordinate system is not as good as that obtained in two different coordinate systems.

[0138] The result obtained by segmentation in the unified coordinate system is better. In addition, the segmentation result is only obtained from the first sub-network, which can better avoid the loss of segmentation accuracy caused by the excessive depth of the network. Finally, we changed the coding structure of the first sub-network of the Joint U-Net (OURS) to AIR and compared it with the coding structures of U-Net++ and DeepLabv3+ whose coding modules are residual modules. The experimental results prove the effectiveness of our model, and the relevant results are shown in Tables 4 and 5. In addition, Figure 9 For the segmentation result on the rectal dataset, where (a) original image (b) label (C) double-Unet (d) Joint U-net; Figure 10 For the segmentation result on the skin lesion dataset, where (a) original image (b) label (C) double-Unet (d) Joint U-net; From Figure 9 and Figure 10 the experimental results, it can be seen that:

[0139] (1) When the boundary of the object to be segmented is not clear, our network has better segmentation accuracy, as shown in the last row of Figure 9 and the second and third rows of Figure 10 .

[0140] (2) When the boundary of the segmented object is discontinuous, our network can better capture this serrated shape, as shown in the first and fourth rows of Figure 10 .

[0141] In summary, medical image segmentation is very important for the diagnosis of related diseases. In order to reduce the annotation work of related medical images, many U-Net-based models have been proposed to achieve automatic segmentation of target regions. However, most of these models are only trained in one coordinate system, ignoring the combined effect of different spatial coordinate systems. In addition, most of the encoding modules of these models do not pay attention to multi-scale spatial information. Therefore, we propose a multi-scale U-net segmentation model based on the joint spatial domain to achieve phased segmentation of medical images. The model uses a self-designed multi-layer dilated convolutional encoding module AIR (Atrous Inception Residual Block) to achieve multi-scale content fusion. In addition, the attention mechanism and rotational invariance are achieved by using the center point and polar coordinates. The output of the polar coordinate network is retrained by the Cartesian coordinate system network to achieve translational invariance of the segmentation result. Our final segmentation result is converted from the polar coordinate network. Compared with the commonly used medical segmentation models, the DICE coefficient of this model is increased by about 2% on the rectal dataset and about 0.5% on the skin lesion dataset.

[0142] This invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide means for implementing the functions specified in Figure 1 one or more of the flows Figure 1Steps of the functions specified in one or more boxes.

[0145] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

[0146] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A multi-scale U-Net medical image segmentation method based on the joint spatial domain, characterized in that It includes the following steps: S1. Obtain the original medical image data; S2. Use the polar coordinate center point prediction network to obtain the center point coordinates of the original medical image; S3. According to the center point coordinates of the medical image, convert the original medical image into a polar coordinate medical image; The step S3 specifically includes the following sub-steps: S31. Reset the target center point to the center point of the image data to determine the converted center point coordinates; S32. Perform a linear coordinate transformation on the corresponding coordinate points in the rectangular coordinate system according to the source sample points and the target regular grid in the polar coordinate system; S33. Construct a sampling grid of the grid generator according to the linear coordinate transformation parameters, sample the original medical image by using the sampler with the sampling grid, and splice it with the original medical image to obtain the final polar coordinate medical image; S4. Construct a multi-scale U-Net network model based on the joint spatial domain, and use the polar coordinate medical image for model training; The multi-scale U-Net network model based on the joint spatial domain specifically includes: A cascaded first segmentation sub-network and second segmentation sub-network, and the first segmentation sub-network specifically includes: An encoder and a decoder that constitute a U-Net network structure; The encoder includes two parallel first dilated convolution channels and second dilated convolution channels, and two parallel first convolution channels and second convolution channels; The calculation formulas for determining the receptive field according to the first continuous dilation rate and determining the receptive field according to the second continuous dilation rate are: Among them, is the determined receptive field size, is the continuous hole rate, is the convolution kernel size; The first dilated convolution channel includes a continuous convolution layer using the first continuous dilation rate, and is used to perform convolution calculation on the input polar coordinate medical image according to the receptive field determined by the first continuous dilation rate; The second dilated convolution channel includes a continuous convolution layer using the second continuous dilation rate, and is used to perform convolution calculation on the input polar coordinate medical image according to the receptive field determined by the second continuous dilation rate, and splice the convolution result with the convolution result of the first dilated convolution channel to obtain a first splicing result; The first convolution channel includes a convolution layer, and is used to splice the convolution result with the first splicing result to obtain a second splicing result; The second convolution channel includes a convolution layer, and is used to perform feature fusion on the convolution result and the second splicing result to obtain the output result of the encoder; The first segmentation sub-network is used to perform image segmentation on the polar coordinate medical image, re-convert the segmentation result into a rectangular coordinate system medical image, and input it into the second segmentation sub-network after feature fusion with the input polar coordinate medical image; The second segmentation sub-network is used to perform image segmentation on the input fused medical image, and perform feature fusion on the segmentation result and the segmentation result of the first segmentation sub-network to obtain the final medical image segmentation result; S5. Use the trained multi-scale U-Net network model based on the joint spatial domain to generate a medical image segmentation result.

2. The multi-scale U-Net medical image segmentation method based on the combined spatial domain according to claim 1, characterized in that, The polar coordinate center point prediction network is specifically an encoder-decoder network based on a stacked hourglass structure, and the output of each stack is fed back as input to the next stack.

3. A multi-scale U-Net medical image segmentation method based on the joint spatial domain according to claim 1, characterized in that, The formula for resetting the target center point to the center point of the image data in the step S31 is: , , where d is the minimum distance between the center point of the target to be segmented and the four sides of the image, is the center point coordinate of the original image, w and h are the width and height of the original image, and (x, y) is the coordinate of any pixel point after conversion.

4. A multi-scale U-Net medical image segmentation method based on the joint spatial domain according to claim 1, characterized in that, The formula for performing linear coordinate transformation on the corresponding coordinate points in the rectangular coordinate system in step S32 is as follows: , wherein, is the coordinate of the source sample point in the polar coordinate system, is the corresponding coordinate point in the rectangular coordinate system.

5. A multi-scale U-Net medical image segmentation method based on the combined spatial domain according to claim 1, characterized in that The multi-scale U-Net network model based on the joint spatial domain uses a loss function based on deep supervision for model training, which is expressed as: Among them, is the loss function of the model, is the loss value of the first segmentation sub-network, is the loss value of the second segmentation sub-network.

Citation Information

Patent Citations

  • Pancreatic cell image segmentation method based on improved U-Net network

    CN109191471A

  • Iris image segmentation, positioning and normalization method based on multi-task neural network

    CN112287872A