A Liver Tumor Segmentation Method and Device Based on a Hybrid Neural Network
By using a hybrid neural network in the CT image segmentation of liver tumors, the 3D module and 2D module are combined to solve the problem of limited accuracy in the prior art, and efficient and accurate segmentation of liver tumor images is achieved.
Patent Information
- Application Number
- CN202210582897.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-26
AI Technical Summary
The prior art has the problem of limited accuracy in the segmentation of CT images of liver tumors. The 2D convolutional network lacks spatial context information, while the 3D convolutional network has a large amount of computation and limited application range.
Using a hybrid neural network-based method, after preprocessing the 3D CT image, continuous multi-layer 2D slices are input into a hybrid neural network model composed of 3D modules and 2D modules for segmentation. Combining the advantages of 3D modules and 2D modules, the shortcomings of a single network structure are overcome.
The precise segmentation of liver tumor images is achieved, which not only overcomes the problem of 2D convolution being unable to obtain context information, but also solves the problem of large amount of pure 3D network computing, and improves the segmentation accuracy and efficiency.
Smart Images

Figure CN115018862B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical imaging, and particularly relates to a liver tumor segmentation method and device based on a hybrid neural network. Background Technique
[0002] The accurate segmentation of the liver tumor contour is an important step in the diagnosis, surgical planning and postoperative evaluation of liver diseases. Computed tomography (CT) can provide more comprehensive information for the diagnosis and treatment of liver tumors. Due to the diverse types and complex structures of liver tumors, computer-aided diagnostic segmentation can help doctors better determine the tumor boundary. The liver tumor CT segmentation method based on deep learning has achieved obvious performance improvement compared with the traditional segmentation method and has obtained rapid development.
[0003] Most of the existing image segmentation algorithms only use the data of one phase as the training set and design 2D convolutional networks or 3D networks for training. Such methods have the following disadvantages: Different types of lesions show different characteristics in different phases. Taking liver cancer as an example, the CT enhancement of liver cancer shows a unique feature of "fast in and fast out", that is, in the early stage (arterial phase), the whole lesion reaches uniform or non-uniform high density, and then rapidly decreases to be close to the density of the liver parenchyma with increasing density. The contour is not obvious in the delayed phase and the plain scan phase. If only one phase is used for training, it is easy to miss the detection of lesions. The liver tumor segmentation methods are basically divided into two types: the segmentation method based on 2D convolutional networks and the segmentation method based on 3D networks. The 2D convolutional network is represented by 2D FCNS (Fully Convolutional Networks), and the network is composed of stacked 2D convolutions. Its input is generally a single slice or adjacent 3 or 5 slices, and the prediction result of the middle slice is output. This method lacks spatial context information for three-dimensional structures, resulting in limited segmentation accuracy. The 3D convolutional network is represented by 3D FCNS, and the network is composed of stacked 3D convolutions. Its input is three-dimensional data. Due to the large GPU video memory consumption and large computational amount of the 3D convolutional network, its application range is relatively limited. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a liver tumor segmentation method and device based on a hybrid neural network.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions.
[0006] In the first aspect, the present invention provides a liver tumor segmentation method based on a hybrid neural network, including the following steps:
[0007] Obtain a 3D CT image and preprocess the image;
[0008] Input the preprocessed continuous multi - layer 2D slices into the trained image segmentation model to segment the liver tumor image. The model is a hybrid neural network mainly composed of 3D modules located at the network input, bottom layer, and output stage, and 2D modules located at the middle layer.
[0009] Perform connected - component processing on the segmented image including the liver to obtain the final segmentation result.
[0010] Furthermore, the pre - processing of the image specifically includes:
[0011] Adjust the window width and window level of the input CT image, and normalize the pixel value of each pixel point to [0, 1] according to the following formula:
[0012]
[0013] In the formula, I(x, y) and I w (x, y) are the pixel values of the pixel point with coordinates (x, y) before and after normalization respectively, W is the window width, and C is the window level.
[0014] Resample the data with different slice thicknesses in the z - axis to 2.5mm, and keep the slice thicknesses in the x - and y - axes unchanged.
[0015] Perform a sliding window on the 3D data in the z - axis, and use adjacent N slice images as the input of the image segmentation model.
[0016] Furthermore, N = 32. Randomly generate a number i between 0 and 31, and use the continuous 32 - layer slice images starting from the i - th slice image I i as the input of the image segmentation model. i 、I i+1 、…、I i+31
[0017] Furthermore, in the test dataset of the image segmentation model, the 32 - layer slice images are I i 、I i+15×1 、I i+15×1 、I i+15×2 、…、I i+15×31 , i ≥ 0.
[0018] Furthermore, the pre - processing of the image also includes data augmentation processing: flipping the image obtained by the sliding window, randomly adding noise, and performing affine transformation.
[0019] Furthermore, the CT images in the training dataset of the image segmentation model include 4 phases, namely the plain - scan phase, arterial phase, venous phase, and delayed phase.
[0020] Further, the 2D module includes a downsampling module with a step size of 2, a 2D convolutional unit, a 2D non-local unit, and a 2D convolutional unit connected in sequence.
[0021] Further, the 3D module includes a 3D convolutional unit, an instance normalization module, and a 3D non-local unit connected in sequence.
[0022] In a second aspect, the present invention provides a liver tumor segmentation device based on a hybrid neural network, including:
[0023] An image acquisition module for acquiring a 3D CT image and preprocessing the image;
[0024] An image segmentation module for inputting the preprocessed continuous multi-layer 2D slices into a trained image segmentation model to segment the liver tumor image, and the model is a hybrid neural network mainly composed of 3D modules located at the network input, bottom layer, and output stages and 2D modules located at the middle layer;
[0025] A post-processing module for performing connected domain processing on the segmented image including the liver to obtain the final segmentation result.
[0026] Further, the preprocessing of the image specifically includes:
[0027] Adjusting the window width and window level of the input CT image, and normalizing the pixel value of each pixel point to [0, 1] according to the following formula:
[0028]
[0029] In the formula, I(x, y) and I w (x, y) are the pixel values of the pixel point with coordinates (x, y) before and after normalization respectively, W is the window width, and C is the window level;
[0030] Resampling the data with different slice thicknesses on the z-axis to 2.5 mm, and keeping the slice thicknesses on the x and y axes unchanged;
[0031] Sliding the window on the z-axis of the 3D data, and using adjacent N slice images as the input of the image segmentation model.
[0032] Furthermore, N = 32, randomly generate a number i between 0 and 31, and use the continuous 32-layer slice images starting from the i-th layer slice image I i as the input of the image segmentation model. i 、I i+1 、…、I i+31
[0033] Furthermore, in the test dataset of the image segmentation model, the 32-layer slice images are Ii , I i+15×1 , I i+15×1 , I i+15×2 , …, I i+15×31 , where \(i\geq0\).
[0034] Furthermore, the preprocessing of the image further includes data augmentation processing: flipping the image obtained by the sliding window, randomly adding noise, and performing affine transformation.
[0035] Further, the CT images in the training dataset of the image segmentation model include 4 phases, namely the plain scan phase, the arterial phase, the venous phase, and the delayed phase.
[0036] Further, the 2D module includes a downsampling module with a stride of 2, a 2D convolutional unit, a 2D non-local unit, and a 2D convolutional unit connected in sequence.
[0037] Further, the 3D module includes a 3D convolutional unit, an instance normalization module, and a 3D non-local unit connected in sequence.
[0038] Compared with the prior art, the present invention has the following beneficial effects.
[0039] The present invention obtains a 3D CT image and performs preprocessing, and inputs the preprocessed continuous multi-layer 2D slices into a trained image segmentation model to segment liver tumor images. The model mainly includes a 3D module located at the network input, bottom layer, and output stages and a 2D module located at the middle layer. The segmented liver image is processed to remove the largest connected component to obtain the final segmentation result. By constructing an image segmentation model using a hybrid neural network structure and fusing 2D convolution and 3D convolution in the same network, the present invention overcomes the problem that 2D convolution cannot obtain context information and solves the problem of large computational complexity of a pure 3D network. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of a method for segmenting liver tumors based on a hybrid neural network according to an embodiment of the present invention.
[0041] Figure 2 is a schematic structural diagram of a hybrid neural network.
[0042] Figure 3 is a schematic structural diagram of a 2D module.
[0043] Figure 4 is a schematic structural diagram of a 2D convolutional unit.
[0044] Figure 5 is a schematic structural diagram of a 2D non-local unit.
[0045] Figure 6It is a schematic structural diagram of a 3D module.
[0046] Figure 7 It is a block diagram of a liver tumor segmentation device based on a hybrid neural network according to an embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Figure 1 It is a flowchart of a liver tumor segmentation method based on a hybrid neural network according to an embodiment of the present invention, including the following steps:
[0049] Step 101: Obtain a 3D CT image and preprocess the image.
[0050] Step 102: Input the preprocessed continuous multi-layer 2D slices into a trained image segmentation model to segment the liver tumor image. The model is a hybrid neural network mainly composed of a 3D module located at the network input, bottom layer and output stage and a 2D module located at the middle layer.
[0051] Step 103: Perform a connected component processing on the segmented image including the liver to obtain the final segmentation result.
[0052] In this embodiment, step 101 is mainly used to obtain the input 3D CT image and perform preprocessing. A CT (Computed Tomography) image, namely computed tomography, is one of the main medical images. The CT image is characterized by high density resolution. To improve the image segmentation accuracy and processing speed, generally, the input 3D CT image needs to be preprocessed first. The preprocessing content mainly includes window width and window level adjustment, resampling, sliding window, etc.
[0053] In this embodiment, step 102 is mainly used for image segmentation. In this embodiment, the preprocessed image is input into a trained image segmentation network to implement image segmentation. There are mainly two types of existing image segmentation networks: one is a segmentation network based on a 2D convolutional network, and the other is a segmentation network based on a 3D network. The 2D convolutional network is represented by 2DFCNS. The network consists of stacked 2D convolutions. Its input is generally a single slice or adjacent 3 or 5 slices, and the prediction result of the middle slice is output. This method lacks spatial context information for three-dimensional structures, resulting in limited segmentation accuracy. The 3D convolutional network is represented by 3D FCNS. The network consists of stacked 3D convolutions, and its input is a whole three-dimensional data. Since the 3D convolutional network consumes a large amount of GPU video memory, the three-dimensional data is generally reduced in size by methods such as interpolation, losing information to a certain extent. In addition, due to the large computational amount of the 3D convolutional network, its application range is relatively limited. Therefore, this embodiment proposes a hybrid neural network composed of a 2D convolutional network and a 3D convolutional network. The overall structure of the hybrid neural network is as shown in Figure 2 shown. Different from the existing solutions that simply combine two complete 2D networks and 3D networks, the 3D module is located at the network input, bottom layer, and output stages, acting on the input and output of the entire network, and is used to analyze and reconstruct three-dimensional features. The remaining convolutional modules are all 2D modules.
[0054] To facilitate the understanding of the technical principle of the hybrid neural network, the following will be described in detail in conjunction with Figure 2 The input of the network is an array of five dimensions (B, 1, D, W, H). Each dimension represents the number of data input into the network each time, the channel size, the number of layers, the width, and the height respectively. After being processed by the 3D module at the input end, the output of the network is a four-dimensional array (B*D, C, W, H). Next, the 4 2D modules of the encoder perform feature extraction, and the network outputs are (B*D / 2, C, W / 2, H / 2), (B*D / 4, C2, W / 4, H / 4), (BD / 8, C4, W / 8, H / 8), (BD / 8, C8, W / 16, H / 16) respectively. The last 2D convolutional module restores the features into a five-dimensional array (B, C16, D / 8, W / 32, H / 32). The outputs of the two 3D modules at the bottom layer of the encoder are (B, C16, D / 8, W / 64, H / 64) and (B, C16, D / 8, W / 128, H / 128) respectively. Similarly, the decoder is the reverse operation of the encoder, and the output feature size of each module in the decoder is the same as the input feature size of the corresponding module in the encoder.
[0055] The hybrid neural network of this embodiment overcomes the shortcomings of the lack of spatial information in the 2D convolutional neural network and the large computational amount of the 3D convolutional neural network at the same time.
[0056] In this embodiment, step 103 is mainly used to obtain the final segmentation result by post-processing the segmented image. Post-processing refers to performing connected component processing on the segmented image that includes the liver. A human has only one liver, which is the largest organ in the human body. However, since the spleen and the liver are somewhat similar to a certain extent, it is possible to predict part of the spleen as the liver. Therefore, in this embodiment, the maximum connected component processing is performed on the image after network segmentation, the volume of each connected component is calculated, and only the part of the connected component with the largest volume is retained.
[0057] As an alternative embodiment, the preprocessing of the image specifically includes:
[0058] Adjust the window width and window level of the input CT image, and normalize the pixel value of each pixel point to [0, 1] according to the following formula:
[0059]
[0060] In the formula, I(x, y) and I w (x, y) are the pixel values of the pixel point with coordinates (x, y) before and after normalization respectively, W is the window width, and C is the window level;
[0061] Resample the data with different slice thicknesses along the z-axis to 2.5 mm, and keep the slice thicknesses of the x and y axes unchanged;
[0062] Perform a sliding window along the z-axis on the 3D data, and use N adjacent slice images as the input of the image segmentation model.
[0063] This embodiment provides a technical solution for image preprocessing. First, perform window width and window level adjustment, set the window width W and window level C. For example, set W = 175 and C = 125, and normalize the value of each pixel point to [0, 1] according to the above formula. Then, perform resampling along the z-axis to make the slice thickness along the z-axis 2.5 mm, and keep the slice thicknesses of the x and y axes unchanged. This is equivalent to listing a series of equally spaced slice images along the z-axis direction. Finally, perform a sliding window on the z-axis and select N adjacent consecutive slice images as the input of the image segmentation model.
[0064] As an alternative embodiment, N = 32, randomly generate a number i between 0 and 31, and use the continuous 32-layer slice images I i starting from the i-th slice image I i 、I i+1 、…、I i+31 as the input of the image segmentation model.
[0065] This embodiment provides a specific technical solution for determining N adjacent slice images. Taking N = 32 as an example (N can also be other integer values), the technical solution is described. Since the starting position of the 32 slices can be freely selected, first, a random integer i between 0 and (32 - 1) is generated; then, 32 consecutive slice images I i starting from the i-th slice image are selected i 、I i+1 、…、I i+31 as the input of the image segmentation model. It is equivalent to moving a sliding window with a window width of 32 to I i , and the 32 slices within the sliding window are the ones required. It should be noted that if the total number of slices is M, it should be ensured that i + 31 ≤ M.
[0066] As an alternative embodiment, in the test dataset of the image segmentation model, the 32-slice images are I i 、I i+15×1 、I i+15×1 、I i+15×2 、…、I i+15×31 , and i ≥ 0.
[0067] This embodiment provides a specific technical solution for determining N adjacent slice images when testing the image segmentation model. The method for determining N slices given in the previous embodiment is suitable for model training and prediction. When testing the model, the N slices are arranged at equal intervals, that is, the slice numbers form an arithmetic sequence. When N = 32, the 32-slice images are I i 、I i+15×1 、I i+15×1 、I i+15×2 、…、I i+15×31 , i ≥ 0, and i + 15×31 ≤ M. Here, the interval is 15. Of course, it can also be selected as other integer values, but generally it should not be too small.
[0068] As an alternative embodiment, the preprocessing of the image further includes data augmentation processing: flipping the image obtained by the sliding window, randomly adding noise, and affine transformation.
[0069] This embodiment expands the preprocessing step. In addition to the preprocessing steps given in the previous embodiments, it also includes data augmentation processing steps, such as flipping the image obtained by the sliding window, randomly adding noise, affine transformation, etc.
[0070] As an alternative embodiment, the CT images in the training dataset of the image segmentation model include 4 phases, namely, non-contrast phase, arterial phase, venous phase, and delayed phase.
[0071] In this embodiment, in order to improve the robustness and generalization ability of the image segmentation model, during model training, the training data set selects CT images of 4 phases. Different types of lesions show different characteristics in different phases. Taking liver cancer as an example, the CT enhancement of liver cancer shows a unique feature of "fast in and fast out", that is, in the early stage (arterial phase), the entire lesion reaches uniform or non-uniform high density, and then rapidly decreases to be close to the density of the liver parenchyma with increasing density. The contour is not obvious in the delayed phase and the plain scan phase. Therefore, if only CT images of one phase are used for training, it is easy to miss lesions. For this reason, this embodiment selects CT images of 4 phases, namely the plain scan phase, the arterial phase, the venous phase, and the delayed phase, for training.
[0072] As an alternative embodiment, the 2D module includes a downsampling module with a stride of 2, a 2D convolutional unit, a 2D non-local unit, and a 2D convolutional unit connected in sequence.
[0073] This embodiment presents a technical solution for the 2D module. As Figure 3 shown, the 2D module is composed of a cascaded downsampling module with a stride of 2, a 2D convolutional unit, a 2D non-local unit (EfficientNonLocal unit), and a 2D convolutional unit. The composition of the 2D convolutional unit is as Figure 4 shown. The 2D convolutional unit draws on the characteristics of mixnet[2] and ghostnet[3] at the same time. It divides the channel dimension of the unit input into two parts, and respectively passes through a grouped convolution with a convolution kernel of 5*5 and a grouped convolution with a convolution kernel of 7*7. The form of grouped convolution is used between features to reduce the number of parameters. The outputs of the two are merged with the input of the unit in the channel dimension, and after passing through a scSE attention mechanism unit and a sub-unit composed of Conv2d+BN+GELU, the module output is obtained. The composition of the 2D non-local unit is as Figure 5 shown. The input of the unit passes through three 2D convolutional operations respectively to obtain three features with halved channels of q, k, and v, and the output of this unit is obtained through transformation and multiplication operations, etc.
[0074] As an alternative embodiment, the 3D module includes a 3D convolutional unit, an instance normalization module, and a 3D non-local unit connected in sequence.
[0075] This embodiment presents a technical solution for the 3D module. As Figure 6 shown, the 3D module is composed of a 3D convolutional unit, an instance normalization module, and a 3D non-local unit connected in sequence. The structure and principle of the 3D module are similar to those of the 2D module, and will not be elaborated in detail here.
[0076] Figure 7 This is a schematic diagram of the composition of an embodiment of the present invention. The device includes:
[0077] An image acquisition module 11, configured to acquire a 3D CT image and preprocess the image;
[0078] An image segmentation module 12, configured to input the preprocessed continuous multi-layer 2D slices into a trained image segmentation model to segment a liver tumor image, where the model is a hybrid neural network mainly composed of 3D modules located at the network input, bottom layer, and output stages and 2D modules located at the middle layer;
[0079] A post-processing module 13, configured to perform a connected component processing of the segmented image including the liver to obtain a final segmentation result.
[0080] The device of this embodiment can be used to execute Figure 1 the technical solution of the method embodiment shown, and its implementation principle and technical effects are similar, so details are not described herein. The same applies to the subsequent embodiments, and no further elaboration will be provided.
[0081] As an optional embodiment, the preprocessing of the image specifically includes:
[0082] Adjust the window width and window level of the input CT image, and normalize the pixel value of each pixel point to [0, 1] according to the following formula:
[0083]
[0084] In the formula, I(x, y) and I w (x, y) are the pixel values of the pixel point with coordinates (x, y) before and after normalization, respectively, W is the window width, and C is the window level;
[0085] Resample the data with different slice thicknesses on the z-axis to 2.5 mm, and keep the slice thicknesses on the x and y axes unchanged;
[0086] Perform a sliding window on the 3D data on the z-axis, and use N adjacent slice images as the input of the image segmentation model.
[0087] As an optional embodiment, N = 32, randomly generate a number i between 0 and 31, and use the continuous 32-layer slice images I i starting from the i-th slice image I i 、I i+1 、…、I i+31 as the input of the image segmentation model.
[0088] As an optional embodiment, in the test dataset of the image segmentation model, the 32-layer slice images are I i 、I i+15×1 、I i+15×1 、I i+15×2 、…、I i+15×31 , i ≥ 0.
[0089] As an optional embodiment, the preprocessing of the image further includes data augmentation processing: flipping the image obtained by the sliding window, randomly adding noise, and affine transformation.
[0090] As an optional embodiment, the CT images in the training dataset of the image segmentation model include 4 phases, namely, plain scan phase, arterial phase, venous phase, and delayed phase.
[0091] As an optional embodiment, the 2D module includes a downsampling module with a stride of 2, a 2D convolution unit, a 2D non-local unit, and a 2D convolution unit connected in sequence.
[0092] As an optional embodiment, the 3D module includes a 3D convolution unit, an instance normalization module, and a 3D non-local unit connected in sequence.
[0093] As described above, the above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for liver tumor segmentation based on a hybrid neural network, characterized in that, Including the following steps: Obtain a 3D CT image and preprocess the image; Input the preprocessed continuous multi-layer 2D slices into a trained image segmentation model to segment the liver tumor image. The model is a hybrid neural network composed of 3D modules located at the network input, bottom layer, and output stage, and 2D modules located at the middle layer. The 2D modules in the middle layer are respectively located between the 3D module at the network input and the 3D module at the bottom layer, and between the 3D module at the bottom layer and the 3D module at the output stage. The 2D module includes a 2D convolution unit and a 2D non-local unit. The 2D convolution unit is used to divide the channel dimension of the input of this unit into two, and respectively pass through a first group of convolutions with a convolution kernel of 5*5 and a second group of convolutions with a convolution kernel of 7*7. Group convolution is used between features to reduce the number of parameters. The outputs of the first group of convolutions and the second group of convolutions are merged with the input of this unit in the channel dimension, and after passing through an scSE attention mechanism unit and a sub-unit composed of Conv2d+BN+GELU, the output of this unit is obtained; The 2D non-local unit is used to respectively pass the input of this unit through three 2D convolution operations to obtain features with the channel halved for q, k, and v, and the output of this unit is obtained through transformation and multiplication operations; Perform connected domain processing on the segmented image including the liver to obtain the final segmentation result.
2. The method for liver tumor segmentation based on a hybrid neural network according to claim 1, characterized in that, The preprocessing of the image specifically includes: Adjust the window width and window level of the input CT image, and normalize the pixel value of each pixel point to [0,1] according to the following formula: wherein, I(x, y) and I w (x, y) are the pixel values of the pixel point with coordinates (x, y) before and after normalization respectively, W is the window width, and C is the window level; Resample the data with different slice thicknesses on the z-axis to 2.5mm, and keep the slice thicknesses on the x and y axes unchanged; Perform sliding window on the 3D data on the z-axis, and use adjacent N slice images as the input of the image segmentation model.
3. The method for liver tumor segmentation based on a hybrid neural network according to claim 2, characterized in that, N = 32, randomly generate a number i between 0 and 31, and use the continuous 32-layer slice images I i starting from the i-th layer slice image I i 、I i+1 、…、I i+31 , as the input of the image segmentation model.
4. The method for liver tumor segmentation based on a hybrid neural network according to claim 3, characterized in that, In the test dataset of the image segmentation model, the 32-layer slice images are I i , I i+15×1 , I i+15×1 , I i+15×2 , …, I i+15×31 , where i ≥ 0.
5. The method for liver tumor segmentation based on a hybrid neural network according to claim 2, characterized in that, The preprocessing of the image further includes data augmentation processing: flip the image obtained by sliding window, randomly add noise, and perform affine transformation.
6. The method for liver tumor segmentation based on a hybrid neural network according to claim 1, characterized in that, The CT images in the training dataset of the image segmentation model include 4 phases, namely the plain scan phase, arterial phase, venous phase, and delayed phase.
7. The method for liver tumor segmentation based on a hybrid neural network according to claim 1, characterized in that, The 2D module includes a downsampling module with a stride of 2, a 2D convolution unit, a 2D non-local unit, and a 2D convolution unit connected in sequence.
8. The method for liver tumor segmentation based on a hybrid neural network according to claim 1, characterized in that, The 3D module includes a 3D convolution unit, an instance normalization module, and a 3D non-local unit connected in sequence.
9. A device for liver tumor segmentation based on a hybrid neural network, characterized in that, Including: An image acquisition module for obtaining a 3D CT image and preprocessing the image; An image segmentation module, which is used to input the preprocessed continuous multi-layer 2D slices into a trained image segmentation model to segment liver tumor images. The model is a hybrid neural network composed of 3D modules located at the network input, bottom layer, and output stage, and 2D modules located at the middle layer. The 2D modules at the middle layer are respectively located between the 3D module at the network input and the 3D module at the bottom layer, and between the 3D module at the bottom layer and the 3D module at the output stage. The 2D module includes a 2D convolutional unit and a 2D non-local unit. The 2D convolutional unit is used to divide the channel dimension of the input of this unit into two parts, and respectively pass through a first group of convolutions with a convolution kernel of 5*5 and a second group of convolutions with a convolution kernel of 7*7. The form of group convolution is used between features to reduce the number of parameters. The outputs of the first group of convolutions and the second group of convolutions are merged with the input of this unit in the channel dimension, and after passing through an scSE attention mechanism unit and a sub-unit composed of Conv2d+BN+GELU, the output of this unit is obtained. The 2D non-local unit is used to respectively pass the input of this unit through three 2D convolution operations to obtain three features with halved channels of q, k, and v, and obtain the output of this unit through transformation and multiplication operations. A post-processing module, which is used to perform connected domain processing on the segmented image including the liver to obtain the final segmentation result.
10. The device for liver tumor segmentation based on a hybrid neural network according to claim 9, characterized in that, The preprocessing of the image specifically includes: Adjust the window width and window level of the input CT image, and normalize the pixel value of each pixel point to [0,1] according to the following formula: wherein, I(x, y) and I w (x, y) are the pixel values of the pixel point with coordinates (x, y) before and after normalization, W is the window width, and C is the window level; Resample the data with different slice thicknesses on the z-axis to 2.5mm, and keep the slice thicknesses on the x-axis and y-axis unchanged. Perform a sliding window on the 3D data on the z-axis, and use adjacent N slice images as the input of the image segmentation model.
Citation Information
Patent Citations
Three-dimensional liver image semantic segmentation method based on context attention strategy
CN112927255A
Image processing method, device and equipment
CN114494442A