A method for segmenting lumbosacral nerve roots in magnetic resonance images
Through the combination of multi-scale amplification-clipping and attention 3D-CNN, the universality and sample imbalance of lumbosacral nerve root segmentation in magnetic resonance images is solved, and high-precision nerve root segmentation is achieved, which is suitable for image data of a variety of imaging devices.
Patent Information
- Application Number
- CN202210429046.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-22
AI Technical Summary
The prior art has insufficient universality and sample imbalance in the segmentation of lumbosacral nerve roots of magnetic resonance images, making it difficult to apply to image data acquired by different imaging devices, resulting in low segmentation accuracy.
Using a method of combining multi-scale amplification-clipping strategy and attention 3D-CNN, pseudo-3D images and masks are generated through preliminary segmentation of 3D coarse segmentation networks, and fine segmentation is used for attention 3D-CNN, the optimal scale is determined and the mask is repaired, and the segmentation accuracy is improved.
It improves the segmentation accuracy of the nerve roots in the lumbosacral area, reduces memory usage, improves network training and testing efficiency, and is suitable for image data of a variety of imaging devices.
Smart Images

Figure CN114723730B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing, and particularly relates to a method for segmenting lumbosacral nerve roots in magnetic resonance images. Background Art
[0002] Medical image segmentation is a process of automatically or semi-automatically detecting the boundaries of organs or tissues in 2D (Two-dimensional, 2D) or 3D (Three-dimensional, 3D) images. MRI (Magnetic resonance imaging) images are widely used in the clinical diagnosis and treatment of lumbosacral diseases due to their advantages such as no radiation and high contrast.
[0003] Currently, the methods for automatically segmenting lumbosacral nerve roots in medical images can be divided into two types: the first is the method based on threshold processing, and the second is the method based on 3D convolutional neural network models. The first method uses thresholding and mathematical morphology methods to segment nerve roots in a special sequence of MRI (Steady-state free precession pulse sequence). There is a very high contrast between the white cerebrospinal fluid and the black nerve roots in this sequence. The second method uses 3D convolutional neural network U-Net with CT image patches as the input of the model to segment nerve roots. This method shows the effectiveness of deep learning methods in automatic nerve root segmentation.
[0004] However, the first method based on threshold processing can only be used for MRI images of specific sequences. Since MRI images of specific sequences are rarely used clinically, this method is difficult to apply to the image data obtained by most existing imaging devices and conventional imaging methods. Therefore, this method has limited universality and a narrow application range. This is because, limited by the specific forms of images that can be obtained by imaging devices, different sequences of images reflect and characterize different tissue physical properties, so there may not be similar gray value differences between the gray values of different sequences of images. The second method based on 3D convolutional neural network models extracts image patches in the entire image range by sliding a window with a fixed step size. Since the nerve roots only occupy a small area of the entire image, most of the cropped image patches do not contain nerve roots, easily resulting in the problem of sample imbalance. The reason for "the disadvantages of the method based on threshold processing" is limited by the specific forms of images that can be obtained by imaging devices. Different sequences of images reflect and characterize different tissue physical properties, so there may not be similar gray value differences between the gray values of different sequences of images. This is because the sizes of different organ tissues are inconsistent, and extracting image patches in the entire image range by sliding a window with a fixed step size may not be suitable for the segmentation task of organ tissues with a small volume ratio.
[0005] Therefore, in view of the deficiencies of the prior art, it is very necessary to provide a method for segmenting the lumbosacral nerve roots in magnetic resonance images to solve the deficiencies of the prior art. Summary of the Invention
[0006] The purpose of the present invention is to avoid the deficiencies of the prior art and provide a method for segmenting the lumbosacral nerve roots in magnetic resonance images. This method for segmenting the lumbosacral nerve roots in magnetic resonance images can improve the segmentation accuracy.
[0007] The above object of the present invention is achieved by the following technical measures:
[0008] Provide a method for segmenting the lumbosacral nerve roots in magnetic resonance images, which is carried out through the following steps:
[0009] Step 1: Obtain the 3D lumbosacral MRI image X and the corresponding ground truth mask Y, and then segment the 3D lumbosacral MRI image X through a 3D coarse segmentation network to obtain a predicted mask.
[0010] Step 2: Perform a multi-scale magnification-cropping strategy on the 3D lumbosacral MRI image X and the ground truth mask Y obtained in Step 1 to obtain a pseudo-3D image and a pseudo-3D mask.
[0011] Step 3: Use the pseudo-3D image and the pseudo-3D mask obtained in Step 2 as the input and label of the attention 3D-CNN respectively. Segment the pseudo-3D image through the attention 3D-CNN to obtain a training pseudo-3D mask. Then determine the optimal scale of the training pseudo-3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo-3D mask to obtain unilateral nerve root mask slices. Finally, stack the unilateral nerve root mask slices on both sides to obtain a nerve root mask.
[0012] Preferably, the above Step 2 includes:
[0013] Step 2.1: Slice the 3D lumbosacral MRI image X and the ground truth mask Y obtained in Step 1 at the same height, respectively, according to the same number j and the same thickness h to obtain j slices of 3D lumbosacral MRI image X slices and j slices of ground truth mask Y slices, where 1 ≤ j ≤ 200 and j is an integer, h > 0, and enter Step 2.2;
[0014] Step 2.2: In all the 3D lumbosacral MRI image X slices and all the ground truth mask Y slices, sequentially select the 3D lumbosacral MRI image X slices and the ground truth mask Y slices at the same position from top to bottom in order. And for each side of the nerve roots in each slice of the ground truth mask Y slice, perform processing respectively, and enter Step 2.3;
[0015] Step 2.3: Define the current 3D lumbosacral MRI image X slice as the i-th slice X i , and define the current true mask Y slice as the i-th slice Y i . Then, respectively fill m i zero pixels along the length direction and n i zero pixels along the width direction for both X *,i and Y *,i so that the unilateral nerve root is centered to obtain X *,i and Y *,i . Proceed to Step 2.4, where 1 ≤ i ≤ j and i is an integer;
[0016] Step 2.4: Use the bilinear method to magnify X *,i with the maximum magnification scale being S and starting from the initial magnification scale of 1. Then, increase the magnification scale by 1 each time until the magnification scale finally reaches the S scale, and obtain the magnified image . Here, S is an integer greater than 1, and the predicted mask after being magnified by the maximum magnification scale S does not exceed the original X x,s range. Use the nearest neighbor interpolation method to magnify Y y,s with the maximum magnification scale being S and starting from the initial magnification scale of 1. Then, increase the magnification scale by 1 each time until the magnification scale finally reaches the S scale, and obtain the magnified mask . Proceed to Step 2.5;
[0017] Step 2.5: Crop the image and the mask to the same size of and respectively, where is represented by Equation (Ⅱ), is represented by Equation (Ⅲ), and correspondingly obtain the image patch and the mask patch . Then proceed to Step 2.6;
[0018]
[0019]
[0020] where (c x , c y ) is the center coordinate of the s-th scale, n i and n i are the radii of the image patch along the length and width directions respectively, and the s-th scale is the s-th scale under the S scale;
[0021] Step 2.6: Introduce a dimension scale, and for the image patch and the mask patch Stacked along the scale axis respectively to form pseudo 3D image patches for training the fine segmentation network and the pseudo 3D mask module wherein the pseudo 3D image patches are represented as a pseudo 3D image The pseudo 3D mask module is represented as a pseudo 3D mask
[0022] Preferably, the above step three includes:
[0023] Step 3.1: Use the pseudo 3D image and the pseudo 3D mask obtained in step two as the input and label of the attention 3D - CNN respectively, segment the pseudo 3D image through the attention 3D - CNN to obtain a training pseudo 3D mask, then determine the optimal scale of the training pseudo 3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo 3D mask. Define the slices of the optimal scale as unilateral nerve root mask slices, and the size of the unilateral nerve root mask slices is the same as the size of the X slice of the 3D lumbosacral MRI image;
[0024] Step 3.2: Stack the unilateral nerve root mask slices on both sides in sequence from top to bottom to obtain a nerve root mask with nerves on both sides;
[0025] Step 3.3: Repair the nerve root mask to obtain a repaired nerve root mask;
[0026] Preferably, the above step 3.1 is specifically to use the pseudo 3D image and the pseudo 3D mask obtained in step 2.6 as the input and label of the 3D U - Net after adding the attention module respectively, segment the pseudo 3D image through the attention 3D - CNN to obtain a training pseudo 3D mask, then determine the optimal scale of the training pseudo 3D mask according to the results of the validation set, and extract the optimal scale layer of the training pseudo 3D mask at the optimal scale as the prediction result of the i - th slice X i and perform an inverse operation transformation to obtain a unilateral nerve root mask target slice with the same size as the i - th slice Y i ;
[0027] Preferably, the above inverse operation transformation includes:
[0028] Step 3.1.1: Define the magnification scale corresponding to the optimal scale layer as k, perform zero - pixel padding on the length and zero - pixel padding on the width in the to obtain a padded mask Y 填,i , so that the length of the mask Y 填,i is the same as that of the enlarged mask {Y when the magnification scale in step 2.4 is the optimal scale k*,i,s} s=k The length of the mask Y 填,i The length is N x,0 + k×r, so that the mask Y 填,i The width is the same as that of the enlarged mask {Y when the enlargement scale in step 2.4 is the optimal scale k *,i,s} s=k The width of the mask Y 填,i The width is N y,0 + k×r;
[0029] Step 3.1.2. Downsample the mask Y 填,i Convert the length from N x,0 + k×r to N x,0 , and convert the width from N y,0 + k×r to N y,0 , and correspondingly obtain the mask Y 缩,i ;
[0030] Step 3.1.3. Cut m 缩,i pixels along the length and n i pixels along the width of the mask Y i to obtain the mask Y 裁,i , and the mask Y 裁,i is the target slice of the unilateral nerve root mask.
[0031] Preferably, the processing steps of the 3D U-Net after adding the attention module include:
[0032] Step A. Process the 3D spatial self-attention module through two "convolution layer + reorganization" operations respectively to obtain the reorganized feature maps as F α and F β , and both the F α and the F β are of size C×N, where N = H×W×S, and N is the number of voxels. The given size of the 3D spatial self-attention module is the feature map F of C×H×W×S, where C, H, W, W, S represent the number of channels of the feature map, the length of the feature map, the width of the feature map, and the total scale number respectively,
[0033] Step B. Perform matrix multiplication between the transpose of F α and F β , and calculate the spatial attention map A of dimension N×N using the softmax layer, which is represented by Equation (IV);
[0034]
[0035] Step C. Perform matrix multiplication between F β and the spatial attention map to obtain the post-attention map P of C×N, which is represented by Equation (V);
[0036]
[0037] Step D: Reorganize the attention map P into an output of size C×H×W×S. The feature map obtained through the attention module will first pass through a transposed convolutional layer, then be juxtaposed with the feature map of the encoder, and then enter the decoder to obtain a 3D U-Net with an added attention module.
[0038] Preferably, in the above step 3.3, the nerve root mask obtained in step 3.2 is dilated using a spherical structure with a radius of R, and then the 3D mask is eroded using a spherical structure with a radius of R to obtain a repaired nerve root mask, and R>0.
[0039] Preferably, the above X *,i The range of is obtained from formula (Ⅰ);
[0040] (N x,s , N y,s )=(N x,0 +s×r, N y,0 ) Formula (Ⅰ);
[0041] Where (N x,0 , N y,0 ) is the size of the starting X *,i , r is the magnification factor, and r is an integer greater than 1.
[0042] In step 2.3, the single nerve root is located at the center of the prediction mask obtained in step one Take the prediction mask at the same position correspondingly Slice, and take the center of the smallest rectangular border covering the prediction mask As the center position.
[0043] Preferably, the above 3D rough segmentation network is constructed as a 3D U-Net.
[0044] Preferably, the above S is 20-60.
[0045] Preferably, the above magnification factor is 10-50.
[0046] Preferably, the number of scales of the feature map of the above 3D U-Net is 3-7.
[0047] In the 3D U-Net, the number of convolutional layers at each scale is 1-3.
[0048] Preferably, the above j is 60-70.
[0049] Preferably, the above h is 0.8 mm.
[0050] Preferably, the above R is 4.
[0051] A method for segmenting the lumbosacral nerve roots of a magnetic resonance image according to the present invention is carried out through the following steps: Step 1, obtain a 3D lumbosacral MRI image X and a corresponding ground truth mask Y, and then segment the 3D lumbosacral MRI image X through a 3D rough segmentation network to obtain a predicted mask Step 2, perform a multi-scale magnification-cropping strategy on the 3D lumbosacral MRI image X and the ground truth mask Y obtained in Step 1 to obtain a pseudo-3D image and a pseudo-3D mask; Step 3, use the pseudo-3D image and the pseudo-3D mask obtained in Step 2 as the input and label of the attention 3D-CNN respectively, segment the pseudo-3D image through the attention 3D-CNN to obtain a training pseudo-3D mask, then determine the optimal scale of the training pseudo-3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo-3D mask to obtain unilateral nerve root mask slices. Finally, stack the unilateral nerve root mask slices on both sides to obtain a nerve root mask. The present invention obtains the nerve root segmentation result through the above three steps. The pseudo-3D image generated by the multi-scale magnification-cropping step in the present invention contains multi-scale information with a compact different target-background ratio compared with the 2D image, so as to improve the segmentation accuracy. At the same time, the structure of the attention 3D-CNN used in the present invention is compact, and it is more compact than the traditional 2D attention module structure, so as to reduce memory occupancy and improve the efficiency of network training and testing. Description of the Drawings
[0052] The present invention is further described with reference to the accompanying drawings, but the content in the drawings does not constitute any limitation to the present invention.
[0053] Figure 1 It is a basic flowchart of the method for segmenting the lumbosacral nerve roots of a magnetic resonance image according to the present invention.
[0054] Figure 2 It is a flowchart of the method for segmenting the lumbosacral nerve roots of a magnetic resonance image.
[0055] Figure 3 It is a structural diagram of the attention module cited in the present invention.
[0056] Figure 4 It is a schematic diagram of the convolutional neural network architecture of Example 2.
[0057] Figure 5 It is a process diagram of Example 2. Detailed Embodiments
[0058] The technical solutions of the present invention are further described in conjunction with the following embodiments.
[0059] Example 1.
[0060] A method for segmenting the lumbosacral nerve roots in magnetic resonance images, as Figure 1 and 2 shown, is carried out through the following steps:
[0061] Step 1: Obtain the 3D lumbosacral MRI image X and the corresponding ground truth mask Y, and then segment the 3D lumbosacral MRI image X through a 3D rough segmentation network to obtain a predicted mask
[0062] Step 2: Perform a multi-scale magnification-cropping strategy on the 3D lumbosacral MRI image X and the ground truth mask Y obtained in Step 1 to obtain a pseudo-3D image and a pseudo-3D mask;
[0063] Step 3: Use the pseudo-3D image and the pseudo-3D mask obtained in Step 2 as the input and label of the attention 3D-CNN respectively. Segment the pseudo-3D image through the attention 3D-CNN to obtain a training pseudo-3D mask. Then determine the optimal scale of the training pseudo-3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo-3D mask to obtain unilateral nerve root mask slices. Finally, stack the unilateral nerve root mask slices on both sides to obtain a nerve root mask.
[0064] It should be noted that the ground truth mask Y of the present invention is obtained by manual delineation, while the 3D lumbosacral MRI image X is obtained by machine acquisition.
[0065] Among them, the 3D rough segmentation network in Step 1 is constructed as a 3D U-Net.
[0066] Step 2 includes:
[0067] Step 2.1: Slice the 3D lumbosacral MRI image X and the ground truth mask Y obtained in Step 1 at the same height, respectively, according to the same number j and the same thickness h to obtain j slices of 3D lumbosacral MRI image X slices and j slices of ground truth mask Y slices, where 1 ≤ j ≤ 200 and j is an integer, h > 0, and enter Step 2.2;
[0068] Step 2.2: In all 3D lumbosacral MRI image X slices and all ground truth mask Y slices, sequentially select the 3D lumbosacral MRI image X slices and the ground truth mask Y slices at the same position from top to bottom, and process each nerve root on each side in each ground truth mask Y slice, and enter Step 2.3;
[0069] Step 2.3: Define the current 3D lumbosacral MRI image X slice as the i-th slice X i , and define the current ground truth mask Y slice as the i-th slice Y i , and process X i and Y iFill along the length m i zero pixels and padding n in width i zero pixels, so that the unilateral nerve root is located in the center and the corresponding X *,i and Y *,i , go to step 2.4, 1≤i≤j, i is an integer;
[0070] Step 2.4: X *,i Use the bilinear method to set the maximum magnification scale to S, and start the magnification with the initial magnification scale of 1, and then increase the magnification scale by 1 each time, so that the magnification scale is finally magnified to S scale, and the magnified image is obtained S is an integer greater than 1, and the predicted mask after being enlarged by the maximum enlargement scale S does not exceed the initial X *,i Range, Y *,i Use the nearest neighbor interpolation method to set the maximum magnification scale to S, and start the magnification with the initial magnification scale of 1, and then increase the magnification scale by 1 each time, so that the magnification scale is finally magnified to S scale, and the magnification mask is obtained. Go to step 2.5; S is 20 to 60;
[0071] Step 2.5: Image and mask Cut to the same size and in It is represented by formula (II): It is expressed by formula (III), and the corresponding image block is and mask module Then proceed to step 2.6;
[0072]
[0073]
[0074] Where (c x,s ,c y,s ) is the center coordinate of the s-th scale, n x and n y are the radii of the image block along the length and width directions, respectively. The s-th scale is the s-th scale under the S scale.
[0075] Step 2.6: Introduce a dimension scale and transform the image block and mask module Stacked along the scale axis to form pseudo 3D image blocks for training fine segmentation networks and pseudo 3D mask module The pseudo 3D image block Represented as a pseudo 3D image Pseudo 3D mask module Represented as a pseudo 3D mask
[0076] It should be noted that in step 2.1, in actual operation, only the middle part of the 3D lumbosacral MRI image X and the real mask Y needs to be sliced, while the upper and lower parts do not need to be sliced, and the thickness of the sliced part is the same as that of the upper and lower parts. The number of slices and the thickness of each slice are preset in the image acquisition machine.
[0077] In step 2.4, in order not to occupy too much computing memory, the total scale is preferably an integer within 20 to 60. The magnification factor is preferably an integer within 10 to 50, and the magnification factor is generally based on the standard that the predicted mask does not exceed the image range. In this embodiment, the total scale number S is set to 40, and the maximum magnification of the image can reach 10 times. This means that the pseudo 3D image contains multi-scale information with a compact ratio of different target backgrounds, which is used to improve the segmentation accuracy.
[0078] In step 2.6,
[0079] The X in step 2.4 of the present invention *,i The range of is obtained from formula (Ⅰ);
[0080] (N x,s , N y,s ) = (N x,0 + s × r, N y,0 + s × r) formula (Ⅰ);
[0081] Where (N x,0 , N y,0 ) is the size of the starting X *,i , r is the magnification factor, and r is an integer greater than 1.
[0082] It can be seen from formula (Ⅰ) that s × r pixels are added in the direction of each axis, as Figure 2 The change of X L,i to X L,M in.
[0083] Step three includes:
[0084] Step 3.1: Use the pseudo 3D image and the pseudo 3D mask obtained in step two as the input and label of the attention 3D-CNN respectively. Segment the pseudo 3D image through the attention 3D-CNN to obtain the training pseudo 3D mask. Then, determine the optimal scale of the training pseudo 3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo 3D mask. Define the slices of the optimal scale as the unilateral nerve root mask slices, and the size of the unilateral nerve root mask slices is the same as the size of the slices of the 3D lumbosacral MRI image X;
[0085] Step 3.2: Stack the unilateral nerve root mask slices on both sides in order from top to bottom to obtain a nerve root mask with nerves on both sides;
[0086] Step 3.3: Repair the nerve root mask to obtain a repaired nerve root mask;
[0087] Specifically, Step 3.1 uses the pseudo 3D image and the pseudo 3D mask obtained in Step 2.6 as the input and label of the 3D U-Net after adding the attention module. Segment the pseudo 3D image through the attention 3D-CNN to obtain a training pseudo 3D mask, then determine the optimal scale of the training pseudo 3D mask according to the results of the validation set, and extract the optimal scale layer of the training pseudo 3D mask at the optimal scale as the prediction result of the i-th slice X i and perform an inverse operation transformation to obtain a unilateral nerve root mask target slice with the same size as the i-th slice Y i .
[0088] It should be noted that the lumbosacral nerve roots consist of two almost symmetrical branches surrounding the lumbar vertebrae. For each slice of the present invention, and for the two lumbosacral nerve roots of each slice, multi-scale magnification-cropping and inverse operation transformation are required.
[0089] The inverse operation transformation of the present invention includes:
[0090] Step 3.1.1: Define the magnification scale corresponding to the optimal scale layer as k. Perform zero-pixel padding on the length and zero-pixel padding on the width in to obtain a padded mask Y 填,i such that the length of the mask Y 填,i is the same as the length of the magnified mask {Y *,i,s} s=k at the optimal scale k in Step 2.4. The length of the mask Y 填,i is N x,0 + k×r, and the width of the mask Y 填,i is the same as the width of the magnified mask {Y *,i,s} s=k at the optimal scale k in Step 2.4. The width of the mask Y 填,i is N y,0 + k×r;
[0091] Step 3.1.2: Convert the length of the mask Y 填,i from N x,0 + k×r to N x,0 through downsampling operation, and convert the width from N y,0 + k×r to N y,0, the corresponding mask Y is obtained 缩,i ;
[0092] Step 3.1.3, in mask Y 缩,i , cut m i pixels along the length and cut n i pixels along the width to obtain mask Y 裁,i . Mask Y 裁,i is the target slice of the unilateral nerve root mask.
[0093] The processing steps of the 3D U-Net after adding the attention module include:
[0094] Step A: Process the 3D spatial self-attention module through two "convolutional layer + reorganization" respectively to obtain the reorganized feature maps F α and F β . And the sizes of F α and F β are both C×N, where N = H×W×S, and N is the number of voxels. The given size of the 3D spatial self-attention module is the feature map F of C×H×W×S, where C, H, W, W, S represent the number of channels of the feature map, the length of the feature map, the width of the feature map, and the total scale number respectively.
[0095] Step B: Perform matrix multiplication between the transpose of F α and F β , and calculate the spatial attention map A of dimension N×N using the softmax layer, which is represented by Equation (IV);
[0096]
[0097] Step C: Perform matrix multiplication between F β and the spatial attention map to obtain the post-attention map P of C×N, which is represented by Equation (V);
[0098]
[0099] Step D: Reorganize the attention map P into an output of size C×H×W×S. The feature map obtained through the attention module will first pass through the deconvolution layer, then be juxtaposed with the feature map of the encoder, and then enter the decoder to obtain the 3D U-Net after adding the attention module, as Figure 3 .
[0100] Using 3D-CNN for fine segmentation in Step 3 means that the neural network can capture the surrounding environment information at different scales within its receptive field from the pseudo 3D image. Because the pseudo 3D mask predicted and generated in the scale-up and crop step in Step 2 has a frustum-like structure, as Figure 2 in step (3) of Multiple small black dots, when stacked in multiple layers, form a frustum-like structure. Since along the scale axis, the mask shapes between adjacent slices in the 3D frustum-like mask are the same, only differing in size. By training on this 3D frustum-like mask target, the network will generate a linear or smooth surface output along the scale axis when predicting this 3D frustum-like mask, imposing a shape constraint between adjacent slices of the 3D frustum-like mask.
[0101] Specifically, in step 3.3, a spherical structure with a radius of R is used to dilate the nerve root mask obtained in step 3.2, and then a spherical structure with a radius of R is used to erode the 3D mask to obtain the repaired nerve root mask, where R > 0.
[0102] It should be noted that through the repair in step 3.3, the missing slices on the z-axis can be repaired, but the existing slices remain unchanged, and only the missing slices are replaced with the slices repaired by dilation-erosion. Such an operation is simple, time-saving, and can effectively solve the discontinuity problem.
[0103] In step 2.3, the single nerve root is located at the center of the predicted mask obtained in step one Take the predicted mask at the same corresponding position Slice, and use the center of the smallest rectangular bounding box covering the predicted mask as the center position.
[0104] It should be noted that when operating with the single nerve root at the center, specifically, the coordinates of the upper left and lower right corner points of the rectangular frame are directly recorded, and the midpoint of the line segment connecting these two corner points is taken as the center position of the rectangular frame. In the present invention, j is 60 - 70, specifically 65 in this embodiment, and h is 0.8 mm.
[0105] For the rough segmentation network in step one and the 3D-CNN in step 3.1, 3D U-Net can be used, or DeepLabV3 or nnUNet, etc. can be used as the localization and segmentation model. The number of scale levels of the feature maps of the 3D U-Net in the present invention is 3 - 7; the number of convolutional layers at each scale level in the 3D U-Net is 1 - 3; R is 4.
[0106] The method for segmenting the lumbosacral nerve roots in magnetic resonance images obtains the nerve root segmentation results through the above three steps. The pseudo 3D images generated by the multi-scale magnification-cropping step in the present invention contain multi-scale information with a compact ratio of different target backgrounds compared to 2D images, so as to improve the segmentation accuracy. At the same time, the structure of the attention 3D-CNN used in the present invention is compact, and is more compact than the traditional 2D attention module structure, thus reducing memory occupancy and improving the efficiency of network training and testing.
[0107] Example 2.
[0108] A method for segmenting lumbosacral nerve roots in magnetic resonance images, such as Figure 5 , including the following steps:
[0109] Step 1: Obtain the 3D lumbosacral MRI image X and the corresponding ground truth mask Y, and then segment the 3D lumbosacral MRI image X through a 3D rough segmentation network to obtain a predicted mask
[0110] The network used in the localization network is a 3D U-Net, with the input being the lumbosacral 3D MRI image and the output being the predicted mask
[0111] Step 2: Apply a multi-scale magnification-cropping strategy to the 3D lumbosacral MRI image X and the ground truth mask Y obtained in Step 1 to obtain a pseudo-3D image and a pseudo-3D mask;
[0112] Step 3: Use the pseudo-3D image and the pseudo-3D mask obtained in Step 2 as the input and label of the attention 3D-CNN respectively, segment the pseudo-3D image through the attention 3D-CNN to obtain a training pseudo-3D mask, then determine the optimal scale of the training pseudo-3D mask according to the results of the validation set, extract the slices of the optimal scale in the training pseudo-3D mask to obtain unilateral nerve root mask slices, and finally stack the unilateral nerve root mask slices on both sides to obtain a nerve root mask.
[0113] In Step 3, the network used in the segmentation network is a 3D U-Net with a 3D spatial attention module added. The input is the pseudo-3D image generated in the multi-scale magnification-cropping step of the present invention, and the output is the training pseudo-3D mask of the present invention. The number of magnification scales used is 40, and the magnification factor is 25.
[0114] In this embodiment, the 3D U-Net with a 3D spatial attention module added is a U-shaped network, which includes an encoding path and a decoding path. Each path contains feature maps of five scales. At each scale, it contains two 3×3×3 convolutional layers, a batch normalization layer, a ReLU non-linear activation function, and a 2×2×2 max-pooling layer with a stride of 2. Short connection is made from the high-resolution layer in the encoding path to the decoding path. In the last layer of the network, a 1×1×1 convolution is used to reduce the number of output channels to the number of labels, which is 2 in our example. A 3D spatial attention model is added between the encoder layer and the decoder layer of the smallest scale of the 3D U-Net. The specific network structure is as Figure 4 .
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for segmenting lumbosacral nerve roots in magnetic resonance images, characterized in that, It is carried out through the following steps: Step 1: Obtain the 3D lumbosacral MRI image X and the corresponding ground truth mask Y, and then segment the 3D lumbosacral MRI image X through a 3D coarse segmentation network to obtain a predicted mask Step 2: Perform a multi-scale magnification-cropping strategy on the 3D lumbosacral MRI image X and the true mask Y obtained in Step 1 to obtain a pseudo-3D image and a pseudo-3D mask; Step 3: Use the pseudo-3D image and the pseudo-3D mask obtained in Step 2 as the input and label of the attention 3D-CNN respectively. Segment the pseudo-3D image through the attention 3D-CNN to obtain a training pseudo-3D mask. Then determine the optimal scale of the training pseudo-3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo-3D mask to obtain unilateral nerve root mask slices. Finally, stack the unilateral nerve root mask slices on both sides to obtain a nerve root mask.
2. The method for segmenting lumbosacral nerve roots from magnetic resonance images according to claim 1, characterized in that, Step 2 includes: Step 2.1: Slice the 3D lumbosacral MRI image X and the true mask Y obtained in Step 1 at the same height, respectively, according to the same number j and the same thickness h to obtain j slices of 3D lumbosacral MRI image X slices and j slices of true mask Y slices, where 1 ≤ j ≤ 200 and j is an integer, h > 0, and enter Step 2.2; Step 2.2: In all 3D lumbosacral MRI image X slices and all true mask Y slices, sequentially select 3D lumbosacral MRI image X slices and true mask Y slices at the same position from top to bottom in order, and process each nerve root on each side in each true mask Y slice, and enter Step 2.3; Step 2.3: Define the current 3D lumbosacral MRI image X slice as the i-th slice X i , and define the current true mask Y slice as the i-th slice Y i . Respectively, fill X i and Y i with m i zero pixels along the length direction and n i zero pixels along the width direction, so that the unilateral nerve root is centered to obtain X *,i and Y *,i . Proceed to Step 2.4, where 1 ≤ i ≤ j and i is an integer; Step 2.4: Take X *,i Use the bilinear method to perform magnification with the maximum magnification scale being S, starting from an initial magnification scale of 1 and then increasing the magnification scale by 1 each time until the magnification scale finally reaches the S scale, and obtain the magnified image where S is an integer greater than 1, and the predicted mask after magnification by the maximum magnification scale S does not exceed the initial X *,i range. Take Y *,i Use the nearest neighbor interpolation method to perform magnification with the maximum magnification scale being S, starting from an initial magnification scale of 1 and then increasing the magnification scale by 1 each time until the magnification scale finally reaches the S scale, and obtain the magnified mask Proceed to Step 2.5; Step 2.
5. Crop the image and the mask to the same size respectively, and where is represented by Equation (II), is represented by Equation (III), and the corresponding image patch and the mask patch are obtained, then go to Step 2.6; where (c x,s , c y,s ) is the center coordinate of the s-th scale, n x and n y are the radii of the image patch along the length and width directions respectively, and the s-th scale is the s-th scale under the S scale; Step 2.6: Introduce a dimension scale, and stack the image patches and the mask module along the scale axis respectively to form pseudo 3D image patches for training the fine segmentation network and pseudo 3D mask modules where the pseudo 3D image patches are represented as pseudo 3D images The pseudo 3D mask modules are represented as pseudo 3D masks 3. The method for segmenting lumbosacral nerve roots from magnetic resonance images according to claim 2, wherein Step 3 includes: Step 3.1: Use the pseudo-3D image and the pseudo-3D mask obtained in Step 2 as the input and label of the attention 3D-CNN respectively. Segment the pseudo-3D image through the attention 3D-CNN to obtain a training pseudo-3D mask. Then determine the optimal scale of the training pseudo-3D mask according to the results of the validation set, and extract the slices of the optimal scale in the training pseudo-3D mask. Define the slices of the optimal scale as unilateral nerve root mask slices, and the size of the unilateral nerve root mask slices is the same as the size of the 3D lumbosacral MRI image X slices; Step 3.2: Stack the unilateral nerve root mask slices on both sides in order from top to bottom to obtain a nerve root mask with nerves on both sides; Step 3.3: Repair the nerve root mask to obtain a repaired nerve root mask.
4. The method for segmenting the lumbosacral nerve roots from magnetic resonance images according to claim 3, wherein: The specific operation of step 3.1 is to use the pseudo 3D image and the pseudo 3D mask as the input and label of the 3D U-Net after adding the attention module respectively. The pseudo 3D image is segmented by the attention 3D-CNN to obtain the training pseudo 3D mask, and then the optimal scale of the training pseudo 3D mask is determined according to the results of the validation set. The optimal scale layer of the training pseudo 3D mask at the optimal scale is extracted as the prediction result of the i-th slice X i , and an inverse operation transformation is performed to obtain a unilateral nerve root mask target slice with the same size as the i-th slice Y i .
5. The method for segmenting the lumbosacral nerve roots of a magnetic resonance image according to claim 4, wherein The inverse operation transformation includes: Step 3.1.1: Define the enlarged scale corresponding to the optimal scale layer as k. Perform zero-pixel padding on the length and zero-pixel padding on the width in the to obtain the padded mask Y 填,i , such that the length of the mask Y 填,i is the same as the length of the enlarged mask {Y *,i,s} s=k in Step 2.4 when the enlarged scale is the optimal scale k. The length of the mask Y 填,i is N x,0 + k×r. The width of the mask Y 填,i is the same as the width of the enlarged mask {Y *,i,s} s=k in Step 2.4 when the enlarged scale is the optimal scale k. The width of the mask Y 填,i is N y,0 + k×r; Step 3.1.2, mask Y 填,i Through downsampling operation, convert the length from N x,0 +k×r to N x,0 , and convert the width from N y,0 +k×r to N y,0 , and correspondingly obtain mask Y 缩, ; Step 3.1.3: Along the length of mask Y 缩,i trim m i pixels, and along the width trim n i pixels to obtain mask Y 裁,i . Mask Y 裁,i is the target slice of the unilateral nerve root mask.
6. The method for segmenting the lumbosacral nerve roots of a magnetic resonance image according to claim 5, characterized in that: The processing steps of the 3D U-Net after adding the attention module include: Step A: Process the 3D spatial self-attention module through two "convolution layer + recombination" operations respectively to obtain the recombined feature maps F α and F β , and both the F α and the F β have a size of C×N, where N = H×W×S, and N is the number of voxels. The given size of the 3D spatial self-attention module is the feature map F of C×H×W×S, where C, H, W, W, S represent the number of channels of the feature map, the length of the feature map, the width of the feature map, and the total scale number respectively. Step B: Perform matrix multiplication between the transpose of F α and F β to calculate the spatial attention map A with dimension N×N using a softmax layer, as represented by Equation (IV); Step C, at F β Perform matrix multiplication between and the spatial attention map to obtain the post-attention map P of C×N, which is represented by Equation (V); Step D: Reorganize the attention map P into an output of size CH × W × S. The feature map obtained through the attention module will first pass through a transposed convolutional layer, then be juxtaposed with the feature map of the encoder, and then enter the decoder to obtain the 3D U-Net after adding the attention module.
7. The method for segmenting the lumbosacral nerve roots in a magnetic resonance image according to claim 6, characterized in that: The specific content of Step 3.3 is to dilate the nerve root mask obtained in Step 3.2 using a spherical structure with a radius of R, and then erode the 3D mask using a spherical structure with a radius of R to obtain a repaired nerve root mask, and R > 0.
8. The method for segmenting the lumbosacral nerve roots in a magnetic resonance image according to claim 7, characterized in that: The said X *,i is obtained from formula (I); (N x,s ,N y,s ) = (N x,0 + s × r, N y,0 + s × r) …… Equation (Ⅰ); where (N x,0 , N y,0 ) is the size of the starting X *,i , r is the magnification factor, and r is an integer greater than 1.
9. The method for segmenting lumbosacral nerve roots from magnetic resonance images according to claim 8, wherein: In the step 2.3, the unilateral nerve root is centered at the predicted mask obtained in the first step and the predicted mask at the same position is correspondingly taken for slicing, and the center of the smallest rectangular border covering the predicted mask is taken as the central position.
10. The method for segmenting the lumbosacral nerve roots in a magnetic resonance image according to claim 9, characterized in that: The 3D rough segmentation network is constructed as a 3D U-Net; S is 20 - 60; The magnification factor is 10 - 50; The number of scales of the 3D U-Net feature map is 3 - 7; The number of convolutional layers at each scale in the 3D U-Net is 1 to 3; The j is 60 to 70; The h is 0.8 mm; The R is 4.
Citation Information
Patent Citations
Novel mammary gland MRI automatic auxiliary diagnosis method based on fusion attention mechanism
CN111401480A
Liver segmentation method of abdomen CT image, and CT imaging method thereof
CN114066901A