Lumbar MRI Image Semantic Segmentation Method and System Based on Divergence Loss
By introducing contour loss and Gaussian divergence loss in semantic segmentation of lumbar spine MRI images, the problem that existing methods are difficult to reflect the smoothness and continuity of segmented images at the pixel level is solved, and a smoother and fuller segmented graphics are achieved, which improves the convenience of graphic index measurement.
Patent Information
- Application Number
- CN202210814704.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The existing semantic segmentation method of lumbar MRI image cannot effectively reflect the smoothness and continuity of the segmented image when processing the pixel level, resulting in misclassified pixel points having an important impact on the quality of the segmentation result, such as forming an "island" area.
A semantic segmentation method of lumbar vertebrae MRI image based on divergence loss is proposed. By adding contour loss and Gaussian divergence loss on the basis of classic cross entropy loss, the boundary smoothness and continuity of the segmented area are optimized and the emergence of "island" areas is reduced.
By introducing contour loss and Gaussian divergence loss, the boundary smoothness and continuity of the segmented area are significantly improved, the emergence of "island" areas is reduced, making the segmented figure more smooth and full, and improving the convenience of subsequent graphic indicator measurements.
Smart Images

Figure CN115393374B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image analysis, and particularly relates to a lumbar MRI image semantic segmentation method and system based on divergence loss. Background Art
[0002] Automatically diagnosing lumbar diseases is of great significance for improving the diagnosis efficiency. Lumbar spinal stenosis (LSS) is a common lumbar disease that causes low back pain and leg pain. Axial lumbar MRI images are of great significance for the diagnosis of lumbar spinal stenosis. The graphical indexes measured from MRI images, such as the anteroposterior diameter of the spinal canal, the distance between intervertebral foramina, and the area of the dural sac, have important guiding effects on clinical diagnosis. Therefore, it is first necessary to perform semantic segmentation on lumbar MRI images, and according to the segmented graphics, measure the required graphical indexes.
[0003] At present, most of the image semantic segmentations based on deep learning classify each pixel point in the image so that the pixel point belongs to a certain region category. The most cutting-edge research results are models represented by U-net. The methods for evaluating such models are generally accuracy and intersection over union (IoU). Although these evaluation methods have achieved relatively good results in most general image segmentation scenarios, these models most commonly use cross-entropy loss or mean absolute error, which are pixel-level loss functions. Essentially, they sum up the differences between the predicted classification and the true classification of all pixel points, without considering the characteristics such as the smoothness and continuity of the segmented image. For lumbar MRI images, the position of the misclassified pixel points has an important impact on the quality of the segmented image results, because this will have a crucial impact on the subsequent analysis work. For example, a background pixel point misclassified as the foreground, if it is not connected to the correct foreground area and forms an isolated "island" area, this will have a great impact on our judgment of the foreground area boundary; conversely, if the background pixel point misclassified as the foreground is connected to the correct foreground area, then this impact will be much smaller, but the difference between these two situations cannot be reflected in the pixel-level loss function. Summary of the Invention
[0004] In view of the above problems, the present invention provides a lumbar MRI image semantic segmentation method based on divergence loss. On the basis of the classic cross-entropy loss, contour line loss and Gaussian divergence loss are added, making the boundary of the segmented region smoother and more continuous, reducing the "island" regions in the segmentation results, and making it easier to obtain smooth and plump segmented graphics, which brings great convenience to the subsequent graphical index measurement work.
[0005] To achieve the above object, the present invention mainly adopts the following technical solutions:
[0006] A semantic segmentation method for lumbar spine MRI images based on divergence loss, comprising the following steps:
[0007] Step S1: Construct a training data set, each sample data in the training data set includes a lumbar spine axial MRI image and its corresponding label, and preprocess the sample data in the training data set;
[0008] Step S2: Construct a first neural network to achieve semantic segmentation of MRI images;
[0009] Step S3: Construct a loss function
[0010] Among them, cross-entropy loss Gaussian divergence loss Contour line loss Denote the probability that pixel point X ij belongs to class t, Denote the variance of the random variable along the t-th axis in the multi-dimensional Gaussian distribution, Denote pixel point X ij as a two-dimensional tensor belonging to class t, Denote pixel point X ij as the difference in the x direction of the t-th segmentation map, Denote pixel point X ij as the difference in the y direction of the t-th segmentation map, and α and β are parameters for controlling the weights of the Gaussian divergence loss and the contour line loss respectively;
[0011] Step S4: Use the preprocessed training data set to train the first neural network, with the minimum value of the loss function as the optimization objective, and gradually update the parameters of the first neural network to obtain an image segmentation neural network.
[0012] In some embodiments, in step S2, an FC-DenseNet network is used to construct the first neural network, including a first convolutional layer, a second convolutional layer, an upsampling unit, and a downsampling unit. The input image is sequentially passed through the first convolutional layer, the downsampling unit, the upsampling unit, and the second convolutional layer to output an image segmentation result.
[0013] In some embodiments, the downsampling unit includes 5 downsampling sub-units, the upsampling unit includes 5 upsampling sub-units, and a dense connection block is connected between the downsampling unit and the upsampling unit.
[0014] In some embodiments, in the step S1, the preprocessing process of the sample data in the training dataset includes:
[0015] Perform a first preprocessing on the sample data in the training dataset to generate a first training dataset;
[0016] Perform a second preprocessing on the sample data in the training dataset to generate a second training dataset.
[0017] In some embodiments, the first preprocessing is to preprocess the sample data in the training dataset by simultaneously using two methods of random flipping and random cropping, and the second preprocessing is to preprocess the sample data in the training dataset only by random flipping.
[0018] In some embodiments, the step S4 includes the following steps:
[0019] Step S41: Batch-train a first neural network using the first training dataset, calculate the loss value according to the loss function and update the parameters of the first neural network to continue training until the value of the loss function no longer decreases;
[0020] Step S42: Batch-train the first neural network using the second training dataset, calculate the loss value according to the loss function and update the parameters of the first neural network to continue training until the value of the loss function no longer decreases, obtaining an image segmentation neural network.
[0021] In some embodiments, the first neural network outputs the MRI image segmentation result and the label segmentation result in the form of 3-D tensors respectively.
[0022] In some embodiments, in the step S3, let α = 1, β = 1, and obtain the loss function
[0023] The present invention also provides a lumbar MRI image semantic segmentation system based on divergence loss, including:
[0024] An image acquisition module, configured to acquire an MRI image to be segmented;
[0025] An image segmentation neural network, configured to output an image segmentation result in the form of 3-D tensors for the input MRI image to be segmented;
[0026] Wherein, the image segmentation neural network is constructed by using an FC-DenseNet network structure.
[0027] The present invention also provides a computer-readable memory, on which a computer program is stored. When the program is executed by a processor, the steps in the lumbar MRI image semantic segmentation method based on divergence loss provided by the present invention are implemented.
[0028] The lumbar MRI image semantic segmentation method based on divergence loss provided by the present invention proposes a new loss function. On the basis of the classical cross-entropy loss, contour loss and Gaussian divergence loss are added. The introduction of contour loss can make the boundary of the segmented region smoother and more continuous. The introduction of Gaussian divergence loss reduces the "island" regions in the segmentation result and makes the segmented region tend to be continuous. The combination of contour loss and Gaussian divergence loss is more likely to obtain a smooth and plump segmentation graph, which brings great convenience to the subsequent determination of graphics metrics. The present invention combines the new loss function with the semantic segmentation model FC-DenseNet and uses a two-stage training method to obtain the Spine-Seg-2phase model. Through experimental verification, compared with the latest model proposed by predecessors, this model can achieve comparable performance in terms of pixel-level accuracy and intersection over union, and is even better than the latter in terms of intersection over union for some regions. Description of the Drawings
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can also be obtained according to these drawings.
[0030] Figure 1 The positions of four indicators concerned when diagnosing spinal stenosis;
[0031] Figure 2 The schematic diagram of the lumbar MRI image semantic segmentation method provided by the present invention;
[0032] Figure 3 The flowchart of the lumbar MRI image semantic segmentation method provided by the present invention;
[0033] Figure 4 A specific embodiment of the MRI image and the corresponding label image in the training dataset S;
[0034] Figure 5 The first neural network structure diagram;
[0035] Figure 6 The first neural network converts the label into a 3-D tensor form;
[0036] Figure 7 The comparison of the change trend diagrams under different α and β parameter values;
[0037] Figure 8 Comparison chart of experimental results of training using four loss functions.
[0038] Explanation of reference numerals:
[0039] 51 - First convolutional layer, 52 - Downsampling unit, 521 - Downsampling sub-unit, 53 - Upsampling unit, 531 - Upsampling sub-unit, 54 - Second convolutional layer; Detailed implementation manner
[0040] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0041] For a T2-weighted lumbar spine axial MRI image, when diagnosing spinal stenosis, radiologists generally focus on four graphic metrics: the anteroposterior diameter of the spinal canal, the width of the intervertebral foramen, the area of the intervertebral disc, and the area of the dural sac, as Figure 1 shown. Line segment 1 represents the anteroposterior diameter of the spinal canal, line segment 2 represents the width of the intervertebral foramen, the white area 3 represents the area of the intervertebral disc, and area 4 represents the area of the dural sac. These 4 graphic metrics can be accurately measured through a stroke algorithm, but first, semantic segmentation of the lumbar MRI image is required, and then measurement is performed based on the segmentation result.
[0042] To achieve semantic segmentation of lumbar MRI images, the present invention provides a method for semantic segmentation of lumbar MRI images based on divergence loss. The principle of the method is as Figure 2 shown, including four parts: constructing a training data set, a first neural network, a loss function, and training the first neural network.
[0043] As Figure 3 shown, the method for semantic segmentation of lumbar MRI images based on divergence loss provided by the present invention specifically includes the following steps:
[0044] Step S1: Construct a training data set S and preprocess the samples in the training data set S;
[0045] Since the MRI images in the transverse plane of the lumbar spine can more specifically reflect detailed information, and when clinically scanning patients' lumbar spines, sometimes there is a preference for only scanning T2-weighted MRI images. Therefore, we use the T2-weighted MRI images in the transverse plane of the lumbar spine to construct the training dataset S. At the same time, each MRI image in the transverse plane of the lumbar spine is annotated. Therefore, each sample data in the training dataset S we constructed includes a transverse lumbar spine MRI image and its corresponding label. In this embodiment, the original MRI images and their corresponding labels in the training dataset S are as Figure 4 shown, Figure 4 (a) is the original MRI image, Figure 4 (b) is its corresponding label image.
[0046] When performing semantic segmentation on the MRI images, we set 5 segmentation categories, represented by 0 - 4 respectively. As Figure 4 (b) shows, 5 segmentation regions are represented by 0 - 4 respectively. 0 represents the background, 1 represents the intermediate region between the front and back tissues (referred to as AAP), 2 represents the dural sac, 3 represents the lamina, and 4 represents the intervertebral disc. The labels in the training dataset S we constructed are two-dimensional tensors stored in the form of pictures, and the value at each pixel position belongs to one of {0, 1, 2, 3, 4}, indicating that the pixel point of the input image at this position belongs to one of the categories.
[0047] Step S2: Construct the first neural network to achieve semantic segmentation of the MRI images;
[0048] The present invention uses the FC-DenseNet network to construct the first neural network to achieve semantic segmentation of the MRI images. As Figure 5 shown, the first neural network includes a first convolutional layer 51, a downsampling unit 52, an upsampling unit 53, and a second convolutional layer 54. The downsampling unit 52 includes 5 downsampling sub-units 521, and each downsampling sub-unit 521 includes a dense connection block (DenseBlock, Figure 5 abbreviated as DB in Figure 5 ) and a downsampling block (Transition Down, Figure 5Abbreviation: TU in Chinese). There is also a dense connection block connected between the downsampling unit 52 and the upsampling unit 53. In this embodiment, the first convolutional layer 51 is a 3x3 convolutional layer; the second convolutional layer is a 1x1 convolutional layer; in the downsampling unit 52, the output of each dense connection block needs to be short-circuited with its input in the channel dimension as the input of the downsampling block; in the upsampling unit 53, each dense connection block has no short circuit connection, but the output of each upsampling block will be short-circuited with the feature map of the same resolution generated during the downsampling process in the channel dimension as the input of the dense connection block in the same upsampling sub-unit 531, as shown in Figure 5 shown.
[0049] As Figure 5 shown, the process of the first neural network performing image segmentation on the input image is as follows: The input image is first processed by a 3x3 convolutional layer to obtain a feature map with 4k channels, and then processed by 5 downsampling sub-units 521, that is, alternately processed by 5 dense connection blocks and downsampling blocks to obtain the feature map with the smallest resolution. After passing through a dense connection block, it enters the upsampling process; during the upsampling process, the feature map will pass through 5 upsampling sub-units 531, that is, alternately processed by 5 dense connection blocks and upsampling blocks to obtain a feature map with the same resolution as the input image. In order to obtain the segmentation result, that is, the classification map at the pixel level, the feature map is further processed by a 1x1 convolutional layer and a soft-max activation function to obtain the final image segmentation result.
[0050] Step S3: Construct a loss function
[0051] For the input image X, considering there are K semantic segmentation categories, the output label Y of the neural network = {p (t) |t = 1, 2,..., K}, p (t) ∈ [0, 1] W×H . Where represents the probability that the pixel point X ij belongs to the category t. Here, we call Y the classification map. The present invention makes the following assumption: For any category t, in its classification map p (t) , the points with high probability should be concentrated around a certain center point, and the probability is higher the closer to the center and lower the closer to the edge. This assumption conforms to the characteristics of the axial lumbar MRI image. Therefore, the present invention normalizes the classification map in the entire two-dimensional space and then fits it with a two-dimensional Gaussian probability distribution:
[0052]
[0053] Then the degree of divergence of its probability can be measured by the covariance matrix, where x is a vector in the two-dimensional space, that is μ is the center point of the probability distribution, and ∑ is the covariance matrix.
[0054] For μ (t) , it is the expectation of the coordinates (i, j) obtained in the entire two-dimensional space of p (t) , that is:
[0055]
[0056] The covariance matrix can be calculated by the following formula:
[0057]
[0058] To facilitate the design of the loss function, we need to use a constant to represent the divergence degree of the probability distribution. First, perform singular value decomposition on the covariance matrix:
[0059] ∑ (t) = U (t) Λ (t) V (t)T (4)
[0060] The covariance matrix of the Gaussian probability distribution must be a symmetric positive semi-definite matrix. Therefore, (4) can also be written as:
[0061] Σ (t) = Q (T) Λ (t) Q (t)T (5)
[0062] Among them, each column of Q (t) is a unit eigenvector of the covariance matrix, and they are pairwise orthogonal. Λ (t) is a diagonal matrix, and the values on the diagonal are the eigenvalues of the covariance matrix, and In fact, is the variance of the random variable along a certain axis in the multi-dimensional Gaussian distribution, that is The magnitude of represents the divergence degree of the probability distribution along a certain direction. The larger, the more evenly the probability distribution diverges. Conversely, the more concentrated. Therefore, the present invention defines the Gaussian divergence loss:
[0063]
[0064] That is, the final calculated result is the average value of the divergence degrees of the classification graphs of each category. Taking the square root is because this can make proportional to the divergence degree of the high-probability points in the classification graph, but this is not necessary.
[0065] The present invention also considers another hypothesis: the contour energy function of the image segmentation result should be as small as possible. Taking the active contour model as an example, the present invention obtains the contour line loss function, that is:
[0066]
[0067] Wherein, represents the difference of the pixel point X ij in the x direction of the t-th segmentation map, represents the pixel point X ij in the y direction of the t-th segmentation map, and ∈ is an arbitrarily very small positive number to ensure that the final result is greater than 0.
[0068] Geometrically speaking, is the boundary line of the predicted segmentation region, and the hypothesis of the present invention is equivalent to making the total length of the boundary line as small as possible.
[0069] Finally, in order to make the predicted segmentation label close to the true label at the pixel level and achieve the highest possible classification accuracy, the present invention still uses the classical cross-entropy loss, that is:
[0070]
[0071] Wherein, represents the two-dimensional tensor of the pixel point X ij belonging to the category t.
[0072] Finally, the loss function of the present invention is obtained, which includes three parts: cross-entropy loss, Gaussian divergence loss, and contour line loss, that is:
[0073]
[0074] Wherein, α and β are parameters for controlling the weights of the Gaussian divergence loss and the contour line loss, respectively, and generally α = 1, β = 1.
[0075] Step S4: Using the optimized loss function as the training objective, training the first neural network, and finally obtaining the image segmentation neural network.
[0076] In the above step S1, the training dataset S includes a large number of lumbar spine axial MRI images and corresponding labels, all with a resolution of 320×320. First, all the MRI images and labels in the training dataset S need to be downsampled to a resolution of 224×224 using the bicubic downsampling method. Then, through data augmentation strategies, the variability of the dataset is increased to combat overfitting in training and train a more robust network model. In this embodiment, two data augmentation methods, random flipping and random cropping, are used to preprocess the training dataset S. When performing random flipping, each MRI image and label has a 50% chance of being flipped left and right. When performing random cropping, each MRI image and label will be randomly cropped to obtain a region, and the cropped region is scaled to a resolution of 224×224 as the input to the model.
[0077] In this embodiment, the preprocessing of the sample data in the training dataset S includes:
[0078] (1) First preprocessing: Both random flipping and random cropping are used to preprocess the MRI images and labels simultaneously to generate the first training dataset S1;
[0079] (2) Second preprocessing: Only random flipping is used to preprocess the MRI images and labels to generate the second training dataset S2.
[0080] The above-generated first training dataset S1 and second training dataset S2 can be respectively used in the two-stage training of the first neural network of the present invention.
[0081] In the above step S4, the training of the first neural network is divided into two stages, and the specific process is as follows:
[0082] Step S41: Use the first training dataset S1 to train the first neural network in batches. Calculate the loss value according to the loss function and update the parameters of the first neural network to continue training until the value of the loss function no longer decreases;
[0083] Step S42: Use the second training dataset S2 to continue training the first neural network in batches. Calculate the loss value according to the loss function and update the parameters of the first neural network to continue training until the value of the loss function no longer decreases, obtaining the image segmentation neural network.
[0084] In this embodiment, when the first neural network performs semantic segmentation on the input MRI images and labels, a 3-D tensor G of size W×H×5 is used, where G = [G (0) , G (1) , G (2) , G (3) , G (4)Output the image segmentation result, such as Figure 6 shown. During each batch of training, the MRI image segmentation result output by the first neural network and the label segmentation result are jointly used as the input of the loss function , calculate the loss value, and gradually update the network parameters of the first neural network to achieve the training goal of optimizing the loss function.
[0085] Using the lumbar MRI image semantic segmentation method based on divergence loss provided by the present invention, a lumbar MRI image semantic segmentation system based on divergence loss is constructed, including an image acquisition module and an image segmentation neural network. The image acquisition module is used to acquire the transverse lumbar MRI image to be segmented, and the image segmentation neural network is used to input the MRI image to be segmented as a 3-D tensor G = [G (0) , G (1) , G (2) , G (3) , G (4) Output the image segmentation result, and then the four graphic medical indicators required for accurately measuring and diagnosing spinal stenosis can be obtained through methods such as the stroke algorithm.
[0086] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the lumbar MRI image semantic segmentation method based on divergence loss provided by the present invention are implemented. Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0087] Next, experiments are used to verify the effect of constructing an image segmentation neural network using the lumbar MRI image semantic segmentation method based on divergence loss provided by the present invention.
[0088] Experiments were conducted using the lumbar MRI segmentation annotation dataset publicly available by Sudirman et al. from Liverpool John Moores University. This dataset contains 1545 axial lumbar MRI images, with each case including the original T1-weighted and T2-weighted images. Sudirman et al. used 5 annotators to label each MRI image, and finally aggregated the results of the 5 annotators and used a voting mechanism to determine the final label for each case. During the experiment, we only used the T2-weighted images and their corresponding labels for the experiment. The entire dataset was divided into three parts: a training dataset, a validation set, and a test set. The training dataset contains 1157 cases, the validation set contains 100 cases, and the test set contains 290 cases. All training processes were completed on the training dataset, the validation set was used to observe the model's performance during the training process to assist in adjusting hyperparameters, and the final performance evaluation was conducted using the test set.
[0089] During the experiment, the Python-3.7.6 environment was used, and the TensorFlow-2.3 framework was used to build the neural network. The training was carried out using a single GTX3090 graphics card. The optimizer selected was the Adaptive Moment Estimation optimizer (Adam), the learning rate was lr, the batch size for batch training was b, and during the training process, L2 regularization was applied to all the weight coefficients in the first neural network, and the regularization weight coefficient was w. The experimental values of all the above hyperparameters are shown in Table 1.
[0090] Table 1 Hyperparameter value settings
[0091] Hyperparameter Value lr <![CDATA[1×10 -3 > b 8 w <![CDATA[1×10 -4 > d 0.8 k 12 α 1 β 1
[0092] During the first-stage training, using the above hyperparameter settings, the number of training epochs was 30 epochs; during the second-stage training, the learning rate lr in the hyperparameters was changed to 1×10 -4 , and the number of training epochs was 40 epochs; this training model is simply referred to as Spine-Seg-2phase.
[0093] To evaluate the performance of the model, first, two parameters, pixel-level accuracy and intersection over union, were used. For class t, the pixel accuracy is defined as follows:
[0094]
[0095] where, TP t is the number of pixel points classified as t and also marked as t in the label; T t is the number of pixel points marked as t in the label. To be able to evaluate the pixel accuracy for all classes, the average pixel accuracy across all classes is defined as:
[0096]
[0097] For class t, the Intersection over Union (IoU) is defined as follows:
[0098]
[0099] where P t represents the number of pixels predicted by the model to belong to class t.
[0100] Similarly, to evaluate the IoU for all classes, the mean Intersection over Union (mIoU) for all classes is defined as follows:
[0101]
[0102] To compare the necessity of two-stage training of Spine-Seg-2phase, we also trained another model, which was trained throughout with the same configuration as the second training stage in Spine-Seg-2phase above, i.e., using a learning rate of 1×10 -4 , and only using the data augmentation strategy of random flipping, and training for 70 epochs. This model is abbreviated as Spine-Seg-1Phase. The performance comparison of the two models is shown in Table 2.
[0103] Table 2 Performance comparison of one-stage training and two-stage training strategies
[0104]
[0105] As can be seen from Table 2, the Spine-Seg-2phase model with two-stage training can achieve better results in terms of pixel accuracy and IoU metrics. This is mainly because the data augmentation strategy in the first training stage of Spine-Seg-2phase is more aggressive, forcing the neural network to learn more semantic information extraction patterns, resulting in better performance in the second training stage and better generalization ability of the entire network compared to Spine-Seg-1Phase.
[0106] Regarding the values of the hyperparameters in the loss function, we selected three schemes: (1) α = β = 0.1 (2) α = β = 1 (3) α = β = 10, and observed the performance of the neural network model during the training process and on the validation set in these three cases. Since the neural network model training basically reaches a stable state after 50 epochs, the contour loss Gaussian divergence loss as well as the average pixel accuracy acc and the mean Intersection over Union iou values on the validation set during the training process from the 50th epoch (about the 7200th step) to the 70th epoch (about the 10000th step) were selected, as shown in Figure 7 shown Figure 7 (a) is the trend chart of the contour lossFigure 7 (b) is the trend chart of the Gaussian divergence loss, Figure 7 (c) is the trend chart of the average pixel accuracy, Figure 7 (d) is the trend chart of the average intersection over union. In the four figures, curve 1 is the trend chart when α = β = 0.1; curve 2 is the trend chart when α = β = 1; curve 3 is the trend chart when α = β = 10.
[0107] From Figure 7 it can be seen that when α = β = 10, compared with the other two cases, the values of the contour line loss and the Gaussian divergence loss are significantly lower, but correspondingly, the cost is that its pixel accuracy and intersection over union also become worse. In the cases of α = β = 1 and α = β = 0.1, the performance of the neural network model in terms of pixel accuracy and intersection over union is not much different, and the difference is not much in terms of the contour line loss and the Gaussian divergence loss either, but the case of α = β = 1 is slightly better. The present invention hopes that the introduction of the contour line loss and the Gaussian divergence loss will not cause a significant degradation in the performance of the neural network. On this premise, the two losses are reduced as much as possible. Therefore, the present invention selects α = β = 1 as the weight value in the final loss function.
[0108] To verify the role of the loss function constructed by the present invention, four loss functions were used to train different Spine - Seg - 2phase models for ablation experiments. The four loss functions used are: (1) only using the classical cross - entropy loss, that is (2) using the cross - entropy loss and the contour line loss, that is (3) using the cross - entropy loss and the Gaussian divergence loss, that is (4) using the cross - entropy loss, the contour line loss and the Gaussian divergence loss, that is The results of the ablation experiment are as Figure 8 shown, Figure 8 (a) is the original input image, Figure 8 (b) is the corresponding input label image, Figure 8 (c) is only using the classical cross - entropy loss, Figure 8 (d) is using the cross - entropy loss and the contour line loss, Figure 8 (e) is using the cross - entropy loss and the Gaussian divergence loss, Figure 8 (f) is using the cross - entropy loss, the contour line loss and the Gaussian divergence loss. From Figure 8It can be clearly seen from lines 2, 3, and 5 in particular that using only the contour line loss or only the Gaussian divergence loss can reduce the "islands" in the segmentation results to a certain extent and make the segmentation results more continuous. Additionally, observing the intervertebral disc part in line 1, the lamina part in line 2, and the dural sac part in line 4, it can be seen that the introduction of the contour line loss can suppress the irregular edges in the segmented figure and make the segmented figure smoother. In the case of using both the contour line loss and the Gaussian divergence loss, it is more balanced and full in all aspects and is closest to the ideal segmentation result.
[0109] The lumbar MRI image semantic segmentation method based on divergence loss provided by the present invention proposes a new loss function. On the basis of the classical cross-entropy loss, the contour line loss and the Gaussian divergence loss are added. The introduction of the contour line loss can make the boundary of the segmented region smoother and continuous. The introduction of the Gaussian divergence loss reduces the "island" regions in the segmentation results and makes the segmented regions tend to be continuous. The combination of the contour line loss and the Gaussian divergence loss is more likely to obtain a smooth and full segmented figure, which brings great convenience to the subsequent determination work of graphics indexes.
[0110] The present invention combines the new loss function with the semantic segmentation model FC-DenseNet and uses a two-stage training method to obtain the Spine-Seg-2phase model. Through experimental verification, compared with the latest model proposed by predecessors, this model can achieve comparable performance in terms of pixel-level accuracy and intersection over union, and even outperforms the latter in terms of intersection over union for some regions.
[0111] It can be understood that the above specific description of the present invention is only used to illustrate the present invention and is not limited to the technical solutions described in the embodiments of the present invention. Those of ordinary skill in the art should understand that the present invention can still be modified or equivalently replaced to achieve the same technical effects; as long as the usage requirements are met, they are all within the protection scope of the present invention.
Claims
1. A lumbar MRI image semantic segmentation method based on divergence loss, characterized in that Including the following steps: Step S1: Construct a training data set, where each sample data in the training data set includes a lumbar spine axial MRI image and its corresponding label, and preprocess the sample data in the training data set; Step S2: Construct a first neural network to achieve semantic segmentation of MRI images; Step S3: Construct a loss function Among them, the cross-entropy loss Gaussian divergence loss Contour line loss Denote the probability that pixel X ij belongs to class t, Denote the variance of the random variable along the t-th axis in the multi-dimensional Gaussian distribution, Denote the two-dimensional tensor that pixel X ij belongs to class t, Denote the difference in the x direction of the t-th segmentation map that pixel X ij belongs to, Denote the difference in the y direction of the t-th segmentation map that pixel X ij belongs to; ∈ is an arbitrarily very small positive number to ensure that the final result is greater than 0; α and β are parameters for controlling the weights of the Gaussian divergence loss and the contour line loss respectively; Step S4: Use the preprocessed training dataset to train the first neural network, aiming to minimize the loss function as the optimization objective, and gradually update the parameters of the first neural network to obtain the image segmentation neural network.
2. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 1, wherein In step S2, use the FC-DenseNet network to construct the first neural network, including a first convolutional layer, a second convolutional layer, an upsampling unit, and a downsampling unit. The input image is sequentially passed through the first convolutional layer, the downsampling unit, the upsampling unit, and the second convolutional layer, and then the image segmentation result is output.
3. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 2, characterized in that, The downsampling unit includes 5 layers of downsampling sub-units, the upsampling unit includes 5 layers of upsampling sub-units, and a dense connection block is connected between the downsampling unit and the upsampling unit.
4. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 1, characterized in that, In step S1, the preprocessing process of the sample data in the training data set includes: Perform a first preprocessing on the sample data in the training data set to generate a first training data set; Perform a second preprocessing on the sample data in the training data set to generate a second training data set.
5. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 4, wherein The first preprocessing is to preprocess the sample data in the training data set by using both random flipping and random cropping at the same time, and the second preprocessing is to preprocess the sample data in the training data set only by using random flipping.
6. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 4, wherein Step S4 includes the following steps: Step S41: Train the first neural network using the first training dataset in batches, calculate the loss value according to the loss function and update the parameters of the first neural network to continue training until the value of the loss function no longer decreases; Step S42: Continuously train the first neural network using the second training dataset in batches, and calculate the loss value according to the loss function and update the parameters of the first neural network to continue training until the value of the loss function no longer decreases, obtaining an image segmentation neural network.
7. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 1, characterized in that The first neural network outputs the MRI image segmentation result and the label segmentation result in the form of 3-D tensors respectively.
8. The lumbar MRI image semantic segmentation method based on divergence loss according to claim 1, characterized in that, In step S3, let α = 1 and β = 1 to obtain the loss function 9. Lumbar MRI image semantic segmentation system based on divergence loss, characterized in that, Used to implement the image semantic segmentation method described in claim 1, including: An image acquisition module for acquiring the MRI image to be segmented; An image segmentation neural network for outputting the image segmentation result in the form of 3-D tensors for the input MRI image to be segmented; Among them, the image segmentation neural network is constructed using the FC-DenseNet network structure.
10. A computer-readable memory having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the lumbar MRI image semantic segmentation method based on divergence loss described in any one of claims 1-8.