A method for semantic segmentation of medical images based on a U-shaped network

The improved U-type network with binary operations addresses the speed and accuracy issues in medical image segmentation by integrating feature extraction and enhancement networks, resulting in efficient and precise medical image analysis.

CN115690413BActive Publication Date: 2025-07-15XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211238982.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-07-15
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

The existing U-shaped network has a long training time, long inference time and low accuracy in medical image semantic segmentation, making it difficult to quickly and accurately complete object detection and recognition in medical image recognition.

Method used

Add binary operations to the classic U-shaped network structure to build an improved U-shaped network, including backbone extraction network and feature enhancement network, strengthen feature extraction through binary convolution and upsampling layers, and combine loss function and Adam algorithm to optimize the model.

Benefits of technology

It realizes fast and accurate semantic segmentation of medical images, improves recognition speed and accuracy, and improves the efficiency of medical image research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690413B_ABST
    Figure CN115690413B_ABST
Patent Text Reader

Abstract

A method for semantic segmentation of medical images based on a U-shaped network, which includes obtaining a sample set and dividing it into a training set and a test set; in the sample set, each medical image and its own mask are used as a group of samples; constructing an identification model and training the identification model using the training set, and testing the training effect using the test set to obtain an optimized model that meets the requirements; the identification model is based on an improved U-shaped network, and the improved U-shaped network includes a backbone extraction network and a feature enhancement network; the improved U-shaped network adds a binarization operation to the classical U-shaped network structure; inputting the medical image to be identified into the optimized model to complete target recognition. The present invention can accurately and quickly complete the semantic segmentation of medical images, with a faster speed and higher accuracy than the traditional U-shaped network, improving the efficiency of medical image research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection and recognition, and particularly relates to a method for semantic segmentation of medical images based on a U-shaped network. Background Art

[0002] The semantic segmentation and recognition of medical images is a very important part of its image processing. Due to the complex internal structure of the human body, it is difficult for traditional machine vision algorithms to completely detect the features of medical images. Moreover, it still relies on rich experience and basic knowledge, and it is difficult for humans to completely extract the information contained in the data. At the same time, its reusability is not high, resulting in a waste of a large amount of human costs.

[0003] The method of deep learning can automatically extract features from data, eliminating the high entry cost and the interference of human factors. In the fields of image processing, speech recognition, medical health diagnosis, etc., machine learning and deep learning have achieved good results in prediction and clustering. However, in the semantic segmentation and recognition of medical images, on the one hand, it is necessary to ensure the complete detection of image features, and on the other hand, it is also necessary to have a relatively fast detection speed. The existing models and algorithms cannot effectively balance these two requirements.

[0004] Among them, the U-shaped network has application prospects in the background of high pixels and small defects due to its pixel-level recognition segmentation. However, at present, the training time of the U-shaped network is long, the inference time is also long, and the accuracy rate is low. After ordinary semantic segmentation, a classification network needs to be applied to identify the image. Summary of the Invention

[0005] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method for semantic segmentation of medical images based on a U-shaped network, in order to quickly and accurately complete the target detection and recognition in medical images.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0007] A method for semantic segmentation of medical images based on a U-shaped network, comprising the following steps:

[0008] Step 1, obtaining a sample set, and dividing it into a training set and a test set;

[0009] In the sample set, each medical image and its own mask are used as a group of samples;

[0010] Step 2, constructing a recognition model, training the recognition model using the training set, testing the training effect using the test set, and obtaining an optimized model that meets the requirements;

[0011] The recognition model is based on an improved U-shaped network, and the improved U-shaped network includes a backbone extraction network and a feature enhancement network; the improved U-shaped network adds a binarization operation to the classical U-shaped network structure.

[0012] Step 3: Input the medical image to be recognized into the optimization model to complete the target recognition.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0014] The model structure draws on the U-shaped network architecture in semantic segmentation and has strong expressive power, enabling it to see very long historical and future information. It is more excellent than ordinary U-shaped networks in terms of robustness and accuracy.

[0015] The advantage of the present invention is that it can accurately and quickly complete the semantic segmentation of medical images, with faster speed and higher accuracy than traditional U-shaped networks, improving the efficiency of medical image research. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is the network structure adopted by the present invention.

[0017] Figure 2 is the image example diagram adopted by the present invention.

[0018] Figure 3 is the mask example diagram adopted by the present invention.

[0019] Figure 4 is the original image used for testing, the mask of the original image used, and the inference result.

[0020] Figure 5 is the result of the model of the present invention for the test set.

[0021] Figure 6 is the result of the ordinary U-shaped network for the test set. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following will describe the implementation plan of the present invention in detail in combination with embodiments. However, those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be regarded as limiting the scope of the present invention.

[0023] As mentioned above, due to the complex internal structure of the human body, existing machine vision algorithms are difficult to completely detect the features of medical images and even more difficult to extract complete information, resulting in the lack of reusability and timeliness of the algorithms.

[0024] To this end, the present invention provides a method for semantic segmentation of medical images based on a U-shaped network. By adding a binarization operation to the classic U-shaped network structure as a semantic segmentation network, it can extract high-level semantic information and shallow features, and can automatically extract features according to the data, with a short prediction time and high accuracy. The present invention combines the U-shaped network with medical image recognition, improving work efficiency.

[0025] Specifically, the method for semantic segmentation of medical images in the present invention includes the following steps:

[0026] Step 1, obtain a sample set and divide it into a training set and a test set.

[0027] Specifically, obtain medical images from an existing medical image database, take a medical image and its own mask as a group of samples, and combine all the sample groups as the sample set.

[0028] To facilitate training the model, the present invention also preprocesses the samples, and the method is as follows:

[0029] The pixel range of each medical image is [0-255]. Divide each pixel value of each medical image by 255 to make its range between [0-1], and convert it into tensor type data. Then standardize the tensor image through the mean and standard deviation, that is, convert the medical image into a standard tensor image. The formula for preprocessing can be expressed as:

[0030] output[channel]=(input[channel]-mean[channel]) / std[channel]

[0031] Among them, channel represents the number of channels. Generally, an image is an RGB (red green blue) image, which is three channels; Output represents the output, input represents the input, mean represents the mean value, and std represents the standard deviation.

[0032] The first obtained (0.5, 0.5, 0.5), that is, the average value of the three channels, corresponds to the three channels of the tensor image respectively, and the mean value is calculated for each channel.

[0033] The second obtained (0.5, 0.5, 0.5), that is, the standard deviation of the three channels, corresponds to the three channels of the tensor image respectively, and the standard deviation is calculated for each channel.

[0034] By taking the maximum and minimum values after dividing the pixel values by 255 on each channel, and then substituting them into the formula, the value after dividing each pixel point by 255 is mapped to the range of [-1, 1].

[0035] Through preprocessing, the recognition of medical images in the present invention is not only limited to ordinary block diagrams, but is improved to the recognition of pixel sets, classifying each pixel in the medical image, which is more accurate.

[0036] Step 2: Construct a recognition model, train the recognition model using a training set, and test the training effect using a test set to obtain an optimized model that meets the requirements.

[0037] In the present invention, the recognition model is based on the improved U-shaped network with the binarization operation added as described above. Refer to Figure 1 , the network structure of the improved U-shaped mainly includes two parts. The first part is the backbone extraction network, and the second part is the feature enhancement network (the upsampling part). Since the network is U-shaped, it is called the U-shaped network. In the recognition model of the present invention, a BinarizeConv2d operation is added after convolution and during the upsampling process, which can enhance the feature extraction.

[0038] Among them, the backbone extraction network is used to extract backbone features. A number of effective feature layers are obtained by using the backbone extraction network, and feature fusion is performed to obtain the backbone features, which successively include: the first convolutional layer, the first binarized convolutional layer, the first max-pooling layer, the second convolutional layer, the second binarized convolutional layer, the second max-pooling layer, the third convolutional layer, the third binarized convolutional layer, the third max-pooling layer, the fourth convolutional layer, the fourth binarized convolutional layer, the fourth max-pooling layer, the fifth convolutional layer, the fifth binarized convolutional layer, the sixth convolutional layer, the seventh convolutional layer, the eighth convolutional layer, the ninth convolutional layer, the tenth convolutional layer, and the sigmoid output layer. That is, in the present invention, five effective feature layers are obtained by using the backbone extraction network.

[0039] The feature enhancement network is used to extract enhanced features. The feature enhancement network performs upsampling on the number of effective feature layers and performs feature fusion to obtain a final effective feature layer that integrates all features, that is, the enhanced features. In the embodiment of the present invention, the feature enhancement network successively includes the sixth binarized upsampling layer, the seventh binarized upsampling layer, the eighth binarized upsampling layer, and the ninth binarized upsampling layer.

[0040] Specifically, the results of the sixth binarized upsampling layer and the fourth binarized convolutional layer are stacked in terms of the number of channels; the results of the seventh binarized upsampling layer and the third binarized convolutional layer are stacked in terms of the number of channels; the results of the eighth binarized upsampling layer and the second binarized convolutional layer are stacked in terms of the number of channels; the results of the ninth binarized upsampling layer and the first binarized convolutional layer are stacked in terms of the number of channels. For example, if the result of the binarized convolutional layer in the backbone extraction network is 128*128*64 and the result of the binarized upsampling is 128*128*64, where 64 is the number of channels, after stacking the number of channels, it is 128*128*128, thereby performing feature fusion.

[0041] Each convolutional layer is used to extract data features. Each convolutional layer contains two convolutional operations. The convolutional operation specifically includes convolution (conv2d), normalization (BN), and non-linear mapping (using the activation function RELU). In a convolutional layer, the two convolutional operations are in a serial relationship, and the order of the two convolutional operations is: conv relu bn conv relu bn, that is, after one convolutional operation is completed, another convolutional operation is performed. The convolutional kernels of each convolutional layer have the same size, with kernel being 3 and padding being 1. The batch normalization unit is used to accelerate the convergence speed of the network prediction model and prevent overfitting. The ReLU activation function is used to complete the non-linear transformation of the data and avoid gradient explosion and gradient disappearance.

[0042] The binarized convolutional layer is used to enhance feature extraction and make local features more obvious. Each binarized convolutional layer contains a binarized convolutional unit (BinarizeConv2d). The convolutional kernels of each binarized convolutional layer have the same size, with kernel being 3, stride being 1, and padding being 1. The sizes of the five binarized convolutional layers in this embodiment are successively: 32*32*3*3, 64*64*3*3, 128*128*3*3, 256*256*3*3, 512*512*3*3.

[0043] The max pooling layer uses max pooling (Maxpool) with kernel being 2, divides the entire image into several non-overlapping small blocks of the same size (pooling size) without overlap, only takes the largest number within each small block, and then discards other nodes, and maintains the original planar structure to obtain the output. Its main function is down-sampling.

[0044] Each binarized upsampling layer includes an upsampling unit (UpSample), a binarized convolutional (BinarizeConv2d) unit, a normalization unit (BN), and an activation function unit. The activation function unit uses the Hardtanh function because the gradient of the Sign(x) function is almost everywhere 0 in binary values, which is not conducive to backpropagation.

[0045] Exemplarily, in the present invention, during training, a loss function is used to measure the matching degree between the predicted value and the expected result, and the network weights are updated to obtain an optimized model; among them, the loss function uses BCEWithLogitsLoss, and the BCEWithLogitsLoss function is more conducive to binary classification. The Adam algorithm is used to update the network weights.

[0046] During training, input the test set data into the trained model, evaluate its prediction accuracy and inference time, and obtain the predicted mask. Among them, the accuracy of a single test data has two metrics, the intersection over union (IOU) of a single test data and the Dice coefficient of a single test data. The accuracy of the test set has three metrics, the mean intersection over union (MIOU) of the overall test set, the average Hausdorff distance (aver_hd) of the overall test set, and the Dice value (dv) of the overall test set.

[0047] Among them, the intersection over union (IOU) of a single test data, that is, the intersection over union between the mask A of the original test data and the mask B inferred by the model, is calculated by the following formula:

[0048]

[0049] The Dice coefficient of a single test data is a measure of set similarity, usually used to calculate the similarity between two samples, with a value range of [0, 1]. When used in image segmentation, the best segmentation result is 1 and the worst is 0. The method is to first obtain the masks of the two images, where the foreground in the mask is represented by 1 and the background image is represented by 0, and then calculate the two masks as arrays, with the formula as follows:

[0050]

[0051] pred is the predicted value and true is the true value;

[0052] The MIOU of the overall test set is obtained by adding up the IOU of each image and then taking the average.

[0053] The aver_hd of the overall test set is obtained by calculating the Hausdorff distance between each image and the inferred result, adding them up, and then taking the average. hd is the Hausdorff distance. H(image_mask, predict) is called the bidirectional Hausdorff distance, and h(image_mask, predict) is called the unidirectional Hausdorff distance from the point set image_mask to the point set predict. Correspondingly, h(predict, image_mask) is called the unidirectional Hausdorff distance from the point set predict to the point set image_mask. The bidirectional Hausdorff distance is the larger of the two unidirectional Hausdorff distances, which measures the maximum mismatch degree between two point sets.

[0054] The dv of the overall test set is obtained by adding up the Dice of each image and then taking the average.

[0055] Step 3: Preprocess the medical image to be recognized as described above, and then input it into the optimized model obtained in Step 2 to complete the target recognition.

[0056] In a specific embodiment of the present invention, the medical image to be recognized is a liver ultrasound image, and the purpose is to recognize the liver information therein. The original image is as Figure 2 shown, where (a) and (b) are the original liver ultrasound images in the divided training set, Figure 3 and (a) and (b) in the middle are the corresponding masks, Figure 4 which shows the test process and results. Among them, (a) and (b) respectively show two different inputs, and both from left to right respectively include the original image used for testing, the inference result of the model of the present invention, and the mask of the original image. It can be seen that the result inferred by the present invention has a relatively high coincidence degree with the mask of the original image.

[0057] Figure 5 shows the test results of the model of the present invention on the test set, Figure 6 shows the test results of the ordinary U-shaped network on the same test set, where the test set contains 20 samples, Figure 5 and Figure 6 both respectively show the IOU, Dice and time of each sample. It can be clearly seen that the results of IOU, Dice, and time of the present invention are better than those of the ordinary U-shaped network, that is, the accuracy and time of the inference of the present invention are significantly better than those of the existing ordinary U-shaped network model.

Claims

1. A method for semantic segmentation of medical images based on a U-shaped network, characterized in that, It includes the following steps: Step 1: Obtain a sample set and divide it into a training set and a test set; In the sample set, each medical image and its own mask image are used as a group of samples; Step 2: Construct a recognition model, train the recognition model using the training set, test the training effect using the test set, and obtain an optimized model that meets the requirements; The recognition model is based on an improved U-shaped network, and the improved U-shaped network includes a backbone extraction network and a feature enhancement network; the improved U-shaped network adds a binarization operation to the classical U-shaped network structure; Among them, the backbone extraction network sequentially includes: a first convolutional layer, a first binary convolutional layer, a first max-pooling layer, a second convolutional layer, a second binary convolutional layer, a second max-pooling layer, a third convolutional layer, a third binary convolutional layer, a third max-pooling layer, a fourth convolutional layer, a fourth binary convolutional layer, a fourth max-pooling layer, a fifth convolutional layer, a fifth binary convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, and an output layer sigmoid; The feature enhancement network sequentially includes a sixth binary upsampling layer, a seventh binary upsampling layer, an eighth binary upsampling layer, and a ninth binary upsampling layer. The results of the sixth binary upsampling layer and the fourth binary convolutional layer are stacked in terms of the number of channels; the results of the seventh binary upsampling layer and the third binary convolutional layer are stacked in terms of the number of channels; the results of the eighth binary upsampling layer and the second binary convolutional layer are stacked in terms of the number of channels; the results of the ninth binary upsampling layer and the first binary convolutional layer are stacked in terms of the number of channels; Step 3: Input the medical image to be recognized into the optimized model to complete target recognition.

2. The method for semantic segmentation of medical images based on the U-shaped network according to claim 1, wherein In Step 1, the samples in the sample set are preprocessed as follows: Divide the pixel value of each medical image by 255 to make its range between [0 - 1], convert it into tensor type data, and then standardize the tensor image through the mean and standard deviation; In Step 3, the medical image to be recognized is preprocessed as described above and then input into the optimized model.

3. The method for semantic segmentation of medical images based on a U-shaped network according to claim 1, characterized in that, The backbone extraction network is used to extract backbone features. A number of effective feature layers are obtained using the backbone extraction network, and backbone features are obtained through feature fusion; the feature enhancement network is used to extract enhanced features. The number of effective feature layers is upsampled using the feature enhancement network, and feature fusion is performed to obtain enhanced features.

4. The method for semantic segmentation of medical images based on the U-shaped network according to claim 1, characterized in that, Each convolutional layer is used to extract data features. Each convolutional layer contains two convolutional operations. Each convolutional operation includes convolution, normalization, and non-linear mapping. In a convolutional layer, the two convolutional operations are in a serial relationship; the convolutional kernels of each convolutional layer have the same size; Each binary convolutional layer contains a binary convolutional unit. The convolutional kernels of each binary convolutional layer have the same size. The sizes of the five binary convolutional layers are successively: 32*32*3*3, 64*64*3*3, 128*128*3*3, 256*256*3*3, 512*512*3*3; For each max-pooling layer, the entire image is divided into several non-overlapping small blocks of the same size (pooling size). Only the largest number is taken within each small block, and after discarding other nodes, the original planar structure is maintained; Each binary upsampling layer includes an upsampling unit, a binary convolution unit, a normalization unit, and an activation function unit, and the activation function unit uses the Hardtanh function.

5. The method for semantic segmentation of medical images based on a U-shaped network according to claim 1, wherein, In step 2 during training, a loss function is used to measure the matching degree between the predicted value and the expected result, and the network weights are updated to obtain an optimized model; the loss function uses BCEWithLogitsLoss, and the Adam optimization algorithm is used to update the network weights.

6. The method for semantic segmentation of medical images based on a U-shaped network according to claim 1, wherein In step 2 during training, the test set data is input into the trained model to evaluate its prediction accuracy and inference time, and a predicted mask is obtained; there are two metrics for the accuracy of a single test data, which are the intersection over union (IOU) of a single test data and the Dice coefficient of a single test data, and there are three metrics for the accuracy of the test set, which are the mean intersection over union (MIOU) of the overall test set, the average Hausdorff distance (aver_hd) of the overall test set, and the Dice value (dv) of the overall test set.

7. The method for semantic segmentation of medical images based on the U-shaped network according to claim 6, wherein The intersection over union (IOU) of a single test data, that is, the intersection over union between the mask A of the original test data and the mask B inferred by the model, is as follows: (1) The Dice coefficient of a single test data is used to calculate the similarity between two samples, and the value threshold is [0, 1]. The best result of segmentation is 1, and the worst result is 0. The method is to first obtain the masks of the two images, where the foreground in the mask is represented by 1 and the background image is represented by 0, and the two masks are converted into arrays for calculation, as follows: (2) pred is the predicted value, and true is the true value; The mean intersection over union (MIOU) of the overall test set is to add up the IOU of each image and then take the average; The average Hausdorff distance (aver_hd) of the overall test set is to calculate the Hausdorff distance between each image and the inferred result, add them up, and then take the average; The Dice value (dv) of the overall test set is to add up the Dice of each image and then take the average.

Citation Information

Patent Citations

  • Damaged two-dimensional code recovery method of convolutional auto-encoder in combination with binary segmentation

    CN111783494A

  • Rain removing method based on image restoration technology

    CN111861935A