Method for Automatically Segmenting Human Tissues in Ultrasonic Images by Deep Learning

The NFNet-based deep learning approach addresses the limitations of batch normalization in ultrasound image segmentation by modifying the network structure for deeper training, achieving improved accuracy and speed in segmenting skin, fat, fascia, muscle, and bone.

CN114119474BActive Publication Date: 2025-07-15ZHEJIANG DE IMAGE SOLUTIONS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111232141.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-07-15
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

Existing ultrasound image segmentation technology cannot quickly and accurately segment tissues such as skin, fat, fascia, muscle and bones, and deep learning models using batch normalized layers are inefficient during training and cannot meet the real-time segmentation needs.

Method used

NFNet without batch normalization layer is used as the basic network to build an encoded and decoded network structure, and real-time segmentation of ultrasonic images is achieved through multiple double-size upsampling and convolution operations, combining data augmentation strategy and loss function optimization.

Benefits of technology

The accuracy and inference speed of ultrasonic image segmentation are improved, the model performance degradation caused by batch normalization layer is solved, and the rapid and accurate organizational segmentation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114119474B_ABST
    Figure CN114119474B_ABST
Patent Text Reader

Abstract

The present invention relates to medical image processing technology, aiming to provide a method for automatically segmenting human tissues in ultrasonic images by deep learning. The method includes: establishing a training and test sample set of ultrasonic images of human tissues; constructing an encoding network structure of a segmentation network with NFNet without batch normalization layer as the basic network; constructing a decoding network structure of the segmentation network; cross-training the segmentation model; multi-model testing and performance evaluation; and performing real-time segmentation after model fusion. The present invention designs a segmentation network structure without batch normalization layer based on NFNet-F0, solves the problem that the network structure with batch normalization layer has poor model performance when using a smaller batch during segmentation training, uses the complete image for training to utilize all context information, and improves the accuracy of the segmentation model and the speed during application inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medical image processing technology, and particularly to a method for automatically segmenting human tissues (such as skin, fat, fascia, muscle, bone, and internal organs, etc.) in ultrasonic images using deep learning, and specifically to the application of a convolutional neural network without a batch normalization layer for automatic segmentation of ultrasonic images. Background Art

[0002] The rapid automatic segmentation of regions of interest in ultrasonic images has important application value, and there have been many research conclusions and practical application results in the identification of human internal organ tissues, tumors, nodules, etc.

[0003] During the process of applying ultrasonic treatment, it is necessary to timely judge the positions of different human tissues during the treatment. Therefore, if various tissues such as skin, fat, fascia, muscle, bone, and internal organs can be quickly segmented from ultrasonic images simultaneously, the thickness, area, etc. changes of various tissues during the treatment can be calculated in real time, reducing the operation time, improving the treatment accuracy, and preventing accidental injury to normal tissues. However, due to the existence of a large number of speckle textures, artifacts, and uneven echoes in ultrasonic images, the existing ultrasonic image segmentation technology can only identify internal organ tissues, tumors, nodules, etc. well, but cannot meet the requirements of quickly and accurately segmenting skin, fat, fascia, muscle, and bone.

[0004] Existing research results show that the segmentation algorithm based on convolutional neural network has a good effect in ultrasonic image segmentation, and generally the deeper the network, the better the segmentation effect. A deeper network structure is mainly composed of a convolutional layer, a pooling layer, an activation layer, and a batch normalization layer (Batch Normalization). The main function of the batch normalization layer is to solve the problem of internal covariate shift, prevent gradient dispersion, and can use a larger learning rate to accelerate convergence, and at the same time has a regularization effect. However, the batch normalization layer also has some disadvantages. The batch normalization layer significantly increases the training time of the model, and the results are different during model training and inference; in addition, the batch normalization layer is very sensitive to the batch size. If the batch size is too small, the performance of the model will deteriorate. Due to limited hardware resources, the size of the maximum model that can be designed is restricted. When using a larger model for training ultrasonic image segmentation, the method of inputting the entire image can learn more context information, improve the segmentation accuracy and application inference speed. When both the model and the input image size are relatively large, only an extremely small batch value can be used for training, so the performance of the final model will be relatively poor. Obviously, such a model cannot meet the requirements of quickly and accurately segmenting skin, fat, fascia, muscle, and bone.

[0005] In order to achieve the functions of the original batch normalization layer while removing the batch normalization layer, NFNet suppresses the activation scale on the residual branch during initialization by introducing small constants or learnable scalars, and uses scale-weight normalization to eliminate the mean shift phenomenon in the hidden activation function. Compared with ResNet with the same number of parameters, NFNet has higher accuracy and can train deeper network structures. Currently, NFNet is mainly used for natural image recognition. If it is used for image segmentation tasks, such as using the conventional method of constructing a segmentation network to simultaneously recognize and segment human tissues such as skin, fat, fascia, muscle, bone, and internal organs in ultrasound images, there will be problems such as difficulties in designing the decoding network and incomplete and inaccurate segmentation of narrow tissue regions, which cause great trouble to researchers. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a method for automatically segmenting human tissues in ultrasound images by deep learning.

[0007] To solve the technical problem, the solution of the present invention is as follows:

[0008] Provide a method for automatically segmenting human tissues in ultrasound images by deep learning, including the following steps:

[0009] (1) Establish a training and test sample set of ultrasound images of human tissues;

[0010] (2) Use NFNet without a batch normalization layer as the basic network to construct the encoding network structure of the segmentation network;

[0011] (3) Construct the decoding network structure of the segmentation network;

[0012] (4) Cross-train the segmentation model;

[0013] (5) Multi-model testing and performance evaluation;

[0014] (6) Perform real-time segmentation after model fusion.

[0015] As a preferred solution of the present invention, the step (1) includes:

[0016] (1.1) Collect ultrasound images containing human tissues, where the human tissues refer to skin, fat, fascia, muscle, bone, and internal organs; crop the ultrasound image area on the image, remove the non-ultrasound image area, and rename the image file;

[0017] (1.2) Outline the area contours of the skin, fat, fascia, muscle, bone, or internal organs on the ultrasound image to generate a mask image as the true label image for ultrasound image segmentation;

[0018] (1.3) Taking the ultrasound images of various human tissues and their corresponding label images as units, all the data are randomly divided into multiple parts, and one of them is taken as the test sample set, and the rest are used as the training sample set.

[0019] As a preferred embodiment of the present invention, the step (2) includes:

[0020] (2.1) Select NFNet-F0 as the basic network, and select ReLU with a fixed scale factor as the activation function; modify the convolution kernel size of the last convolutional layer of NFNet-F0 to 3×3, adjust the number of output channels to 2048, and keep the size of the output feature map unchanged; remove the last global pooling layer, Dropout layer and fully connected layer of NFNet-F0;

[0021] (2.2) Train the modified NFNet-F0 network on the public dataset. During the training process, randomly crop image regions with random area and aspect ratio, and use random data augmentation, label smoothing, MixUp and CutMix as data input regularization strategies, and select the model parameters with the highest accuracy on the validation set matching the public dataset;

[0022] (2.3) Re-initialize the NFNet-F0 network modified in step (2.1) with the model parameters trained in step (2.2) as the encoding network structure of the segmentation network.

[0023] As a preferred embodiment of the present invention, in the step (2.2), the modified NFNet-F0 network is trained on the ImageNet dataset of the ILSVRC2012 competition, and the accuracy of the model parameters is verified on the ImageNet validation set.

[0024] As a preferred embodiment of the present invention, the step (3) includes:

[0025] (3.1) Connect a double-size upsampling layer after the output layer of the encoding part of the segmentation network, and use the double-size upsampling method to increase the size of the feature map;

[0026] (3.2) Connect a 3×3 convolutional layer after the double-size upsampling layer to learn the segmentation features after upsampling for the decoding network. After the convolutional layer, use a double-size upsampling layer to adjust the size of the output feature map; then use another 3×3 convolutional layer to strengthen the feature learning of the decoding network; finally, use a double-size upsampling layer to enlarge the size of the feature map, and use a convolutional layer to output 6 feature channels, and output the segmentation probability map after softmax mapping; there is no skip connection between the decoding network and the encoding network;

[0027] (3.3) Complete the decoding network structure of the segmentation network according to steps (3.1) and (3.2). During each training, downsample the ground truth label image corresponding to the input image by a factor of 4 using nearest neighbor interpolation. The downsampled label image has the same size as the output probability map of the decoding network to calculate the loss function and perform backpropagation of gradients.

[0028] As a preferred embodiment of the present invention, step (4) includes:

[0029] (4.1) Cross-train the segmentation model on the training set. First, randomly and evenly divide the ultrasound image training set into at least 4 parts. Each time, select 1 part as the validation set, and the remaining parts as the training set.

[0030] (4.2) Since the sizes of ultrasound images are not uniform, randomly set the scaling factor of the longest side of the image during each training. If the width or height of the image is less than the preset value, use random translation for padding. Randomly crop an image area of a specified size within the padded image, and perform data augmentation using random transformations such as horizontal mirroring, brightness, contrast, and sharpness, and then input it into the network.

[0031] (4.3) Use softmax cross-entropy as the loss function, and adopt the stochastic gradient descent method to optimize the loss function, and perform several trainings on the training set. After cross-training is completed, select at least 4 models with the highest average intersection over union on the validation set as the segmentation models for each training.

[0032] As a preferred embodiment of the present invention, step (5) includes: testing the trained model on the test set to evaluate the model performance; specifically including:

[0033] After outputting the segmentation probability map, perform the inverse operation on the probability map in the same way as the input image is processed to restore the segmentation probability map to the original image size, that is, obtain the segmentation prediction value corresponding to each pixel value. On the test set, each image is segmented and predicted using at least 4 models, and the average of their predictions is used as the final segmentation result. Select an appropriate threshold to calculate the average intersection over union, and select the training result with the highest average intersection over union as the segmentation model.

[0034] As a preferred embodiment of the present invention, step (6) includes:

[0035] (6.1) Repeat steps (4) and (5) after adjusting the parameters. Use grouped convolution to connect the convolutional layer parameters of multiple models, and the total amount of parameters remains unchanged after model merging to obtain the final segmentation model.

[0036] (6.2) Real-time obtain ultrasound images, input them into the final segmentation model, and real-time obtain the segmentation results for different human tissues in the ultrasound images.

[0037] The present invention also provides an implementation method for a deep learning automatic segmentation model, which uses NFNet without batch normalization layer as the basic network structure, and modifies the convolution kernel size and output channel number of the last convolutional layer as the encoding network structure of the model; the decoding part of the model uses a hybrid operation of multiple double-size upsamplings and convolutions without skip connections; the decoding network is upsampled 8 times in total. During model training, the real label image is downsampled 4 times, and then the loss function is calculated with the output value of the decoding network and gradient backpropagation is performed to update the parameters to obtain the final segmentation model.

[0038] Description of the invention principle:

[0039] By researching the use of NFNet and constructing an ultrasonic image segmentation network, the present invention innovatively proposes to use the modified NFNet without batch normalization layer as the encoding part of the segmentation network, and designs a decoding network structure without skip connections and batch normalization layer, which can reduce the video memory consumption and computational amount while retaining all the output information of the encoding network, and can use a smaller batch size to train the entire image. Therefore, the present invention can perfectly solve the problems of the batch normalization layer during the training of a deeper network model, and improve the accuracy of ultrasonic image segmentation and the inference speed during application.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] The present invention designs a segmentation network structure without batch normalization layer based on NFNet-F0, solves the problem that the network structure with batch normalization layer has poor model performance when using a smaller batch size during segmentation training, uses the complete image for training to utilize all context information, and improves the accuracy of the segmentation model and the speed during application inference. Description of the drawings

[0042] Figure 1 It is a segmentation network structure diagram without batch normalization layer used in the embodiment of the present invention.

[0043] Figure 2 It is the original ultrasonic image and the contour images of different tissue regions outlined in the embodiment of the present invention. Detailed implementation manners

[0044] The present invention will be further described in detail below in conjunction with the drawings and specific implementation manners. The embodiments can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any way.

[0045] This example provides a method for implementing a deep learning automatic segmentation model, which uses NFNet without batch normalization layer as the basic network structure, modifies the convolution kernel size and output channel number of the last convolutional layer as the encoding network structure of the model; the decoding part of the model uses a hybrid operation of multiple double-size upsamplings and convolutions, without skip connections; the decoding network upsamples a total of 8 times, downsamples the real label image by 4 times during model training, then calculates the loss function with the output value of the decoding network and performs gradient backpropagation to update the parameters to obtain the final segmentation model.

[0046] Based on this deep learning automatic segmentation model, the method for deep learning automatic segmentation of human tissues in ultrasonic images according to the present invention specifically comprises the following steps:

[0047] Step 1, establish an ultrasonic image training and test sample set;

[0048] (1.1) Collect ultrasonic images containing human tissues, crop the ultrasonic image area on the image, remove the non-ultrasonic image area, and rename the image file.

[0049] Some of the ultrasonic images used in this embodiment are as Figure 2 shown on the left. The ultrasonic image includes skin, fat, fascia, muscle and bone. The following ultrasonic image drawing operations will also be described by taking these human tissues as examples. Since the recognition and segmentation technology of internal organs is relatively mature and similar to the following operations, the present invention will not give separate examples.

[0050] (1.2) Outline the regional contours of skin, fat, fascia, muscle and bone on the ultrasonic image according to conventional standards to generate a mask image as the real label image for ultrasonic image segmentation. In the generated label image, the background is represented by 0, the skin area is represented by 1, the fat area is represented by 2, the fascia area is represented by 3, the muscle area is represented by 4, and the bone area is represented by 5, a total of 6 categories. The different tissue areas outlined are as Figure 2 shown on the right, from top to bottom are skin, fat, fascia, muscle and bone, respectively represented by different colors.

[0051] The outlining process is carried out manually according to the conventional medical image recognition rules. Usually, licensed doctors in the ultrasonic examination departments of each hospital can complete this work.

[0052] (1.3) Take the ultrasonic images of various human tissues (such as skin, fat, fascia, muscle, bone and internal organs) and their corresponding label images as units, randomly divide all the data into multiple parts, take 1 part as the test sample set, and the remaining as the training sample set.

[0053] Step 2, construct the encoding network structure of the segmentation network

[0054] (2.1) Select NFNet-F0 as the base network and choose ReLU with a fixed scale factor as the activation function. For NFNet-F0 for the classification task, the kernel size of the last convolutional layer is 1×1 and the number of output feature channels is 3072. Since the main role of a convolutional layer with a kernel of 1 is to adjust the number of output channels and does not have strong feature expression ability, and it occupies a large amount of video memory in the segmentation network, the kernel size of the last convolutional layer is modified to 3×3, the number of output channels is adjusted to 2048, and the size of the output feature map remains unchanged. Remove the last global pooling layer, Dropout layer, and fully connected layer of NFNet-F0.

[0055] (2.2) Train the modified NFNet-F0 network on the ImageNet dataset of the ILSVRC2012 competition. The input image size of the network is 192×192. During the training process, randomly crop image regions by area and aspect ratio, and use random data augmentation, label smoothing, MixUp, and CutMix as data input regularization strategies. Train a total of 350 times on the training set, and select the model parameters with the highest accuracy on the ImageNet validation set.

[0056] (2.3) Re-initialize the NFNet-F0 network modified in step (2.1) with the model parameters trained in step (2.2) as the encoding network structure of the segmentation network. Set the input image size of the segmentation network to 896×896. The size of the feature map of the encoding network is downsampled by a total of 32 times, that is, the size of the output feature map of the last convolutional layer is 28×28.

[0057] Process three, construct the decoding network structure of the segmentation network

[0058] (3.1) Use the double-size upsampling method to increase the size of the feature map. The double-size upsampling layer is a network layer that replaces the bilinear interpolation and transposed convolution upsampling layers. This network layer doubles the size of the input feature map, reduces the number of channels to 1 / 4 of the original, and keeps the total data volume of the feature map unchanged. Connect a double-size upsampling layer after the output layer of the encoding part of the segmentation network. The output feature map size becomes 56×56, and the number of feature channels becomes 512. Since the double-size upsampling utilizes all features and does not introduce duplicate information, it greatly reduces the computational complexity and video memory consumption.

[0059] (3.2) Connect one 3×3 convolutional layer after the double-size upsampling layer for the decoding network to learn the segmentation features after upsampling. After the convolutional layer, use one double-size upsampling layer to adjust the output feature map size to 112×112, and the number of feature channels becomes 128; then use another 3×3 convolutional layer with an output channel number of 128 to strengthen the feature learning of the decoding network. Finally, use one double-size upsampling layer to enlarge the feature map size to 224×224, and use one convolutional layer to output 6 feature channels. After softmax mapping, the segmentation probability map is output. There is no skip connection between the decoding network and the encoding network.

[0060] (3.3) Complete the decoding network structure of the segmentation network according to steps (3.1) and (3.2). The output segmentation probability map of the final decoding network of the segmentation network is 224×224, and the size is 1 / 4 of the input image size. During each training, the corresponding ground truth label image of the input image is downsampled 4 times using nearest neighbor interpolation. The downsampled label image has the same size as the output probability map of the decoding network, and the loss function can be calculated and gradient backpropagation can be performed.

[0061] Process Four, Cross-Training the Segmentation Model

[0062] (4.1) Train the segmentation model on the training set using 4-fold cross-validation. First, randomly divide the ultrasound image training set into 4 equal parts. Each time, select 3 of them as the training set and the other 1 as the validation set.

[0063] (4.2) Since the sizes of ultrasound images are not uniform, randomly set the scaling scale of the longest side of the image during each training. The scaling range is [627, 1164]. If the width or height of the image is less than 896, use random translation padding to ensure that the width or height is not less than 896. Then, randomly crop an 896×896 image area from the padded image, and perform data augmentation using random transformations such as horizontal mirroring, brightness, contrast, and sharpness before inputting it into the network.

[0064] (4.3) Use softmax cross-entropy as the loss function, and use the stochastic gradient descent method to optimize the loss function. The initial learning rate is 0.01, and the learning rate is adjusted using polynomial decay. The weight decay is set to 2e-5, the momentum is set to 0.9, the single-device training batch size is 2, and the total batch size is 40. Train on the training set for 50 times in total. After cross-training, select the 4 models with the highest average intersection over union on the validation set as the segmentation models for each training.

[0065] Process Five, Multi-Model Testing and Performance Evaluation

[0066] After the model training is completed, test it on the test set to evaluate the model performance.

[0067] During the test, the longest side of the image is scaled to 896, the aspect ratio is kept unchanged, and the short sides are evenly filled. After filling, the length of each side is 896. After the segmentation probability map is output, the probability map is inversely operated according to the processing method of the input image, and the segmentation probability map is restored to the original image size, that is, the segmentation prediction value corresponding to each pixel value is obtained. In the test set, each image uses 4 models for segmentation prediction, and the average of their predictions is used as the final segmentation result. The appropriate threshold is selected to calculate the average intersection-over-union ratio, and the training result with the highest average intersection-over-union ratio is selected as the segmentation model.

[0068] Process 6: Real-time segmentation after model fusion

[0069] (6.1) Adjust the parameters and repeat steps (4) and (5) to obtain the final segmentation model.

[0070] In actual applications, although the inference speed of a single segmentation network is very fast, it takes a long time to execute the four models in sequence, which cannot meet the requirements of real-time applications. Therefore, the four models need to be fused. Since the convolutional layer output feature maps of multiple models for inference are independent of each other, group convolution can be used to connect the convolutional layer parameters of the four models. After the models are merged, the total number of parameters remains unchanged, and only one forward inference is required to obtain the final segmentation result.

[0071] (6.2) Acquire the ultrasound image to be segmented in real time, input it into the final segmentation model, and obtain the segmentation results of the skin, fat, fascia, muscle, bone and internal organs in the ultrasound image in real time (such as Figure 2 shown).

[0072] Finally, it should be noted that the above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and there are many variations and application scenarios. All variations that can be directly derived or associated with the content disclosed by ordinary technicians in this field should be considered as the protection scope of the present invention.

Claims

1. A method for automatically segmenting human tissues in ultrasonic images by deep learning, characterized in that, It includes the following steps: (1) Establish an ultrasonic image training and test sample set of human tissues; (2) Using NFNet without batch normalization layer as the basic network, construct the encoding network structure of the segmentation network; This step specifically includes: (2.1) Select NFNet-F0 as the basic network, and choose ReLU with a fixed scale factor as the activation function; Modify the convolution kernel size of the last convolutional layer of NFNet-F0 to 3×3, adjust the number of output channels to 2048, and keep the size of the output feature map unchanged; Remove the last global pooling layer, Dropout layer and fully connected layer of NFNet-F0; (2.2) Train the modified NFNet-F0 network on the public dataset. During the training process, randomly crop image regions by area and aspect ratio, and use random data augmentation, label smoothing, MixUp and CutMix as data input regularization strategies, and select the model parameters with the highest accuracy on the validation set matching the public dataset; (2.3) Use the model parameters trained in step (2.2) to re-initialize the NFNet-F0 network modified in step (2.1) as the encoding network structure of the segmentation network; (3) Construct the decoding network structure of the segmentation network; This step specifically includes: (3.1) Connect a double-size upsampling layer after the output layer of the encoding part of the segmentation network, and use the double-size upsampling method to increase the size of the feature map; (3.2) Connect a 3×3 convolutional layer after the double-size upsampling layer to enable the decoding network to learn the segmentation features after upsampling. After the convolutional layer, use a double-size upsampling layer to adjust the size of the output feature map; Then use another 3×3 convolutional layer to strengthen the feature learning of the decoding network; Finally, use a double-size upsampling layer to enlarge the size of the feature map, and use a convolutional layer to output 6 feature channels, and output the segmentation probability map after softmax mapping; There is no skip connection between the decoding network and the encoding network; (3.3) Complete the decoding network structure of the segmentation network according to steps (3.1) and (3.2). During each training, downsample the real label image corresponding to the input image by 4 times using nearest neighbor interpolation. The size of the downsampled label image is the same as the output probability map of the decoding network to calculate the loss function and perform gradient backpropagation; (4) Cross-train the segmentation model; This step specifically includes: (4.1) Cross-train the segmentation model on the training set. First, randomly and evenly divide the ultrasonic image training set into at least 4 parts. Each time, select 1 part as the validation set, and the remaining parts as the training set; (4.2) Since the sizes of ultrasonic images are not uniform, randomly set the scaling scale of the longest side of the image during each training; If the width or height of the image is less than the preset value, use random translation to fill it; Randomly crop an image region of a specified size within the filled image, and perform data augmentation using random transformations such as horizontal mirroring, brightness, contrast, and sharpness, and then input it into the network; (4.3) Use softmax cross-entropy as the loss function, and adopt the stochastic gradient descent method to optimize the loss function, and perform several trainings on the training set; after the cross-training is completed, select at least 4 models with the highest average intersection over union on the validation set as the segmentation models for each training; (5) Multi-model testing and performance evaluation; (6) Perform real-time segmentation after model fusion.

2. The method according to claim 1, characterized in that, The step (1) includes: (1.1) Collect ultrasound images of human tissues, crop the ultrasound image areas on the images, remove non-ultrasound image areas, and rename the image files; the human tissues refer to skin, fat, fascia, muscle, bone, and internal organs; (1.2) Outline the area contours of the skin, fat, fascia, muscle, bone, or internal organs on the ultrasound image, and generate a mask image as the ground truth label image for ultrasound image segmentation; (1.3) Take all data in units of various human tissue ultrasound images and their corresponding label images, randomly divide them into multiple parts, take 1 part as the test sample set, and the remaining as the training sample set.

3. The method according to claim 1, characterized in that In the step (2.2), the modified NFNet-F0 network is trained on the ImageNet dataset of the ILSVRC2012 competition, and the accuracy of the model parameters is verified on the ImageNet validation set.

4. The method according to claim 1, characterized in that The step (5) includes: testing the trained model on the test set and evaluating the model performance; specifically including: After outputting the segmentation probability map, perform the inverse operation on the probability map according to the processing method of the input image to restore the segmentation probability map to the original image size, that is, obtain the segmentation prediction value corresponding to each pixel value; on the test set, each image is segmented and predicted using at least 4 models, and the average of their predictions is used as the final segmentation result; select an appropriate threshold to calculate the average intersection over union, and select the training result with the highest average intersection over union as the segmentation model.

5. The method according to claim 1, wherein The step (6) includes: (6.1) Repeat steps (4) and (5) after adjusting the parameters, use group convolution to connect the convolution layer parameters of multiple models, and the total amount of parameters remains unchanged after model merging to obtain the final segmentation model; (6.2) Obtain ultrasound images in real time, input them into the final segmentation model, and obtain the segmentation results for different human tissues in the ultrasound images in real time.

Citation Information

Patent Citations

  • Fundus image retinal vessel segmentation method and system based on deep learning

    CN106408562A