An Ultrasonic Image Semantic Segmentation Method Based on Joint Supervision of the Subject and the Boundary

By building a dual-branch supervision network for subject and boundary and training in combination with boundary and subject labels, the problem of inaccurate prediction of pixels near boundaries in ultrasound images is solved, the accuracy of segmented images is improved, and the better ultrasound image segmentation effect is achieved.

CN116012356BActive Publication Date: 2025-06-20WUHAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310086638.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-06-20
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

The prior art is difficult to correctly predict pixels near the boundary in ultrasonic images, resulting in low accuracy of segmented images, especially in cases of poor contrast and blurred boundary.

Method used

The semantic segmentation method of ultrasonic image based on joint supervision of subject and boundary is adopted. By building a dual-branch supervision network of subject and boundary, the image input layer, encoder, decoder and prediction result output layer are used, and the boundary and subject labels are combined for training, to achieve more accurate segmentation of ultrasonic images.

Benefits of technology

Through this method, pixels near the boundary can be correctly predicted, the accuracy of segmented images can be improved, the inaccuracy problem of existing methods in ultrasonic image boundary segmentation is improved, and the overall segmentation results are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012356B_ABST
    Figure CN116012356B_ABST
Patent Text Reader

Abstract

The present invention provides an ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary. After building a semantic segmentation network including an image input layer, an encoder, a decoder, and a prediction result output layer, and completing training for a specified number of rounds using the training set data, the test image is input into the trained segmentation network to obtain the final segmentation result and evaluation metrics, realizing the function of correctly predicting pixels near the boundary and improving the accuracy of the segmented image. The dual-branch boundary and main body supervision network constructed by the present invention effectively alleviates the problems of inaccurate segmentation caused by acoustic artifacts, poor contrast between lesions and surrounding tissues, and blurred boundaries existing in current ultrasonic images, and improves the segmentation performance of the neural network on ultrasonic images; it solves the problem that the existing semantic segmentation methods have unsatisfactory segmentation results when segmenting the boundaries of ultrasonic images due to the poor contrast and blurred boundaries of ultrasonic images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to an ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary. Background Art

[0002] Ultrasound (US) has been widely used in clinical diagnosis and treatment of tissues and lesions, including the breast, thyroid, kidney, lymph nodes, and liver, etc., due to its advantages of low-level radioactivity, low cost, and wide availability. Generally, clinical applications require experienced radiologists to review the collected ultrasound images and mark the size, location, and shape of the lesions for further application, which requires a great deal of effort from doctors. On the other hand, the rapid development of automatic medical image segmentation technology has solved this problem. However, there is a lot of noise in the directly obtained images, which results in a large number of acoustic shadows in the ultrasound images, poor image contrast, and blurred boundaries, posing a huge challenge to existing methods.

[0003] In the past few years, automatic medical segmentation based on deep learning has become the dominant direction. Among these deep learning-based methods, the encoder-decoder structure is one of the most popular architectures for realizing image pixel classification, such as U-Net, LinkNet, and DeepLabV3+. Although these methods can achieve good segmentation accuracy, they still have some limitations. Since the poor contrast and blurred boundaries in ultrasound images are not fully considered, these models still lack the ability to correctly predict pixels near the boundary. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary, which is used to correctly predict pixels near the boundary and improve the accuracy of the segmented image.

[0005] The technical solution adopted by the present invention to solve the above technical problem is: an ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary, comprising the following steps:

[0006] S1: Collect ultrasonic image data, process and annotate it to obtain an adjusted annotation map G final , form a processed ultrasonic image data set, and divide the ultrasonic image data set into a training set and a test set;

[0007] S2: Process the annotation map G final , and obtain a boundary label and a main body label;

[0008] S3: Build a dual-branch supervision network for the main body and the boundary, which sequentially includes an image input layer, an encoder, a decoder, and a prediction result output layer;

[0009] The encoder consists of five stages, each stage includes two convolutional layers, and after each convolutional layer, there are Batch Normalization and ReLU activation functions; in the two convolutional layers, the first convolutional layer is a 1x1 convolution, and the second convolutional layer is a regular convolution with dilation. From top to bottom, the dilation rates in the regular convolution are 1, 2, 3, 5, and 7 respectively;

[0010] The decoder includes two double-branch modules DB, two main bodies, and a boundary supervision module DBS connected in sequence;

[0011] The double-branch module DB includes three parts: the first part is a 1x1 convolution for adjusting the dimension of the input features; the second part is two branches with the same structure connected to the first part respectively, and each branch includes two regular convolutions for learning rich features; the third part includes two adaptive parameters that are continuously updated through backpropagation, which are multiplied by the feature maps output by the two branches respectively for feature fusion, and then added to obtain a more representative feature map as the output of the double-branch module;

[0012] The internal structure of the main body and the boundary supervision module DBS is the same as that of the double-branch module DB, which is used to splice the feature map of the encoder and the feature map of the previous layer of the decoder, and input it into a 1x1 convolution for feature dimensionality reduction processing; then input it into the double-branch block, one branch is used to generate the main body feature map F Body ,and the other branch is used to generate the boundary feature map F BD ,and the two branches each include two regular convolutions;

[0013] S4: Use the training set to train the main body and boundary double-branch supervision network;

[0014] S5: Input the test set into the trained main body and boundary double-branch supervision network to obtain the segmentation result of the ultrasonic medical image to be segmented.

[0015] According to the above scheme, in step S1, the specific steps are as follows:

[0016] S11: Clean the collected ultrasonic image data to remove unqualified ultrasonic image data;

[0017] S12: Annotate the ultrasonic image data according to the expert experience of radiology professionals, and divide the annotated ultrasonic image data into a training set and a test set according to a ratio of 7:3.

[0018] Furthermore, in step S12, the specific steps are as follows:

[0019] The images in the training set are augmented by scaling, random cropping, and random flipping to achieve better performance and robustness; the labeled images are adjusted accordingly according to the operations of the training set.

[0020] According to the above scheme, in step S2, the specific steps are as follows:

[0021] S21: Divide the adjusted labeled graph G final into the foreground part G fg and the background part G bg , perform a distance transform on the foreground part G fg to obtain the shortest distance from each pixel in the foreground part G fg to the background part G bg , which is called the distance map;

[0022] S22: Classify the corresponding pixels in the foreground with distances greater than the specified threshold in the distance map as the main body label, and classify the corresponding pixels in the foreground with distances less than or equal to the threshold as the boundary label: Let represent a single pixel point in the foreground, represent the shortest distance from the pixels in the foreground to the background, G Body represent the main body label, G BD represent the boundary label, and α represent the manually set threshold; then the classification process is:

[0023]

[0024] S23: Set the values of the unprocessed parts in the main body label and the boundary label to 0, and the relationship between the main body label, the boundary label, and the labeled graph is:

[0025] G final = G Body + G BD .

[0026] According to the above scheme, in step S3,

[0027] Let L i represent the total loss of supervising the i-th main body and boundary supervision module; and respectively represent the losses obtained by taking the loss of the feature maps of the corresponding channels after the main body feature map and the boundary feature map of the i-th main body and boundary supervision module are restored to the original resolution through the upsampling operation and then passing through a 1x1 convolution and then taking the loss with their respective corresponding labels; represent the boundary feature map of the i-th main body and boundary supervision module; represent the main body feature map of the i-th main body and boundary supervision module; λ1 and λ2 are hyperparameters, and the default settings are both 1; then the main body feature map F Body and the boundary feature map FBD Supervised by the subject label and the boundary label respectively, which is expressed by the following formula:

[0028]

[0029] Let λ3 and λ4 be the adaptive weight parameters set during the network training process, and the initial values are set to 0.5; then the subject feature map and the boundary feature map are fused in terms of features under the adaptive weight parameters to obtain more representative features D i :

[0030]

[0031] The subject and boundary double-branch supervised network realizes multiple supervision methods by designing different loss functions; let L i be the total loss for supervising the i-th subject and boundary supervision module, and L final be the loss obtained from the finally obtained global feature map and the annotated map, then the total loss L is:

[0032]

[0033] Furthermore, in the step S4, the specific steps are as follows:

[0034] S41: Input the training set from the image input layer into the encoder, and obtain feature maps F1~F5 with different semantic information and different resolution sizes from top to bottom in different layers of the encoder;

[0035] S42: Perform bilinear interpolation on F5 to obtain a feature map with the same resolution size as F4; splice the processed F5 and F4 and input them into the first double-branch module DB of the decoder to obtain the decoded feature map D4; perform bilinear interpolation on D4 to obtain a feature map with the same resolution size as F3; splice the bilinear-interpolated D4 and F3 and input them into the second double-branch module DB of the decoder to obtain the decoded feature map D3;

[0036] S43: Perform bilinear interpolation on D3 to obtain a feature map with the same resolution size as F2; splice the bilinear-interpolated D3 and F2 and then input them into the first subject and boundary supervision module DBS of the decoder to obtain the decoded feature map D2; perform bilinear interpolation on D2 to obtain a feature map with the same resolution size as F1; splice the bilinear-interpolated D2 and F1 and then input them into the second subject and boundary supervision module DBS of the decoder to obtain the decoded feature map D1;

[0037] S44: After processing D1 through the classifier, obtain pixel-level classification predictions to get the final ultrasonic image segmentation result.

[0038] A main body and boundary double-branch supervision network successively includes an image input layer, an encoder, a decoder, and a predicted result output layer; the encoder includes five stages, each stage includes two convolutional layers, and a Batch Normalization and a ReLU activation function are provided after each convolutional layer; among the two convolutional layers, the first convolutional layer is a 1x1 convolution, and the second convolutional layer is a conventional convolution with dilation. From top to bottom, the dilation rates in the conventional convolution are 1, 2, 3, 5, and 7 respectively; the decoder includes two double-branch modules DBs and two main body and boundary supervision modules DBSs connected in sequence; the double-branch module DB includes three parts: the first part is a 1x1 convolution for dimension adjustment of the input features; the second part is two branches with the same structure respectively connected to the first part, and each branch includes two conventional convolutions for learning rich features; the third part includes two adaptive parameters that are continuously updated through backpropagation, which are multiplied by the feature maps respectively output by the double branches for feature fusion, and then added to obtain a more representative feature map as the output of the double-branch module; the internal structure of the main body and boundary supervision module DBS is the same as that of the double-branch module DB, which is used to splice the feature map of the encoder and the feature map of the previous layer of the decoder, input it into a 1x1 convolution for feature dimensionality reduction processing; then input it into the double-branch block, one branch is used to generate the main body feature map F Body , and the other branch is used to generate the boundary feature map F BD , and the two branches each include two conventional convolutions.

[0039] A computer storage medium stores a computer program executable by a computer processor, and the computer program executes an ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary.

[0040] The beneficial effects of the present invention are as follows:

[0041] 1. For the ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary of the present invention, by building a semantic segmentation network including an image input layer, an encoder, a decoder, and a predicted result output layer, training the built semantic segmentation network and testing the semantic segmentation network, after using the training set data to complete the training of the specified number of rounds, inputting the test image into the trained segmentation network to obtain the final segmentation result and evaluation index, the function of correctly predicting the pixels near the boundary and improving the accuracy of the segmented image is realized.

[0042] 2. At the feature level, the boundary serves as a constraint for the main body, and the main body part implies the boundary. The present invention utilizes the correlation between the boundary and the main body, effectively improves the problem of inaccurate segmentation of the ultrasonic image boundary in the existing method, and achieves better segmentation accuracy in the overall segmentation result.

[0043] 3. The encoder module of the present invention realizes feature extraction through repeated 1x1 convolutions and conventional convolutions with dilation. Compared with the encoder in U-Net, the present invention has fewer parameters and faster speed.

[0044] 4. The present invention selects to use backpropagation to enable the network to learn a set of parameters as the weights of the main features and boundary features, so as to better obtain representative features.

[0045] 5. The dual-branch boundary and main body supervision network constructed by the present invention effectively alleviates the problems existing in current ultrasound images, such as acoustic artifacts, poor contrast between lesions and surrounding tissues, and inaccurate segmentation caused by blurred boundaries, improving the segmentation performance of the neural network on ultrasound images; it solves the problem that the existing semantic segmentation methods have unsatisfactory segmentation results when segmenting the boundaries of ultrasound images due to the poor contrast and blurred boundaries of ultrasound images. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart of an embodiment of the present invention.

[0047] Figure 2 is an example diagram of generating boundary and main body labels in an embodiment of the present invention.

[0048] In the figure: (a) is the original image; (b) is the annotated image; (c) is the boundary label; (d) is the main body label.

[0049] Figure 3 is a schematic diagram of the network structure in an embodiment of the present invention.

[0050] Figure 4 is a schematic diagram of the dual-branch module DB in an embodiment of the present invention.

[0051] Figure 5 is a schematic diagram of the main body and boundary supervision module DBS in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0053] See Figure 1 , a method for semantic segmentation of ultrasound images based on joint supervision of main body and boundary in an embodiment of the present invention, includes the following steps:

[0054] S1: Collect ultrasound image data, process and annotate the data set to obtain the adjusted annotated map G final , and form the processed ultrasound image data set, and divide the data set into a training set and a test set;

[0055] First, clean the collected ultrasonic image data to remove unqualified ultrasonic image data. Then, annotate the images by professional radiologists. The annotated images are divided into a training set and a test set according to a ratio of 7:3. The images in the training set are augmented by means such as scaling, random cropping, and random flipping to achieve better performance and robustness, and the annotated images need to be adjusted accordingly.

[0056] S2: Process the annotated images to obtain boundary labels and body labels, see Figure 2 ; The specific steps are as follows:

[0057] S21: Divide the adjusted annotated image G final into a foreground part G fg and a background part G bg . Perform a distance transform on the foreground part G fg to obtain the shortest distance from each pixel in the foreground part G fg to the background part G bg , which is called the distance map;

[0058] S22: Classify the corresponding pixels in the foreground with distances greater than a specified threshold in the distance map as body labels, and classify the corresponding pixels in the foreground with distances less than or equal to the threshold as boundary labels: Let represent the shortest distance from a pixel in the foreground to the background, represent a single pixel point in the foreground, G Body represent the body label, G BD represent the boundary label, and α represent a manually set threshold; then the classification process is:

[0059]

[0060] S23: Set the values of the unprocessed parts in the body label and the boundary label to 0 to obtain the final body label and boundary label; the body label, the boundary label, and the annotated image satisfy the following formula:

[0061] G final =G Body +G BD ;

[0062] S3: Build a dual-branch supervised network for the body and the boundary, see Figure 3 , including an image input layer, an encoder, a decoder, and a prediction result output layer; the specific steps are as follows:

[0063] S31: Input the training images into the encoder, and obtain feature maps F1~F5 with different semantic information and different resolution sizes from top to bottom in different layers of the encoder;

[0064] The encoder consists of five stages, each stage containing two convolutional layers, each followed by Batch Normalization and ReLU activation function. Among the two convolutions, the first convolution is a 1x1 convolution, and the second convolution is a regular convolution with dilation. From top to bottom, the dilation rates in the regular convolution are 1, 2, 3, 5, and 7 respectively.

[0065] S32: Perform bilinear interpolation on F5 to obtain a feature map with the same resolution as F4. After the processed F5 and F4 are concatenated, they are input into the dual-branch network (DB) in the decoder to obtain the decoded feature map D4. Further perform bilinear interpolation on D4 to obtain a feature map with the same resolution as F3. After the bilinear-interpolated D4 and F3 are concatenated, they are input into the DB in the decoder to obtain the decoded feature map D3;

[0066] See Figure 4 , the dual-branch module consists of three parts. The first is a 1x1 convolution; the second is two identical dual-branches, each containing two regular convolutions; the third is to perform feature fusion through two adaptive parameters.

[0067] The input features are first adjusted in dimension through a 1x1 convolution, then rich features are obtained through dual-branch learning, and finally, more representative feature maps are obtained through learnable parameter fusion of the input.

[0068] More specifically, the feature maps from the dual-branch are multiplied by two parameters that are continuously updated by the network through backpropagation, and finally added together as the output of the dual-branch module.

[0069] S33: Perform bilinear interpolation on D3 to obtain a feature map with the same resolution as F2. After the bilinear-interpolated D3 and F2 are concatenated, they are input into the first main body and boundary supervision block (DBS) in the decoder to obtain the decoded feature map D2. Further perform bilinear interpolation on D2 to obtain a feature map with the same resolution as F1. After the bilinear-interpolated D2 and F1 are concatenated, they are input into the second DBS in the decoder to obtain the decoded feature map D1;

[0070] See Figure 5 , the main body and boundary supervision module first concatenates the feature map from the decoder and the feature map of the previous layer of the decoder, inputs them into a 1x1 convolution for feature dimensionality reduction, and then inputs them into a dual-branch block. One branch is used to generate the main body feature map F Body , and the other branch is used to generate the boundary feature map F BD , and each of the two branches contains two regular convolutions. The main body feature map and the boundary feature map are supervised by the main body label and the boundary label respectively, as expressed by the following formula:

[0071]

[0072] where L i represents the total loss of supervising the i-th subject and the boundary supervision block, and respectively represent the losses obtained by taking the dot product of the feature maps of the corresponding channels, which are obtained by performing 1x1 convolution on the subject feature map and the boundary feature map of the i-th subject and the boundary supervision block after upsampling them to the original resolution, with their respective corresponding labels; represents the boundary feature map in the i-th subject and the boundary supervision block, represents the subject feature map in the i-th subject and the boundary supervision block, and λ1 and λ2 are hyperparameters, with default settings both being 1.

[0073] Finally, the subject feature map and the boundary feature map are fused under learnable weight parameters to obtain more representative features. The above process is expressed by the following formula:

[0074]

[0075] where λ3 and λ4 are learnable weight parameters set during the network training process, and the initial values are set to 0.5.

[0076] S34: After processing D1 by the classifier, pixel-level classification predictions are obtained to get the final ultrasonic image segmentation result;

[0077] S4: Use the training set to train the constructed network;

[0078] The subject and boundary double-branch supervised network realizes multiple supervision methods by designing different loss functions. The total loss is denoted as L, and its loss is as follows:

[0079]

[0080] where L i represents the total loss of supervising the i-th subject and the boundary supervision block, and L final represents the loss obtained by taking the dot product of the finally obtained global feature map and the annotation map.

[0081] S5: Input the test set into the trained subject and boundary double-branch supervised network to get the segmentation result of the ultrasonic medical image to be segmented.

[0082] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. An ultrasonic image semantic segmentation method based on joint supervision of the main body and the boundary, characterized in that: It includes the following steps: S1: Collect ultrasonic image data, process and annotate it to obtain an adjusted annotated map G final , form a processed ultrasonic image data set, and divide the ultrasonic image data set into a training set and a test set; S2: Process the labeled graph G final to obtain the boundary label and the body label; S3: Build a main body and boundary double-branch supervision network, which sequentially includes an image input layer, an encoder, a decoder, and a prediction result output layer; The encoder includes five stages, each stage includes two convolutional layers, and after each convolutional layer, there are BatchNormalization batch normalization and ReLU activation function; among the two convolutional layers, the first convolutional layer is a 1x1 convolution, and the second convolutional layer is a conventional convolution with dilation. From top to bottom, the dilation rates in the conventional convolution are 1, 2, 3, 5, and 7 respectively; The decoder includes two double-branch modules DB and two main body and boundary supervision modules DBS connected in sequence; the double-branch module DB includes three parts: the first part is a 1x1 convolution for dimension adjustment of the input features; The second part is two branches with the same structure respectively connected to the first part. Each branch includes two conventional convolutions for learning rich features; the third part includes two adaptive parameters that are continuously updated through backpropagation, which are multiplied by the feature maps respectively output by the double branches for feature fusion, and then added to obtain a more representative feature map as the output of the double-branch module; The internal structure of the main body and the boundary supervision module DBS is the same as that of the double-branch module DB, which is used to splice the feature map of the encoder and the feature map of the previous layer decoder, and input it into a 1x1 convolution for feature dimension reduction processing; then input it into the double-branch block, one branch is used to generate the main body feature map F Body , and the other branch is used to generate the boundary feature map F BD , and the two branches respectively include two conventional convolutions; S4: Use the training set to train the main body and boundary double-branch supervision network; S5: Input the test set into the trained main body and boundary double-branch supervision network to obtain the segmentation result of the ultrasonic medical image to be segmented.

2. The ultrasonic image semantic segmentation method according to claim 1, characterized in that: In the step S1, the specific steps are: S11: Clean the collected ultrasonic image data to remove unqualified ultrasonic image data; S12: Label the ultrasonic image data according to the expert experience in radiology, and divide the labeled ultrasonic image data into a training set and a test set according to a ratio of 7:

3.

3. The ultrasonic image semantic segmentation method according to claim 2, characterized in that: In the step S12, the specific steps are: The images in the training set are augmented by methods including scaling, random cropping, and random flipping to achieve better performance and robustness; the labeled images are adjusted accordingly according to the operations of the training set.

4. The ultrasonic image semantic segmentation method according to claim 1, characterized in that: In the step S2, the specific steps are: S21: Divide the adjusted labeled graph G final into a foreground part G fg and a background part G bg . Perform a distance transformation on the foreground part G fg to obtain the nearest distance from each pixel in the foreground part G fg to the background part G bg , which is called the distance map; S22: Classify the corresponding pixels in the foreground with distances greater than the specified threshold in the distance map as the main body label, and classify the corresponding pixels in the foreground with distances less than or equal to the threshold as the boundary label: Let represent a single pixel point in the foreground, represent the shortest distance from the pixels in the foreground to the background, G Body represent the main body label, G BD represent the boundary label, and α represents the manually set threshold; then the classification process is: S23: Set the values of the unprocessed parts in the main body label and the boundary label to 0, and the relationship among the main body label, the boundary label, and the annotation map is: G final = G Body + G BD .

5. The ultrasonic image semantic segmentation method according to claim 1, characterized in that: In the step S3, Let L i denote the total loss for supervising the i-th subject and the boundary supervision module; and respectively denote the losses obtained by taking the losses of the subject feature map and the boundary feature map of the i-th subject and the boundary supervision module after being upsampled to the original resolution and then passing through a 1x1 convolution to obtain the feature maps of the corresponding channels, and then taking the losses with their respective corresponding labels; denote the boundary feature map of the i-th subject and the boundary supervision module; denote the subject feature map of the i-th subject and the boundary supervision module; λ1 and λ2 are hyperparameters, and their default settings are both 1; then the subject feature map F Body and the boundary feature map F BD are respectively supervised by the subject label and the boundary label, and are expressed by the following formula: Let λ3 and λ4 be the adaptive weight parameters set during network training, with the initial value set to 0.5; then the main feature map and the boundary feature map perform feature fusion under the adaptive weight parameters to obtain more representative features D i : The main body and boundary dual-branch supervision network realizes multiple supervision methods by designing different loss functions; let L i be the total loss for supervising the i-th main body and boundary supervision module, and L final be the loss obtained from the finally obtained global feature map and the annotation map. Then the total loss L is:

6. The ultrasonic image semantic segmentation method according to claim 5, characterized in that: In the step S4, the specific steps are: S41: Input the training set from the image input layer into the encoder, and obtain feature maps F1~F5 with different semantic information and different resolution sizes from top to bottom in different layers of the encoder; S42: Perform bilinear interpolation on F5 to obtain a feature map with the same resolution size as F4; splice the processed F5 and F4 and input them into the first double-branch module DB of the decoder to obtain the decoded feature map D4; perform bilinear interpolation on D4 to obtain a feature map with the same resolution size as F3; splice the bilinear interpolated D4 and F3 and input them into the second double-branch module DB of the decoder to obtain the decoded feature map D3; S43: Perform bilinear interpolation on D3 to obtain a feature map with the same resolution as F2; concatenate the bilinearly interpolated D3 and F2 and input them into the first main body and boundary supervision module DBS of the decoder to obtain the decoded feature map D2; perform bilinear interpolation on D2 to obtain a feature map with the same resolution as F1; concatenate the bilinearly interpolated D2 and F1 and input them into the second main body and boundary supervision module DBS of the decoder to obtain the decoded feature map D1; S44: After processing D1 by the classifier, obtain pixel-level classification predictions to get the final ultrasonic image segmentation result.

7. A main body and boundary double-branch supervision network for the ultrasonic image semantic segmentation method based on the joint supervision of the main body and the boundary according to any one of claims 1 to 6, characterized in that: It sequentially includes an image input layer, an encoder, a decoder, and a prediction result output layer; The encoder includes five stages, each stage includes two convolutional layers, and after each convolutional layer, there are BatchNormalization batch normalization and ReLU activation functions; among the two convolutional layers, the first convolutional layer is a 1x1 convolution, and the second convolutional layer is a regular convolution with dilation. From top to bottom, the dilation rates in the regular convolution are 1, 2, 3, 5, and 7 respectively; The decoder includes two double-branch modules DB, two main bodies, and a boundary supervision module DBS connected in sequence; the double-branch module DB includes three parts: the first part is a 1x1 convolution for adjusting the dimension of the input features; The second part is two branches with the same structure respectively connected to the first part. Each branch includes two regular convolutions for learning rich features; the third part includes two adaptive parameters that are continuously updated through backpropagation, which are multiplied by the feature maps respectively output by the double-branch and then added to obtain a more representative feature map as the output of the double-branch module; The internal structure of the main body and the boundary supervision module DBS is the same as that of the double-branch module DB, which is used to splice the feature maps of the encoder and the feature maps of the previous layer decoder, and input them into a 1x1 convolution for feature dimensionality reduction processing; then input into the double-branch block, one branch is used to generate the main body feature map F Body , and the other branch is used to generate the boundary feature map F BD , and the two branches each include two conventional convolutions.

8. A computer storage medium, characterized in that: It stores a computer program executable by a computer processor, and the computer program executes a method for ultrasonic image semantic segmentation based on joint supervision of the main body and the boundary as described in any one of claims 1 to 6.