End-to-End Segmentation Method of Breast Ultrasound Nodules Based on Multi-Scale and Cross-Spatial Fusion
By introducing multi-scale feature extraction and fusion module, receptive field adaptive aggregation module and cross-spatial residual fusion module in U-Net network, the problem of insufficient accuracy and boundary accuracy in phacodile nodule segmentation is solved, and more efficient breast nodule segmentation and diagnostic auxiliary effects are achieved.
Patent Information
- Application Number
- CN202211177184.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-09-26
AI Technical Summary
The prior art is difficult to meet the clinical needs of ultrasound physicians in phacodile nodules segmentation, and the imaging characteristics of ultrasound imaging and the diversification of breast nodules shapes increase the difficulty of segmentation.
The end-to-end segmentation method of phacopad nodules based on multi-scale and cross-space fusion is adopted. By introducing multi-scale feature extraction and fusion modules, receptive field adaptive aggregation modules and cross-space residual fusion modules in traditional U-Net networks, the network's feature extraction and information fusion capabilities are enhanced.
It improves the accuracy and boundary accuracy of phacodile nodules segmentation, reduces the misdiagnosis rate and missed diagnosis rate, and can assist doctors in quickly and accurately diagnosing mammary nodules.
Smart Images

Figure CN115578559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field, and in particular to an end-to-end ultrasound breast nodule segmentation method based on multi-scale and cross-space fusion. Background Art
[0002] According to statistics from the World Health Organization, breast cancer has become one of the most common malignant tumors in women worldwide, seriously threatening women's physical and mental health. Clinical experience shows that the exact pathogenesis of breast cancer is still unclear, related high-risk factors are difficult to control, and primary etiology prevention is difficult to achieve. Therefore, the current prevention and control of breast cancer is mainly based on secondary prevention of "early detection, early diagnosis, and early treatment". Therefore, early screening and diagnosis are key factors in reducing breast cancer mortality. Nowadays, non-invasive breast diagnosis has developed rapidly, including X-ray, magnetic resonance imaging, ultrasound imaging and other technologies. Ultrasound imaging has the advantages of no radiation damage, easy to use, imaging can be observed at any angle, fast imaging speed, and low price, and has become the most important method and means for early diagnosis of breast cancer. However, ultrasound imaging has the disadvantages of high noise, uneven grayscale, and low contrast. In addition, breast nodules have different morphologies and textures, and benign and malignant nodules are difficult to distinguish with the naked eye, which brings certain difficulties to ultrasound breast nodule detection. In order to solve the above problems, a computer-aided diagnosis (CAD) system based on artificial intelligence algorithm is used in the diagnosis of ultrasound breast images.
[0003] In early algorithm research, scholars mostly used traditional machine learning methods to segment ultrasound breast nodules. For example, in 2012, Jiang et al. [1] First, the Adaboost+Haar framework is used to detect the initial ultrasound breast nodule area, then the support vector machine is used to further filter the set of detected nodules, and finally the random walk algorithm is used to refine the segmentation of the nodules. Shan et al. [2] proposed the NLM (neutrosophic L-means) generalized clustering method for ultrasound breast single nodule fuzzy boundary segmentation. In 2016, Luo et al. [3] The particle swarm algorithm and the optimized graph theory algorithm are combined to achieve the segmentation of ultrasound breast nodules, but two diagonal lines need to be manually selected to determine the nodule area. In 2018, Liu et al. [4] Adaptive thresholding and morphological filtering methods were used to locate and initialize the nodule contour. The nodule contour was further refined by improving the active contour method. However, since only the largest area was retained as the initial nodule contour, it was not suitable for the segmentation of multiple nodules. Lotfollahi et al. [5]An improved active contour method is proposed, and combined with neutrosophic theory for ultrasound breast image segmentation, which overcomes the inherent speckle noise of ultrasound images. The disadvantage is that the nodule contour requires manual operation. All of the above methods can achieve ultrasound breast nodule segmentation, but they do not get rid of the tedious multi-step processing or manual initial contour. Therefore, researchers consider taking a different approach to seek a relatively simple and efficient segmentation method.
[0004] In recent years, with the development of deep learning in the field of image processing, its role in the medical field has gradually emerged. In the task of ultrasound breast nodule segmentation, in view of the shortcomings of traditional machine methods, many scholars have begun to study the use of deep learning methods. In 2019, Han et al. [6] The multi-scale feature extraction network (BUS-S) and the dual attention fusion network (BUS-E) were used for ultrasound breast nodule segmentation, but the training load was too large. Zhuang et al. [7] Based on the original UNet model, each ordinary convolutional layer introduces cross-layer connections to alleviate the vanishing gradient, uses dilated convolutions in the bottleneck layer to obtain more features, and uses attention gate modules to replace the cross-layer connections of the encoder and decoder to suppress background information and improve the segmentation performance of breast nodules. In 2020, Byra et al. [8] The standard convolution in the encoding and decoding path is replaced by a selective kernel convolution, which can adaptively adjust the receptive field and solve the segmentation problem of breast nodules with variable morphology. Vakanski et al. [9] Adding attention to the pooling operation is used to focus on the ultrasound breast nodule area, but it is not effective for segmenting the fuzzy boundaries of nodules. In 2021, Xue et al.
[10] Combining multi-layer contextual information with breast lesion boundary detection, the boundary quality is refined to improve the segmentation results. Iqbal et al.
[11] A multi-scale dual attention network based on ultrasound breast nodule image segmentation is proposed. Multi-scale convolution replaces all standard convolutions in Unet for feature extraction. Dual attention is introduced after upsampling to adaptively learn high-level features, but the segmentation effect of nodules with fuzzy boundaries is poor. In 2022, Punn et al.
[12] By fusing four different scale convolutions to adapt to ultrasound breast nodule regions of different sizes, mixed pooling is used to retain more nodule feature information. In addition, the adjacent spatial features of the encoding path and the spatial features of the equivalent layer of the decoding path are combined with attention to focus on the correlation in different spatial dimensions to better identify the nodule region. Chen et al.
[13] It is proposed to improve the Unet network with hybrid adaptive attention, so that the network has the ability to adaptively adjust the receptive field in channels and space to capture features of different dimensions. This method achieves better segmentation results.
[0005] Although deep learning methods achieve better segmentation results than traditional machine learning methods, their segmentation accuracy still cannot meet the clinical needs of ultrasound physicians.
[0006] References:
[0007] [1]JIANG P,PENG J,ZHANG G,etal.Learning-based automatic breast tumor detection and segmentation in ultrasound images[C] / / 2012 9th IEEEInternational Symposium on Biomedical Imaging(ISBI).IEEE,2012:1587-1590.
[0008] [2]SHAN J, CHENG HD, WANG YA novel segme-ntation method for breastultrasound images based on neutrosophic l-means clustering[J]. Me-dicalphysics, 2012, 39(9):5669-5682.
[0009] [3]LUO Y,HAN S,HUANG QA novel graph-based segmentation method forbreast ultrasound images[C] / / 2016International Conference on Digital ImageComputing:Techniques and Applications(DICTA).IEEE,2016:1-6.
[0010] [4]LIU L, LI K, QIN W, et al. Automated breast tumor detection and segmentation with a novel computati-onal framework of whole ultrasound images[J]. Medical&biological engineering&computing, 2018, 56(2):183-199.
[0011] [5]LOTFOLLAHI M,GITY M,YE J Y,etal.Segme-ntation of breast ultrasoundimages based on active contours using neutrosophic theory[J].Journal ofMedical Ultrasonics,2018,45(2):205-212.
[0012] [6]HAN L,HUANG Y,DOU H,et al.Semi-supervised segmentation of lesionfrom breast ultrasound images with attentional generative adversarial network[J].Computer methods and programs in biomedicine,2020,189:105275.
[0013] [7]ZHUANG Z,LI N,JOSEPH RAJ A N,et al.An RDAU-NET model for lesionsegmentation in breast ultrasound images[J].PloS one,2019,14(8):e0221535.
[0014] [8]BYRA M,JAROSIK P,SZUBERT A,et al.Breast mass segmentation inultrasound with selective kernel U-Ne t convolutional neural network[J].Biomedical Signal Processing and Control,2020,61:102027.
[0015] [9]VAKANSKI A,XIAN M,FREER P E.Attention-enriched deep learning modelfor breast tumor segmentation in ultrasound images[J].Ultrasoun-d inMedicine&Biology,2020,46(10):2819-2833.
[0016]
[10] XUE C,ZHU L,Fu H,et al.Global guidance network for breast lesionsegmentation in ultraso-und images[J].Medical image analysis,2021,70:101989.
[0017]
[11] IQBAL A,SHARIF M.MDA-Net:Multiscale dual attention-based networkfor breast lesion segmentation using ultrasound images[J].Journal of KingSaud University-Computer and Informa-tion Sciences,2021.
[0018]
[12] PUNN N S,AGARWAL S.RCA-IUnet:a residual cross spatial attention-guided inception U-Net model for tumor segmentation in breast ultrasou-ndimaging[J].Machine Vision and Applications,2022,33(2):1-10.
[0019]
[13] CHEN G,DAI Y,ZHANG J,et al.AAU-net:An Adaptive Attention U-netfor Breast Lesions Segmentation in Ultrasound Images[J].arXiv preprint arXiv:2204.12077,2022.
[0020]
[14] HAN K,WANG Y,TIAN Q,et al.Ghostnet:More features from cheapoperations[C] / / Procee-dings of the IEEE / CVF conference on computer vision andpattern recognition.2020:1580-1589.
[0021]
[15] Wu H,Wang W,Zhong J,et al.Scs-net:A scale and context sensitivenetwork for retinal vessel segmentation[J].Medical Image Analysis,2021,70:102025.
[0022]
[18] OKTAY O,SCHLEMPER J,FOLGOC L L,et al.Attention u-net:Learningwhere to look for the pancreas[J].arXiv preprint arXiv:1804.03999,2018.
[0023]
[19] JHA D,SMEDSRUD P H,RIEGLER M A,et al.Resunet++:An advancedarchitecture for medical image segmentation[C] / / 2019 IEEE InternationalSymposium on Multimedia(ISM).IEEE,2019:225-2255.
[0024]
[20] LI X,WANG W,HU X,et al.Selective kernel networks[C] / / Proceedingsof the IEEE / CVF conference on computer vision and pattern recognition.2019:510-519.
[0025]
[21] LOU A, GUAN S, LOEW M. Cfpnet-m: A light-weight encoder-decoder based network for multi-modal biomedical image real-time segmentation[J]. arXiv preprint arXiv:2105.04075, 2021.
[0026]
[22] NING Z, WANG K, ZHONG S, et al. CF2-Net: Coarse-to-fine fusion convolutional network for breast ultrasound image segmentation[J]. arXiv preprint arXiv:2003.10144, 2020.
[0027]
[23] SHAREEF B, VAKANSKI A, Xian M, et al. Estan: Enhanced small tumor-aware network for breast ultrasound image segmentation[J]. arXiv preprint arXiv:2009.12894, 2020.
[0028]
[24] PI J, QI Y, LOU M, et al. FS-UNet: Mass segment-ation in mammograms using an encoder-decoder architecture with feature strengthening[J]. Computers in Biology and Medicine, 2021, 137:104800 Summary of the Invention
[0029] According to the above proposal, the segmentation accuracy still cannot meet the clinical needs of ultrasound doctors. In addition, the imaging characteristics of ultrasound images and the diversity of breast nodule shapes objectively increase the technical difficulty of breast nodule segmentation. A method for end-to-end segmentation of ultrasound breast nodules based on multi-scale and cross-space fusion is provided. The present invention mainly adopts the encoding-decoding structure of the traditional U-Net as the main framework. Its encoder replaces the convolution operation of the traditional Unet with the Multi-scale Feature Extraction and Fusion module (MFEF) to obtain contextual information of different receptive fields to improve the extraction and expression capabilities of shallow network features; in the bottleneck layer, the Receptive-fields Adaptive Aggregation module (RAA) is used to generate a weight matrix by fusing information under the multi-scale receptive field, and the deep semantic channel is feature-screened by weights to highlight the semantic features related to the segmentation results to enhance the extraction and expression capabilities of deep features in the encoding stage; on the jump connection between the encoder and the decoder, a Cross-spatial Residuals Fusion module (CRF) is added to alleviate the semantic differences between the codec peer layers, thereby better compensating for the information loss in the decoding stage. Through ablation experiments, it is verified that the combined use of the above three modules can achieve the best segmentation effect.
[0030] The technical means adopted by the present invention are as follows:
[0031] An end-to-end ultrasound breast nodule segmentation method based on multi-scale and cross-space fusion, comprising:
[0032] Step 1: The encoding-decoding structure of the traditional U-Net is used as the main framework, wherein the encoder replaces the convolution operation of the traditional U-Net with a multi-scale feature extraction and fusion module to obtain context information of different receptive fields;
[0033] Step 2: At the bottleneck layer, the information under different receptive fields obtained by the multi-scale feature extraction and fusion module is fused to generate a weight matrix through the receptive field adaptive aggregation module, and the deep semantic channel is feature screened through the weights to highlight the semantic features related to the segmentation results;
[0034] Step 3: In the skip connection between the encoder and decoder, a cross-space residual fusion module is added to alleviate the semantic differences between the encoder and decoder equivalent layers;
[0035] Step 4: Verify the segmentation effect through ablation experiments.
[0036] Compared with the prior art, the present invention has the following advantages:
[0037] The end-to-end ultrasound breast nodule segmentation method based on multi-scale and cross-space fusion invented in this paper has made the following innovative improvements based on U-Net:
[0038] First, by adding a multi-scale feature extraction and fusion module, the feature extraction receptive field is expanded, the nonlinearity of the module is enhanced, and the network's ability to capture target features is enhanced;
[0039] Secondly, the bottleneck layer is at the deep stage of the encoder and contains more advanced context information. Therefore, its feature extraction should consider the importance differences between different semantic features, but existing methods rarely consider this. The present invention adopts a receptive field adaptive aggregation module in the bottleneck layer to weight the convolution features under different receptive fields, so that important semantic features play a more significant role in the segmentation process, thereby improving the segmentation effect;
[0040] Finally, there are differences in semantic features between the encoder and decoder layers of the traditional U-Net network. The present invention adds a cross-space residual fusion module, changes the direct connection and splicing jump connection mode in the U-Net network, and constructs a nonlinear jump connection mode between the encoding path and the decoding path, which not only realizes the information complementarity between different encoding layers, but also alleviates the semantic differences between the encoder and decoder layers.
[0041] The present invention can automatically extract the nodule lesion area of ultrasound breast images, has a high segmentation accuracy and a relatively precise segmentation boundary, can assist doctors in making rapid and accurate diagnosis of nodule lesions, reduce the misdiagnosis rate and missed diagnosis rate, and alleviate the current situation of a relative shortage of excellent ultrasound doctors in primary hospitals. The extracted area can also serve as the basis for subsequent automatic discrimination of benign and malignant nodule lesions, and has very important research value and application prospects for promoting computer-assisted ultrasound medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0043] Figure 1 Schematic diagram of the overall structure of the segmentation method of the present invention.
[0044] Figure 2 Schematic diagram of the multi-scale feature extraction and fusion module of the present invention.
[0045] Figure 3 Schematic diagram of the receptive field adaptive aggregation module of the present invention.
[0046] Figure 4 Schematic diagram of the cross-space residual fusion module of the present invention.
[0047] Figure 5 The diagram is a comparative diagram of the ablation experiment of the present invention. (a) is the original image; (b) is segmentation using only the original Unet network; (c) is adding the MFEF module, RAA module, and CRF module separately on the basis of the Unet network architecture; (d) is adding two modules in a mixed manner on the basis of the Unet network architecture; (e) is adding all three modules on the basis of the Unet network architecture.
[0048] Figure 6 Schematic diagram of the segmentation results of different models of the present invention. Among them, (a) is the original image; (b) AttUnet; (c) ResUNet++; (d) AKUNet; (e) CFPNet; (f) the method of the present invention. DETAILED DESCRIPTION
[0049] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0051] like Figure 1-4 As shown, the present invention provides an end-to-end ultrasound breast nodule segmentation method based on multi-scale and cross-space fusion, comprising:
[0052] Step 1: The encoding-decoding structure of the traditional U-Net is used as the main framework, wherein the encoder replaces the convolution operation of the traditional U-Net with a multi-scale feature extraction and fusion module to obtain contextual information of different receptive fields.
[0053] Breast nodules under ultrasound images often show complex morphological features, such as different sizes and shapes (malignant nodules often present various irregular shapes), unclear boundaries between the edges of some nodules and surrounding tissues, and low contrast of ultrasound images and the presence of artifact interference, all of which pose great challenges to the accurate segmentation of nodules. The encoding layer of the traditional UNet network uses two consecutive 3×3 convolutions when extracting features. Therefore, it can only extract feature information under a single-scale receptive field, and has poor adaptability to irregular changes in nodule features. As an embodiment of the present application, a multi-scale feature extraction and fusion module (MFEF) in the present application is based on four convolution branches with different receptive fields as the main line, and its structure is as follows: Figure 2 shown.
[0054] Among them, the first branch is a 1×1 convolution for feature extraction, with a 1×1 minimum receptive field; the second to fourth branches are 3×3 convolutions for feature extraction, with a 3×3 receptive field;
[0055] The third branch and the fourth branch are both provided with a multi-scale convolution block, namely the SplitB module; and the multi-scale feature extraction and fusion module splices and fuses the back ends of the other three branches in pairs while keeping the convolution result of the first branch unchanged. This design enables the entire module to form a nested effect of multi-scale convolution, which not only ensures multi-scale feature extraction, but also enhances the nonlinearity of the module. In addition, while keeping the convolution result of the first branch unchanged, the module splices and fuses the back ends of the other three branches in pairs, in order to retain more feature information under adjacent receptive fields.
[0056] Preferably, the SplitB module is inspired by Ghost-net to extract features and expand the receptive field of input features through grouped convolution and dilated convolution; the SplitB module extracts features and expands the receptive field of input features through grouped convolution and dilated convolution; the SplitB module divides the input feature channels evenly and multiple times into two groups, group1 and group2, as shown in the figure; and then the two groups, i.e., group1 or group2, are grouped again, i.e., group3 and group4; and then any group after the secondary grouping, i.e., group3 or group4, is subjected to feature extraction through dilated convolution, such as group3 and group4 in the figure, and then group4 is spliced with group2, and feature extraction is performed using dilated convolution, and finally the obtained feature map is spliced with group3, and finally the obtained feature map is spliced with a group that is not subjected to secondary grouping, i.e., group2 or group1, to form the final module output. This form of convolution and merging after grouping can effectively reduce model parameters and reduce the amount of calculation, while increasing the complexity of model connection and improving the nonlinearity of the module.
[0057] The multi-scale feature extraction and fusion module in Section 1 is only applicable to shallow feature extraction. The reason is that when extracting features, the model splices channel features of different scales with equal weights. This processing method can make every detail of shallow features be paid attention to. However, after entering the bottleneck layer, the shallow features are further abstracted into high-level semantic features. The abstractness of the features is enhanced, making the influence of semantic features of different channels on the final segmentation results different. Therefore, in addition to considering the multi-scale receptive field, feature extraction at the bottleneck layer should also consider the importance differences between different semantic features.
[0058] Therefore, step 2: at the bottleneck layer, the information under different receptive fields obtained by the multi-scale feature extraction and fusion module is fused to generate a weight matrix through the receptive field adaptive aggregation module, and the deep semantic channel is feature-screened through the weights to highlight the semantic features related to the segmentation results. The receptive field adaptive aggregation module in step 2 performs weight screening on the convolution features under different receptive fields; the receptive field adaptive aggregation module has 4 convolution branches, and the receptive fields of each convolution branch are 1×1, 3×3, 5×5 and 7×7 respectively. The feature maps F1, F2, and F3 of the 3×3, 5×5, and 7×7 convolution branches are cascaded and fused in pairs. The fused information generates a weight matrix through the softmax function and then performs channel-by-channel weighted addition with the feature maps F1, F2, and F3, thereby realizing the adaptive selection of convolution features under different receptive fields, thereby realizing the adaptive selection of convolution features under different receptive fields. The essence of this adaptive selection is to highlight the useful information and weaken the useless information in the multi-scale receptive field through the weight matrix. The selected result is superimposed with the original 1×1 convolution branch feature to ensure that information is not lost while highlighting the strengthened information.
[0059] To compensate for the information loss caused by the decoding stage, the classic U-Net network uses skip connections to directly transfer the low-level spatial information of the encoding part and splice it to the decoding part. However, the decoder is characterized by the formation of higher-level context information after multiple convolutions and pooling in the encoding stage, which makes the features between the encoder and decoder layers have large semantic differences. If the lower-level features are directly transferred to the higher level through skip connections, a semantic gap will be generated.
[0060] Step 3: In the jump connection between the encoder and the decoder, a cross-space residual fusion module is added to alleviate the semantic differences between the codec peer layers; this structure of first parallel, then series, and then parallel allows the information on the encoder side to be fully fused during the jump transmission process to the decoder side. This fusion, on the one hand, realizes the information complementarity between different coding layers, and on the other hand, further extracts features from the coding layer information, thereby alleviating the semantic differences between the codec peer layers. In addition, the introduction of residual branches makes the network easier to optimize. The cross-space residual fusion module includes five dual-space fusion modules; the five dual-space fusion modules form a cross-fusion network structure in the form of series connection and cross-parallel connection;
[0061] Features F of four different encoding layers s, s∈{1,2,3,4} are input into the first two parallel DSFs in CRF in pairs according to the adjacent order, so that the two adjacent codes in the encoder are fused; the fused outputs are then input into three serial-parallel DSF modules for further fusion. The dual-space fusion module includes: 4 convolution branches, the first branch uses 3×3 convolution to extract features from F3, the second branch uses 3×3, stride 2 downsampling operation to extract features from F3, and at the same time, the third branch uses upsampling operation to extract features from F4, and the fourth branch uses 3×3 convolution to extract features from F4;
[0062] The first branch is cascaded and merged with the third branch, and the second branch is cascaded and merged with the fourth branch.
[0063] Step 4: Verify the segmentation effect through ablation experiments.
[0064] Embodiment 1:
[0065] As an embodiment of the present application, the effectiveness of each designed module and its combination for the ultrasonic breast nodule segmentation task is verified by ablation experiments. In order to fully verify, this ablation experiment conducted 4 groups of experiments. The first group used only the original Unet network for segmentation; the second group added MFEF module, RAA module, and CRF module separately on the basis of the Unet network architecture; the third group added two mixed modules on the basis of the Unet network architecture; the fourth group added all three modules on the basis of the Unet network architecture. The results of the ablation experiment are shown in Table 1.
[0066] As can be seen from the table, the values of each index of the second group of models have been improved on the basis of the UNet model, which shows that the use of each module is effective. However, compared with the third group of models, it is found that adding two modules improves the index value more than using one module alone, which shows that the mixed use of modules can produce an advantage accumulation, and this advantage accumulation effect is more obvious in the fourth group of experiments. Except for the precision index, the other indicators of the fourth group of experimental results have been significantly improved compared with the results of the third group, with Dice, Recall, and IOU increasing by an average of 0.258, 0.307, and 0.243 respectively. Although the precision index of the fourth group is slightly lower than that of the second and third groups, the gap is not very obvious, and the comprehensive improvement of all indicators is more important. In addition, in the process of medical image segmentation, the recall rate is a very important indicator, which shows how many diseased tissues are found. Therefore, when the precision index is similar, the experiment pays more attention to the improvement of the recall rate. The above analysis shows that the best evaluation index is obtained by using MFEF module, RAA module and CRF module based on Unet network architecture. This is also reflected in the visual effect comparison of segmentation results. Figure 5 shown.
[0067] Figure 5 This is a visual comparison of the segmentation results of the 1st to 4th group of experiments. The green outline is the nodule area annotation given by the expert, and the red outline represents the segmentation result of the model in this paper. As can be seen from the figure, there are many false positives in the segmentation results using the Unet model alone, as shown in the areas a, b, and c in the figure. Adding the MFEF module on the basis of the Unet network significantly eliminates the false positives, but there is still a certain gap between the segmentation boundary and the expert-annotated boundary; after adding the MFEF, RAA, and CRF modules, there is a more complete outline, retaining more breast nodule boundary details, and the segmentation result is morphologically closer to the annotation image, with better performance.
[0068] Table 1 Network structure ablation comparison
[0069] Table 1 Comparison of network structure ablation
[0070]
[0071] To further verify the effectiveness of the proposed method, we first compare the proposed model with existing classic segmentation models such as AttUNet
[18] 、ResUnet++
[19] SKUnet
[20] 、CFPNet
[21] This paper uses these classic models to achieve ultrasound breast nodule segmentation and evaluates the segmentation results, as shown in Table 2. As can be seen from the table, the evaluation index of the proposed method is the best among all the compared models, especially the Dice and IOU indicators are significantly improved, indicating that the segmentation results of this method are closest to the gold standard given by doctors. This is also reflected in the visual effect comparison of the segmentation results, such as Figure 6 As shown in the figure, the green outline is the nodule area annotation given by the expert, and the red one is the segmentation result of the model in this paper. It can be seen from the figure that the proposed method can locate the nodule area better than other comparison methods, and identify more lesion areas. For nodules with blurred boundaries, the segmentation boundary is closer to the boundary marked by experts. From a comprehensive comparison, the method proposed in this paper is better than other comparison networks in terms of the ability to distinguish nodules from the background and the morphology of the segmentation results.
[0072] In addition, this paper also compares with 4 papers on ultrasound breast nodule segmentation. The comparison results are shown in Table 3. The papers involved in the comparison all use the ultrasound breast nodule public dataset provided by Yap et al., just like this paper. As can be seen from Table 3, the evaluation indicators of the proposed method are all higher than those of other literature methods.
[24] The index of this method is second only to that of this paper. The dice index and IOU index of the two are very similar, but the recall rate of this paper is significantly higher than that of the literature.
[24] , indicating that compared with the literature
[24] , the false negative segmentation area of this paper is reduced, and more lesions are segmented out.
[0073] Table 2 Comparison of different segmentation algorithms
[0074] Table 2 Comparison of different segmentation algorithms
[0075]
[0076] Table 3 Comparison of different segmentation algorithms for the same dataset
[0077] Table 3 Comparison of different segmentation algorithms for the same dataset
[0078]
[0079] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An end-to-end segmentation method for ultrasonic breast nodules based on multi-scale and cross-space fusion, characterized in that, it includes: Step 1: Use the encoding-decoding structure of the traditional U-Net as the main framework. Among them, the encoder replaces the convolution operation of the traditional Unet through a multi-scale feature extraction and fusion module to obtain context information with different receptive fields; the multi-scale feature extraction and fusion module contains 4 convolutional branches with different receptive fields; among them, the first branch performs feature extraction through 1×1 convolution and has a minimum receptive field of 1×1; the second to fourth branches perform feature extraction through 3×3 convolution and have a receptive field of 3×3; However, after the 3×3 convolution of the third branch and the fourth branch, a multi-scale convolution block, namely the SplitB module, is set; and the multi-scale feature extraction and fusion module performs pairwise splicing and fusion on the back ends of the other three branches respectively while keeping the convolution result of the first branch unchanged; the SplitB module performs feature extraction and expands the receptive field on the input features through grouped convolution and dilated convolution; the SplitB module divides the input feature channels into two groups, group1 and group2; then any one of the two groups, that is, group1 or group2, is grouped again, that is, divided into group3 and group4; then any one of the groups after the secondary grouping, that is, group3 or group4, is subjected to feature extraction through dilated convolution, and finally the obtained feature map is spliced with the group that has not been subjected to secondary grouping, that is, group2 or group1, to form the final module output; Step 2: At the bottleneck layer, the information under different receptive fields obtained through the multi-scale feature extraction and fusion module is fused by the receptive field adaptive aggregation module to generate a weight matrix, and the deep semantic channels are feature-screened through the weights to highlight the semantic features related to the segmentation result; Step 3: On the skip connection between the encoder and the decoder, a cross-space residual fusion module is added to alleviate the semantic differences between the equivalent layers of the encoder and decoder; the cross-space residual fusion module includes five dual-space fusion modules; the five dual-space fusion modules form a cross-fusion network structure in the form of series connection and cross-parallel connection; For the features F of four different coding layers s , s ∈ {1, 2, 3, 4} are input into the first two parallel dual - space fusion modules DSF in the cross - space residual fusion module CRF in pairs in adjacent order, so as to fuse the information of two adjacent encodings in the encoder; the fused outputs are then respectively input into three series - parallel dual - space fusion modules DSF for further fusion; the dual - space fusion module includes: 4 convolutional branches. The first branch uses a 3×3 convolution to extract features from F3, the second branch uses a 3×3 downsampling operation with a stride of 2 to extract features from F3. At the same time, the third branch uses an upsampling operation to extract features from F4, and the fourth branch uses a 3×3 convolution to extract features from F4; The first branch and the third branch are cascaded and fused, and the second branch and the fourth branch are cascaded and fused; Step 4: Through ablation experiments, verify the segmentation effect.
2. The end-to-end segmentation method for ultrasonic breast nodules based on multi-scale and cross-space fusion according to claim 1, characterized in that, In the step 2, the receptive field adaptive aggregation module performs weight screening on the convolution features under different receptive fields; the receptive field adaptive aggregation module has 4 convolution branches, and the receptive fields of each convolution branch are 1×1, 3×3, 5×5, and 7×7 respectively. The feature maps F1, F2, and F3 of the 3×3, 5×5, and 7×7 convolution branches are cascaded and fused pairwise. The fused information generates a weight matrix through the softmax function and then performs channel-wise weighted addition with the feature maps F1, F2, and F3, so as to realize the adaptive selection of convolution features under different receptive fields.