Mammary gland ultrasound image segmentation method based on neural network

By combining the cross attention mechanism and the convolution attention module, using the U-shaped network skeleton in parallel to use the convolution layer and Swin Transformer Block, the problems of local receptive field limitation and noise interference in breast ultrasound image segmentation are solved, and high-precision breast ultrasound image segmentation is achieved.

CN119991689AActive Publication Date: 2025-05-13SICHUAN UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510126909.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-13
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

The prior art has problems of local receptive field limitation and noise interference in breast ultrasound image segmentation, resulting in inaccurate segmentation results, especially when dealing with objects of different scales and complex boundaries.

Method used

The cross attention mechanism is used to combine with the convolution attention module, and the convolution layer is used in parallel with the Swin Transformer Block through the U-shaped network framework, and the combination of pipeline attention and spatial attention is used to reduce noise interference in the convolutional low-level feature map.

Benefits of technology

The ultrasonic breast image segmentation taking into account both global and local details is realized, which significantly improves the segmentation accuracy and high-precision prediction capabilities of boundaries, and reduces the impact of noise on the segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991689A_ABST
    Figure CN119991689A_ABST
Patent Text Reader

Abstract

The invention discloses a breast ultrasound image segmentation method based on a neural network, and belongs to the technical field of image processing, and the method comprises the following steps: S1, collecting a breast ultrasound image, and carrying out the convolution processing and flattening processing of the breast ultrasound image through a convolution layer; s2, performing feature extraction to obtain an original feature map; s3, carrying out remodeling processing and flattening processing to obtain a flattened feature map; s4, inputting the flattened feature map into a convolution feature extraction module to obtain a low-layer feature map; s5, inputting the low-layer feature map into a convolution attention module to obtain a weighted feature map; s6, inputting the original feature map and the weighted feature map into a cross attention module to obtain a final feature map; and S7, completing breast ultrasound image segmentation. According to the method, noise interference in the convolutional low-level feature map is reduced by combining pipeline attention and space attention, introduction of excessive noise to a deep network is avoided, and the influence of the noise on a segmentation result is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a breast ultrasound image segmentation method based on a neural network. Background Art

[0002] At present, convolutional neural network (CNN) has achieved great success in the field of medical image segmentation, promoting the rapid development of intelligent diagnosis technology. In the field of breast ultrasound image segmentation, CNN makes full use of prior knowledge such as the local continuous distribution of tumor lesions in ultrasound images, extracts the morphological features of the tumor area through the local receptive field of the convolution kernel, and effectively separates it from the background to achieve automatic segmentation of tumor lesions. Specifically, each convolution kernel will focus on the local area of ​​the image for feature extraction instead of processing the entire image at one time. This local processing method can effectively capture local features such as texture, edge and shape at different positions in the image. Through the stacking of multiple layers of convolution and pooling, CNN can gradually learn features from low-level to high-level. In breast ultrasound images, CNN will first learn local features such as texture and edge, which will have more significant differences between the tumor area and the background area; as the network goes deeper, the features learned by the convolution kernel will become more abstract and can identify larger structures, such as the shape and boundary of the tumor. However, due to the local receptive field of the convolution kernel, the extracted features often lack globality. In addition, compared with other images, ultrasound images often have low contrast, blurred edges and local noise. Building a CNN ultrasound image segmentation network often faces local noise interference and difficulty in feature extraction, resulting in inaccurate segmentation results.

[0003] Transformer was first successful in the field of natural language processing. It uses the self-attention mechanism to establish global relationships between elements in the input sequence and has powerful global modeling capabilities. With the continuous development of the field of computer vision, Transformer has been applied to image processing. It captures the dependencies between distant pixels in the image through the self-attention mechanism, overcomes the limitations of the local receptive field in traditional CNNs, and effectively suppresses the impact of image noise on feature extraction. Although Transformer has great advantages in image processing, due to the large amount of computation, directly performing self-attention calculations on the entire image often requires expensive computing power. Therefore, the Transformer based on the window self-attention mechanism was proposed to ensure globality while reducing computational complexity.

[0004] Swin Transformer is an efficient visual Transformer architecture. Its core is to introduce the window self-attention mechanism and the sliding window mechanism. It retains the advantages of Transformer in capturing global context information while greatly reducing the computational complexity. The window self-attention mechanism reduces the amount of computation by dividing the image into multiple small windows and performing self-attention calculations inside each window, so that the self-attention calculations in each window only involve pixels in the local area, significantly reducing the computational complexity. In order to solve the limitations of information propagation caused by window self-attention, each window can only learn local information and cannot effectively model the long-range dependencies between different windows. The sliding window mechanism is introduced. By "sliding" the window, the relationship between windows between multiple Transformer layers can be gradually captured. Specifically, at each layer, Swin Transformer first divides the input image into windows of fixed size and performs self-attention calculations; then, in the next layer, the position of the window is offset in a fixed direction (such as left or up). In this way, the intersection between windows changes at each layer, so that information transfer across windows can be gradually realized. In this way, information can be gradually converged between multiple layers, and finally the modeling of global information is completed. Through multi-level window offsets, Swin Transformer can not only extract local information within each level, but also gradually capture long-range dependencies between levels, thereby enhancing the global modeling capability.

[0005] Swin Transformer reduces the computational complexity of Transformer through window self-attention mechanism and sliding window mechanism, while achieving the grasp of global and local information, alleviating the problems of local receptive field and noise interference of CNN in breast ultrasound image segmentation. However, in the task of breast ultrasound image segmentation, tumor lesions often have different sizes and complex boundaries, which require fine segmentation. Swin Transformer uses a fixed window size. When processing objects of different scales, the window size may not be able to effectively adapt to the size of different objects. Especially for smaller local areas, window self-attention cannot fully extract local detail information such as small tumor lesions or complex texture boundaries. In addition, the design of Swin Transformer focuses on global modeling, and too much global modeling may lead to local features that are not fine enough, especially when facing edges, boundaries or tiny objects, making it difficult to achieve fine segmentation of tumor lesions in breast ultrasound images.

[0006] In breast ultrasound image segmentation, CNN can effectively capture the difference between tumor lesions and background due to its local receptive field characteristics, and gradually learn features from low-level to high-level through multi-layer convolution. However, the local receptive field of CNN limits its extraction of global information, and ultrasound images usually have problems such as low contrast, blurred edges and noise interference, which makes CNN face the problem of difficult feature extraction and inaccurate segmentation results when processing details; Transformer can capture long-distance pixel dependencies in the image through the self-attention mechanism, overcoming the local receptive field limitation of CNN. SwinTransformer further reduces the computational complexity through the window self-attention mechanism and the sliding window mechanism, ensuring the modeling ability of global information. However, Swin Transformer may not be able to effectively adapt to various sizes when processing objects of different scales, and excessive attention to global information may lead to the loss of local details, especially when processing complex boundaries or small objects. Therefore, how to fully combine the advantages of CNN and Swin Transformer to build a breast ultrasound image segmentation network that can effectively model global information and accurately capture local detail features has become an important issue. Summary of the invention

[0007] In order to solve the above problems, the present invention proposes a breast ultrasound image segmentation method based on neural network.

[0008] The technical solution of the present invention is: a neural network-based breast ultrasound image segmentation method comprises the following steps:

[0009] S1, collecting breast ultrasound images, and using convolution layers to perform convolution processing and flattening processing on the breast ultrasound images in sequence;

[0010] S2, extracting features from the processed breast ultrasound image to obtain an original feature map;

[0011] S3, reshape and flatten the original feature map to obtain a flattened feature map;

[0012] S4, inputting the flattened feature map into the convolutional feature extraction module to obtain a low-level feature map;

[0013] S5, input the low-level feature map into the convolutional attention module to obtain a weighted feature map;

[0014] S6, input the original feature map and the weighted feature map into the cross attention module to obtain the final feature map;

[0015] S7. Input the final feature map into the downsampling convolution layer and the upsampling convolution layer in sequence to complete the breast ultrasound image segmentation.

[0016] Further, in S4, the convolutional feature extraction module includes a first standard convolutional layer, a first dilated convolutional layer, a depthwise separable convolutional layer, a second dilated convolutional layer, a second standard convolutional layer, a residual convolutional layer, and an adder;

[0017] The input end of the first standard convolution layer serves as the input end of the convolution feature extraction module, and the input end of the first standard convolution layer is also connected to the input end of the residual convolution layer; the output end of the first standard convolution layer is connected to the input end of the first dilated convolution layer; the output end of the first dilated convolution layer is connected to the input end of the depthwise separable convolution layer; the output end of the depthwise separable convolution layer is connected to the input end of the second dilated convolution layer; the output end of the second dilated convolution layer is connected to the input end of the second standard convolution layer; the output end of the second standard convolution layer is connected to the first input end of the adder; the output end of the residual convolution layer is connected to the second input end of the adder; the output end of the adder serves as the output end of the convolution feature extraction module.

[0018] Furthermore, S5 includes the following sub-steps:

[0019] S51, using the channel attention layer of the convolutional attention module to perform global average pooling and global maximum pooling on the low-level feature map to obtain a first channel descriptor and a second channel descriptor, and using the first channel descriptor and the second channel descriptor to process the low-level feature map;

[0020] S52. Use the spatial attention layer of the convolutional attention module to perform global average pooling and global maximum pooling on the processed low-level feature map to obtain a first spatial descriptor and a second spatial descriptor, and use the first spatial descriptor and the second spatial descriptor to obtain a weighted feature map.

[0021] Further, in S51, the first channel descriptor M avg The calculation formula is:

[0022] M avg =AvgPool(F);

[0023] Where F represents the low-level feature map, AvgPool(·) represents the global average pooling operation;

[0024] In S51, the second channel descriptor M max The calculation formula is:

[0025] M max =MaxPool(F);

[0026] Where MaxPool(·) represents the global maximum pooling operation;

[0027] In S51, the expression of the processed low-level feature map F′ is:

[0028] F′=F·A c ;

[0029] A c =σ(FC(ReLU(FC(M avg +M max ))));

[0030] In the formula, A c represents the attention weight of each channel in the multi-layer perceptron network, ReLU(·) represents the activation function, σ(·) represents the Sigmoid activation function, and FC represents the fully connected layer.

[0031] Further, in S52, the first space descriptor M avg The calculation formula of ′ is:

[0032] M avg ′ = AvgPool(F′);

[0033] Where F′ represents the processed low-level feature map, and AvgPool(·) represents the global average pooling operation;

[0034] In S52, the second space descriptor M max The calculation formula of ′ is:

[0035] M max ′ = MaxPool(F′);

[0036] Where MaxPool(·) represents the global maximum pooling operation;

[0037] In S52, the expression of the weighted feature map F″ is:

[0038] F″=F′·A s ;

[0039] A s =σ(Conv([M avg ′;M max ′]));

[0040] In the formula, A s represents the spatial attention weight map, Conv represents the convolution operation, σ(·) represents the Sigmoid activation function, and [·;·] represents the concatenation operation.

[0041] Furthermore, S6 includes the following sub-steps:

[0042] S61, using the linear layer of the cross attention module to map the original feature map into a query matrix, and map the weighted feature map into a key matrix and a value matrix;

[0043] S62, calculating the attention score matrix between the query matrix and the key matrix, and obtaining the normalized attention weight of each key to the query matrix;

[0044] S63, calculating weighted output features according to the normalized attention weights of each key to the query matrix and the value matrix;

[0045] S64, obtaining a fused feature map by adding the weighted output features and the query matrix;

[0046] S65. Map the fused feature map back to the original feature space through linear transformation to obtain the final feature map.

[0047] Furthermore, in S62, the expression of the attention score matrix A is:

[0048]

[0049] Where Q represents the query matrix, K represents the key matrix, T represents the matrix transpose, and C represents the number of channels;

[0050] In S62, the calculation formula of the normalized attention weight α of each key-to-query matrix is:

[0051] α = Softmax(A);

[0052] Where Softmax(·) represents the Softmax activation function.

[0053] Furthermore, in S63, the calculation formula of the weighted output feature O is:

[0054] O = αV;

[0055] Where α represents the normalized attention weight of each key to the query matrix, and V represents the value matrix;

[0056] Further, in S64, the fusion feature map Q fused The calculation formula is:

[0057] Q fused =Q+O;

[0058] Where O represents the weighted output feature and Q represents the query matrix.

[0059] The beneficial effects of the present invention are as follows: the present invention applies the cross attention mechanism and the convolution attention module to the design of the segmentation network architecture, relies on the U-shaped network skeleton, uses the convolution layer and the Swin Transformer Block to extract features in parallel, and uses the cross attention mechanism to fuse the output feature maps of the two to achieve both global and local detail features. The present invention uses the combination of pipeline attention and spatial attention to reduce noise interference in the convolution low-level feature map, avoids introducing too much noise into the deep network, and alleviates the impact of noise on the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of a neural network-based breast ultrasound image segmentation method;

[0061] Figure 2 Schematic diagram of the structure of the convolutional feature extraction module. DETAILED DESCRIPTION

[0062] The embodiments of the present invention will be further described below in conjunction with the accompanying drawings.

[0063] like Figure 1 As shown, the present invention provides a method for segmenting breast ultrasound images based on a neural network, comprising the following steps:

[0064] S1, collecting breast ultrasound images, and using convolution layers to perform convolution processing and flattening processing on the breast ultrasound images in sequence;

[0065] S2, extracting features from the processed breast ultrasound image to obtain an original feature map;

[0066] S3, reshape and flatten the original feature map to obtain a flattened feature map;

[0067] S4, inputting the flattened feature map into the convolutional feature extraction module to obtain a low-level feature map;

[0068] S5, input the low-level feature map into the convolutional attention module to obtain a weighted feature map;

[0069] S6, input the original feature map and the weighted feature map into the cross attention module to obtain the final feature map;

[0070] S7. Input the final feature map into the downsampling convolution layer and the upsampling convolution layer in sequence to complete the breast ultrasound image segmentation.

[0071] In an embodiment of the present invention, the image is converted into the input format expected by the Swin Transformer block through image block division and linear embedding operations. Since there are differences in the input formats expected by the convolution module and the Swin Transformer block, the feature map is reshaped and flattened to adapt to the convolution module input. The image feature extraction part is completed by the Swin Transformer Block in parallel with the convolution module, and the feature maps of the two are fused through the cross attention module. Downsampling aims to gradually reduce the spatial resolution of the image while increasing the depth of the feature, while upsampling aims to restore a high-resolution image or feature map from a low-resolution feature map for further processing or accurate prediction. At the same time, the jump connection structure of the U-shaped architecture is retained, and the feature map of each layer of the encoder is passed to the corresponding decoder layer, so that the decoding process not only depends on the upsampled feature map passed from the previous layer, but also can use the original features of the lower layer.

[0072] In the embodiment of the present invention, in S1, before inputting the image into the Swin Transformer block, the image needs to be divided into image blocks and converted into a form suitable for the Swin Transformer block input through linear embedding. The specific process is as follows:

[0073] Convolution operation: Use a convolution layer. The convolution kernel size of this convolution layer is 4×4, the step size is 4, and the number of output channels is C. The input image I∈R H×W×3 (where H and W are the height and width pixels of the image, respectively, and 3 represents the image color channel) After convolution processing, the image block division and linear embedding can be completed, specifically:

[0074] Flatten output: Flatten the output of the convolution operation to get the image tensor (Width is height is 1, the number of channels is C), which is used as the input of the Swin Transformer block.

[0075] In S2, the Swin Transformer Block in the existing Swin Transformer architecture is used for feature extraction.

[0076] In S3, in order to convert the image into the input format required by the convolutional feature extraction module and restore the original format after feature extraction, feature map reshaping and feature map flattening operations are performed.

[0077] Feature map reshaping: The input feature map (W n ×H n ,1,C n) is input to the nth layer of the network, such as C 1 =C) into a feature map (W n ,H n ,C n ).

[0078] Feature map flattening: Flatten the output of the convolutional feature extraction module, that is, flatten the feature map (W n ,H n ,C n ) is converted into a feature map (W n ×H n ,1,C n ).

[0079] like Figure 2 As shown, in S4, the convolutional feature extraction module includes a first standard convolutional layer, a first dilated convolutional layer, a depthwise separable convolutional layer, a second dilated convolutional layer, a second standard convolutional layer, a residual convolutional layer, and an adder;

[0080] The input end of the first standard convolution layer serves as the input end of the convolution feature extraction module, and the input end of the first standard convolution layer is also connected to the input end of the residual convolution layer; the output end of the first standard convolution layer is connected to the input end of the first dilated convolution layer; the output end of the first dilated convolution layer is connected to the input end of the depthwise separable convolution layer; the output end of the depthwise separable convolution layer is connected to the input end of the second dilated convolution layer; the output end of the second dilated convolution layer is connected to the input end of the second standard convolution layer; the output end of the second standard convolution layer is connected to the first input end of the adder; the output end of the residual convolution layer is connected to the second input end of the adder; the output end of the adder serves as the output end of the convolution feature extraction module.

[0081] For the 3×3 standard convolution layer, a convolution kernel with a kernel size of 3, a stride of 1, and a padding of 1 is used to achieve feature extraction without changing the size of the feature map. For the two dilated convolution layers in the module, a dilated convolution with a kernel size of 3, a stride of 1, a padding of 2, a dilation factor of 2 and a kernel size of 3, a stride of 1, a padding of 3, and a dilation factor of 3 are used respectively. There are 2 or 3 pixels of gap between the convolution kernels, which gradually increases the effective receptive field of the convolution kernel. Depthwise separable convolution is used to calculate the features of each channel using depthwise convolution, and then point-by-point convolution is used to process the feature map of each position and synthesize the output to obtain finer-grained detail features. 1×1 convolution is used as a bridge for residual connection, allowing the original feature map and the convolution output feature map to be added to avoid gradient disappearance and promote information flow.

[0082] In this embodiment of the present invention, S5 includes the following sub-steps:

[0083] S51, using the channel attention layer of the convolutional attention module to perform global average pooling and global maximum pooling on the low-level feature map to obtain a first channel descriptor and a second channel descriptor, and using the first channel descriptor and the second channel descriptor to process the low-level feature map;

[0084] S52. Use the spatial attention layer of the convolutional attention module to perform global average pooling and global maximum pooling on the processed low-level feature map to obtain a first spatial descriptor and a second spatial descriptor, and use the first spatial descriptor and the second spatial descriptor to obtain a weighted feature map.

[0085] In the embodiment of the present invention, in S51, the first channel descriptor M avg The calculation formula is:

[0086] M avg =AvgPool(F);

[0087] Where F represents the low-level feature map, AvgPool(·) represents the global average pooling operation;

[0088] In S51, the second channel descriptor M max The calculation formula is:

[0089] M max =MaxPool(F);

[0090] Where MaxPool(·) represents the global maximum pooling operation;

[0091] In S51, the expression of the processed low-level feature map F′ is:

[0092] F′=F·A c ;

[0093] A c =σ(FC(ReLU(FC(M avg +M max ))));

[0094] In the formula, A c represents the attention weight of each channel in the multi-layer perceptron network, ReLU(·) represents the activation function, σ(·) represents the Sigmoid activation function, and FC represents the fully connected layer.

[0095] In the embodiment of the present invention, in S52, the first space descriptor M avg The calculation formula of ′ is:

[0096] M avg ′ = AvgPool(F′);

[0097] Where F′ represents the processed low-level feature map, and AvgPool(·) represents the global average pooling operation;

[0098] In S52, the second space descriptor M max The calculation formula of ′ is:

[0099] M max ′ = MaxPool(F′);

[0100] Where MaxPool(·) represents the global maximum pooling operation;

[0101] In S52, the expression of the weighted feature map F″ is:

[0102] F″=F′·A s ;

[0103] A s =σ(Conv([M avg ′;M max ′]));

[0104] In the formula, A s represents the spatial attention weight map, Conv represents the convolution operation, σ(·) represents the Sigmoid activation function, and [·;·] represents the concatenation operation.

[0105] In this embodiment of the present invention, S6 includes the following sub-steps:

[0106] S61, using the linear layer of the cross attention module to map the original feature map into a query matrix, and map the weighted feature map into a key matrix and a value matrix;

[0107] S62, calculating the attention score matrix between the query matrix and the key matrix, and obtaining the normalized attention weight of each key to the query matrix;

[0108] S63, calculating weighted output features according to the normalized attention weights of each key to the query matrix and the value matrix;

[0109] S64, obtaining a fused feature map by adding the weighted output features and the query matrix;

[0110] S65. Map the fused feature map back to the original feature space through linear transformation to obtain the final feature map.

[0111] In this embodiment of the present invention, in S62, the expression of the attention score matrix A is:

[0112]

[0113] Where Q represents the query matrix, K represents the key matrix, T represents the matrix transpose, and C represents the number of channels;

[0114] In S62, the calculation formula of the normalized attention weight α of each key-to-query matrix is:

[0115] α = Softmax(A);

[0116] Where Softmax(·) represents the Softmax activation function.

[0117] In the embodiment of the present invention, in S63, the calculation formula of the weighted output feature O is:

[0118] O = αV;

[0119] Where α represents the normalized attention weight of each key to the query matrix, and V represents the value matrix;

[0120] In the embodiment of the present invention, in S64, the fusion feature map Q fused The calculation formula is:

[0121] Q fused =Q+O;

[0122] Where O represents the weighted output feature and Q represents the query matrix.

[0123] In an embodiment of the present invention, in order to verify the effectiveness of the breast ultrasound image segmentation network based on CNN and Swin Transformer, model training and evaluation are performed on the BUSI breast ultrasound dataset. The BUSI dataset has a total of 780 ultrasound images, of which 437 are benign tumors, 210 are malignant tumors, and 133 are tumor-free. The tumor-free images in the BUSI dataset are removed, and the BUSI dataset is processed using data enhancement technology. In order to balance the number of benign tumor and malignant tumor images, 2100 benign tumor data enhanced images and 2100 malignant tumor enhanced images were generated for the original dataset, for a total of 4847 images. 2906 images were randomly selected as training sets, 1294 images as test sets, and 647 original images as verification sets. There are three main reasons for using the above method to select and construct the dataset: 1) Since the BUSI dataset is relatively small, in order to avoid network overfitting, data enhancement methods (such as rotation, flipping, scaling, etc.) are used to increase the diversity of images, thereby improving the performance of the model in different scenarios and improving the generalization ability and performance of the model; 2) The BUSI dataset is challenging in terms of tumor boundaries and image noise; 3) The BUSI dataset has a wide range of applications in breast ultrasound image segmentation tasks.

[0124] Task evaluation indicators:

[0125] In terms of evaluating model performance, Dice coefficient (Dice), accuracy (Acc), intersection over union (IoU) and Hausdorff distance (HD) are used as evaluation indicators. The Dice coefficient is used to measure the similarity between the predicted and true segmentation, the accuracy is used to measure the proportion of correctly classified pixels, the intersection over union is used to measure the overlap between the predicted area and the true area, and the Hausdorff distance is used to measure the maximum error between the predicted and true boundaries, which is used for boundary accuracy evaluation.

[0126] Algorithm parameter settings:

[0127] For the designed segmentation network architecture, the loss function combines Dice loss and cross entropy loss, with weights of 0.5. The optimizer uses stochastic gradient descent (SGD) to train the model. In the training parameter setting, the initial learning rate is set to 0.0005 and gradually decreases through the CosineAnnealing LR scheduler. The model is trained on the BUSI enhanced dataset, with batch size set to 8, input resolution of 224x224, image block size of 4x4, embedding dimension of 96, the number of Transformer heads decreases with depth to [3, 6, 12, 24], the window size is 7, and 120 cycles of training are performed to obtain the final results.

[0128] Comparison of results:

[0129] The results of our method (Ours) comparing the current advanced network architectures (HCT-Net, MF-Net) and classic structures (Unet, DeepLabV3+, SwinUet) on the same dataset (BUSI) are shown in Table 1.

[0130] Table 1

[0131]

[0132] Result analysis:

[0133] In the comparative experiment, the breast ultrasound image segmentation network (Ours) based on the combination of CNN and Swin Transformer performed significantly better than other existing advanced networks and classic segmentation networks on the BUSI dataset. Specifically, the Dice coefficient, accuracy, intersection over union (IoU) and Hausdorff distance (HD) of the present invention achieved the best results in four important evaluation indicators. The Dice coefficient reached 84.59, 1.48 percentage points higher than the second place, showing stronger segmentation accuracy; the accuracy reached 98.11, 1.17 percentage points higher than the second place, indicating that the model made correct classification decisions on most pixels; the intersection over union reached 76.43, 2.64 percentage points higher than the closest MF-Net, indicating that the model can accurately capture the shape and position of the tumor; the Hausdorff distance was 17.13, 3.52 lower than MF-Net, showing that it has a significant advantage in the accuracy of the tumor boundary and can effectively reduce the boundary error.

[0134] Compared with traditional networks such as Unet and DeepLabV3+, the method of the present invention not only performs better in all indicators, but also has obvious improvements in boundary accuracy and regional overlap. For example, the Hausdorff distance is 35.63 lower than that of Unet, showing high-precision prediction ability of boundaries. Overall, the breast ultrasound segmentation network based on CNN and Swin Transformer demonstrates extremely outstanding segmentation capabilities, can more effectively handle noise interference in breast ultrasound images, and provide more accurate and reliable segmentation results.

[0135] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A method for segmenting breast ultrasound images based on neural network, characterized in that: The following steps are involved: S1, collecting breast ultrasound images, and using convolution layers to perform convolution processing and flattening processing on the breast ultrasound images in sequence; S2, extracting features from the processed breast ultrasound image to obtain an original feature map; S3, reshape and flatten the original feature map to obtain a flattened feature map; S4, inputting the flattened feature map into the convolutional feature extraction module to obtain a low-level feature map; S5, input the low-level feature map into the convolutional attention module to obtain a weighted feature map; S6, input the original feature map and the weighted feature map into the cross attention module to obtain the final feature map; S7. Input the final feature map into the downsampling convolution layer and the upsampling convolution layer in sequence to complete the breast ultrasound image segmentation.

2. The neural network-based breast ultrasound image segmentation method according to claim 1, characterized in that: In S4, the convolution feature extraction module includes a first standard convolution layer, a first dilated convolution layer, a depthwise separable convolution layer, a second dilated convolution layer, a second standard convolution layer, a residual convolution layer and an adder; The input end of the first standard convolutional layer serves as the input end of the convolutional feature extraction module, and the input end of the first standard convolutional layer is also connected to the input end of the residual convolutional layer; the output end of the first standard convolutional layer is connected to the input end of the first dilated convolutional layer; the output end of the first dilated convolutional layer is connected to the input end of the depthwise separable convolutional layer; the output end of the depthwise separable convolutional layer is connected to the input end of the second dilated convolutional layer; the output end of the second dilated convolutional layer is connected to the input end of the second standard convolutional layer; the output end of the second standard convolutional layer is connected to the first input end of the adder; the output end of the residual convolutional layer is connected to the second input end of the adder; the output end of the adder serves as the output end of the convolutional feature extraction module.

3. The neural network-based breast ultrasound image segmentation method according to claim 1, characterized in that: The S5 comprises the following sub-steps: S51, using the channel attention layer of the convolutional attention module to perform global average pooling and global maximum pooling on the low-level feature map to obtain a first channel descriptor and a second channel descriptor, and using the first channel descriptor and the second channel descriptor to process the low-level feature map; S52. Use the spatial attention layer of the convolutional attention module to perform global average pooling and global maximum pooling on the processed low-level feature map to obtain a first spatial descriptor and a second spatial descriptor, and use the first spatial descriptor and the second spatial descriptor to obtain a weighted feature map.

4. The neural network-based breast ultrasound image segmentation method according to claim 3, characterized in that: In S51, the first channel descriptor M avg The calculation formula is: M avg =AvgPool(F); Where F represents the low-level feature map, AvgPool(·) represents the global average pooling operation; In S51, the second channel descriptor M max The calculation formula is: M max =MaxPool(F); Where MaxPool(·) represents the global maximum pooling operation; In S51, the processed low-level feature map F ′ The expression is: F ′ =F·A c ; And c =σ(FC(ReLU(FC(M avg +M max )))); In the formula, A c represents the attention weight of each channel in the multi-layer perceptron network, ReLU(·) represents the activation function, σ(·) represents the Sigmoid activation function, and FC represents the fully connected layer.

5. The neural network-based breast ultrasound image segmentation method according to claim 3, characterized in that: In S52, the first space descriptor M avg The calculation formula of ′ is: M avg ′=AvgPool(F′); In the formula, F ′ represents the processed low-level feature map, AvgPool(·) represents the global average pooling operation; In S52, the second space descriptor M max The calculation formula of ′ is: M max ′=MaxPool(F′); Where MaxPool(·) represents the global maximum pooling operation; In S52, the expression of the weighted feature map F″ is: F″=F′·A s ; A s =σ(Conv([M avg ′;M max ′])); In the formula, A s represents the spatial attention weight map, Conv represents the convolution operation, σ(·) represents the Sigmoid activation function, and [·;·] represents the concatenation operation.

6. The neural network-based breast ultrasound image segmentation method according to claim 1, characterized in that: The S6 comprises the following sub-steps: S61, using the linear layer of the cross attention module to map the original feature map into a query matrix, and map the weighted feature map into a key matrix and a value matrix; S62, calculating the attention score matrix between the query matrix and the key matrix, and obtaining the normalized attention weight of each key to the query matrix; S63, calculating weighted output features according to the normalized attention weights of each key to the query matrix and the value matrix; S64, obtaining a fused feature map by adding the weighted output features and the query matrix; S65. Map the fused feature map back to the original feature space through linear transformation to obtain the final feature map.

7. The neural network-based breast ultrasound image segmentation method according to claim 6, characterized in that: In S62, the expression of the attention score matrix A is: Where Q represents the query matrix, K represents the key matrix, T represents the matrix transpose, and C represents the number of channels; In S62, the calculation formula of the normalized attention weight α of each key-to-query matrix is: α = Softmax(A); Where Softmax(·) represents the Softmax activation function.

8. The neural network-based breast ultrasound image segmentation method according to claim 6, characterized in that: In S63, the calculation formula of the weighted output feature O is: O = αV; Where α represents the normalized attention weight of each key on the query matrix, and V represents the value matrix.

9. The neural network-based breast ultrasound image segmentation method according to claim 6, characterized in that: In S64, the fusion feature map Q fused The calculation formula is: Q fused =Q+O; Where O represents the weighted output feature and Q represents the query matrix.

Citation Information

Patent Citations

  • Ultrasonic automatic mammary gland full-volume image identification method and device, terminal and medium

    CN115187530A

  • Medical image depth segmentation method based on fuzzy logic

    CN116188435A

  • Breast cancer pathological image classification device and method based on deep learning

    CN116524226A

  • Medical image automatic segmentation method of U-shaped network based on fusion convolution and attention mechanism

    CN117474866A

  • Mammary gland ultrasound image segmentation method based on denoising diffusion probability model

    CN118196121A