Breast mass detection method and device for enhancing feature extraction

By introducing the AE-YOLO network of aggregated dynamic convolution and visual enhancement block modules into the YOLOv8 network, the problem of insufficient accuracy of breast mass detection is solved, and higher detection performance and robustness are achieved.

CN120451042APending Publication Date: 2025-08-08SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510366206.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing deep learning-based breast mass detection methods have insufficient accuracy, especially when the treatment is irregular, small in size and severe interference with surrounding tissues.

Method used

Build an AE-YOLO network and enhance feature extraction capabilities by introducing aggregation dynamic convolution (ADC) module and visual enhancement block (VEB) module into the YOLOv8 network. The ADC module dynamically adjusts the convolution kernel weights through channel attention, filter attention and nuclear attention mechanisms. The VEB module captures global information through the TFormer module and reduces feature redundancy through the FRC module.

Benefits of technology

It improves the accuracy and robustness of breast mass detection, enhances the adaptability to different mass characteristics, and significantly improves detection performance, including recall, accuracy and average accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451042A_ABST
    Figure CN120451042A_ABST
Patent Text Reader

Abstract

The invention provides a breast mass detection method and device for enhancing feature extraction, and relates to the field of image recognition, and the method comprises the steps: S1, constructing an AE-YOLO network through an ADC module, a VEB module and a YOLOv8 network; s2, obtaining a training image set, training the AE-YOLO network through the training image set, and obtaining a trained AE-YOLO network; and S3, carrying out breast mass detection through the trained AE-YOLO network. The invention provides an AE-YOLO network. The architecture of the AE-YOLO network is established on the basis of an original YOLOv8 architecture; the improvements enhance the fine-grained sampling ability of breast mass features and the ability of the model to capture global information and utilize redundant features. The AE-YOLO network is enabled to have better performance in breast mass detection, and higher detection performance is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a method and device for detecting breast masses with enhanced feature extraction. Background Art

[0002] The continuous development of deep learning has prompted researchers to apply it to medical image analysis. In breast mass detection, deep neural network models can automatically analyze and process mammograms to detect suspected mass areas, thereby assisting doctors in diagnosis and treatment and improving work efficiency. Compared with traditional CAD systems, deep learning-based CAD systems have better generalization capabilities and can provide faster and more accurate detection results in clinical practice. Teare et al. used dual-depth convolutional neural networks of different scales and combined them with a random forest network to detect masses in mammograms. The proposed algorithm achieved sensitivity and specificity similar to those of professional radiologists. Maqsood et al. proposed a transferable texture CNN scheme for effective breast malignancy detection using mammograms. In 2021, Meraj et al. proposed a quantized UNet-based approach for breast lesion detection. Deep features extracted by DenseNet were fused with features extracted by independent component analysis to improve recognition performance. Al-Masni et al. proposed a neural network-based CAD system that can effectively detect breast masses and distinguish between benign and malignant masses. Chen et al. proposed a network model for breast mass detection based on RetinaNet. By modifying the structure of the backbone and detection head and adding enhancements such as PAN, the performance and detection speed of RetinaNet for breast mass detection are improved.

[0003] Although some studies have achieved encouraging results, existing deep learning-based detection methods still have much room for improvement in improving the accuracy of breast mass detection. Breast masses are usually small and irregular in shape, and interference from surrounding tissues makes them prone to missed detection and misdetection. Summary of the Invention

[0004] In view of this, the object of the present invention is to provide a method and device for breast lump detection with enhanced feature extraction, so as to solve the problem of insufficient accuracy of existing detection methods based on deep learning.

[0005] The present invention provides a method for detecting breast masses by enhancing feature extraction, comprising the steps of:

[0006] S1: Build the AE-YOLO network through the ADC module, VEB module and YOLOv8 network;

[0007] S2: Obtain a training image set, train the AE-YOLO network using the training image set, and obtain a trained AE-YOLO network;

[0008] S3: Detect breast masses using the trained AE-YOLO network.

[0009] Preferably, step S1 is specifically as follows:

[0010] S11: Obtain the YOLOv8 network, which includes: the initial backbone network, the initial Neck network, and the initial Head network;

[0011] S12: adding an ADC module after the first CspLayer layer, the second CspLayer layer, and the third CspLayer layer of the initial backbone network, and adding a VEB module at the end of the initial backbone network to obtain an optimized backbone network;

[0012] S13: Construct the AE-YOLO network by optimizing the backbone network, initial Neck network and initial Head network.

[0013] Preferred:

[0014] The ADC module includes: GFR module, first FA module, FS module, second FA module, channel attention module, kernel attention module and filter attention module;

[0015] The GFR module is connected to the first FA module, the FS module, and the second FA module;

[0016] The first FA module is connected to the channel attention module, the FS module is connected to the kernel attention module, and the second FA module is connected to the filter attention module.

[0017] Preferred:

[0018] The output characteristic of the ADC module is expressed as:

[0019] y=(α c1 ☉α f1 ☉α k1 ☉W1+…+α ci ☉α fi ☉α ki ☉W i +…+α cn ☉α fn ☉α kn ☉W n )*x (1)

[0020] Among them, x and y represent the input features and output features of the ADC module respectively, n represents the maximum dimension of the ADC module, and W i represents the convolution kernel of the i-th dimension, a ci 、a fi and a ki They represent the channel attention weight, filter attention weight and kernel attention weight in the i-th dimension respectively, the symbol * represents the convolution operation, and the symbol ⊙ represents the aggregation operation.

[0021] Preferred:

[0022] VEB modules include: Stem module, TFormer module, FRC module and SE module;

[0023] The Stem module is connected to the TFormer module and the FRC module, and the TFormer module and the FRC module are connected to the SE module.

[0024] Preferred:

[0025] The expression of the output feature of the VEB module is:

[0026] Xout=SE(cat(TFormer(Xin);FRC(Xin)))

[0027] Among them, Xout is the output feature of the VEB module, cat(·) represents the feature concatenation along the channel dimension, Xin represents the output feature of the Stem module, TFormer(Xin) represents the output feature of the TFormer module, FRC(Xin) represents the output feature of the FRC module, and SE() represents the processing by the SE module.

[0028] Preferred:

[0029] The TFormer module includes a first residual module and a second residual module connected in sequence;

[0030] The first residual module includes a PConv module, and the second residual module includes an MLp module.

[0031] Preferred:

[0032] The FRC module includes a GN layer and a Sigmoid function layer connected in sequence;

[0033] The workflow of the FRC module is:

[0034] Input the initial feature X into the GN layer, and obtain the weight of each channel in the initial feature X through the GN layer;

[0035] Map each weight to the range of 0 to 1 through the Sigmoid function layer to obtain the corresponding weight mapping value;

[0036] Set a threshold, and use weight mapping values greater than or equal to the threshold as information-rich mapping values, and weight mapping values less than the threshold as information-poor mapping values;

[0037] The information-rich mapping value and the information-poor mapping value are multiplied by the initial feature X to form the information-rich feature X1 and the information-poor feature X2 respectively;

[0038] Split the information-rich feature X1 into feature X 11 and feature X 12 , split the information-deficient feature X2 into feature X 21 and feature X 22 ;

[0039] The feature X 11 and feature X 21 Perform the inversion operation to obtain the inverted feature X 11 and the inverted feature X 21 ;

[0040] The inverted feature X 11 With feature X 22 Add to get feature X w1 , the reversed feature X 21 With feature X 12 Add to get feature X w2 , the feature X w1 and feature X w2 The output features of the FRC module are obtained by splicing.

[0041] Preferably, step S2 is specifically as follows:

[0042] S21: constructing a training image set by acquiring a breast mass image set;

[0043] S22: Input the training image into the AE-YOLO network, obtain the predicted detection image and calculate the loss value;

[0044] S23: Compare the predicted detection image with the real detection image to obtain a comparison result, and adjust the parameters of the AE-YOLO network according to the comparison result;

[0045] S24: Repeat steps S22-S23 until the maximum number of iterations is reached or the loss value converges to obtain a trained AE-YOLO network.

[0046] A breast mass detection device with enhanced feature extraction comprises: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the breast mass detection method with enhanced feature extraction.

[0047] The present invention has the following beneficial effects:

[0048] (1) An AE-YOLO network is proposed. The architecture of the AE-YOLO network is built on the original YOLOv8 architecture. This enhancement increases the fine-grained sampling of mass features and the model's ability to capture global information and utilize redundant features. This makes the AE-YOLO network perform better in breast mass detection and achieves stronger detection performance.

[0049] (2) To improve the model’s adaptability to different tumor characteristics and enhance feature extraction capabilities, an aggregated dynamic convolution (ADC) module is proposed, which applies three parallel structures to the feature map. These structures allow the convolution kernel to dynamically acquire weights for the kernel dimension, input channel dimension, and output channel dimension. The aggregation operation then applies these three weights to the convolution kernel, enabling it to simultaneously adjust the weights of the kernel space, input channel, and output channel according to the feature information, thereby improving the robustness of the model.

[0050] (3) It is proposed to place the Visual Enhancement Block (VEB) module at the end of the backbone network. VEB consists of two parallel modules: one is the TFormer module that extracts global information, and the other is the Feature Reconstruction Center (FRC) module, which reduces feature redundancy and improves the quality of feature information. When the features enter VEB, they are divided into two parts: one part enters TFormer and the other part enters FRC. After processing the two parts, the feature maps are connected. In order to fuse spatial and channel information, a SE (Squeeze-and-Excitation) module is inserted at the end, which effectively fuses the information from TFormer and FRC, further improving the network performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of a method according to an embodiment of the present invention;

[0052] Figure 2 This is the YOLOv8 network structure diagram;

[0053] Figure 3 This is the ADC module structure diagram;

[0054] Figure 4 This is the VEB module structure diagram;

[0055] Figure 5 This is the structure diagram of the TFormer module;

[0056] Figure 6 It is the structure diagram of the residual module;

[0057] Figure 7 This is the visual feature map of the middle layer of the TFormer module;

[0058] Figure 8 This is the FRC module structure diagram;

[0059] Figure 9 This is a comparison chart of FROC curves of different detection methods;

[0060] Figure 10 It is a visualization of the detection results of different detection methods;

[0061] Figure 11 Visual feature map of each layer of ADC;

[0062] Figure 12 This is a visual heat map of VEB;

[0063] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0064] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0065] Reference Figure 1 The present invention provides a method for detecting breast masses by enhancing feature extraction, comprising the steps of:

[0066] S1: Build the AE-YOLO network through the ADC module, VEB module and YOLOv8 network;

[0067] As an example, AE-YOLO is built on the YOLOv8 framework. Figure 2 As shown in the figure, AE-YOLO is a model designed for breast mass detection. To improve detection performance based on the characteristics of breast masses, it introduces aggregated dynamic convolution (ADC) and visual enhancement blocks (VEB). ADC is added after the first, second, and third CspLayer layers of the backbone to further extract richer mass feature information. Adding VEB at the end of the backbone provides two benefits: first, it allows the model to capture global information about breast masses, aiding detection; second, VEB helps reduce the high redundancy of mass features, thereby improving the quality of feature information.

[0068] Step S1 is specifically as follows:

[0069] S11: Obtain the YOLOv8 network, which includes: the initial backbone network, the initial Neck network, and the initial Head network;

[0070] S12: adding an ADC module after the first CspLayer layer, the second CspLayer layer, and the third CspLayer layer of the initial backbone network, and adding a VEB module at the end of the initial backbone network to obtain an optimized backbone network;

[0071] S13: Construct the AE-YOLO network by optimizing the backbone network, initial Neck network and initial Head network.

[0072] As an example:

[0073] The ADC module includes: GFR module, first FA module, FS module, second FA module, channel attention module, kernel attention module and filter attention module;

[0074] The GFR module is connected to the first FA module, the FS module, and the second FA module;

[0075] The first FA module is connected to the channel attention module, the FS module is connected to the kernel attention module, and the second FA module is connected to the filter attention module.

[0076] Specifically, breast tumors exhibit varying shapes, attributes, and other characteristics across patients, posing a significant challenge to the generalization of the model. To improve the model's adaptability to diverse tumor characteristics and enhance feature extraction capabilities, we propose Aggregation Dynamic Convolution (ADC).

[0077] During the operation, the kernel weights remain fixed and are uniformly applied to all input samples in traditional convolution. However, the same weights cannot be applied to all breast masses. In dynamic convolution [15,16], a linear combination of n convolution kernels is used, and an attention mechanism is used to dynamically adjust the kernel weights based on the input feature information. Introducing dynamic convolution into breast mass detection allows the kernel weights to be adjusted according to different masses, improving the convolution's ability to capture mass-related information.

[0078] Standard dynamic convolution can adjust the convolution kernel weights according to feature information, but usually only dynamically operates on the number of convolution kernels. That is, it adapts to different input features through a weighted combination of multiple kernels, while ignoring the effects of the other two dimensions on the convolution kernel. Due to the unique characteristics of breast mass features, the number of kernels alone is not enough to extract sufficient information; additional information must also be collected from the context and different channels. The Aggregated Dynamic Convolution (ADC) proposed in the present invention extracts mass information from three dimensions (input channel dimension, output channel dimension, kernel dimension), and uses three different attention mechanisms in parallel: channel attention, filter attention, and kernel attention. These three attention mechanisms are used through aggregation operations to adjust the weights of the three dimensions of the convolution kernel.

[0079] As an example:

[0080] The output characteristic of the ADC module is expressed as:

[0081] y=(α c1 ☉α f1 ☉α k1 ☉W1+…+α ci ☉α fi ☉α ki ☉W i +…+α cn ☉α fn ☉α kn ☉W n )*x (1)

[0082] Among them, x and y represent the input features and output features of the ADC module respectively, n represents the maximum dimension of the ADC module, and W i represents the convolution kernel of the i-th dimension, a ci 、a fi and a ki They represent the channel attention weight, filter attention weight and kernel attention weight in the i-th dimension respectively, the symbol * represents the convolution operation, and the symbol ⊙ represents the aggregation operation.

[0083] Specifically, channel attention adjusts the weights of convolution kernels in different input channels to enhance the network's focus on salient feature channels. Filter attention dynamically adjusts the weights of convolution kernels in various output channels, allowing the network to prioritize important output features. Kernel attention adaptively modifies the number and weights of convolution kernels based on the input's signature, facilitating the capture of diverse local features.

[0084] The structure of the ADC module is as follows Figure 3 As shown, Figure 3The operational process of the three types of attention parallel generation and aggregation of the present invention is described: the input feature x first passes through the GFR module, where the global average pooling operation compresses the input feature x into a feature vector of length Cin. After passing through the fully connected layer and the ReLU activation layer, the feature vector is processed by three branch structures. Each branch uses the FA or FS module. In the FA and FS modules, a fully connected layer maps the compressed feature vector to a low-dimensional space, followed by a Sigmoid activation function or Softmax to generate the normalized channel attention weights a respectively. ci , filter attention weight a fi and kernel attention weight a ki After generating the three weights, the aggregation process begins. First, the channel attention weight a ci Multiply the input feature map to generate a weighted intermediate feature map; predefine a learnable weight and set the filter attention weight c ki Multiply this weight to form a convolution kernel. Then, the kernel is convolved with the intermediate feature. Finally, the feature map obtained by the convolution operation is multiplied by the filter attention weight a fi , to produce the final output feature map.

[0085] As an example:

[0086] VEB modules include: Stem module, TFormer module, FRC module and SE module;

[0087] The Stem module is connected to the TFormer module and the FRC module, and the TFormer module and the FRC module are connected to the SE module.

[0088] Specifically, such as Figure 4 As shown in the figure, the Visual Enhancement Block (VEB) proposed in the present invention mainly consists of two parallel connected blocks. Among them, TFormer is used to capture global information. At the same time, in order to reduce feature redundancy and improve feature quality, a feature reconstruction center (FRC) is proposed. The output feature maps of the two blocks are connected together along the channel dimension, and then the information is fused using the SE block (squeeze-and-excitation block). In the implementation of the present invention, the Stem block is used to smooth the features before the features are input to the VEB. The Stem block consists of a 7x7 convolution kernel, a batch normalization layer, and an activation function layer. The process of the Stem block can be described as follows:

[0089] Xin=σ(BN(Conv7×7(X))) (2)

[0090] Where Conv7x7(·) represents a 7x7 convolution with a stride of 1, BN(·) represents a batch normalization layer, and σ(·) represents a ReLU activation function layer.

[0091] As an example:

[0092] The expression of the output feature of the VEB module is:

[0093] Xout=SE(cat(TFormer(Xin);FRC(Xin))) (3)

[0094] Among them, Xout is the output feature of the VEB module, cat(·) represents the feature concatenation along the channel dimension, Xin represents the output feature of the Stem module, TFormer(Xin) represents the output feature of the TFormer module, FRC(Xin) represents the output feature of the FRC module, and SE() represents the processing by the SE module.

[0095] As an example:

[0096] The TFormer module includes a first residual module and a second residual module connected in sequence;

[0097] The first residual module includes a PConv module, and the second residual module includes an MLp module.

[0098] Specifically, the global information in deep features is particularly important for breast mass recognition. Usually, at deeper levels of the model, the mass features may become less obvious, and the model needs to rely on contextual information to collect feature information. In order to enable the model to learn global information, the present invention designs a lightweight MLP structure TFormer. Figure 5 shown.

[0099] The proposed TFormer consists of two residual blocks. In the first residual block that performs spatial operations, partial convolution (PConv) is introduced. Figure 6 As shown. Compared to regular convolution, partial convolution applies a standard convolution with a kernel size of k to spatially sample a subset of channels, while another subset of channels uses point-by-point convolution. The shape of partial convolution is like a T, which is why the present invention named the proposed MLpTFormer as TFormer. The feature maps between different channels show a high degree of similarity. Therefore, the present invention re-examined the mammograms in the dataset and found that most images contained only one mass, and only a few images contained two to three masses. In addition, the mass accounts for a very small proportion of the total image pixels. As Figure 7 As shown, there is little useful feature information in the image and there is significant redundancy between channels.

[0100] The introduction of PCony in the task of breast mass detection is effective. Not only does it extract sufficient feature information, but it also reduces the number of parameters, reduces memory access, and improves the overall calculation speed. In the second residual block that processes channel operations, the present invention uses channel MLP. The normalization layer adopts group normalization (GN); compared with batch normalization, the calculation of GN does not depend on the batch size. In addition, the present invention also adds channel scaling and Drop Path operations to improve the generalization ability and robustness of the model by adjusting each channel and randomly discarding the convolution path. In general, the output features from the Stem module first undergo spatial interaction through the GN layer and PConv, followed by channel scaling and drop path. Finally, channel interaction is performed through channel MLP, and the residual connection ensures the gradient flow and alleviates the gradient disappearance problem.

[0101] As an example:

[0102] The FRC module includes a GN layer and a Sigmoid function layer connected in sequence;

[0103] The workflow of the FRC module is:

[0104] Input the initial feature X into the GN layer, and obtain the weight of each channel in the initial feature X through the GN layer;

[0105] Map each weight to the range of 0 to 1 through the Sigmoid function layer to obtain the corresponding weight mapping value;

[0106] Set a threshold, and use weight mapping values greater than or equal to the threshold as information-rich mapping values, and weight mapping values less than the threshold as information-poor mapping values;

[0107] The information-rich mapping value and the information-poor mapping value are multiplied by the initial feature X to form the information-rich feature X1 and the information-poor feature X2 respectively;

[0108] Split the information-rich feature X1 into feature X 11 and feature X 12 , split the information-deficient feature X2 into feature X 21 and feature X 22 ;

[0109] The feature X 11 and feature X 21 Perform the inversion operation to obtain the inverted feature X 11 and the inverted feature X 21 ;

[0110] The inverted feature X 11 With feature X 22Add to get feature X w1 , the reversed feature X 21 With feature X 12 Add to get feature X w2 , the feature X w1 and feature X w2 The output features of the FRC module are obtained by splicing.

[0111] Specifically, during breast X-ray photography, tissue factors such as lighting and mass shadows will affect the image quality. The backbone network may also mistakenly extract irrelevant features when extracting mass features. This leads to redundancy in the extracted features. In order to reduce the interference of these redundant features on breast mass detection, a Feature Reconstruction Center (FRC) module is designed. FRC uses separation and reconstruction operations. The separation operation separates the feature map with rich mass information from the feature map with less mass information to ensure that the main features are not interfered by noise. The reconstruction operation further reduces the redundancy of features on the basis of separation, and adds the feature map with rich information to the feature map with less information, thereby generating a feature map with richer information. Figure 8 Specifically, in the separation stage, the present invention uses the learnable scaling factors in the GN layer to evaluate the importance of different channel feature maps and converts these scaling factors into a series of weights through normalization. Higher weights reflect richer features. Next, the present invention applies the Sigmoid function to map the weighted feature values to the range of 0 to 1. By setting a threshold, the weighted features are divided into information-rich features X i And feature X2 with less information. Then the separated features are reconstructed. In the reconstruction stage, in order to fully integrate the features and strengthen the information flow between the features, the present invention divides the feature X1 into X 11 and X 12 Two parts, also split X2 into X 21 and X 22 The present invention combines feature X 11 and feature X 21 Perform the inversion operation and then cross-combine, that is, the inverted X 11 With X 22 Add up to get feature X w1 , the inverted X 21 With X 12 Add up to get feature X w2 , and finally X w1 and X w2 Splicing to get more informative features Figure X w .

[0112] As an example:

[0113] The SE layer (Squeeze-and-Excitation) acts as a channel attention module, adjusting the weight of each channel based on the input, thereby enhancing feature representation and improving performance. In VEB, the input features are processed by two parallel branches, one through TFormer and the other through FRC, and the output features of the two structures are then concatenated. The result is that the two parts of the feature have different types of information, but there is no direct connection between them, making it difficult to share information across the entire channel. To fuse the feature information of the two parts, the present invention places the SE layer at the end of VEB.

[0114] S2: Obtain a training image set, train the AE-YOLO network using the training image set, and obtain a trained AE-YOLO network;

[0115] As an example:

[0116] Step S2 is specifically as follows:

[0117] S21: constructing a training image set by acquiring a breast mass image set;

[0118] Specifically, the data used in the present invention comes from two public datasets: DDSM (Digital Database for Screening Mammography) and MIAS (The Mammographic Image Analysis Society).

[0119] The DDSM dataset contains 2,610 cases, including 695 normal cases, 1,011 benign cases, and 914 malignant cases. These cases include various types of abnormalities, such as calcifications, masses, architectural deformations, and asymmetry. Each case includes cranial (CC) and mediolateral oblique (MLO) views of the left breast, as well as CC and MLO views of the right breast, for a total of 4 images. In this study, the present invention selected images containing masses. Due to the high resolution of the original images and their LJPEG format, for the convenience of training, the images were downsampled 8 times and converted to JPG format.

[0120] The MIAS dataset contains 208 cases, all of which are from MLO views. We selected images containing masses, converted the original PGM format images into JPG format, and standardized the resolution to 1024x1024 for training.

[0121] S22: Input the training image into the AE-YOLO network, obtain the predicted detection image and calculate the loss value;

[0122] S23: Compare the predicted detection image with the real detection image to obtain a comparison result, and adjust the parameters of the AE-YOLO network according to the comparison result;

[0123] S24: Repeat steps S22-S23 until the maximum number of iterations is reached or the loss value converges to obtain a trained AE-YOLO network.

[0124] S3: Detect breast masses using the trained AE-YOLO network.

[0125] 1. Specific experimental details:

[0126] All experiments were implemented in PyTorch using an NVIDIA GeForce RTX 4080 GPU (with 16GB of memory). The initial learning rate was 0.01, the batch size was set to 16, and the number of epochs was 200. The experimental dataset consisted of 1930 images, with a training set, validation set, and test set split ratio of 6:2:2.

[0127] 2. Evaluation indicators:

[0128] The evaluation metrics used in this paper are Recall, Precision, F1-Score, mAP50, and mAP50:95. The Intersection over Union (IoU) measures the overlap between the predicted bounding box and the ground-truth bounding box. The IoU is calculated as follows:

[0129]

[0130] Recall represents the number of true positive examples that the model can correctly predict. Precision represents the number of correct positive example predictions among all positive example predictions made by the model

[23] . The calculation formulas for recall and precision are as follows:

[0131]

[0132] Where TP represents true positive, FP represents false positive, and FN represents false negative. By plotting Recall on the x-axis and Precision on the y-axis, the present invention can obtain a precision-recall (PR) curve. The F1 score is the harmonic mean of precision and recall and is used to provide a comprehensive assessment of detection performance. The calculation formula is as follows:

[0133]

[0134] The recall rate, precision rate and F1 in the table of the present invention are all calculated under the condition that the IoU (Intersection over Union) threshold is 0.7. When different confidence thresholds are set, the recall rate and precision rate will change. The Free-response Receiver Operating Characteristic (FROC) curve can be used to evaluate the performance of the model under different thresholds. By plotting the false positive rate (False Positive Per Image, FPPI) of each image on the x-axis and the true positive rate (True Positive Rate, TPR), i.e., the recall rate, on the y-axis, the FROC curve can be obtained. In order to reduce the impact of excessive false positives, the study limits the number of false positives per image to 2

[24] . AFROC is used as one of the evaluation indicators. mAP (mean Average Precision) is the average value of AP (Average Precision) of each category. AP is obtained by recall rate and precision rate, which is used to measure the detection performance of the model for a specific category. AP is the area under the PR (Precision-Recall) curve, and the formula is as follows:

[0135]

[0136] mAP50 refers to the average AP of each category when the IoU is set to 0.5. mAP50:95 refers to the average mAP at different IoU thresholds (from 0.5 to 0.95, with a step size of 0.05).

[0137] 3. Results:

[0138] A comparative experiment with different detection methods

[0139] To verify the performance of the proposed method in breast mass detection, the present invention compared it with other detection methods, and the results are shown in Table 1. Compared with the original YOLOv8, AE-YOLO showed significant improvements in detection accuracy, with mAP50 increasing from 81.7% to 84.9%, mAP50:95 increasing from 44.9% to 48.4%, and recall increasing from 74.7% to 77.2%. Compared with other detection methods, the proposed method also showed significant advantages. Compared with other classic detection methods (YOLOv9, Gelan, Rtdetr, YOLOv10), the proposed method achieved better detection results. To further verify the advantages of the proposed method in breast mass detection, it was compared with ERtinaNet, the latest method in this field. The proposed method achieved leading results in all indicators, indicating that the proposed method not only detected more masses but also located them more accurately.

[0140] Table 1 Comparison of detection performance of different detection methods

[0141]

[0142] When considering other confidence thresholds, the present invention draws a FROC curve with FPPI as the x-axis and the true positive rate as the y-axis, such as Figure 9 As shown. The FROC curve shows that AE-YOLO has the highest TPR for most FPPI values on the axis. In order to numerically illustrate the difference in detection performance, the present invention calculates the true positive rate values at 0.25FPPI, 0.75FPPI, 1.25FPPI and 1.75FPPI, as shown in Table 2. For example, when FPPI is set to 0.25, which means that one false positive is allowed for every 4 images, AE-YOLO achieves a high TPR of 95.4%. This method outperforms other methods in terms of true positive rate for all FPPI values. The present invention also calculates the p-value between the true positive rate of AE-YOLO and other methods when the FPPI value range is 0-2. A p-value less than 0.05 indicates that the difference is statistically significant. As shown in Table 2, the method of the present invention is significantly different from other methods, indicating that the method of the present invention is more accurate.

[0143] Table 2 Comparative experiment of TPR performance under different FPPI values

[0144]

[0145] The paper provides a visual representation of the results of different detection methods. Figure 10 As shown in the example, Figure 10Column (a) shows Ground Truth, column (b) shows the detection results of AE-YOLO, column (c) shows the detection results of YOLOv8, column (d) shows the detection results of Gelan, column (e) shows the detection results of ERetinaNet, column (f) shows the detection results of YOLOv9, and column (g) shows the detection results of RtDetr.

[0146] Most tumors are small, and surrounding tissue often interferes with detection. For smaller tumors or those with significant surrounding tissue interference, other methods either fail to detect the tumor or produce false positives, while AE-YOLO successfully detects the tumor. This demonstrates AE-YOLO's superior performance in tumor detection. Thanks to the proposed aggregated dynamic convolution and visual enhancement blocks, AE-YOLO can adapt to various tumor types and accurately capture feature information even in the presence of surrounding tissue interference.

[0147] B. Ablation experiment

[0148] The present invention proposes two modules, ADC and VEB, where VEB includes TFormer and FRC. To demonstrate their improvement in network performance, the present invention conducts ablation experiments.

[0149] Table 3 Ablation experiment results of different modules

[0150]

[0151] Ablation test results are shown in Table 3. Compared to the baseline model YOLOv8, the methods of the present invention significantly improve various network metrics both when used individually and in combination. When the three methods are used together, the model achieves the best results in recall, F1, mAP50, and mAP50:95. Compared to the baseline model YOLOv8, the accuracy is improved by 5.2%, the recall is improved by 2.5%, the mAP50 is improved by 3.2%, and the mAP50:95 is improved by 3.5%. These results demonstrate that the methods of the present invention can achieve significant performance improvements in various aspects.

[0152] The role of ADC is to enhance the adaptability of the model to different tumor characteristics and improve the ability of feature extraction. The present invention places it in different layers of the model backbone to sample the characteristics of breast tumors.

[0153] Table 4 Effect of inserting ADC at different layers of the backbone

[0154]

[0155]

[0156] As shown in Table 4, when ADC is placed only in the first layer of the backbone, the model has achieved significant improvements in accuracy, recall, mAP50, and mAP50:95 compared to the baseline model. The present invention believes that it is necessary to place ADC in the first layer, and extracting richer features in the shallow layer is conducive to deeper feature extraction. When ADC is placed in the 1st and 2nd layers, the performance of the model is not much different from that of the first layer. When ADC is placed in all four layers, the model performance is the worst among the four placement methods. When placed in the 1st, 2nd, and 3rd layers, the effect of the model is significantly improved, reaching the highest level in accuracy, F1, mAP50, and mAP50:95. In order to further analyze these phenomena, the present invention selected a picture and then visualized the feature map of each layer of ADC.

[0157] The visual feature map of each layer of ADC is as follows Figure 11 As shown, Figure 11 (a) is the gold standard image, (b) is the feature image of the first layer ADC, (c) is the feature image of the second layer ADC, (d) is the feature image of the third layer ADC, and (e) is the feature image of the fourth layer ADC.

[0158] In the feature maps of the ADC in the first and second layers, the characteristic information of the mass is not yet clear. However, in the feature map of the third layer, the characteristic information of the mass is very clear, and the black dot at the mass location is more prominent. In the fourth layer, after multiple layers of deep convolution, the resolution of the feature map is only 20x20, and the mass occupies only a few pixels. Therefore, despite the attention of the ADC, the characteristic information of the mass is still interfered with by the surrounding tissue and is aggregated into a larger surrounding area. Based on the above analysis, the present invention uses ADC in the first, second, and third layers of the backbone.

[0159] Next, we analyze the impact of VEB on the model. In VEB, TFormer acquires global information, FRC reduces feature redundancy and improves feature quality, and finally, SE layers fuse the output features of TFormer and FRC. To verify the impact of the three different architectures on the model, we conducted ablation experiments.

[0160] Table 5 Effects of different modules in VEB

[0161]

[0162] As can be seen from Table 5, using either TFormer or FRC alone improves model performance. Adding SE at the end of VEB further improves performance, reaching optimal levels for accuracy, F1, mAP50, and mAP50:95. This demonstrates the effectiveness of SE in fusing the output features of TFormer and FRC. To illustrate the connection between TFormer, FRC, and SE, we visualized a heatmap of VEB.

[0163] The VEB visualization heat map is as follows Figure 12 As shown, Figure 12 (a) is the gold standard, (b) is the heat map of TFormer, (c) is the heat map of VEB without SE, and (d) is the heat map of VEB with SE.

[0164] like Figure 12 As shown in (b), after the input features are processed by TFormer, thanks to TFormer's ability to obtain global information, the model effectively learns the characteristic information of the mass, but it is also interfered by the surrounding tissues. After the introduction of FRC, as shown in Figure 12 As shown in (c), FRC filters out the interference of surrounding tissues, and the model is more effective in extracting tumor feature information. Figure 12 As shown in (d), the SE structure is introduced at the end of VEB to fully integrate the output results of TFormer and FRC, enabling the model to more fully filter interference and extract global information.

[0165] Finally, the images in the test set were classified based on tumor size, and the detection performance of AE-YOLO was demonstrated. After processing the DDSM dataset, the pixel resolution was increased from the initial 0.05 mm to 0.4 mm, while the pixel resolution of the MIAS dataset was 0.2 mm. In breast mass research, tumors with a radius less than 7 mm are classified as small. After calculation, the present invention considers tumors smaller than 908 pixels in the DDSM dataset to be small, while the rest are non-small. Similarly, in the MIAS dataset, tumors smaller than 3847 pixels are considered small, while the rest are non-small. The present invention tested both small and non-small tumors separately. As shown in Table 6, compared to the original YOLOv8, the present invention's method performs better in detecting small tumors. Precision increased by 4.7% and recall by 2.3%, reducing the risk of the model missing small tumors. The improvement in F1 further demonstrates the improved overall performance of the model. Furthermore, the improvements in mAP50 and mAP50.95 also demonstrate better detection of small tumors. When detecting non-small masses, the precision is improved by 3.3%, the recall is improved by 5.8%, the mAP50 is improved by 5%, and the mAP50:95 is improved by 0.9%.

[0166] Table 6 Comparison of the effects of AE-YOLO and baseline YOLOv8 on small masses and non-small masses

[0167]

[0168] Experiments show that AE-YOLO achieves a recall rate of 77.2%, an mAP50:95 of 84.9%, and an mAP50:95 of 48.4%, which are all state-of-the-art results. Compared with the original YOLOv8, the detection accuracy has been significantly improved.

[0169] A breast mass detection device with enhanced feature extraction comprises: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the breast mass detection method with enhanced feature extraction.

[0170] It should be noted that, in the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0171] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same item of hardware. The use of the terms first, second, and third, etc., does not denote any order and should be construed as identifiers.

[0172] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for detecting breast masses by enhancing feature extraction, characterized in that: Including steps: S1: Build the AE-YOLO network through the ADC module, VEB module and YOLOv8 network; S2: Obtain a training image set, train the AE-YOLO network using the training image set, and obtain a trained AE-YOLO network; S3: Detect breast masses using the trained AE-YOLO network.

2. The breast mass detection method with enhanced feature extraction according to claim 1, wherein: Step S1 is specifically as follows: S11: Obtain the YOLOv8 network, which includes: the initial backbone network, the initial Neck network, and the initial Head network; S12: adding an ADC module after the first CspLayer layer, the second CspLayer layer, and the third CspLayer layer of the initial backbone network, and adding a VEB module at the end of the initial backbone network to obtain an optimized backbone network; S13: Construct the AE-YOLO network by optimizing the backbone network, initial Neck network and initial Head network.

3. The method for detecting breast masses by enhanced feature extraction according to claim 2, wherein: The ADC module includes: GFR module, first FA module, FS module, second FA module, channel attention module, kernel attention module and filter attention module; The GFR module is connected to the first FA module, the FS module, and the second FA module; The first FA module is connected to the channel attention module, the FS module is connected to the kernel attention module, and the second FA module is connected to the filter attention module.

4. The method for detecting breast masses by enhanced feature extraction according to claim 3, wherein: The output characteristic of the ADC module is expressed as: y=(a c1 ⊙a f1 ⊙a k1 ⊙W1+…+α ci ⊙a fi ⊙a ki ⊙W i +…+a cn ⊙a fn ⊙a kn ⊙W n )*x (1) Among them, x and y represent the input features and output features of the ADC module respectively, n represents the maximum dimension of the ADC module, and W i represents the convolution kernel of the i-th dimension, a ci 、a fi and a ki They represent the channel attention weight, filter attention weight and kernel attention weight in the i-th dimension respectively, the symbol * represents the convolution operation, and the symbol ⊙ represents the aggregation operation.

5. The method for detecting breast masses by enhanced feature extraction according to claim 2, wherein: VEB modules include: Stem module, TFormer module, FRC module and SE module; The Stem module is connected to the TFormer module and the FRC module, and the TFormer module and the FRC module are connected to the SE module.

6. The method for detecting breast masses by enhanced feature extraction according to claim 5, wherein: The expression of the output feature of the VEB module is: Xout=SE(cat(TFormer(Xin);FRC(Xin))) Among them, Xout is the output feature of the VEB module, cat(·) represents the feature concatenation along the channel dimension, Xin represents the output feature of the Stem module, TFormer(Xin) represents the output feature of the TFormer module, FRC(Xin) represents the output feature of the FRC module, and SE() represents the processing by the SE module.

7. The method for detecting breast masses by enhanced feature extraction according to claim 5, wherein: The TFormer module includes a first residual module and a second residual module connected in sequence; The first residual module includes a PConv module, and the second residual module includes an MLp module.

8. The method for detecting breast masses by enhanced feature extraction according to claim 5, wherein: The FRC module includes a GN layer and a Sigmoid function layer connected in sequence; The workflow of the FRC module is: Input the initial feature X into the GN layer, and obtain the weight of each channel in the initial feature X through the GN layer; Map each weight to the range of 0 to 1 through the Sigmoid function layer to obtain the corresponding weight mapping value; Set a threshold, and use weight mapping values greater than or equal to the threshold as information-rich mapping values, and weight mapping values less than the threshold as information-poor mapping values; The information-rich mapping value and the information-poor mapping value are multiplied by the initial feature X to form the information-rich feature X1 and the information-poor feature X2 respectively; Split the information-rich feature X1 into feature X 11 and feature X 12 , split the information-deficient feature X2 into feature X 21 and feature X 22 ; The feature X 11 and feature X 21 Perform the inversion operation to obtain the inverted feature X 11 and the inverted feature X 21 ; The inverted feature X 11 With feature X 22 Add to get feature X w1 , the reversed feature X 21 With feature X 12 Add to get feature X w2 , the feature X w1 and feature X w2 The output features of the FRC module are obtained by splicing.

9. The method for detecting breast masses by enhanced feature extraction according to claim 1, wherein: Step S2 is specifically as follows: S21: constructing a training image set by acquiring a breast mass image set; S22: Input the training image into the AE-YOLO network, obtain the predicted detection image and calculate the loss value; S23: Compare the predicted detection image with the real detection image to obtain a comparison result, and adjust the parameters of the AE-YOLO network according to the comparison result; S24: Repeat steps S22-S23 until the maximum number of iterations is reached or the loss value converges to obtain a trained AE-YOLO network.

10. A breast mass detection device with enhanced feature extraction, characterized by: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the breast mass detection method with enhanced feature extraction according to any one of claims 1 to 9.

Citation Information

Cited By

  • Surface defect detection method based on deep learning

    CN121616523A