Multi-scale adaptive lesion detection method based on breast ultrasound

Through the improved YOLOv12 model and CA attention mechanism, the multi-scale features in breast ultrasound images are adaptively extracted, solving the problem of insufficient flexibility and key area identification in the prior art, and achieving high-precision detection of breast lesions.

CN120374631AActive Publication Date: 2025-07-25SHANDONG FUTURE NETWORK RES INST (PURPLE MOUNTAIN LAB IND INTERNET INNOVATION APPL BASE)

Patent Information

Application Number
CN202510883971.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-25
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing breast lesion detection technology has insufficient flexibility, generalization ability, multi-scale lesion capture and key area identification, and it is particularly difficult to accurately identify key features such as smaller calcification points, cyst wall changes, and fibroblastoma boundaries.

Method used

Using a multi-scale adaptive lesion detection method based on breast ultrasound, the improved YOLOv12 model and CA attention mechanism are combined with convolutional blocks and multi-scale convolutional branches to adaptively extract feature information from different scales to enhance the capture ability of key areas.

Benefits of technology

It improves the accuracy and specificity of breast lesions detection, can clearly identify different tissue boundaries and multi-scale lesion characteristics, and enhances the ability to accurately judge breast lesions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374631A_ABST
    Figure CN120374631A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale adaptive lesion detection method based on mammary gland ultrasound. The method comprises the following steps: pre-processing an image; the input module is used for extracting features through convolution blocks to clearly display boundaries, and then deconvolution blocks are used for refining boundary features to complete image expression; a trunk module; in the neck network, the CA generates feature maps in two directions, the feature maps in the two directions are spliced, features are extracted through convolution operation, and attention weights in the two directions are further generated; and the detection head network performs positioning prediction, classification prediction and loss calculation, and finally outputs a result. According to the method, the precision and richness of feature extraction are improved, and the clear recognition capability of the model on different tissue boundaries is enhanced; the capacity of capturing multi-scale lesion features is improved, and the lesion features are better captured; important areas such as small calcification points, changes of cyst walls and boundaries of fibroadenoma can be highlighted, and the accuracy and specificity of detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of breast ultrasound lesion detection, and particularly to a multi-scale adaptive lesion detection method based on breast ultrasound. Background Art

[0002] Many breast lesions may not have obvious symptoms in the early stage. Early detection of lesions can gain precious time for treatment, improve the cure rate. Especially for malignant lesions such as breast cancer, early diagnosis and treatment can significantly improve the survival rate of patients. Breast lesions detected early usually have less treatment difficulty and relatively lower treatment costs. As one of the most common female problems, the detection of breast lesions is an important topic in clinical practice.

[0003] Currently, for the detection of breast lesions, it is mostly based on breast ultrasound, mammography, breast magnetic resonance imaging, and uses waveform feature rules, machine learning or deep learning methods, mainly as follows: The method based on waveform feature rules uses technologies such as expert systems to classify and diagnose breast lesions according to predefined rules. These rules are usually based on medical knowledge and clinical experience. Some studies judge the nature of lesions according to rules such as the mass shape, edge features, and calcification distribution in mammography. The advantage of this method is strong interpretability, but the features of breast lesions are complex, and fixed rules are difficult to adapt to all situations. Especially for some atypical lesions, it is difficult to automatically adjust and optimize according to new data or cases.

[0004] The method based on machine learning first uses ultrasonic elastography technology to evaluate tissue hardness. Malignant tumors are usually harder than the surrounding normal tissues, and then applies machine learning algorithms to analyze ultrasonic elastography data to help doctors more accurately judge the nature of lesions. It is necessary to manually extract features, which not only increases the workload of preprocessing, but may also lead to the omission of key features or inaccurate feature selection. In addition, the performance of machine learning models highly depends on the quality and quantity of training data. If the training data is insufficient or unbalanced, the generalization ability of the model will be limited.

[0005] Based on deep learning methods, often through deep learning networks such as convolutional neural networks, image data is analyzed to automatically extract features in the image and detect breast lesions. For example, some studies have used CNN to identify benign and malignant masses in breast ultrasound images, with an accuracy of 90%, and the diagnostic sensitivity and specificity are 86% and 96% respectively. In addition, CNN can also be combined with other technologies, such as transfer learning, multi-instance learning, etc., to improve the performance of the model. In mammography, CNN automatically identifies tiny calcifications and masses by training a large number of labeled image data. However, the problem of a large scale span of lesions is often not considered, and the feature extraction is not accurate enough, resulting in limited ability to capture key features such as smaller calcification points, changes in cyst walls, and the boundaries of fibroadenomas.

[0006] Therefore, there is an urgent need for a multi-scale adaptive lesion detection method based on breast ultrasound to solve the above problems. Summary of the Invention

[0007] The purpose of the present invention is to solve the problems existing in the prior art, such as poor flexibility, insufficient generalization ability, and deficiencies in multi-scale lesion capture and key area recognition. A multi-scale adaptive lesion detection method based on breast ultrasound is proposed, and an artificial intelligence algorithm is developed. Based on non-invasive and non-radiative highly adaptable breast ultrasound, accurate judgment of four common different types of breast lesions, namely cysts, fibroadenomas, suspicious mass groups, and calcifications, can be achieved, so as to timely remind these lesions and protect women's health.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions: A multi-scale adaptive lesion detection method based on breast ultrasound includes the following specific steps: S1: Image preprocessing: First, based on the letterbox method, the breast ultrasound is converted into a size of 1280×1280×3. This method can simultaneously maintain the original aspect ratio of the image and the integrity of the breast ultrasound content; determine the scaling ratio according to the original length and width of the breast ultrasound and the target length and width, and adjust the image size. The scaling ratio is selected as the smaller ratio of the target size to the original size to ensure that the image does not exceed the target size after scaling; add 0-pixel padding at the short sides of the top, bottom, left, and right of the image to make the long and short sides match the target size; S2: Input module: Receive the preprocessed image, extract features such as edges and textures through a convolutional block, clearly display the boundaries of fat, glands, fibrous tissues, etc., and then refine these boundary features through a deconvolutional block to complete the image expression. After passing through the input module, the image is converted into a size of 640×640×3 that is most suitable for processing by the backbone module; S3: Backbone Module: Input the ultrasound image into the backbone module of the improved YOLOv12 model based on the MSCA module. Its core is the residual efficient layer aggregation network. First, it passes through the first convolutional layer and the second convolutional layer, then through the first C3k2 module, then through the third convolutional layer. After the output passes through the second C3k2 module and the fourth convolutional layer in sequence, it is input into the first A2C2f module. After passing through the fifth convolutional layer, it is input into the A2C2f module, and finally input into the MSCA module; S4: Neck Network: CA aggregates features in the horizontal and vertical directions respectively through 1D pooling operations to generate feature maps in two directions. Specifically, use average pooling with a pooling kernel size of [pooling kernel size 1] and [pooling kernel size 2] to process the input feature map to obtain two direction-aware feature maps. Concatenate the feature maps in two directions, then extract features through convolutional operations to further generate attention weights in two directions. Use the h-swish non-linear activation function to increase the expressive ability of the model. Apply the generated attention weights to the horizontal and vertical directions of the input feature map respectively to achieve direction-aware and position-sensitive attention enhancement; S5: Detection Head Network: The detection head network receives multi-scale feature maps from the neck network. These feature maps are fused and adjusted in features, containing semantic information and spatial details at different levels, and perform localization prediction, classification prediction, loss calculation, and finally output the results.

[0009] As a further technical solution of the present invention, in the S2, the convolutional block is composed of a convolutional layer, a normalization layer, and an activation function. The convolutional layer uses a 3×3 convolutional kernel and the stride is set to 2. Then, the convolutional output is further normalized and activated to improve its expressive ability; extract low-level features in breast ultrasound through the first convolutional block. The first convolutional block selects the GELU activation function to adaptively adjust the output according to the input data and enhance the feature expression; then extract more abstract features through the second convolutional block. The activation function selects SiLU to further enhance the non-linearity of the feature expression.

[0010] As a further technical solution of the present invention, in the S3, the MSCA module is suitable for detecting targets with diverse sizes. First, perform feature extraction of different scales on the feature map through the multi-scale convolution branch to adaptively obtain multiple feature maps with different scale information; then calculate the channel weights among these feature maps through the attention mechanism and perform weighted fusion on the feature maps according to the importance of the channels; the fused feature map then passes through the spatial attention mechanism to generate a spatial attention weight map and multiply it with the original feature map to make the model pay more attention to the key feature regions.

[0011] As a further technical solution of the present invention, in step S3, the improved YOLOv12 model adopts a new type of convolutional block, which better meets the requirements of the medical scenario for lightweight operations and high parallelism. The Yolov12 model utilizes a series of smaller kernels, generally expressed as: , where: is the output feature, is the input feature, is the weight of the i-th convolutional kernel, is the bias of the i-th convolutional kernel.

[0012] As a further technical solution of the present invention, step S3 specifically includes: S31: The first convolutional layer uses a 3×3 convolutional kernel, with a stride set to 2 and padding of 1, to perform basic feature extraction on the input image. After this convolutional operation, the model initially captures basic features such as the edges and textures of the image; S32: The output feature is input into the second convolutional layer. This layer uses the same 3×3 convolutional kernel, but the stride becomes 1 and the padding remains 1. After the convolutional operation, the size of the feature map remains unchanged, and the number of channels doubles. The second convolutional layer further explores deeper feature information and enriches the feature expression while maintaining the stability of the feature map size; S33: The output of the second convolutional layer then enters the first C3k2 module. The input feature map is divided into two parts. One part is directly passed to retain shallow features, and the other part processes deep features through a standard Bottleneck and finally is spliced and fused; The C3k2 module assists in feature extraction. The C3k2 module retains multiple bottleneck layers, similar to CSPNet. Through the stacking of multiple bottleneck layers, features are gradually extracted and fused to improve the feature expression ability; S34: The output feature map of the first C3k2 module flows into the third convolutional layer, which uses the same filter as the first convolutional layer. The main function of the third convolutional layer is to further downsample, compress the size of the feature map, extract deeper and more abstract feature information, and at the same time expand the number of channels to meet the requirements of subsequent complex feature extraction; S35: The output feature map of the third convolutional layer enters the second C3k2 module and the fourth convolutional layer, which follows the structure of the first C3k2 module and the second convolutional layer; S36: The output feature map of the fourth convolutional layer is then input into the first A2C2f module. The A2C2f module strengthens the network's perception through an adaptive channel and spatial attention mechanism and improves the capture efficiency of objects at different scales; S37: The output feature map of the first A2C2f module is fed into the fifth convolutional layer, which uses a 3×3 convolutional kernel with a stride of 2 and a padding of 1. The role of the fifth convolutional layer is to further downsample, compress the size of the feature map, extract deeper and more semantic feature information, and at the same time increase the number of channels to meet the requirements of subsequent feature extraction; S38: The output feature map of the fifth convolutional layer is input into the second A2C2f module, which follows the architecture of the first A2C2f module. After the feature map is divided into regions, the self-attention mechanism is applied to each region, and it is dynamically adjusted according to the associations between regions to highlight the key regions and weaken the interference items; the features processed in each region are fused, superimposed on the input feature map, and output through the ReLU activation function; the size of the output feature map of the second A2C2f module remains unchanged, and the feature map is optimized by attention, and the features of the key regions are further strengthened, and the accuracy of feature expression is improved; S39: The output feature map of the second A2C2f module is input into the MSCA module, which is mainly composed of a multi-scale convolutional module and an attention module. The multi-scale convolutional module uses convolutional kernels of different sizes to perform convolution on the feature map, so as to capture different scale features from fine-grained to coarse-grained, integrate the features extracted by multi-scale convolution, and generate attention weights; local information is aggregated through depth convolution, then multi-branch depth convolution is used to capture multi-scale context information, and finally 1×1 convolution is used to simulate the relationship between different channels in the features to generate the weights of convolutional attention.

[0013] As a further technical solution of the present invention, in S36, the A2C2f module processes the input features using two convolutional layers cv1 and cv2 to obtain feature maps with two different numbers of channels. Cv1 reduces the number of input channels by half, while cv2 keeps the number of input channels unchanged. Each convolutional layer consists of a Conv2d layer, a BatchNorm2d layer, and an activation function SiLU or Identity, which are used to perform convolutional operations, normalization, and non-linear transformation on the input features. In each ABlock module, there is an attention mechanism module AAttn and a multi-layer perceptron MLP module. First, the qkv layer in AAttn generates query, key, and value vectors, and the Flash Attention mechanism is used to calculate the attention weights to achieve the spatial attention mechanism for features, highlighting the features in important spatial regions. At the same time, position encoding is performed through the pe convolutional layer to enhance the model's perception ability of position information. Then, the attention result is projected back to the original dimension through the proj convolutional layer and connected to the input features through a residual connection to achieve feature fusion and enhancement. Next, the MLP module performs further non-linear transformation on the fused features to further extract and fuse features. Among them, qkv is a convolutional layer used to generate query, key, and value vectors, and the number of output channels is three times that of the input. Proj is another convolutional layer used to project the attention result back to the original dimension. Pe is the position encoding, which is implemented through a depthwise separable convolution with a kernel size of 7x7. By stacking two ABlock modules, multiple attention calculations and feature fusions can be performed on the features, thereby enhancing the model's ability to capture objects of different scales.

[0014] As a further technical solution of the present invention, in S4, the neck network is improved based on the CA attention mechanism, which combines the channel attention and spatial position information contained in the previous feature map to improve the model's expression ability for the feature map.

[0015] As a further technical solution of the present invention, S4 specifically includes: S41: The output feature map of the MSCA in the backbone module is upsampled and fused with the first A2C2f layer, and then passed into the third A2C2f layer. The A2C2f layer has the same structure as the first A2C2f layer. The output feature map is upsampled and fused with the output features of the second A2C2f layer, and then output to the fourth A2C2f layer. S42: The features processed by the fourth A2C2f layer are input into the embedded CA attention module and then output to the multi-scale module to the detection module. S43: The features processed by the fourth A2C2f layer are also output to the sixth convolutional layer, and the features processed by the sixth convolutional layer are fused with the output of the third A2C2f layer. S44: After fusion, input the fifth A2C2f layer, and the output feature map is input into the CA attention module for processing, and the multi-scale feature map is output to the detection module; S45: The feature map output by the fifth A2C2f layer is simultaneously input into the seventh convolutional layer, the output result is fused with the output of the MSCA module, and then output to the third C3k2 layer. Its output feature map flows into the CA attention module, and then the output of the multi-scale module three is given to the detection module.

[0016] As a further technical solution of the present invention, in S5, the detection head network of the YOLOv12 model is mainly responsible for performing the object detection tasks of cyst, fibroadenoma, suspicious mass, and calcification detection.

[0017] As a further technical solution of the present invention, S5 specifically includes: S51: Location prediction: Use the anchor box mechanism to predict the bounding box. Define anchor boxes with different scales and aspect ratios in advance, and then predict the offset and confidence of the bounding box for each anchor box to determine the position and size of the target. Through the combination of fully connected layers and convolutional layers, predict the coordinates of the bounding box for each feature point, usually including the center coordinates (x, y), width, and height of the bounding box; S52: Classification prediction: Perform classification prediction on each feature point or anchor box, and output the probability distribution that the position belongs to the categories of cyst, fibroadenoma, suspicious mass, and calcification. Use the softmax activation function to convert the output of the classification prediction into probability values to make the results more interpretable; S53: Loss calculation: Use the cross-entropy loss function to measure the difference between the classification prediction result and the true label, calculate the difference between the predicted bounding box and the true bounding box based on the mean square error and CIoU loss, and output the bounding box coordinates, corresponding class labels, and confidence scores after non-maximum suppression as the final detection results.

[0018] The beneficial effects of the present invention are: 1. The input layer can capture the boundary features of different tissue components in breast ultrasound images, including fat, glandular, and fibrous tissues, etc. This not only improves the accuracy and richness of feature extraction but also enhances the model's ability to clearly identify different tissue boundaries, solving the problem of inaccurate feature extraction in the prior art.

[0019] 2. Improve the backbone network of YOLOv12 based on the MSCA module, enhancing the ability to capture multi-scale lesion features. The improved backbone network can adaptively extract feature information of different scales, thus better capturing lesion features from fine-grained to coarse-grained. Compared with the prior art, it improves the feature extraction ability for multi-scale targets.

[0020] 3. The neck network improved by combining the CA attention mechanism enhances the model's ability to capture key feature regions. The CA attention mechanism can highlight important regions such as smaller calcification points, changes in the cyst wall, and the boundaries of fibroadenomas through the combination of channel attention and spatial location information, improving the accuracy and specificity of detection and solving the deficiencies of the prior art in key region recognition. Brief Description of the Drawings

[0021] Figure 1 It is a flowchart of a multi-scale adaptive lesion detection method based on breast ultrasound proposed by the present invention.

[0022] Figure 2 It is a schematic diagram of the MSCA attention module.

[0023] Figure 3 It is a schematic diagram of steps S3 - S5. Detailed Embodiments

[0024] To make the technical means, creative features, achieved purposes, and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0025] Please refer to the attached Figure 1 , a multi-scale adaptive lesion detection method based on breast ultrasound, including the following specific steps: S1: Image preprocessing: First, based on the letterbox method, the breast ultrasound is converted into a size of 1280×1280×3. This method can simultaneously maintain the original aspect ratio of the image and the integrity of the breast ultrasound content; determine the scaling ratio according to the original length and width of the breast ultrasound and the target length and width, and adjust the image size. The scaling ratio is selected as the smaller ratio of the target size to the original size to ensure that the image does not exceed the target size after scaling; add 0-pixel padding at the short sides of the top, bottom, left, and right of the image to make the long and short sides match the target size; S2: Input module: Receive the preprocessed image, extract features such as edges and textures through a convolutional block, clearly display the boundaries of fat, glands, fibrous tissues, etc., and then refine these boundary features through a deconvolutional block to complete the image expression. After passing through the input module, the image is converted into a size of 640×640×3 that is most suitable for processing by the backbone module; The convolutional block consists of a convolutional layer, a normalization layer, and an activation function. The convolutional layer uses a 3×3 convolutional kernel with a stride of 2, and then further normalizes and activates the convolutional output to enhance its expression ability; extract low-level features in the breast ultrasound through the first convolutional block, where the first convolutional block selects the GELU activation function to adaptively adjust the output according to the input data and enhance the feature expression; then extract more abstract features through the second convolutional block, where the activation function selects SiLU to further enhance the non-linearity of the feature expression.

[0026] This input module can retain important features while reducing computational complexity, and the deconvolution block can also help detect subtle lesions such as microcalcifications.

[0027] S3: Backbone module: Input the ultrasound image into the backbone module of the improved YOLOv12 model based on the MSCA module. Its core is the residual efficient layer aggregation network. First, it passes through the first convolutional layer and the second convolutional layer, then through the first C3k2 module, then through the third convolutional layer. After the output passes through the second C3k2 module and the fourth convolutional layer in sequence, it is input into the first A2C2f module. After passing through the fifth convolutional layer, it is input into the A2C2f module, and finally input into the MSCA module; The MSCA module is suitable for detecting targets with diverse sizes. First, it performs feature extraction on the feature map at different scales through the multi-scale convolution branch, and adaptively obtains multiple feature maps with different scale information; then, it calculates the channel weights for these feature maps through the attention mechanism, and performs weighted fusion on the feature maps according to the importance of the channels; the fused feature map then passes through the spatial attention mechanism to generate a spatial attention weight map, which is multiplied by the original feature map to make the model pay more attention to the key feature regions; The improved YOLOv12 model adopts a new type of convolutional block, which better meets the requirements of the medical scenario for lightweight operations and high parallelism. The Yolov12 model uses a series of smaller kernels, generally expressed as: , where: is the output feature, is the input feature, is the weight of the i-th convolutional kernel, is the bias of the i-th convolutional kernel; S31: The first convolutional layer uses a 3×3 convolutional kernel with a stride of 2 and a padding of 1 to perform basic feature extraction on the input image. After this convolutional operation, the model initially captures basic features such as the edges and textures of the image; S32: Input the output feature into the second convolutional layer. This layer continues to use a 3×3 convolutional kernel, but the stride becomes 1 and the padding remains 1. After the convolutional operation, the size of the feature map remains unchanged, and the number of channels doubles. The second convolutional layer further explores deeper feature information and enriches the feature expression while maintaining the stability of the feature map size; S33: The output of the second convolutional layer immediately enters the first C3k2 module. The input feature map is divided into two parts. One part is directly passed to retain shallow features, and the other part processes deep features through the standard Bottleneck and finally spliced and fused; The C3k2 module assists in feature extraction. The C3k2 module retains multiple bottleneck layers, similar to CSPNet. Through the stacking of multiple bottleneck layers, features are gradually extracted and fused to improve the feature expression ability; S34: The output feature map of the first C3k2 module flows into the third convolutional layer, which uses the filters of the first convolutional layer. The main function of the third convolutional layer is to further downsample, compress the size of the feature map, extract deeper and more abstract feature information, and at the same time expand the number of channels to meet the requirements of subsequent complex feature extraction; S35: The output feature map of the third convolutional layer enters the second C3k2 module and the fourth convolutional layer, which follows the structures of the first C3k2 module and the second convolutional layer; S36: The output feature map of the fourth convolutional layer is then input into the first A2C2f module. The A2C2f module enhances the network's perception through an adaptive channel and spatial attention mechanism and improves the capture efficiency of objects at different scales; The A2C2f module uses two convolutional layers cv1 and cv2 to process the input features, obtaining two feature maps with different numbers of channels. Cv1 reduces the number of input channels by half, while cv2 keeps the number of input channels unchanged; each convolutional layer consists of a Conv2d layer, a BatchNorm2d layer, and an activation function SiLU or Identity, which are used to perform convolutional operations, normalization, and non-linear transformation on the input features; in each ABlock module, there is an attention mechanism module AAttn and a multi-layer perceptron MLP module. First, the qkv layer in AAttn generates query, key, and value vectors, and the Flash Attention mechanism is used to calculate the attention weights to implement the spatial attention mechanism for features, highlighting the features in important spatial regions; at the same time, position encoding is performed through the pe convolutional layer to enhance the model's perception of position information; then, the attention result is projected back to the original dimension through the proj convolutional layer and connected with the input features through a residual connection to achieve feature fusion and enhancement; next, the MLP module performs further non-linear transformation on the fused features to further extract and fuse features; among them, qkv is a convolutional layer used to generate query, key, and value vectors, and the output number of channels is three times that of the input; proj is another convolutional layer used to project the attention result back to the original dimension; pe is the position encoding, which is implemented through a depthwise separable convolution with a kernel size of 7x7; by stacking two ABlock modules, multiple attention calculations and feature fusions can be performed on the features, thereby enhancing the model's ability to capture objects at different scales; S37: The output feature map of the first A2C2f module is fed to the fifth convolutional layer, which uses a 3×3 convolutional kernel with a stride of 2 and a padding of 1. The function of the fifth convolutional layer is to further downsample, compress the size of the feature map, extract deeper and more semantic feature information, and at the same time increase the number of channels to meet the requirements of subsequent feature extraction; S38: The output feature map of the fifth convolutional layer is input into the second A2C2f module. This module follows the architecture of the first A2C2f module. After the feature map is divided into regions, the self-attention mechanism is applied to each region. Based on the correlations between regions, it is dynamically adjusted to highlight key regions and weaken interference items. The features processed in each region are fused, and the input feature map is superimposed, and then output after passing through the ReLU activation function. The size of the output feature map of the second A2C2f module remains unchanged. After the feature map is optimized by attention, the features in the key regions are further strengthened, and the accuracy of feature expression is improved. S39: The output feature map of the second A2C2f module is input into the MSCA module. This module is mainly composed of a multi-scale convolutional module and an attention module. The multi-scale convolutional module uses convolutional kernels of different sizes to perform convolution on the feature map, so as to capture different scale features from fine-grained to coarse-grained. The features extracted by multi-scale convolution are integrated to generate attention weights. Local information is aggregated through depth convolution, then multi-scale context information is captured by using multi-branch depth convolution, and finally the relationship between different channels in the features is simulated through 1×1 convolution to generate the weights of convolutional attention.

[0028] S4: Neck network: CA aggregates features in the horizontal and vertical directions respectively through 1D pooling operations to generate feature maps in two directions. Specifically, average pooling with a pooling kernel size of and is used to process the input feature map to obtain two direction-aware feature maps. The two direction-aware feature maps are concatenated, and then features are extracted through convolution operations to further generate attention weights in two directions. The h-swish non-linear activation function is used to increase the expression ability of the model. The generated attention weights are applied to the horizontal and vertical directions of the input feature map respectively to achieve direction-aware and position-sensitive attention enhancement. The neck network is improved based on the CA attention mechanism. It combines the channel attention and spatial position information contained in the previous feature map to improve the model's expression ability for the feature map. S41: The output feature map of the MSCA of the backbone module is upsampled and then fused with the first A2C2f layer, and then passed into the third A2C2f layer. The A2C2f layer has the same structure as the first A2C2f layer. The output feature map is upsampled and then fused with the output features of the second A2C2f layer, and then output to the fourth A2C2f layer. S42: The features processed by the fourth A2C2f layer are input into the embedded CA attention module and then output to the multi-scale module to the detection module. S43: The features processed by the fourth A2C2f layer are also output to the sixth convolutional layer. The features processed by the sixth convolutional layer are fused with the output of the third A2C2f layer. S44: Input into the fifth A2C2f layer after fusion. The output feature map is input into the CA attention module for processing, and the output multi-scale feature map is sent to the detection module; S45: The feature map output by the fifth A2C2f layer is simultaneously input into the seventh convolutional layer. The output result is fused with the output of the MSCA module and then sent to the third C3k2 layer. Its output feature map flows into the CA attention module, and then the output of the multi-scale module three is given to the detection module.

[0029] S5: Detection head network: The detection head network receives the multi-scale feature maps from the neck network. These feature maps are subjected to feature fusion and adjustment, containing semantic information and spatial details at different levels, for location prediction, classification prediction, loss calculation, and finally the output result; The detection head network of the YOLOv12 model is mainly responsible for performing object detection tasks for cyst, fibroadenoma, suspicious masses, and calcification detection; S51: Location prediction: Use the anchor box mechanism to predict the bounding box. Define anchor boxes with different scales and aspect ratios in advance, and then predict the offsets and confidences of the bounding boxes for each anchor box to determine the position and size of the object. Through the combination of fully connected layers and convolutional layers, predict the coordinates of the bounding box for each feature point, usually including the center coordinates (x, y), width, and height of the bounding box; S52: Classification prediction: Perform classification prediction for each feature point or anchor box, and output the probability distribution of the position belonging to the categories of cyst, fibroadenoma, suspicious masses, and calcification. Use the softmax activation function to convert the output of the classification prediction into probability values to make the results more interpretable; S53: Loss calculation: Use the cross-entropy loss function to measure the difference between the classification prediction result and the true label. Calculate the difference between the predicted bounding box and the true bounding box based on the mean square error and CIoU loss. Output the bounding box coordinates, corresponding class labels, and confidence scores after non-maximum suppression as the final detection results.

[0030] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects: Design a brand-new input layer. For the boundaries of different tissue components, its structure can extract features such as edges and textures, clarify the boundaries of clear fat, glands, fibrous tissues, etc. For the case of limited resolution of ultrasound images, the input layer reduces the computational complexity while retaining important features, improving the accuracy and richness of feature extraction.

[0031] In breast ultrasound images, breast lesions such as cysts, fibroadenomas, and calcifications present a variety of different sizes and shapes. Based on the MSCA module, the backbone network of YOLOv12 is improved, and feature maps of different scale information are adaptively extracted through multi-scale convolutional features, enabling the improved YOLOv12 model to better capture lesion features from fine-grained to coarse-grained. By using a simple element-wise multiplication operation to weight the input features, while achieving low computational complexity, the spatial information encoding of breast ultrasound is efficiently extracted.

[0032] The neck network is improved by combining the CA attention mechanism. Through the combination of channel attention and spatial position information, irrelevant background information is suppressed, interference is reduced, and the model's ability to capture key feature regions is enhanced, highlighting regions crucial for detecting lesions, including small calcification points, changes in the cyst wall, and the boundaries of fibroadenomas.

[0033] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present invention is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity.

Claims

1. A multi-scale adaptive lesion detection method based on breast ultrasound, characterized in that It includes the following specific steps: S1: Image preprocessing: Based on the letterbox method, convert the breast ultrasound image to a size of 1280×1280×3, determine the scaling ratio according to the original length and width and the target length and width of the breast ultrasound, and adjust the image size; S2: Input module: Receive the preprocessed image, extract features through a convolutional block to clearly display the boundaries, and then refine these boundary features through a deconvolutional block to complete the image representation; S3: Backbone module: Input the ultrasound image into the backbone module of the YOLOv12 model improved based on the MSCA module. After passing through the first convolutional layer and the second convolutional layer, then through the first C3k2 module, then through the third convolutional layer, the output passes through the second C3k2 module and the fourth convolutional layer in sequence, and then is input into the first A2C2f module. After passing through the fifth convolutional layer, it is input into the A2C2f module, and finally input into the MSCA module; S4: Neck network: CA aggregates features in the horizontal and vertical directions respectively through 1D pooling operations to generate feature maps in two directions. Concatenate the feature maps in two directions, and then extract features through convolutional operations to further generate attention weights in two directions; S5: Detection head network: The detection head network receives the multi-scale feature maps from the neck network, performs localization prediction, classification prediction, loss calculation, and finally outputs the results.

2. The multi-scale adaptive lesion detection method based on breast ultrasound according to claim 1, wherein In S2, the convolutional block consists of a convolutional layer, a normalization layer, and an activation function. The convolutional layer uses a 3×3 convolutional kernel with a stride of 2, and then further normalizes and activates the convolutional output; extract low-level features in the breast ultrasound through the first convolutional block, where the first convolutional block selects the GELU activation function and adaptively adjusts the output according to the input data; then extract more abstract features through the second convolutional block, where the activation function selects SiLU.

3. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 1, characterized in that In S3, the MSCA module is suitable for detecting targets with diverse sizes. First, perform feature extraction on the feature map at different scales through a multi-scale convolution branch to adaptively obtain multiple feature maps with different scale information; then calculate the weights between channels of these feature maps through an attention mechanism, and perform weighted fusion on the feature maps according to the importance of the channels; the fused feature map then passes through a spatial attention mechanism to generate a spatial attention weight map, which is multiplied by the original feature map to make the model pay more attention to key feature regions.

4. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 3, characterized in that, In S3, the improved YOLOv12 model adopts a new type of convolutional block. The Yolov12 model utilizes a series of smaller kernels, generally expressed as: , where: is the output feature, is the input feature, is the weight of the i-th convolutional kernel, is the bias of the i-th convolutional kernel.

5. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 4, wherein, S3 specifically includes: S31: The first convolutional layer uses a 3×3 convolutional kernel with a stride of 2 and a padding of 1 to perform basic feature extraction on the input image; S32: Input the output features into the second convolutional layer. This layer still uses a 3×3 convolutional kernel, but the stride becomes 1 and the padding remains 1. After the convolutional operation, the size of the feature map remains unchanged, and the number of channels doubles; S33: The output of the second convolutional layer immediately enters the first C3k2 module. The input feature map is divided into two parts. One part is directly passed to retain shallow features, and the other part processes deep features through a standard Bottleneck, and finally they are concatenated and fused; S34: The output feature map of the first C3k2 module flows into the third convolutional layer, which uses the filters of the first convolutional layer. The main function of the third convolutional layer is to further downsample, compress the size of the feature map, extract deeper and more abstract feature information, and at the same time expand the number of channels; S35: The output feature map of the third convolutional layer enters the second C3k2 module and the fourth convolutional layer, which follows the structures of the first C3k2 module and the second convolutional layer; S36: The output feature map of the fourth convolutional layer is then input into the first A2C2f module, and the A2C2f module enhances the network's perception through an adaptive channel and spatial attention mechanism and improves the capture efficiency of objects at different scales; S37: The output feature map of the first A2C2f module is fed to the fifth convolutional layer, which uses a 3×3 convolutional kernel with a stride of 2 and a padding of 1. The role of the fifth convolutional layer is to further downsample, compress the size of the feature map, extract deeper and more semantic feature information, and at the same time increase the number of channels; S38: The output feature map of the fifth convolutional layer is input into the second A2C2f module, which follows the architecture of the first A2C2f module. After the feature map is divided into regions, the self-attention mechanism is applied to each region, dynamically adjusted according to the correlation between regions to highlight key regions and weaken interference items; fuse the features after processing each region, superimpose the input feature map, and output through the ReLU activation function; the size of the output feature map of the second A2C2f module remains unchanged, the feature map is optimized by attention, the features of key regions are further strengthened, and the accuracy of feature expression is improved; S39: The output feature map of the second A2C2f module is input into the MSCA module, which is mainly composed of a multi-scale convolutional module and an attention module. The multi-scale convolutional module convolves the feature map with convolutional kernels of different sizes to capture different scale features from fine-grained to coarse-grained, integrates the features extracted by multi-scale convolution to generate attention weights; aggregates local information through depth convolution, then captures multi-scale context information using multi-branch depth convolution, and finally simulates the relationship between different channels in the features through 1×1 convolution to generate the weights of convolutional attention.

6. The multi-scale adaptive lesion detection method based on breast ultrasound according to claim 5, wherein, In S36, the A2C2f module processes the input features using two convolutional layers cv1 and cv2 to obtain feature maps with two different numbers of channels. Cv1 reduces the number of input channels by half, while cv2 keeps the number of input channels unchanged. Each convolutional layer consists of a Conv2d layer, a BatchNorm2d layer, and an activation function SiLU or Identity. In each ABlock module, there is an attention mechanism module AAttn and a multi-layer perceptron MLP module. First, the qkv layer in AAttn generates query, key, and value vectors, and the Flash Attention mechanism is used to calculate the attention weights to implement the spatial attention mechanism for features, highlighting the features in important spatial regions. At the same time, position encoding is performed through the pe convolutional layer to enhance the model's perception ability of position information. Then, the attention result is projected back to the original dimension through the proj convolutional layer and residual connection is made with the input features to achieve feature fusion and enhancement. Next, the MLP module performs further non-linear transformation on the fused features to further extract and fuse features. By stacking two ABlock modules, multiple attention calculations and feature fusions can be performed on the features, thereby enhancing the model's ability to capture objects of different scales.

7. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 1, characterized in that, In S4, the neck network is improved based on the CA attention mechanism, which combines the channel attention and spatial position information contained in the previous feature map to improve the model's expression ability for the feature map.

8. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 7, characterized in that S4 specifically includes: S41: The output feature map of the MSCA in the backbone module is upsampled and fused with the first A2C2f layer, and then passed into the third A2C2f layer. The A2C2f layer has the same structure as the first A2C2f layer. The output feature map is upsampled and fused with the output features of the second A2C2f layer, and then output to the fourth A2C2f layer. S42: The features processed by the fourth A2C2f layer are input into the embedded CA attention module and then output to the multi-scale module to the detection module. S43: The features processed by the fourth A2C2f layer are also output to the sixth convolutional layer, and the features processed by the sixth convolutional layer are fused with the output of the third A2C2f layer. S44: After fusion, it is input into the fifth A2C2f layer, and the output feature map is input into the CA attention module for processing, and the multi-scale feature map is output to the detection module. S45: The feature map output by the fifth A2C2f layer is simultaneously input into the seventh convolutional layer, and the output result is fused with the output of the MSCA module, and then output to the third C3k2 layer. Its output feature map flows into the CA attention module, and then the multi-scale module three is output to the detection module.

9. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 1, characterized in that In S5, the detection head network of the YOLOv12 model is mainly responsible for performing the object detection tasks of cyst, fibroadenoma, suspicious masses, and calcification detection.

10. A multi-scale adaptive lesion detection method based on breast ultrasound according to claim 9, characterized in that, S5 specifically includes: S51: Location prediction: Use the anchor box mechanism to predict the bounding box. Define anchor boxes with different scales and aspect ratios in advance, and then predict the offsets and confidences of the bounding boxes for each anchor box to determine the location and size of the target. Through the combination of fully connected layers and convolutional layers, predict the coordinates of the bounding box for each feature point; S52: Classification prediction: Perform classification prediction on each feature point or anchor box, and output the probability distribution that the location belongs to categories such as cysts, fibroadenomas, suspicious masses, and calcifications. Use the softmax activation function to convert the output of the classification prediction into probability values to make the results more interpretable; S53: Loss calculation: Use the cross-entropy loss function to measure the difference between the classification prediction result and the true label. Calculate the difference between the predicted bounding box and the true bounding box based on the mean squared error and CIoU loss. Output the coordinates of the bounding box after non-maximum suppression, the corresponding class label, and the confidence score as the final detection result.

Citation Information

Patent Citations

  • Breast mass detection method based on multi-scale cross-path feature fusion

    CN115423806A

  • Pulmonary nodule detection and identification system based on YOLOv5

    CN115661029A

  • Tumor lesion area detection method and device based on position prior and feature perception

    CN117392119A

  • Ultrasonic image breast tumor classification method based on feature fusion and attention mechanism

    CN117746119A

  • Ultrasonic mammary gland lesion area automatic segmentation method based on multi-scale hybrid convolution

    CN119600041A

Cited By

  • Breast ultrasonic automatic pressurization method and device based on image feedback

    CN121081018A

  • Breast ultrasound automatic compression method and device based on image feedback

    CN121081018B

  • Ultrasonic contrast video analysis method, system, equipment, medium and product

    CN121599950A