Deep learning-based breast mass focus detection method and device, and storage medium
By adopting the deep learning method of adaptive cross attention pyramid network, breast density information perception module and uncertainty boundary modeling module in breast mass lesions detection, the accuracy challenge of breast mass detection in DBT images is solved, and higher detection accuracy and distinction capabilities are achieved.
Patent Information
- Application Number
- CN202510133510.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art presents accuracy challenges in breast mass lesions detection, especially in the case of low contrast between the masses and surrounding tissues, variable morphology, blurred edges, and highly similar to dense glands in DBT images, resulting in false detection.
Using the deep learning-based breast mass lesion detection method, the three-dimensional DBT image is preprocessed, and an adaptive cross attention pyramid network, a breast density information perception module and an uncertainty boundary modeling module are established to improve the detection network and realize the accurate detection of breast mass lesion.
It improves the detection accuracy of breast mass lesions in DBT images, reduces error detection, enhances the ability to distinguish dense breast areas, and effectively deals with blurred or obstructed mass boundaries.
Smart Images

Figure CN120107169A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing, and in particular to a method and device for detecting breast mass lesions based on deep learning, and a storage medium. Background Art
[0002] Medical image detection is the basis of medical image analysis and a key step in computer-aided diagnosis. Medical image detection refers to the process of identifying and locating specific anatomical structures or lesion areas in medical images, such as breast mass detection in digital breast tomosynthesis (DBT) images. The accuracy of the detection results is of great significance for the early diagnosis of breast cancer. Although the convolutional neural network-based method has achieved excellent performance in medical image detection tasks, the accurate detection of mass lesions is still challenging due to the low contrast between the mass lesion area and the surrounding tissue in the DBT image, the existence of morphological changes, blurred edges, and the high similarity between the mass lesion and the dense gland, which can lead to false detection.
[0003] Therefore, there is an urgent need for a breast mass lesion detection method based on DBT images. Summary of the invention
[0004] In order to achieve the above-mentioned purpose and other advantages of the present invention, the first object of the present invention is to provide a method for detecting breast mass lesions based on deep learning, comprising the following steps:
[0005] Preprocess the three-dimensional DBT image dataset;
[0006] Establish a detection network architecture based on convolutional neural network;
[0007] The detection network is improved by constructing an adaptive cross-attention pyramid network, a breast density information perception module, an uncertainty boundary modeling module, and merging 2D detection results into 3D detection results to obtain a breast mass lesion detection network;
[0008] Training the breast mass lesion detection network by using the preprocessed data;
[0009] The DBT image to be detected is input into the breast mass lesion detection network to detect the breast mass lesion area.
[0010] Furthermore, the step of preprocessing the three-dimensional DBT image data set includes:
[0011] The three-dimensional DBT images were uniformly transferred to the same side of the view;
[0012] The three-dimensional DBT images were subjected to window width and window position adjustment, breast area segmentation based on threshold method, image mean normalization, and image scaling.
[0013] Furthermore, the step of preprocessing the three-dimensional DBT image data set also includes:
[0014] The three adjacent slices form a three-channel two-dimensional image as the input of the model;
[0015] In the training phase, the center slice and two adjacent slices are selected as training data according to the coordinates of the label box;
[0016] The images are processed using online data augmentation.
[0017] Furthermore, the step of establishing a detection network architecture based on a convolutional neural network includes:
[0018] The FCOS network is used as the baseline network structure;
[0019] Improvements to the detection network include:
[0020] The backbone network of the baseline network is kept unchanged, the feature pyramid network is replaced by an adaptive cross-attention pyramid network with a breast density information perception module, and the regression branch in the detection regression head is replaced by an uncertainty boundary modeling module.
[0021] Furthermore, the construction of the adaptive cross-attention pyramid network includes:
[0022] Using convolution to unify the number of feature map channels of the multi-scale feature map obtained by the backbone network;
[0023] Use bilinear interpolation based upsampling to achieve uniform resolution.
[0024] The output feature maps from the second to fourth layers are used as the input of the adaptive criss-cross attention pyramid network.
[0025] Furthermore, the step of using the output feature maps of the second to fourth layers as the input of the adaptive cross attention pyramid network includes:
[0026] For each layer of feature maps, they are first upsampled to the resolution of the second layer of feature maps, and then mapped to three feature spaces of the same dimensions, namely query space, key space, and value space;
[0027] Perform matrix multiplication on the query space feature map of this layer and the key space of all layers respectively to obtain three correlation weight matrices;
[0028] The output feature map of this layer is obtained by adding feature maps, adjusting the number of channels through convolution, and upsampling.
[0029] Furthermore, the construction of the breast density information perception module includes:
[0030] For each layer of multi-scale feature maps obtained from the backbone network, convolution is used to unify the number of feature map channels;
[0031] Use bilinear interpolation based upsampling to achieve uniform resolution.
[0032] All feature maps are sent to the breast density information perception module.
[0033] Furthermore, the step of sending all feature maps to the breast density information perception module includes:
[0034] The number of channels of feature maps at each scale compressed by convolution;
[0035] The feature maps of each scale are spliced and then convolved to obtain the perception result of the breast density map;
[0036] The obtained breast density map and the input feature map are unified in number of channels, and then the Hadamard matrix product is performed, and then the weighted result of the breast density map of the input feature map of this layer is obtained after pixel-by-pixel addition;
[0037] The breast density supervision map is obtained by processing the DBT image slices corresponding to the input image.
[0038] Furthermore, the loss function used in the breast density information perception module is:
[0039]
[0040] Among them, N pos Indicates the number of pixels involved in the calculation in the breast density map, D x,y and Represent the value of each position in the breast density map and the breast density supervision map respectively.
[0041] Furthermore, the construction of the uncertainty boundary modeling module includes:
[0042] By predicting the distance offset from the current pixel position to the four boundaries of the label box, the discrete boundary distribution function is predicted.
[0043] Furthermore, the step of predicting the discrete boundary distribution function by predicting the distance offset from the current pixel position to the four boundaries of the label box includes:
[0044] For each pixel position, the uncertainty boundary modeling module predicts 4k results, which represent discrete distribution vectors of length k in four directions;
[0045] The correlation between the elements in the distribution vector is further modeled through grouped convolution; the value in the vector represents the probability that the label box in the corresponding direction is located in the current pixel interval, and the final predicted box position is obtained by weighted summation.
[0046] Furthermore, the step of merging the 2D detection results into 3D detection results includes:
[0047] In the testing phase, all slices of each DBT data are fed into the model one by one to obtain the corresponding two-dimensional detection box;
[0048] The two-dimensional detection frames are fused along the z-axis to obtain the final three-dimensional detection result.
[0049] Furthermore, the step of fusing the two-dimensional detection frames along the z-axis direction to obtain the final three-dimensional detection result includes:
[0050] Calculate the intersection-and-union ratio of two-dimensional detection boxes of adjacent slices;
[0051] If it is greater than the threshold, it is considered that the same lesion is detected, and the bounding boxes will be merged only when the lesion is detected in three consecutive slices or more. The final 3D bounding box is the union of the 2D bounding boxes, and the maximum confidence value in the 2D result is selected as the predicted confidence of the 3D bounding box.
[0052] Furthermore, the training of the breast mass lesion detection network includes:
[0053] The original image and the gold standard are sent to the entire network for supervised learning. The overall loss function L total It consists of five parts: classification loss L cls , prediction box quality evaluation loss L qlr , uncertainty boundary modeling loss L mod , the bounding box regression loss L reg and density-aware loss L dm , the overall loss function is defined as follows:
[0054]
[0055] Among them, N pos Represents the number of positive samples; x and y represent the position coordinates on the full scale feature map; They are the prediction results and labels for the classification task, prediction box quality assessment task, bounding box modeling task, bounding box regression task, and density perception task; is the positive sample labeling function, which is 1 only when there is a target to be detected at the current position.
[0056] Furthermore, the classification loss L cls Select Focal Loss, the prediction box quality assessment loss Lqlt Using binary cross entropy loss, the uncertainty boundary modeling loss L mod Select cross entropy loss, the bounding box regression loss L reg The GIoU loss is selected, and the density-aware loss L dm Use smooth L1 loss.
[0057] A second object of the present invention is to provide a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0058] A third object of the present invention is to provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] The present invention provides a medical image detection method for breast mass lesions, which can effectively realize breast mass lesion detection in DBT images.
[0061] The present invention improves the problems existing in the existing detection network in the breast mass detection task: in order to solve the problem that the complex background of the mass in DBT affects the detection accuracy, the present invention proposes an adaptive cross-attention pyramid to realize adaptive multi-scale feature fusion and improve the feature interaction capability between multi-scale feature maps; in order to solve the problem of false detection caused by dense breast areas, the present invention proposes to use breast density distribution for weighting to guide the network to strengthen the distinction between breast cancer lesions and dense glands; in order to solve the phenomenon of blurred or occluded mass boundaries, the present invention proposes to use a discrete distribution modeling method to obtain a discrete distribution function for predicting the position of the detection box, thereby reducing the difficulty of model optimization.
[0062] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. The specific implementation of the present invention is given in detail by the following embodiments and their accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0064] Figure 1 This is a flow chart of a method for detecting breast mass lesions based on deep learning according to Example 1;
[0065] Figure 2 This is a schematic diagram of the deep learning-based breast mass lesion detection method of Example 1;
[0066] Figure 3 Schematic diagram of the adaptive cross-attention pyramid network structure of Example 1;
[0067] Figure 4 This is a schematic diagram of the structure of a breast density information sensing module according to Example 1;
[0068] Figure 5 This is a schematic diagram of the process of aggregating the 2D detection results of Example 1 into 3D detection results;
[0069] Figure 6 A schematic diagram showing the comparison of different detection methods in Example 1 on DBT images;
[0070] Figure 7 This is a schematic diagram of a computer device according to Embodiment 2;
[0071] Figure 8 Schematic diagram of a computer-readable storage medium of Example 3. DETAILED DESCRIPTION
[0072] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. It should be noted that, under the premise of no conflict, the embodiments or technical features described below can be arbitrarily combined to form a new embodiment.
[0073] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0074] The figure numbers in this application are only used to distinguish the various steps in the scheme, and are not used to limit the execution order of the various steps. The specific execution order is subject to the description in the specification.
[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0076] Example 1
[0077] A breast mass lesion detection method based on deep learning, such as Figure 1 , Figure 2 As shown, the following steps are included:
[0078] S1. Preprocessing the 3D DBT image dataset;
[0079] According to the difference in shooting position and shooting left and right breasts, DBT images can be divided into left breast axial position (LCC), left breast oblique position (LMLO), right breast axial position (RCC), and right breast oblique position (RMLO), and each image occupies a different area of the DBT view. To facilitate the subsequent unified processing of images, further, the preprocessing step of the three-dimensional DBT image data set includes:
[0080] Rotate the 3D DBT images uniformly to the same side of the view; for example, rotate the images uniformly to the left side of the view.
[0081] Common preprocessing operations were performed on the three-dimensional DBT images, including: window width and window position adjustment, breast region segmentation based on threshold method, image mean normalization, and the processed images were scaled to a resolution of 1000×500 to reduce the algorithm calculation amount.
[0082] In order to ensure the spatial position correlation of the slices in the z-axis direction, this embodiment uses three adjacent slices to form a three-channel two-dimensional image as the input of the model. In the training stage, the center slice and two adjacent slices are selected as training data according to the coordinates of the label box. In order to expand the data diversity, online data enhancement methods such as random horizontal flipping and vertical flipping are used.
[0083] S2. Establish a detection network architecture based on convolutional neural network;
[0084] In this embodiment, the detection network uses the FCOS network as the baseline network structure. FCOS (Fully Convolutional One-Stage Object Detection) is a target detection method based on convolutional neural networks. The network consists of three parts: a backbone network (ResNet50) for image feature extraction, a feature pyramid network (FPN) layer for feature fusion and enhancement, and a set of detection regression heads for decoding detection results from features.
[0085] S3, improving the detection network by constructing an adaptive cross-attention pyramid network, a breast density information perception module, an uncertainty boundary modeling module, and merging 2D detection results into 3D detection results to obtain a breast mass lesion detection network;
[0086] In this embodiment, the improvement of the detection network includes:
[0087] The backbone network of the baseline network is kept unchanged, and the feature pyramid network is replaced by an adaptive cross-attention pyramid network with a breast density information perception module to enhance the cross-layer feature interaction capability and breast lesion differentiation capability. The regression branch in the detection regression head is replaced by an uncertainty boundary modeling module to obtain a more accurate bounding box position.
[0088] This embodiment constructs an adaptive cross-attention pyramid network for global correlation modeling of multi-scale feature maps to improve the effectiveness of cross-scale feature interaction and the directness of feature fusion.
[0089] Furthermore, if Figure 3 As shown, the construction of the adaptive cross attention pyramid network includes:
[0090] In terms of network structure design, the multi-scale feature map obtained by the backbone network is firstly convolved with a convolution kernel of 1×1 to unify the number of feature map channels, and then the upsampling operation based on bilinear interpolation is used to achieve the unification of resolution. Considering the feature extraction effect of ResNet50, this embodiment uses the output feature maps of the second to fourth layers as the input of the adaptive cross attention pyramid network. Specifically, for each layer of feature maps, it is first upsampled to the resolution of the second layer feature map, and then mapped to three feature spaces of the same dimension, namely the query space, key space and value space. The query space feature map of this layer is matrix multiplied with the key space of all layers respectively to obtain three correlation weight matrices. Then multiply the weight matrix with the corresponding value space matrix to obtain three multi-scale fused feature maps. Finally, the output feature map of this layer can be obtained by adding feature maps, adjusting the number of channels by 1×1 convolution and upsampling.
[0091] The feature fusion in the adaptive cross-attention pyramid network designed in this embodiment is bidirectional, that is, the prediction branch of the top-level feature map can obtain guidance from the underlying detail features, which can improve the accuracy of boundary positioning. In addition, feature maps between different scales can be directly fused, avoiding the problem of layer-by-layer attenuation of high-level features in traditional feature pyramids when they are passed to the bottom layer. Cross-attention realizes global correlation modeling through matrix multiplication. The feature vector of each position in the correlation weight map is weighted by the feature vectors of all positions in the input feature map, which also provides a global receptive field to facilitate the acquisition of additional global features.
[0092] This embodiment constructs a breast density information perception module to use breast density information to weight the intermediate feature map, so as to enhance the network's attention to dense breast areas and guide the network to strengthen the distinction between breast lumps and lesions and dense glands.
[0093] Furthermore, if Figure 4As shown, the construction of the breast density information perception module includes:
[0094] In terms of network structure design, for each layer of multi-scale feature maps obtained from the backbone network, the convolution kernel of 1×1 is first used to unify the number of feature map channels, and the upsampling operation based on bilinear interpolation is used to achieve resolution unification, and then all feature maps are sent to the breast density information perception module. Specifically, each scale feature map is first compressed to 3 channels by 1×1 convolution, and then the channels of each scale feature map are spliced and then convolved by 1x1 to obtain the perception result of the breast density map. The obtained breast density map is unified with the input feature map, and the Hadamard matrix product is performed, and then pixel-by-pixel addition is performed to obtain the weighted result of the breast density map of the input feature map of this layer. During training, the breast density distribution perception process is supervised by the artificially extracted breast density supervision map. The breast density supervision map is obtained by processing the DBT image slices corresponding to the input image. Specifically, the original slices are first gamma filtered to suppress the low-brightness area in the image and retain the high-brightness area, that is, the information of the high-density breast area is retained. The obtained image is then subjected to Gaussian filtering to smooth the breast area and remove the detailed texture information in the breast density map. Finally, it is normalized to obtain the breast density supervision map.
[0095] Furthermore, the loss function used in the breast density information perception module is:
[0096]
[0097] Among them, N pos Indicates the number of pixels involved in the calculation in the breast density map, D x,y and Represent the value of each position in the breast density map and the breast density supervision map respectively.
[0098] The breast density information perception module designed in this embodiment is placed before the cross attention calculation in the adaptive cross attention pyramid network. The breast density map in the module is directly learned from the network feature map. While the module is supervised by the module loss function, it is also supervised and optimized by the main network loss function. On the one hand, this saves the step of manually extracting breast density features in the reasoning process, and on the other hand, it also alleviates the negative impact of noise in the manually extracted breast density map on the network.
[0099] This embodiment constructs an uncertainty boundary modeling module to predict the discretized coordinate probability distribution function, so as to flexibly implement probability distribution modeling of the boundary box.
[0100] Furthermore, if Figure 5 As shown, the construction of the uncertainty boundary modeling module includes:
[0101] In terms of network structure design, the present invention predicts the discrete boundary distribution function by predicting the distance offset from the current pixel position to the four boundaries of the label box. Specifically, for each pixel position, the uncertainty boundary modeling module predicts 4k results, representing discrete distribution vectors of length k in four directions. The correlation between the elements in the distribution vector is further modeled through 1x1 grouped convolution. The value in the vector represents the probability that the label box in the corresponding direction is located within the current pixel interval, and the final predicted box position is obtained by weighted summation.
[0102] The uncertainty boundary modeling module designed in this embodiment is used to replace the boundary modeling method based on Dirac distribution in classical boundary regression. The traditional detection model predicts a single position coordinate instead of predicting a set of data, which is used to represent the probability distribution function of the discretized predicted coordinates. This can obtain a more accurate prediction of the fuzzy boundary position of the mass lesion.
[0103] This embodiment uses a two-dimensional slice-by-slice detection and fusion method to obtain the final three-dimensional prediction frame, so a post-processing operation of merging the detection frames is required. In the test phase, all slices of each DBT data are sent to the model one by one to obtain the corresponding two-dimensional detection frame. The two-dimensional detection frames are then fused along the z-axis to obtain the final three-dimensional detection result. Specifically, the intersection and union ratio of the two-dimensional detection frames of adjacent slices is first calculated. If it is greater than the threshold, it is considered that the same lesion is detected, and the bounding box will only be merged if the lesion is detected in three consecutive slices or more. The final three-dimensional bounding box is the union of the two-dimensional bounding boxes to ensure that all two-dimensional detection results can be contained to the greatest extent, and the maximum confidence value in the two-dimensional result is selected as the prediction confidence of the three-dimensional bounding box.
[0104] S4, training the breast mass lesion detection network using the preprocessed data;
[0105] Furthermore, the training of the breast mass lesion detection network includes:
[0106] To train the entire detection network, the original image and the detection gold standard need to be sent to the entire network for supervised learning. The overall loss function L total It consists of five parts: classification loss L cls , prediction box quality evaluation loss L qlt , uncertainty boundary modeling loss L mod , the bounding box regression loss L reg and density-aware loss L dm , the overall loss function is defined as follows:
[0107]
[0108] Among them, Npos Represents the number of positive samples; x and y represent the position coordinates on the full scale feature map; They are the prediction results and labels for the classification task, prediction box quality assessment task, bounding box modeling task, bounding box regression task, and density perception task; is the positive sample labeling function, which is 1 only when there is a target to be detected at the current position.
[0109] In this embodiment, the classification loss L cls Select Focal Loss, the prediction box quality assessment loss L qlt Using binary cross entropy loss, the uncertainty boundary modeling loss L mod Select cross entropy loss, the bounding box regression loss L reg The GIoU loss is selected, and the density-aware loss L dm Use smooth L1 loss.
[0110] S5. Input the DBT image to be detected into the breast mass lesion detection network to detect the breast mass lesion area.
[0111] In this embodiment, when testing the detection network, only the image to be tested needs to be input, and there is no need to detect the gold standard result. The detection network will automatically detect the breast mass lesion area based on the test image.
[0112] Comparison of different detection methods on DBT images Figure 6 As shown in the figure, the green box represents the gold standard, the yellow line represents the wrong detection result, and the blue line represents the correct detection result.
[0113] This embodiment adopts a detection network architecture based on a convolutional neural network, uses an adaptive cross-attention pyramid to guide the network to adaptively fuse multi-scale feature maps, and directly and efficiently extracts mass lesion features; uses a breast density information perception module to weight the intermediate feature map to improve the network's distinction between dense breasts and mass lesions; uses an uncertainty boundary modeling method to predict the distribution function of the bounding box position to improve the accuracy of locating the fuzzy edges of mass lesions. Through the above design, the performance of the convolutional neural network in the DBT image breast mass lesion detection task is improved.
[0114] Example 2
[0115] A computer device 600, such as Figure 7 As shown, it includes a memory 610, a processor 620, and a computer program 630 stored in the memory and executable on the processor. When the processor executes the computer program, the steps of a method for detecting breast lesions based on deep learning are implemented. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiment, which will not be repeated here.
[0116] Example 3
[0117] A computer readable storage medium such as Figure 8 As shown, a computer program is stored thereon, and when the computer program is executed by the processor, the steps of a method for detecting breast lesions based on deep learning are implemented. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiment, which will not be repeated here.
[0118] The number of devices and processing scales described here are used to simplify the description of the present invention. Applications, modifications and variations of the present invention will be obvious to those skilled in the art.
[0119] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and implementation modes. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.
[0120] The apparatus, computer device, non-volatile computer storage medium and method provided in the embodiments of this specification correspond to each other, and therefore, the apparatus, computer device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, computer device and non-volatile computer storage medium will not be repeated here.
[0121] Those skilled in the art also know that, in addition to implementing the controller in a purely computer-readable program code, the controller can be made to implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered as a hardware component, and the devices for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software units for implementing the method and structures within the hardware component.
[0122] The systems, devices or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above devices are described separately by functions in various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or more software and / or hardware.
[0123] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may be in the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of this specification may be in the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0124] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0125] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0127] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0128] The specification may be described in the general context of computer-executable instructions executed by a computer, such as program units. Generally, program units include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program units may be located in local and remote computer storage media, including storage devices.
[0129] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0130] The above description is only an embodiment of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of one or more embodiments of this specification.
Claims
1. A method for detecting breast mass lesions based on deep learning, characterized in that: The following steps are involved: Preprocess the three-dimensional DBT image dataset; Establish a detection network architecture based on convolutional neural network; The detection network is improved by constructing an adaptive cross-attention pyramid network, a breast density information perception module, an uncertainty boundary modeling module, and merging 2D detection results into 3D detection results to obtain a breast mass lesion detection network; Training the breast mass lesion detection network by using the preprocessed data; The DBT image to be detected is input into the breast mass lesion detection network to detect the breast mass lesion area.
2. A method for detecting breast lesions based on deep learning as claimed in claim 1, characterized in that: The step of preprocessing the three-dimensional DBT image data set comprises: The three-dimensional DBT images were uniformly transferred to the same side of the view; The three-dimensional DBT images were subjected to window width and window position adjustment, breast area segmentation based on threshold method, image mean normalization, and image scaling.
3. A method for detecting breast mass lesions based on deep learning as claimed in claim 2, characterized in that: The step of preprocessing the three-dimensional DBT image data set also includes: The three adjacent slices form a three-channel two-dimensional image as the input of the model; In the training phase, the center slice and two adjacent slices are selected as training data according to the coordinates of the label box; The images are processed using online data augmentation.
4. A method for detecting breast mass lesions based on deep learning as claimed in claim 1, characterized in that: The steps of establishing a detection network architecture based on a convolutional neural network include: The FCOS network is used as the baseline network structure; Improvements to the detection network include: The backbone network of the baseline network is kept unchanged, the feature pyramid network is replaced by an adaptive cross-attention pyramid network with a breast density information perception module, and the regression branch in the detection regression head is replaced by an uncertainty boundary modeling module.
5. A method for detecting breast mass lesions based on deep learning as claimed in claim 4, characterized in that: The construction of the adaptive cross-attention pyramid network includes: Using convolution to unify the number of feature map channels of the multi-scale feature map obtained by the backbone network; Use bilinear interpolation based upsampling to achieve uniform resolution; The output feature maps from the second to fourth layers are used as the input of the adaptive criss-cross attention pyramid network.
6. A method for detecting breast mass lesions based on deep learning as claimed in claim 5, characterized in that: The step of using the output feature maps of the second to fourth layers as the input of the adaptive cross attention pyramid network includes: For each layer of feature maps, they are first upsampled to the resolution of the second layer of feature maps, and then mapped to three feature spaces of the same dimensions, namely query space, key space, and value space; Perform matrix multiplication on the query space feature map of this layer and the key space of all layers respectively to obtain three correlation weight matrices; The output feature map of this layer is obtained by adding feature maps, adjusting the number of channels through convolution, and upsampling.
7. A method for detecting breast mass lesions based on deep learning as claimed in claim 4, characterized in that: The construction of the breast density information perception module includes: For each layer of multi-scale feature maps obtained from the backbone network, convolution is used to unify the number of feature map channels; Use bilinear interpolation based upsampling to achieve uniform resolution; All feature maps are sent to the breast density information perception module.
8. A method for detecting breast mass lesions based on deep learning as claimed in claim 7, characterized in that: The step of sending all feature maps to the breast density information perception module includes: The number of channels of feature maps at each scale compressed by convolution; The feature maps of each scale are spliced and then convolved to obtain the perception result of the breast density map; The obtained breast density map and the input feature map are unified in number of channels, and then the Hadamard matrix product is performed, and then the weighted result of the breast density map of the input feature map of this layer is obtained after pixel-by-pixel addition; The breast density supervision map is obtained by processing the DBT image slices corresponding to the input image.
9. A method for detecting breast mass lesions based on deep learning as claimed in claim 8, characterized in that: The loss function used in the breast density information perception module is: Among them, N pos Indicates the number of pixels involved in the calculation in the breast density map, D x,y and Represent the value of each position in the breast density map and the breast density supervision map respectively.
10. A method for detecting breast lesions based on deep learning as claimed in claim 1, characterized in that: The construction of the uncertainty boundary modeling module includes: By predicting the distance offset from the current pixel position to the four boundaries of the label box, the discrete boundary distribution function is predicted.
11. A method for detecting breast mass lesions based on deep learning as claimed in claim 10, characterized in that: The step of predicting the discrete boundary distribution function by predicting the distance offset from the current pixel position to the four boundaries of the label box includes: For each pixel position, the uncertainty boundary modeling module predicts 4k results, which represent discrete distribution vectors of length k in four directions; The correlation between the elements in the distribution vector is further modeled through grouped convolution; the value in the vector represents the probability that the label box in the corresponding direction is located in the current pixel interval, and the final predicted box position is obtained by weighted summation.
12. A method for detecting breast lesions based on deep learning as claimed in claim 1, characterized in that: The step of merging the 2D detection results into the 3D detection results comprises: In the testing phase, all slices of each DBT data are fed into the model one by one to obtain the corresponding two-dimensional detection box; The two-dimensional detection frames are fused along the z-axis to obtain the final three-dimensional detection result.
13. A method for detecting breast mass lesions based on deep learning as claimed in claim 12, characterized in that: The step of fusing the two-dimensional detection frames along the z-axis direction to obtain the final three-dimensional detection result includes: Calculate the intersection-and-union ratio of two-dimensional detection boxes of adjacent slices; If it is greater than the threshold, it is considered that the same lesion is detected, and the bounding boxes will be merged only when the lesion is detected in three consecutive slices or more. The final 3D bounding box is the union of the 2D bounding boxes, and the maximum confidence value in the 2D result is selected as the predicted confidence of the 3D bounding box.
14. A method for detecting breast mass lesions based on deep learning as claimed in claim 9, characterized in that: The training of the breast mass lesion detection network includes: The original image and the gold standard are sent to the entire network for supervised learning. The overall loss function L total It consists of five parts: classification loss L cls , prediction box quality assessment loss L qlr , uncertainty boundary modeling loss L mod , the bounding box regression loss L reg and density-aware loss L dm , the overall loss function is defined as follows: Among them, N pos Represents the number of positive samples; x and y represent the position coordinates on the full scale feature map; They are the prediction results and labels for the classification task, prediction box quality assessment task, bounding box modeling task, bounding box regression task, and density perception task; is the positive sample labeling function, which is 1 only when there is a target to be detected at the current position.
15. A method for detecting breast mass lesions based on deep learning as claimed in claim 14, characterized in that: The classification loss L cls Select Focal Loss, the prediction box quality assessment loss L qlt Using binary cross entropy loss, the uncertainty boundary modeling loss L mod Select cross entropy loss, the bounding box regression loss L reg The GIoU loss is selected, and the density-aware loss L dm Use smooth L1 loss.
16. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 15 are implemented.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.