Real-time detection method and system for breast nodules in ultrasonic image
The deep convolutional neural network model uses the deep convolutional neural network model to process breast ultrasound imaging data, real-time automatic detection and tracking of breast nodules, solving the problems of long examination time and poor repeatability caused by relying on physician experience in the existing technology, and improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510335320.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
AI Technical Summary
During the existing breast ultrasound diagnosis, the identification of breast nodules depends on the experience of the doctor, resulting in long examination time, poor repeatability, and difficulty in meeting the growing demand for imaging diagnosis and treatment.
The deep convolutional neural network model is used to preprocess and feature extraction of breast ultrasound imaging data, and the location of breast nodules is automatically captured to reduce the workload of the sonicator.
Real-time automatic detection and tracking of breast nodules is realized, which reduces the examination time of the sonographer, improves the efficiency and accuracy of the detection, and can assist the sonographer in completing the diagnosis of breast nodules.
Smart Images

Figure CN120219348A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of ultrasonic examination, and more specifically, relates to a real-time detection method and system for breast nodules in ultrasonic images. Background Art
[0002] Due to the advantages of low consumption, real-time imaging, and no radiation, ultrasonic imaging has been widely used in the screening and diagnosis of diseases. During the process of ultrasonic diagnosis using ultrasonic imaging, clinicians first need to obtain scanned images of the breast position, identify breast nodules, check whether the main anatomical structures are abnormal, and make analysis and diagnosis.
[0003] In the current process of breast ultrasonic diagnosis, the identification of breast nodules mainly relies on the experience of physicians. However, this identification method has some non-negligible defects: First, since it depends on the clinical experience of ultrasonic physicians and anatomical structure knowledge, the identification and measurement of breast nodules need to be processed by ultrasonic physicians. However, the traditional manual film reading mode has been difficult to meet the growing needs of diagnosis and treatment in the imaging department. According to statistics, the annual growth rate of medical imaging data in China is about 30%, while the annual growth rate of imaging department physicians is only 4.1%. The growth rate of imaging data far exceeds that of imaging department doctors, which means that imaging department doctors need to read a large number of medical images overload every day. It usually takes 5 - 30 minutes for an imaging department physician to diagnose a patient, which cannot meet the actual needs of many patients queuing for consultation. In addition, under the existing medical imaging technology, disease diagnosis still heavily relies on the professional knowledge and clinical experience of individual doctors, with strong subjectivity and poor repeatability. Especially in some relatively complex situations in medical imaging, such as small early lesion tissues and low contrast between lesions and surrounding anatomical structures, it strongly depends on doctors' rich clinical diagnosis experience, otherwise it is easy to miss diagnosis and misdiagnosis. Summary of the Invention
[0004] The object of the present invention is to provide a real-time detection method and system for breast nodules in ultrasonic images, which can assist ultrasonic physicians to complete the automatic positioning and diagnosis of breast nodules and reduce the workload of ultrasonic physicians.
[0005] To achieve the above object, in a first aspect, the present application provides a real-time detection method for breast nodules in ultrasonic images, including:
[0006] Step 1, obtain breast ultrasonic image data;
[0007] Step 2, perform preprocessing operations on the breast ultrasonic image data obtained in Step 1 to obtain preprocessed breast ultrasonic image data;
[0008] Step 3: Input the preprocessed breast ultrasound image data obtained in Step 2 into the trained deep convolutional neural network model to capture the positions of breast nodules contained in the breast ultrasound image data;
[0009] The trained deep convolutional neural network model has the function of outputting the positions of breast nodules contained in the input breast ultrasound image data according to the input breast ultrasound image data;
[0010] The deep convolutional neural network model includes a feature extraction network and a detection subnet connected in sequence; the feature extraction network takes the preprocessed breast ultrasound image data as input data and obtains the feature map corresponding to the input breast ultrasound image data; the detection subnet takes the feature map as input data and captures the positions of breast nodules contained in the corresponding breast ultrasound image data according to the input feature map.
[0011] In a possible implementation, the breast ultrasound image data is a breast ultrasound image dataset; the breast ultrasound image dataset includes a continuous multi-frame breast ultrasound images obtained from an ultrasound device.
[0012] In a possible implementation, the continuous multi-frame breast ultrasound images are extracted from a breast ultrasound video.
[0013] In Step 3, capturing the positions of breast nodules contained in the breast ultrasound image data includes: capturing the positions of breast nodules contained in each frame of breast ultrasound image in the breast ultrasound image data.
[0014] In a possible implementation, the preprocessing operation on the obtained dataset in Step 2 includes the following sub-steps:
[0015] Step 2-1: For each frame of image in the obtained dataset, scale the image to a preset size and normalize the scaled image using a linear function to obtain a normalized image;
[0016] In a possible implementation, the preset size is 800×600 pixels. The resulting image is an image matrix of 800×600×3 pixels;
[0017] Step 2-2: Perform a random enhancement operation on each normalized image obtained in Step 2-1 to obtain a randomly enhanced image.
[0018] In this step, by scaling the images to a preset size and performing normalization, it can be ensured that all images have the same size, and the pixel values of the images are scaled to a fixed range to ensure the consistency and standardization of the input images for the subsequent model, facilitating the training and convergence of the model. By performing random augmentation operations on the normalized images, the diversity of the data is increased, and the robustness and generalization ability of the model are improved.
[0019] In a possible implementation, the feature extraction network is based on the ResNet-50 architecture and includes a common convolutional layer, a conv1 layer, a conv2.x layer, a conv3.x layer, a conv4.x layer, and a conv5.x layer connected in sequence;
[0020] The input data of the common convolutional layer is the input data of the feature extraction network;
[0021] The common convolutional layer adjusts the size and number of channels of the input data and then outputs; the conv1 layer, conv2.x layer, conv3.x layer, conv4.x layer, and conv5.x layer are used for feature extraction;
[0022] The conv2.x layer, conv3.x layer, conv4.x layer, and conv5.x layer respectively output feature maps M2, M3, M4, and M5;
[0023] Take the feature maps M2, M3, M4, and M5 as the output data of the feature extraction network, that is, the feature maps corresponding to the input breast ultrasound image data.
[0024] The so-called connection in sequence means that the output data of the previous layer / module is used as the input data of the next layer / module.
[0025] The detection subnet includes a Spatial Pyramid Pooling Fast Module (SPPF), a feature extraction module, a path enhancement module, and a prediction head; the feature map M5 passes through a Spatial Pyramid Pooling Fast Module (SPPF), and then its output result is upsampled to the same size as the feature map M4, and is concatenated with the feature map M4 through channels and input into the feature extraction module to obtain the feature map M6; then the feature map M6 is upsampled, and its feature channels are concatenated with the feature map M3, and the concatenated result is input into the feature extraction module to obtain the feature map M7; similarly, M7 is upsampled, and its feature channels are concatenated with the feature map M2 extracted in the second layer, and the concatenated result is input into the path enhancement module to obtain the feature V2, V2 is upsampled, and its feature channels are concatenated with M7, and the obtained result is input into the path enhancement module to obtain the feature V3, V3 is upsampled, and its feature channels are concatenated with the feature map M6, and the obtained result is input into the path enhancement module to obtain the feature V4, and it is concatenated with the output result of the Spatial Pyramid Pooling Fast Module (SPPF) and then input into the path enhancement module to obtain the feature V5. Finally, the path-enhanced features V2, V3, V4, and V5 are respectively output to the prediction head, and the prediction head respectively extracts the class feature cls, the location feature bbox, and the confidence from V2, V3, V4, and V5, which are respectively denoted as D2, D3, D4, and D5; the location feature bbox is the generated detection box; the class feature cls is used to mark the target type in the detection box, and the target types include breast nodules and non-nodules (such as normal glands); the confidence is used to represent the probability of the target appearing in the detection box; based on the confidence, the non-maximum suppression (NMS) method is used for D2, D3, D4, and D5 to obtain the final detection result.
[0026] In a possible implementation, the deep convolutional neural network model further includes a tracking subnet; the tracking subnet takes the preprocessed breast ultrasound image data and the positions of breast nodules included in each frame of breast ultrasound image in the breast ultrasound image data as input data, tracks the trajectories of each breast nodule in each frame of breast ultrasound image in the breast ultrasound image data, and counts the total number of breast nodules in the breast ultrasound image data.
[0027] Although the static object detection model has reached a quite high level of accuracy, when it is directly applied to frame-by-frame video detection, the detection results of consecutive frames often show inconsistencies or the target movement trajectories jump. To alleviate this phenomenon, in this application, the tracking subnet uses a bidirectional temporal propagation tracking algorithm. The Kalman filter is used to predict the position of the current tracking trajectory in the next frame and generate a prediction box. The intersection over union (IoU) between the prediction box and the actual detection box in the next frame is used as the similarity measurement criterion for the two matches. The Hungarian algorithm is used to find the best-matched detection boxes in the consecutive frame images, and the best-matched detection boxes are associated with data, assigned a unique tracking identification number (ID), and the corresponding tracking trajectory is generated. When there is a high-confidence prediction box in the next frame, it is directly matched with the existing tracking trajectories. If no high-confidence detection box is matched or there is no high-confidence prediction box in the next frame, the Kalman filter is used to predict the position of the current tracking trajectory in the next frame to generate a prediction box, and a detection box with the highest confidence value is selected from the prediction boxes in the next frame to match the tracking trajectory of the detection box that was not matched with a high confidence for the first time. For a high-confidence detection box in the next frame that does not match the existing tracking trajectories, a new tracking trajectory is created for it.
[0028] The high-confidence prediction box refers to a detection box with a confidence greater than the first threshold.
[0029] In a possible implementation, this application aligns the detection results that are successfully matched in several consecutive frames within a segment (sliding window), and then uses segment consistency to retain the high-confidence object detection results. The detection results with a low IoU (false detections and low-quality detections) are fused with the aligned detection results to obtain the finally generated detection results. This can not only ensure the real-time performance of the model but also ensure the consistency of the tracking results and improve the detection accuracy.
[0030] In a possible implementation, each prediction head has two parallel branches with three cascaded convolutional layers, which are respectively used to extract class features cls, location features bbox, and confidence. The classification and localization tasks are completed with three layers of convolution in a decoupled form to avoid feature conflicts caused by different tasks and improve the accuracy and efficiency of object detection.
[0031] In a possible implementation, the input data of the feature extraction module is split into two parts after passing through a convolutional layer. One part is input into n sequentially connected deformable large kernel convolutional modules DN, and the other part is feature concatenated with the results output by the n deformable large kernel convolutional modules DN. The result after feature concatenation passes through a convolutional layer to obtain the output data of the feature extraction module. It should be noted that the deformable large kernel convolutional module includes a convolutional layer, a deformable depth convolution, a deformable dilated convolution, and two convolutional layers connected in sequence. The input data of the deformable large kernel convolutional module DN serves as the input data of the first convolutional layer. The input data of the first convolutional layer is added to the result output by the penultimate convolutional layer and then serves as the input data of the last convolutional layer, and the output data of the deformable large kernel convolutional module DN is obtained by the last convolutional layer.
[0032] In a possible implementation, the input data of the fast spatial pyramid pooling module SPPF is input into three cascaded max pooling layers after passing through a convolutional layer. The output results of these four layers are concatenated by channels and then continue to pass through a convolutional layer to obtain the output data of the fast spatial pyramid pooling module SPPF.
[0033] In a possible implementation, the overall architecture of the path enhancement module is similar to that of the feature extraction module, except that the deformable large kernel convolutional module DN is replaced by the visual state space module VN. Specifically, the input data of the path enhancement module is split into two parts after passing through a convolutional layer. One part is input into n sequentially connected visual state space modules VN, and the other part is feature concatenated with the results output by the visual state space modules VN. The result after feature concatenation passes through a convolutional layer to obtain the output data of the path enhancement module. The visual state space module VN integrates the state space model SSM in the Mamba architecture, which can perform global visual context modeling and position embedding, and can more fully extract feature context information and model it to achieve feature enhancement and filtering processing. The visual state space module VN includes two branches. The first branch includes a linear layer L, a depth convolution, an activation function AF, a state space model SSM, and a convolutional layer connected in sequence. The second branch includes a linear layer L and an activation function AF connected in sequence. The input data of the visual state space module VN serves as the input of the two branches. The output of the second branch is multiplied by the output of the state space model SSM in the first branch and then input into the convolutional layer in the first branch. The input of the convolutional layer in the first branch is added to the input data of the visual state space module VN and then serves as the output data of the visual state space module VN.
[0034] In a possible implementation, training a deep convolutional neural network model includes:
[0035] Step a1: Obtain the breast ultrasound image data annotated by experts;
[0036] In a possible implementation, the breast ultrasound image data is selected by an ultrasound physician based on clinical experience;
[0037] In a possible implementation, the unannotated breast ultrasound image data can be sent to an ultrasound expert, and the dataset annotated by the ultrasound expert is obtained to get the breast ultrasound image data annotated by experts.
[0038] In a possible implementation, the breast ultrasound image data is several breast ultrasound videos obtained from ultrasound devices manufactured by mainstream manufacturers in the market (including Samsung, Siemens, Kaili, etc.), and these images are parsed into continuous image frames.
[0039] Step a2: Perform a preprocessing operation on the breast ultrasound image data obtained in step a1 to obtain the preprocessed breast ultrasound image data;
[0040] Among them, the preprocessing operation method is the same as the preprocessing operation method in step 2.
[0041] Step a3: Use a batch of data in the training set part of the preprocessed breast ultrasound image data obtained in step a2, input it into the deep convolutional neural network model to obtain an inference output, and input the inference output and the breast ultrasound image data annotated by experts in step a1 into the loss function of the deep convolutional neural network model to obtain a loss value.
[0042] Step a4: According to the Adam algorithm and use the loss value obtained in step a3 to optimize the loss function of the deep convolutional neural network model for optimization;
[0043] Through this step, the gradual update of the parameters in the deep convolutional neural network model is realized;
[0044] Step a5: For the remaining batches of data in the preprocessed breast ultrasound image data obtained in step a2, repeat steps a3 and a4 above in sequence until the number of iterations is reached, so as to obtain a trained deep convolutional neural network model.
[0045] In a possible implementation, the training process in this step includes 300 epochs, and the number of iterations per epoch is 100 times.
[0046] In a possible implementation, the loss value used in the deep convolutional neural network model is calculated by the following loss function Calculate:
[0047]
[0048] Among them, is the classification loss, is the detection box regression loss.
[0049] In a possible implementation, training the deep convolutional neural network model includes: Step a6, using the validation set part of the preprocessed breast ultrasound image data obtained in Step a2 to verify the trained deep convolutional neural network model.
[0050] The images for breast ultrasound examination are all input into the trained model, and the model automatically identifies and outputs the category of breast nodules and gives the segmentation result, as Figure 3 shown.
[0051] The present invention can not only monitor breast nodules in the ultrasound video in real time, but also count the number of breast nodules in the ultrasound video, which greatly reduces the workload of ultrasound physicians and lays a foundation for subsequent computer-aided diagnosis.
[0052] Embodiment 2
[0053] In a first aspect, the present application provides a real-time detection system for breast nodules in ultrasound images, including:
[0054] A first module for acquiring breast ultrasound image data;
[0055] A second module for performing preprocessing operations on the breast ultrasound image data acquired by the first module to obtain preprocessed breast ultrasound image data;
[0056] A third module for inputting the preprocessed breast ultrasound image data obtained by the second module into the trained deep convolutional neural network model to capture the positions of breast nodules included in the breast ultrasound image data;
[0057] The trained deep convolutional neural network model has the function of outputting the positions of breast nodules included in the input breast ultrasound image data according to the input breast ultrasound image data;
[0058] The deep convolutional neural network model includes a feature extraction network and a detection subnet connected in sequence; the feature extraction network takes the preprocessed breast ultrasound image data as input data and obtains the feature map corresponding to the input breast ultrasound image data; the detection subnet takes the feature map as input data and captures the positions of breast nodules included in the corresponding breast ultrasound image data according to the input feature map.
[0059] In a possible implementation, the breast ultrasound image data is a breast ultrasound image dataset; the breast ultrasound image dataset includes a series of consecutive breast ultrasound images obtained from an ultrasound device.
[0060] The third module captures the positions of breast nodules included in the breast ultrasound image data, including: capturing the positions of breast nodules included in each frame of the breast ultrasound image data.
[0061] In a possible implementation, the deep convolutional neural network model in the third module further includes a tracking subnet; the tracking subnet takes the preprocessed breast ultrasound image data and the positions of breast nodules included in each frame of the breast ultrasound image data as input data, tracks the trajectories of each breast nodule in each frame of the breast ultrasound image data, and counts the total number of breast nodules in the breast ultrasound image data.
[0062] In a second aspect, the present application provides an electronic device, including: a memory and a processor;
[0063] The memory is used to store a computer program;
[0064] The processor is used to call the computer program to execute the method described above.
[0065] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on an electronic device, the electronic device is enabled to implement the method described above.
[0066] In a fourth aspect, the present application provides a computer program product, including a computer program. When the computer program runs on an electronic device, the electronic device is enabled to implement the method described above.
[0067] The specific implementation manners of the second to fourth aspects of the present application may refer to the implementation manner of the first aspect above, and will not be elaborated here.
[0068] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0069] (1) The present application can make the entire process of a doctor's examination of a patient be programmed and automated, thereby being able to solve the problems in the existing methods for automatic detection and tracking of breast nodules in an ultrasound video, which overly rely on doctors' clinical experience and anatomical structure knowledge, and can also assist ultrasound doctors in screening and diagnosing breast diseases, reducing the workload of ultrasound doctors.
[0070] (2) Since the present invention is fully automated and programmed, during the detection process, the ultrasound doctor does not need to stop to measure, nor does he need to check the accuracy of breast nodules. The detection and tracking are both real-time. Therefore, it can solve the technical problem of the overly long examination time existing in the current ultrasound doctors during disease screening.
[0071] (3) Since the data set used in the learning process of the present invention can adopt the data selected and accurately labeled by the ultrasound doctor according to clinical experience, the knowledge of the ultrasound doctor can be obtained through machine learning in the present invention. It can not only make full use of the professional knowledge and clinical experience of doctors, but also avoid over-reliance on doctors and achieve automated and efficient detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a flowchart of an embodiment of the present application.
[0073] Figure 2 It is an architecture diagram of a deep convolutional neural network model in an embodiment of the present application.
[0074] Figure 3 It is a schematic diagram of a bidirectional time series propagation tracking algorithm in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution of the present application will be further described in detail below in conjunction with the embodiments and drawings of the present application. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0076] Embodiment 1
[0077] This embodiment provides a real-time detection method for breast nodules in ultrasound images, including:
[0078] Step 1, obtain breast ultrasound image data;
[0079] Step 2, perform preprocessing operations on the breast ultrasound image data obtained in Step 1 to obtain preprocessed breast ultrasound image data;
[0080] Step 3, input the preprocessed breast ultrasound image data obtained in Step 2 into a trained deep convolutional neural network model to capture the positions of breast nodules contained in the breast ultrasound image data;
[0081] The trained deep convolutional neural network model has the function of outputting the positions of breast nodules contained in the input breast ultrasound image data according to the input breast ultrasound image data;
[0082] The deep convolutional neural network model includes a feature extraction network and a detection subnet connected in sequence; the feature extraction network takes the preprocessed breast ultrasound image data as input data and obtains a feature map corresponding to the input breast ultrasound image data; the detection subnet takes the feature map as input data and captures the positions of breast nodules contained in the corresponding breast ultrasound image data according to the input feature map.
[0083] Preferably, the breast ultrasound image data is a breast ultrasound image dataset; the breast ultrasound image dataset includes a series of consecutive breast ultrasound images obtained from an ultrasound device.
[0084] Preferably, the series of consecutive breast ultrasound images are extracted from a breast ultrasound video.
[0085] In step 3, capturing the positions of breast nodules contained in the breast ultrasound image data includes: capturing the positions of breast nodules contained in each frame of breast ultrasound image in the breast ultrasound image data.
[0086] Preferably, the preprocessing operation on the obtained dataset in step 2 includes the following sub-steps:
[0087] Step 2-1: For each frame of image in the obtained dataset, scale the image to a preset size and normalize the scaled image using a linear function to obtain a normalized image;
[0088] In some embodiments, the preset size is 800×600 pixels. The resulting image is an image matrix of 800×600×3 pixels;
[0089] Step 2-2: Perform a random augmentation operation on each normalized image in step 2-1 to obtain a randomly augmented image.
[0090] In this step, by scaling the image to the preset size and performing normalization processing, it can be ensured that all images have the same size, and the pixel values of the images are scaled to a fixed range to ensure the consistency and standardization of the subsequent model input images, facilitating the training and convergence of the model. By performing a random augmentation operation on the normalized images, the diversity of the data is increased, and the robustness and generalization ability of the model are improved.
[0091] Preferably, the feature extraction network is based on the ResNet-50 architecture and includes a common convolutional layer, a conv1 layer, a conv2.x layer, a conv3.x layer, a conv4.x layer, and a conv5.x layer connected in sequence;
[0092] The input data of the common convolutional layer is the input data of the feature extraction network;
[0093] In some embodiments, the input data of the ordinary convolutional layer is an image matrix of 800×600×3 pixels;
[0094] The ordinary convolutional layer adjusts the size and number of channels of the input data and then outputs; the conv1 layer, conv2.x layer, conv3.x layer, conv4.x layer, and conv5.x layer are used for feature extraction;
[0095] The conv2.x layer, conv3.x layer, conv4.x layer, and conv5.x layer respectively output feature maps M2, M3, M4, and M5;
[0096] The feature maps M2, M3, M4, and M5 are used as the output data of the feature extraction network, that is, the feature maps corresponding to the input breast ultrasound image data.
[0097] The sequential connection means using the output data of the previous layer / module as the input data of the next layer / module.
[0098] The detection subnet includes a fast spatial pyramid pooling module SPPF, a feature extraction module, a path enhancement module, and a prediction head; the feature map M5 passes through a fast spatial pyramid pooling module SPPF, and then its output result is upsampled to the same size as the feature map M4 and concatenated with the feature map M4 through channels, and then input into the feature extraction module to obtain the feature map M6; then the feature map M6 is upsampled, and its feature channels are concatenated with the feature map M3, and the concatenated result is input into the feature extraction module to obtain the feature map M7; similarly, M7 is upsampled, and its feature channels are concatenated with the feature map M2 extracted in the second layer, and the concatenated result is input into the path enhancement module to obtain the feature V2, V2 is upsampled, and its channels are concatenated with M7, and the obtained result is input into the path enhancement module to obtain the feature V3, V3 is upsampled, and its channels are concatenated with the feature map M6, and the obtained result is input into the path enhancement module to obtain the feature V4, and it is concatenated with the output result of the fast spatial pyramid pooling module SPPF and then input into the path enhancement module to obtain the feature V5. Finally, the path-enhanced features V2, V3, V4, and V5 are respectively output to the prediction head, and the prediction head respectively extracts the class feature cls, the location feature bbox, and the confidence from V2, V3, V4, and V5, which are respectively denoted as D2, D3, D4, and D5; the location feature bbox is the generated detection box; the class feature cls is used to mark the target type in the detection box, and the target types include breast nodules and non-nodules (such as ordinary glands); the confidence is used to represent the probability of the target appearing in the detection box; based on the confidence, the non-maximum suppression (NMS) method is used for D2, D3, D4, and D5 to obtain the final detection result.
[0099] Furthermore, the deep convolutional neural network model further includes a tracking subnet; the tracking subnet takes the preprocessed breast ultrasound image data and the positions of breast nodules contained in each frame of breast ultrasound images in the breast ultrasound image data as input data, tracks the trajectories of each breast nodule in each frame of breast ultrasound images in the breast ultrasound image data, and counts the total number of breast nodules in the breast ultrasound image data.
[0100] Although the static object detection model has reached a quite high level in terms of accuracy, when it is directly applied to video frame-by-frame detection, there often occur phenomena such as inconsistent detection results between consecutive frames or jumping of the target movement trajectory. To alleviate this phenomenon, in this application, the tracking subnet uses the bidirectional temporal propagation tracking algorithm, adopts the Kalman filter to predict the position of the current tracking trajectory in the next frame and generate a prediction box, and uses the intersection over union between the prediction box and the actual detection box in the next frame as the similarity measurement criterion for the two matches; finds the optimally matched detection boxes in the images of consecutive frames through the Hungarian algorithm, associates the optimally matched detection boxes, assigns a unique tracking identification number (ID), and generates corresponding tracking trajectories. When there is a high-confidence prediction box in the next frame, it is directly matched with the existing tracking trajectories; if no high-confidence detection box is matched or there is no high-confidence prediction box in the next frame, the Kalman filter is used to predict the position of the current tracking trajectory in the next frame to generate a prediction box, and a detection box with the highest confidence value is selected from the prediction boxes in the next frame to be matched with the tracking trajectory of the detection box that was not matched with a high-confidence detection box for the first time; for a high-confidence detection box in the next frame that is not matched with the existing tracking trajectories, a new tracking trajectory is created for it;
[0101] The high-confidence prediction box refers to a detection box with a confidence greater than the first threshold.
[0102] Furthermore, in this application, within a segment (sliding window), the detection results that are successfully matched between consecutive frames are aligned, and then the fragment consistency is used to retain the high-confidence object detection results, and the detection results with a lower intersection over union (false detections and low-quality detections) are fused with the aligned detection results to obtain the finally generated detection results, thereby not only ensuring the real-time performance of the model but also ensuring the consistency of the tracking results and improving the detection accuracy. As Figure 3 shown, the bidirectional temporal propagation tracking algorithm can remove the misdetected target box C of the detection subnet and can also fill in the missed target box B in the third frame.
[0103] Preferably, the prediction head is as Figure 2As shown, each prediction head has two parallel branches with three cascaded convolutional layers, which are used to extract class features cls, location features bbox, and confidence respectively. The classification and localization tasks are completed with three convolutional layers in a decoupled form, so as to avoid feature conflicts caused by different tasks and improve the accuracy and efficiency of object detection.
[0104] Preferably, the feature extraction module is as Figure 2 As shown, the input data of the feature extraction module is split into two parts after passing through a convolutional layer. One part is input into n sequentially connected deformable large kernel convolutional modules DN, and the other part is feature concatenated with the results output by the n deformable large kernel convolutional modules DN. The result after feature concatenation passes through a convolutional layer to obtain the output data of the feature extraction module. It should be noted that the deformable large kernel convolutional module includes a convolutional layer, a deformable depth convolution, a deformable dilated convolution, and two convolutional layers connected in sequence; the input data of the deformable large kernel convolutional module DN is used as the input data of the first convolutional layer, and the input data of the first convolutional layer is added to the result output by the penultimate convolutional layer and used as the input data of the last convolutional layer, and the output data of the deformable large kernel convolutional module DN is obtained by the last convolutional layer.
[0105] Preferably, the fast spatial pyramid pooling module SPPF is as Figure 2 As shown, the input data of the fast spatial pyramid pooling module SPPF is input into three cascaded max pooling layers after passing through a convolutional layer. The output results of these four layers are concatenated through channels and then continue to pass through a convolutional layer to obtain the output data of the fast spatial pyramid pooling module SPPF.
[0106] Preferably, the path enhancement module is as Figure 2As shown, the overall architecture is similar to the feature extraction module, except that the deformable large kernel convolution module DN is replaced by the visual state space module VN. Specifically, the input data of the path enhancement module is split into two parts after passing through a convolutional layer. One part is input into n sequentially connected visual state space modules VN, and the other part is feature concatenated with the result output by the visual state space module VN. The result after feature concatenation passes through a convolutional layer to obtain the output data of the path enhancement module. Among them, the visual state space module VN integrates the state space model SSM in the Mamba architecture. This model can perform global visual context modeling and position embedding, and can more fully extract feature context information and model it to achieve feature enhancement and filtering processing. The visual state space module VN includes two branches. The first branch includes a linear layer L, a depth convolution, an activation function AF, a state space model SSM, and a convolutional layer connected in sequence; the second branch includes a linear layer L and an activation function AF connected in sequence; the input data of the visual state space module VN serves as the input of the two branches. The output of the second branch is multiplied by the output of the state space model SSM in the first branch and then input into the convolutional layer in the first branch. The input of the convolutional layer in the first branch is added to the input data of this visual state space module VN to serve as the output data of the visual state space module VN.
[0107] Preferably, training the deep convolutional neural network model includes:
[0108] Step a1, obtaining breast ultrasound image data annotated by experts;
[0109] In some embodiments, the breast ultrasound image data is selected by an ultrasound physician according to clinical experience;
[0110] In some embodiments, unannotated breast ultrasound image data can be sent to an ultrasound expert, and a dataset annotated by the ultrasound expert can be obtained to obtain breast ultrasound image data annotated by experts.
[0111] In some embodiments, the breast ultrasound image data is several (such as 200) breast ultrasound videos obtained from ultrasound devices manufactured by mainstream manufacturers in the market (including Samsung, Siemens, Kaili, etc.), and these images are parsed into continuous image frames.
[0112] In some embodiments, the dataset can be randomly divided into three parts: a training set (Train set), a validation set (Validation set), and a test set (Test set). Exemplarily, 80% of the dataset is used as the training set, 10% as the validation set, and 10% as the test set.
[0113] Step a2: Preprocess the breast ultrasound image data obtained in step a1 to obtain the preprocessed breast ultrasound image data;
[0114] The preprocessing operation method is the same as the preprocessing operation method in step 2.
[0115] Step a3: Use a batch of data in the training set part of the preprocessed breast ultrasound image data obtained in step a2, input it into the deep convolutional neural network model to obtain an inference output, and input the inference output and the breast ultrasound image data annotated by experts in step a1 into the loss function of the deep convolutional neural network model to obtain a loss value.
[0116] In some embodiments, a batch of data is 16 images.
[0117] Step a4: According to the Adam algorithm and using the loss value obtained in step a3, optimize the loss function of the deep convolutional neural network model ;
[0118] Through this step, the gradual update of the parameters in the deep convolutional neural network model is realized;
[0119] During the optimization process, set the learning rate lr = 0.001, the momentum ξ = 0.9, and the weight decay ψ = 0.004.
[0120] Step a5: For the remaining batches of data in the preprocessed breast ultrasound image data obtained in step a2, repeat steps a3 and a4 above in sequence until the number of iterations is reached, so as to obtain a trained deep convolutional neural network model.
[0121] In some embodiments, the training process in this step includes 300 epochs, and the number of iterations per epoch is 100 times.
[0122] Preferably, the loss value used in the deep convolutional neural network model is calculated by the following loss function :
[0123]
[0124] where is the classification loss, is the detection box regression loss.
[0125] In some embodiments, training the deep convolutional neural network model includes: Step a6: Use the validation set part of the preprocessed breast ultrasound image data obtained in step a2 to validate the trained deep convolutional neural network model.
[0126] The images for breast ultrasound examinations are all input into the trained model, and the model automatically identifies and outputs the categories of breast nodules and gives the segmentation results, such as Figure 3 as shown.
[0127] The present invention can not only monitor breast nodules in the ultrasound video in real time, but also count the number of breast nodules in the ultrasound video, which greatly reduces the workload of ultrasound physicians and lays a foundation for subsequent computer-aided diagnosis.
[0128] Embodiment 2
[0129] In a first aspect, the present application provides a real-time detection system for breast nodules in ultrasound images, including:
[0130] A first module for acquiring breast ultrasound image data;
[0131] A second module for performing preprocessing operations on the breast ultrasound image data acquired by the first module to obtain preprocessed breast ultrasound image data;
[0132] A third module for inputting the preprocessed breast ultrasound image data obtained by the second module into the trained deep convolutional neural network model to capture the positions of breast nodules contained in the breast ultrasound image data;
[0133] The trained deep convolutional neural network model has the function of outputting the positions of breast nodules contained in the input breast ultrasound image data according to the input breast ultrasound image data;
[0134] The deep convolutional neural network model includes a feature extraction network and a detection subnet connected in sequence; the feature extraction network takes the preprocessed breast ultrasound image data as input data and obtains the feature map corresponding to the input breast ultrasound image data; the detection subnet takes the feature map as input data and captures the positions of breast nodules contained in the corresponding breast ultrasound image data according to the input feature map.
[0135] Preferably, the breast ultrasound image data is a breast ultrasound image dataset; the breast ultrasound image dataset includes a continuous multi-frame breast ultrasound image acquired from an ultrasound device.
[0136] The third module captures the positions of breast nodules contained in the breast ultrasound image data, including: capturing the positions of breast nodules contained in each frame of the breast ultrasound image in the breast ultrasound image data.
[0137] Further, the deep convolutional neural network model in the third module further includes a tracking subnet; the tracking subnet takes the preprocessed breast ultrasound image data and the positions of breast nodules contained in each frame of breast ultrasound images in the breast ultrasound image data as input data, tracks the trajectories of each breast nodule in each frame of breast ultrasound images in the breast ultrasound image data, and counts the total number of breast nodules in the breast ultrasound image data.
[0138] Embodiment III
[0139] This embodiment provides an electronic device, including: a memory and a processor;
[0140] The memory is used to store a computer program;
[0141] The processor is used to call the computer program to execute the method as described in Embodiment I.
[0142] Embodiment IV
[0143] This embodiment provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on an electronic device, the electronic device is enabled to implement the method as described in Embodiment I.
[0144] Embodiment V
[0145] This embodiment provides a computer program product, including a computer program. When the computer program runs on an electronic device, the electronic device is enabled to implement the method as described in Embodiment I.
[0146] The specific implementation manners of a system, an electronic device, a computer-readable storage medium, and a computer program product provided in the embodiments of the present application may refer to the specific embodiments of the above method, and will not be elaborated herein.
[0147] Obviously, those skilled in the art should understand that the above units or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0148] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A real-time detection method for breast nodules in ultrasound images, characterized in that: include: Step 1, obtaining breast ultrasound image data; Step 2: preprocessing the breast ultrasound image data obtained in step 1 to obtain preprocessed breast ultrasound image data; Step 3: input the preprocessed breast ultrasound image data obtained in step 2 into the trained deep convolutional neural network model to capture the location of the breast nodules contained in the breast ultrasound image data; The trained deep convolutional neural network model has the function of outputting the positions of breast nodules contained in the breast ultrasound image data according to the input breast ultrasound image data; The deep convolutional neural network model includes a feature extraction network and a detection subnet connected in sequence; the feature extraction network takes the preprocessed breast ultrasound image data as input data to obtain a feature map corresponding to the input breast ultrasound image data; the detection subnet takes the feature map as input data and captures the location of breast nodules contained in the corresponding breast ultrasound image data based on the input feature map.
2. The method according to claim 1, characterized in that The breast ultrasound image data is a breast ultrasound image data set; the breast ultrasound image data set includes continuous multi-frame breast ultrasound images acquired from an ultrasound device; In the step 3, capturing the position of the breast nodules contained in the breast ultrasound image data includes: capturing the position of the breast nodules contained in each frame of the breast ultrasound image data.
3. The method according to any one of claims 1 to 3, characterized in that The feature extraction network is based on the ResNet-50 architecture, including a common convolutional layer, a conv1 layer, a conv2.x layer, a conv3.x layer, a conv4.x layer, and a conv5.x layer connected in sequence; The input data of the common convolutional layer is the input data of the feature extraction network; The common convolution layer adjusts the size and number of channels of the input data and then outputs it; the conv1 layer, conv2.x layer, conv3.x layer, conv4.x layer and conv5.x layer are used for feature extraction; The conv2.x layer, conv3.x layer, conv4.x layer and conv5.x layer output feature maps M2, M3, M4 and M5 respectively; The feature maps M2, M3, M4 and M5 are used as output data of the feature extraction network, that is, the feature maps corresponding to the input breast ultrasound image data.
4. The method according to claim 3, characterized in that The detection subnet includes a fast spatial pyramid pooling module SPPF and a feature extraction module, a path enhancement module and a prediction head; The feature map M5 passes through a fast spatial pyramid pooling module SPPF, and then its output result is upsampled to the same size as the feature map M4, and is spliced with the feature map M4 through a channel, and input into the feature extraction module to obtain the feature map M6; The feature map M6 is upsampled, and the feature channel is spliced with the feature map M3, and the spliced result is input into the feature extraction module to obtain the feature map M7; M7 is upsampled and concatenated with the feature map M2 extracted from the second layer through channels, and input into the path enhancement module to obtain feature V2; V2 is upsampled and channel-concatenated with M7. The result is input into the path enhancement module to obtain feature V3. V3 is upsampled and channel-joined with the feature map M6. The result is input into the path enhancement module to obtain feature V4, which is then joined with the result output by the fast spatial pyramid pooling module SPPF and input into the path enhancement module to obtain feature V5. Finally, features V2, V3, V4, and V5 are respectively output into the prediction head, which extracts category feature cls, position feature bbox, and confidence from V2, V3, V4, and V5, respectively, and are denoted as D2, D3, D4, and D5. The position feature bbox is the generated detection frame. The category feature cls is used to mark the target type in the detection frame, including breast nodules and non-nodules. The confidence is used to indicate the probability of the target appearing in the detection frame. Based on the confidence, the non-maximum suppression method is used to obtain the final detection result for D2, D3, D4, and D5.
5. The method according to claim 4, characterized in that The deep convolutional neural network model also includes a tracking subnet; the tracking subnet takes the preprocessed breast ultrasound image data and the position of the breast nodules contained in each frame of the breast ultrasound image data as input data, tracks the trajectory of each breast nodule in each frame of the breast ultrasound image data, and counts the total number of breast nodules in the breast ultrasound image data.
6. The method according to claim 5, characterized in that The tracking subnet uses a bidirectional time propagation tracking algorithm and a Kalman filter to predict the position of the current tracking trajectory in the next frame and generate a prediction frame. The intersection-and-union ratio between the prediction frame and the actual detection frame of the next frame is used as the similarity measure for the two matches. The Hungarian algorithm is used to find the best matching detection frame in the previous and next frame images, and the best matching detection frame is associated with the data, and a unique tracking identification number (ID) is assigned to generate the corresponding tracking trajectory. When there is a high-confidence prediction frame in the next frame, it is directly matched with the existing tracking trajectory; if a high-confidence detection frame is not matched or there is no high-confidence prediction frame in the next frame, the Kalman filter is used to predict the position of the current tracking trajectory in the next frame to generate a prediction frame, and a detection frame with the highest confidence value is selected from the prediction frame of the next frame to match the tracking trajectory of the detection frame that did not match the high-confidence detection frame the first time; For the detection box in the next frame that does not match the existing tracking track but has high confidence, a new tracking track is created for it; The high-confidence prediction box refers to a detection box whose confidence is greater than a first threshold.
7. A real-time detection system for breast nodules in ultrasound images, characterized in that: include: The first module is used to obtain breast ultrasound image data; The second module is used to perform a preprocessing operation on the breast ultrasound image data acquired by the first module to obtain preprocessed breast ultrasound image data; The third module is used to input the preprocessed breast ultrasound image data obtained in the second module into the trained deep convolutional neural network model to capture the location of breast nodules contained in the breast ultrasound image data; The trained deep convolutional neural network model has the function of outputting the positions of breast nodules contained in the breast ultrasound image data according to the input breast ultrasound image data; The deep convolutional neural network model includes a feature extraction network and a detection subnet connected in sequence; the feature extraction network takes the preprocessed breast ultrasound image data as input data to obtain a feature map corresponding to the input breast ultrasound image data; the detection subnet takes the feature map as input data and captures the location of the breast nodules contained in the corresponding breast ultrasound image data according to the input feature map; The system realizes real-time detection of breast nodules in ultrasound images based on the method described in any one of claims 1 to 6.
8. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor is configured to call the computer program to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on an electronic device, the electronic device implements the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed on an electronic device, the electronic device implements the method according to any one of claims 1 to 7.