A method for detecting abnormal signal intensity in medical images
By using deep learning technology to extract and fuse multi-scale features from diffusion-weighted digital medical images, the accuracy and efficiency issues of abnormal signal intensity detection in existing medical image evaluation methods are solved. This enables automated and accurate abnormal signal intensity detection, reducing misdiagnosis and missed diagnosis, and lowering the consumption of medical resources.
Patent Information
- Application Number
- CN202210538446.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-05-18
AI Technical Summary
Existing medical image evaluation methods are subject to subjective factors, making it difficult to accurately and promptly detect abnormal signal intensity in diffusion-weighted digital medical images, leading to misdiagnosis or missed diagnosis, which is particularly prominent in primary hospitals.
Using deep learning technology, convolutional neural networks are used to extract and fuse multi-scale features from diffusion-weighted digital medical images. Combined with inversion operation and spatial attention extraction, regions with abnormal signal intensity are automatically detected.
It improves the accuracy and efficiency of abnormal signal intensity detection, reduces misdiagnosis and missed diagnosis, lowers the workload of medical staff, and saves medical resources.
Smart Images

Figure CN115115576B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of medical engineering, and in particular relates to a new method for detecting abnormal signal intensity of medical images. BACKGROUND
[0002] At present, many diseases have seriously affected human health, and early detection of lesions and evaluation are beneficial to diagnosis and treatment. For example, stroke is an acute cerebrovascular disease that seriously threatens people's life and health. Every year, about 795,000 people experience new or recurrent stroke. Among them, about 610,000 are first-time strokes, and 185,000 are recurrent episodes. Stroke is divided into two types: cerebral vascular obstruction caused by ischemic stroke and cerebral vascular hemorrhage caused by hemorrhagic stroke (ischemic stroke accounts for 87% of all strokes). Stroke often causes damage to brain tissue, resulting in loss of local function, brain tissue necrosis, and even death. Since the clinical outcome is directly related to the duration of treatment, accurate and timely detection and diagnosis is extremely important. With the continuous development of magnetic resonance imaging (MRI) in clinical diagnosis, diffusion-weighted imaging technology has recently become a standard and routinely used clinical tool. Diffusion-weighted medical images can be well used to examine patients with diseases such as stroke, especially patients with suspected acute ischemic stroke symptoms, due to short acquisition time and high sensitivity. These problems and difficulties include: (1) Different types of lesions have different sizes, shapes, and locations in diffusion-weighted digital medical images. In addition, the area of abnormal signal intensity of the lesion is extremely small, the boundary is blurred, and it is easy to miss diagnosis, even for experienced radiologists, it is difficult to find, often leading to misdiagnosis or missed diagnosis, which is an unavoidable objective factor for radiologists, especially during the fatigue period, which is more obvious and common. (2) Clinicians and radiologists often face the dilemma of a large amount of image interpretation pressure and limited diagnostic time window, and it is difficult to cope with a large amount of data and meet clinical requirements in a short period of time, and sometimes there is a lag and deviation in image data interpretation and evaluation, which may delay the opportunity for treatment, which may result in high morbidity and mortality, and will inevitably increase the medical burden and social burden. (3) Medical image interpretation is often affected by the subjective factors of doctors, and their diagnostic experience and business level often determine their judgment results, which is an unavoidable subjective factor for radiology departments at all levels of hospitals, especially in remote areas and primary hospitals, this problem is particularly prominent and obvious, which greatly limits the diagnosis and treatment level of primary hospitals. SUMMARY
[0003] The application provides a medical image abnormal signal intensity detection method, which aims to solve the subjective and objective factors of the existing medical image evaluation method, avoid occupying a large amount of time and energy of radiologists, overcome the subjective factors of the image interpretation evaluation method, avoid the objective factors caused by the lesion factors, achieve an efficient, objective and intelligent inspection method, avoid incorrect judgment, improve the image interpretation work efficiency, relieve the labor intensity of radiologists and clinicians, and reduce errors.
[0004] Technical scheme
[0005] A medical image abnormal signal intensity detection method, characterized in that it comprises the following steps.
[0006] Obtain a digital medical image, and perform a cross-sectional operation on the digital medical image to obtain a plurality of cross-sectional feature maps;
[0007] Perform a multi-scale feature extraction operation on the plurality of cross-sectional feature maps to obtain a multi-scale feature map; and fuse the multi-scale feature map and a global feature map to be updated to obtain an updated global feature map;
[0008] After each multi-scale feature extraction operation, perform an inversion operation on the updated global feature map to obtain background information corresponding to a multi-scale abnormal signal intensity feature; and fuse the background information corresponding to the multi-scale abnormal signal intensity feature and the multi-scale feature map to obtain a multi-scale boundary feature map;
[0009] According to the multi-scale feature map and the boundary feature map, the abnormal signal intensity region of the digital medical image is detected.
[0010] Optionally, the method for performing a multi-scale feature extraction operation on the plurality of cross-sectional feature maps to obtain a multi-scale feature map comprises: performing a convolution operation on the plurality of cross-sectional feature maps to obtain a multi-scale feature map.
[0011] And / or, the method for fusing the multi-scale feature map and the global feature map to be updated to obtain an updated global feature map comprises the following steps.
[0012] Perform a convolution operation on the multi-scale feature map to obtain a feature map to be fused;
[0013] Scale the feature map to be fused or the global feature map to be updated to the same size;
[0014] Add the scaled feature map to be fused or the global feature map to be updated to obtain an updated global feature map.
[0015] Optionally, before the background information corresponding to the multi-scale abnormal signal strength feature is obtained by performing the reverse operation on the updated global feature map, the method further comprises: performing a spatial attention extraction operation on the updated global feature map to obtain a spatial attention feature; and performing a reverse operation on the spatial attention feature to obtain the background information corresponding to the multi-scale abnormal signal strength feature.
[0016] Optionally, the method for fusing the background information corresponding to the multi-scale abnormal signal strength feature and the multi-scale feature map to obtain the multi-scale boundary feature map comprises: multiplying the background information corresponding to the multi-scale abnormal signal strength feature and the multi-scale feature map to obtain the multi-scale boundary feature map.
[0017] Optionally, the method for detecting the abnormal signal strength region of the digital medical image based on the multi-scale feature map and the boundary feature map comprises:
[0018] fusing the multi-scale feature map and the boundary feature map to obtain a feature bounding box with the highest confidence, and using the feature bounding box to detect the abnormal signal strength region of the digital medical image.
[0019] Optionally, the method for fusing the multi-scale feature map and the boundary feature map to obtain a feature bounding box with the highest confidence comprises:
[0020] splicing the multi-scale feature map and the boundary feature map to obtain a plurality of output feature maps consistent with the number of the multi-scale feature maps;
[0021] determining the maximum size of the plurality of output feature maps;
[0022] up-sampling the output feature maps of other sizes to the maximum size, and then splicing the plurality of output feature maps to obtain a feature bounding box with the highest confidence.
[0023] Optionally, before the updated global feature map is obtained by fusing the multi-scale feature map and the global feature map to be updated, the multi-scale feature map to be fused is determined, and the method comprises:
[0024] obtaining a set size;
[0025] if the size of the multi-scale feature map is less than or equal to the set size, the multi-scale feature map is determined to be fused.
[0026] Advantages and effects: With the development of deep learning, convolutional neural networks (CNNs) have gradually been applied to a wide range of medical digital medical image tasks due to their powerful ability to automatically extract high-discriminative features determined by deep neural networks. These tasks include identifying musculoskeletal joint injuries, classifying common tumors and skin cancers, detecting lung nodules and retinal lesions, and segmenting and evaluating brain tumors and stroke diseases. Due to the powerful ability of convolutional neural networks to learn invariant features, robust and discriminative features that can help automatic detection systems accurately locate lesion areas can be learned from medical digital medical images. Although deep learning techniques have been applied to detect various diseases, few works have attempted to use such techniques to address the detection of related diseases such as stroke in diffusion-weighted digital medical images. Although the current method can detect most disease abnormal signal intensity regions such as stroke, some lesions such as stroke abnormal signal intensity are difficult to detect or missed due to small area, unclear texture, and fuzzy boundary, which poses a major challenge to accurately detecting and identifying missed lesions such as stroke abnormal signal intensity. This method is beneficial for concise, accurate, and rapid examination of abnormal changes in lesions, improves the level of finding abnormal changes in lesions, reduces missed and false judgments, reduces medical errors and mistakes, reduces the intensity of medical workers, promotes the improvement of human health level, reduces the corresponding medical costs and saves medical resources, greatly reduces the heavy workload and intensity of medical workers, has far-reaching economic and social benefits, and has incomparable advantages in terms of current medical resource allocation and social professional human resource use and optimization.
[0027] In the embodiments of the present disclosure, real-time and accurate abnormal signal intensity detection can be performed, multi-scale feature extraction operations are performed on the multiple cross-sectional feature maps to obtain multi-scale feature maps, the multi-scale feature maps and the global feature map to be updated are fused to obtain an updated global feature map, so as to continuously update high-level features, and at the same time, after each multi-scale feature extraction operation, an inverse operation is performed on the updated global feature map to obtain background information corresponding to the multi-scale abnormal signal intensity features; the background information corresponding to the multi-scale abnormal signal intensity features and the multi-scale feature maps are fused to obtain multi-scale boundary feature maps, so as to timely track boundary relationships. Through the detection method provided by the present disclosure, after re-diagnosis according to the automatic detection result, the misdiagnosis or missed diagnosis of radiologists is effectively reduced, so as to solve the problem that abnormal signal intensity cannot be detected due to small area, unclear texture, and fuzzy boundary, and to reduce the misdiagnosis or missed diagnosis.
[0028] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure.
[0029] Other features and aspects of the present disclosure will become apparent from a detailed description of exemplary embodiments with reference to the following drawings.
[0030] BRIEF DESCRIPTION OF DRAWINGS (EXPLAINING BY STROKE)
[0031] The drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description, serve to explain the principles of the present disclosure.
[0032] Figure 1 Flow chart of the abnormal signal strength detection method of the present application;
[0033] Figure 2 Network structure diagram of the present application;
[0034] Figure 3 Stroke disease data set diagram of the present application;
[0035] Figure 4 Digital medical image instance generated by data augmentation of the present application;
[0036] Figure 5 Structure diagram of the present application for updating the global feature map aggregation pool (AP) module;
[0037] Figure 6 Structure diagram of the reverse attention (RA) module for implicitly learning edge feature maps;
[0038] Figure 7 Distribution of the position, width and height of the lesions in the stroke training set;
[0039] Figure 8 Result graph of P, R, F1 and PR curves of the independent stroke test set using the present application;
[0040] Figure 9 Result graph of the FROC curve of the independent stroke test set using the present application;
[0041] Figure 10 Example graph of stroke lesions detected by different models;
[0042] Figure 11 Typical example graph of lesion detection results of the TE-YOLOv5-S model of the present application;
[0043] Figure 12 Comparison graph between manually labeled targets and automatically detected targets;
[0044] Figure 13 Block diagram of the abnormal signal strength detection unit of the present application;
[0045] Figure 14 a block diagram of an electronic device 800 of the present application;
[0046] Figure 15 a block diagram of an electronic device 1900 of the present application. DETAILED DESCRIPTION
[0047] The present application is a new detection method based on medical image data processing and improving the identification of difficult lesions. In view of the objective reasons such as small area of lesions, unclear texture and blurred boundary, this method can completely change the difficulty of detecting lesions, and reduce the high rate of misdiagnosis or missed diagnosis. Its application can be extended to related medical image data lesion identification, which can cover image data such as CT images, magnetic resonance images and X-ray images, and is convenient for popularization and application.
[0048] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings represent functionally the same or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.
[0049] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0050] The term "and / or", merely used to describe the associated relationship of the associated objects, means that there can be three relationships, for example, A and / or B, which can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. In addition, the term "at least one" herein means any one of the plurality or any combination of at least two of the plurality, for example, including at least one of A, B and C, which can mean including any one or more elements selected from the set consisting of A, B and C.
[0051] In addition, in order to better illustrate the present disclosure, a large number of specific details are given in the specific embodiments below. Those skilled in the art should understand that without certain specific details, the present disclosure can also be implemented. In some examples, methods, means, elements and circuits well known to those skilled in the art are not described in detail, in order to highlight the main idea of the present disclosure.
[0052] It can be understood that the above-mentioned various method embodiments of the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to the limited space, the present disclosure will not be described again.
[0053] Figure 1 a flowchart of the abnormal signal strength detection method of the present application, as shown in Figure 1As shown, the abnormal signal intensity detection method comprises the following steps: S101, acquiring a digital medical image and performing cross-section operation on the digital medical image to obtain a plurality of cross-section feature maps; S102, performing multi-scale feature extraction operation on the plurality of cross-section feature maps to obtain a multi-scale feature map, and fusing the multi-scale feature map and a global feature map to be updated to obtain an updated global feature map; S103, after each multi-scale feature extraction operation, performing reverse operation on the updated global feature map to obtain background information corresponding to the multi-scale abnormal signal intensity feature; and fusing the background information corresponding to the multi-scale abnormal signal intensity feature and the multi-scale feature map to obtain a multi-scale boundary feature map; and S104, detecting the abnormal signal intensity region of the digital medical image according to the multi-scale feature map and the boundary feature map. The abnormal signal intensity can be detected in real time and accurately. The multi-scale feature map is obtained by performing multi-scale feature extraction operation on the plurality of cross-section feature maps; the updated global feature map is obtained by fusing the multi-scale feature map and the global feature map to be updated, so as to continuously update the high-level features; meanwhile, the background information corresponding to the multi-scale abnormal signal intensity feature is obtained by performing reverse operation on the updated global feature map after each multi-scale feature extraction operation; and the multi-scale boundary feature map is obtained by fusing the background information corresponding to the multi-scale abnormal signal intensity feature and the multi-scale feature map, so as to timely track the boundary relationship. Through the detection method provided by the present disclosure, after re-diagnosis according to the automatic detection result, the misdiagnosis or missed diagnosis of the radiologists is effectively reduced, so as to solve the problem that the abnormal signal intensity cannot be detected due to small area, unclear texture and fuzzy boundary, etc., and the problem of misdiagnosis or missed diagnosis.
[0054] Figure 2 The network structure diagram of the abnormal signal intensity detection method of the present application is shown. The specific embodiments of the present disclosure are described in conjunction with Figure 2 .
[0055] In the embodiments and other possible embodiments of the present disclosure, the YOLOv5 target detection network (YOLOv5-S, YOLOv5-L) is applied as a basic network to automatically detect the abnormal signal intensity of diseases such as stroke on the collected data. These can simultaneously detect the positions of objects in the input digital medical image and classify them into different categories. Unlike traditional R-CNN and Fast-RCNN, YOLOv5 uses a single convolutional network for detection and classification tasks, providing the coordinates of the bounding box and its class prediction. Therefore, YOLO can encode global context information by looking at the entire input digital medical image once.
[0056] The input digital medical image is divided into a regular grid with intervals slightly smaller than the minimum object to be detected. Each grid square has an associated label (detection class) and pixel coordinates of the bounding box. The grid square where the object center is responsible for detecting the object. When there are multiple objects in the same grid square, the network selects the object with the most pixels in the grid to cover. Features obtained from the entire digital medical image are used to predict each bounding box so that the network can learn the objects in the entire digital medical image.
[0057] Step S101: Obtain a digital medical image and operate on the cross section of the digital medical image to obtain a plurality of cross section feature maps.
[0058] In the embodiments and other possible embodiments of the present disclosure, the digital medical image can be a diffusion weighted (DWI) digital medical image. In order to deeply study the automatic detection process of images of diseases such as stroke by the research team of the present application (Northern Theater General Hospital), 500 cases were retrospectively collected, and each case had 40 layers of original DICOM digital medical images. For ethical reasons, the patient's identity has been omitted, and the hospital ethics committee has approved it. The present application excludes some cases with poor quality, including tumors, metal artifacts (such as metal scalp clips, ventricular external drainage or aneurysm clips after skull resection) or any other significant lesions. Finally, 319 cases (12,760 digital medical images) were reserved for research. The medical ethics committee approved the research of Northern Theater General Hospital. The Declaration of Helsinki (1964) and its later amendments or similar ethical standards were followed. Since this is a retrospective study, informed consent was waived. All diffusion weighted digital medical images were collected using GE Discovery MR750 3.0T with a 16-channel phased array head coil.
[0059] During the scanning process, the subject was asked to lie on his back, close his eyes, keep his head still, remain awake, and first put his head into the scanner. The following settings were used to obtain the diffusion-weighted digital medical images: matrix size of 256x256, cross-sectional thickness of 6mm, cross-sectional interval of 7.5mm, repetition time (TR) of 4097ms, echo time (TE) of 84ms, flip angle (FA) of 90 degrees, field of view (FOV) of 240x240mm, and b value of 0 and 1000s / mm2. Two radiologists independently reviewed each diffusion-weighted digital medical image using the LabelImg tool, accurately labeled the boundary box coordinates of the abnormal signal intensity types of diseases such as stroke in the diffusion-weighted digital medical images, and saved them in YOLO format. After giving their own labels for the controversial abnormal signal intensity, they reached a consensus after discussion. The annotations were stored in a text file with the same name as the diffusion-weighted digital medical image, which included the class of the boundary box, the center coordinates of the lesion area, and the relative width and height of the boundary box. Finally, the present application obtained 1,681 diffusion-weighted digital medical images (including images of diseases such as ischemic stroke and hemorrhagic stroke), each of which may have one or more labeled abnormal signal intensity regions. The total number of labeled abnormal signal intensity regions is 2,180. There are irregular patterns, speckle noise, and fuzzy boundaries in the dataset, which are very suitable for verifying the effectiveness of the method of the present application in handling the above-mentioned stroke disease detection task.
[0060] Figure 3 As shown in the schematic diagram of the stroke dataset of the present application, as shown in Figure 3 two randomly selected labeled diffusion-weighted digital medical images of stroke and other diseases are shown; wherein (a) is a single-target diffusion-weighted digital medical image, and (b) is a multi-target diffusion-weighted digital medical image.
[0061] Like the COCO detection dataset, the 1,681 diffusion-weighted digital medical images were randomly divided into a training set and a test set in a ratio of about 8:2. This resulted in 1,338 digital medical images (2,180 labeled abnormal signal intensities) and 343 digital medical images (532 labeled abnormal signal intensities), respectively. In other words, 252 patient digital medical images were used to train the network of the present application, and 67 patient digital medical images were used for testing and detection. In addition, the present application splits all the layers of the patient diffusion-weighted digital medical images from the original DICOM digital medical images for testing in subsequent patient-level detection. In addition, for YOLO stage-based input, all digital medical images are adjusted to 640x640 pixels through adaptive scaling to reduce black borders.
[0062] In the medical field, it is still difficult to obtain training digital medical images for diseases such as stroke, because they rely on manual annotation, which costs doctors a lot of time and effort. However, in order to avoid overfitting, we need to use data augmentation to obtain more data while training the model. This method effectively expands the data set and improves the robustness of the model. In this work, in order to further improve the performance of TE-YOLOv5 (the tracking edge YOLOv5 proposed in the present disclosure, denoted as TE-YOLOv5), the present application uses a series of data augmentation techniques. Image enhancement is to emphasize the overall or local characteristics of the image, make the originally unclear image clear or emphasize certain features of interest, expand the differences between different object features in the image, suppress features that are not of interest, improve image quality, rich information, and strengthen image interpretation and recognition effect, to meet the needs of some special analysis. Due to the problem of high sample annotation cost of medical images in the present application, resulting in less data, it is necessary to perform data augmentation on the data set to expand the data set. Specifically, the present application uses tone, saturation, numerical value, horizontal flip, mosaic and other data enhancement methods to expand the data set, and then obtains more useful data. Horizontal flip is to flip around the X-axis, which expands the data of different positions of the left and right brain. At the same time, due to the diversity of detection equipment, some equipment has problems such as insufficient saturation and unclear lesions in the image during image acquisition, which is particularly prone to occur during detection. These problems can be improved through post-processing adjustment of tone, saturation, numerical value and other algorithms. Tone enhancement combined with saturation adjustment, numerical value enhancement and other technologies can greatly improve the image. The Mosaic method is based on the CutMix data enhancement method proposed by Sangdoo Yun et al. in 2019. The CutMix method only processes the merging of two pictures, while the Mosaic method reads four pictures, then performs flip, scaling and other operations, and then splices them into one picture. This method can enrich the picture background and greatly expand the training data set. Moreover, the random scaling operation increases a lot of small targets, making the trained model more robust. The Mosaic enhancement method is also a new data enhancement technology introduced in YOLOv5, which combines training digital medical images according to a certain proportion to detect small objects, and promotes the adaptation to the small size changes of objects in digital medical images.
[0063] Figure 4 The digital medical image instances generated by the data augmentation of the present application are shown in Figure 4 As shown, the diffusion-weighted digital medical images enhanced in various forms are Mosaic spliced four random digital medical images, respectively (a)-(d), which greatly enrich the background of the detection object.
[0064] As shown in Figure 2As shown, the improved TE-YOLOv5 network TE-YOLOv5 of the present disclosure is based on the architecture of YOLOv5, including three parts: backbone, neck and output. The backbone extracts features as an encoder layer, and the neck combines features as a decoder layer. The backbone network YOLOv5 has four modules. They are respectively the cross-sectional module (Focus), the Conv module, the convolution module (C3 module) and the pooling pyramid (SSP) module. The input digital medical image with a resolution of 640x640x3 is first changed into a feature map of 320x320x12 through the Focus module using the cross-sectional operation, and then through the convolution operation of 32 convolution kernels, finally changed into a feature map of 320x320x32, making the feature extraction more robust than the traditional down-sampling. Among them, the convolution module C3 includes a convolution layer, a BN layer and a nonlinear activation layer.
[0065] Step S102: performing a multi-scale feature extraction operation on the multiple cross-sectional feature maps to obtain a multi-scale feature map; and fusing the multi-scale feature map and the global feature map to be updated to obtain an updated global feature map.
[0066] In an embodiment of the present disclosure, the method of performing a multi-scale feature extraction operation on the multiple cross-sectional feature maps to obtain a multi-scale feature map includes: performing a convolution operation on the multiple cross-sectional feature maps to obtain a multi-scale feature map.
[0067] For example, in the embodiment, first, the two C3 modules in the Low-level output a feature map of 64x64x128, and then the two C3 modules in the High-level output a feature map of 32x32x256. The output result is sent into the SSP module, and the output feature map is finally output as a feature map of 16x16x512 after passing through the C3 module. Figure 2
[0068] In an embodiment of the present disclosure, before the multi-scale feature map and the global feature map to be updated are fused to obtain an updated global feature map, the multi-scale feature map to be fused is determined, and the determination method includes: acquiring a set size; and if the size of the multi-scale feature map is less than or equal to the set size, determining the multi-scale feature map as the multi-scale feature map to be fused.
[0069] For example, the set size is 24x24. At this time, the multi-scale feature map of the embodiment includes a feature map of 64x64x128, a feature map of 32x32x256 and a feature map of 16x16x512.
[0070] In the embodiments of the present disclosure, the method for fusing the multi-scale feature map and the global feature map to be updated to obtain an updated global feature map comprises: performing convolution operation on the multi-scale feature map to obtain a feature map to be fused; scaling the feature map to be fused or the global feature map to be updated to the same size; and adding the scaled feature map to be fused or the global feature map to be updated to obtain an updated global feature map.
[0071] Figure 5 The structure diagram for updating the global feature map aggregation pool (AP) module of the present application. Since the detection accuracy of YOLOv5 depends on the output features, it is important to extract more effective information from the backbone network and the neck network. Features are divided into low-level and high-level categories: low-level features refer to some small details in an image, such as edges, corners, colors, pixels, and gradients. High-level features are built on low-level features and can be used to identify and detect target or object shapes in an image, with more semantic information. Compared with high-level features, low-level features require more computing resources due to their larger spatial resolution, but they provide less useful information. Based on this observation, the present application proposes to use an aggregation pool to update high-level features in a timely manner.
[0072] Figure 6 The structure diagram for the reverse attention (RA) module to implicitly learn the edge feature map. Although the neck of YOLOv5 is based on PANet, which collects features from bottom-up and top-down paths, it is difficult to extract the location and boundary of ambiguous features. In clinical practice, radiologists usually browse the entire image to locate suspicious lesion areas, and then accurately determine these areas by examining the location of the lesion or the tissue structure of the edge. Therefore, in Figure 6 , the network of the present application mimics the doctor's diagnosis process in two steps to detect ambiguous lesion areas.
[0073] Step S103: After each multi-scale feature extraction operation, performing an inverse operation on the updated global feature map to obtain background information corresponding to the multi-scale abnormal signal intensity feature; and fusing the background information corresponding to the multi-scale abnormal signal intensity feature and the multi-scale feature map respectively to obtain a multi-scale boundary feature map.
[0074] Step S104: Detecting the abnormal signal intensity area of the digital medical image according to the multi-scale feature map and the boundary feature map.
[0075] Figure 7The location, width, height distribution of the stroke training set lesions, and the relative size of the bounding boxes were annotated. Most stroke lesions are symmetrically distributed around the X-axis (right-left) and Y-axis (anterior-posterior). Most lesions are concentrated in the middle of the left and right hemispheres along the X-axis and in the brain along the Y-axis. In addition, it can be found that most lesions are very small, indicating that the DWI stroke dataset collected from the hospital is very suitable for testing the improved algorithm of the present application.
[0076] Figure 8 The results of P, R, F1 and PR curves for the independent stroke test set using the present application, where the method of the present application is compared with YOLOv5-S and YOLOv5-L. The precision, recall and F1 score curves of YOLOv5-S, YOLOv5-L, TE-YOLOv5-S and TE-YOLOv5-L as a function of confidence are shown in Figures (a), (b) and (c). The P-R curves of the four models are shown in Figure (d). The present application finally confirms that TE-YOLOv5-L produces the highest performance among the four models. Figure 8 Figure 8 (d). The present application finally confirms that TE-YOLOv5-L produces the highest performance among the four models.
[0077] Figure 9 The results of FROC curves for the independent stroke test set using the present application, where the method of the present application is compared with YOLOv5-S and YOLOv5-L. The FROC curves of YOLOv5-S, YOLOv5-L, TE-YOLOv5-S and TE-YOLOv5-L are shown in Figure Figure 9 . The first model to reach a sensitivity of 0.8 was TE-YOLOv5-L. TE-YOLOv5-S and TE-YOLOv5-L were more sensitive than YOLOv5-S and YOLOv5-L when the value of the average number of false positives per scan was less than 1.
[0078] Figure 10 Examples of stroke lesions detected by different models. Figure 10 The intuitive results of these four models compared to the ground truth in detecting stroke lesion areas are given. In the stroke detection task, TE-YOLOv5 effectively identifies those lesion areas with fuzzy boundaries. In the first image (first row), the lesions that are not detected by YOLOv5-S can be detected by TE-YOLOv5-S with a confidence of 0.4 (higher than 0.3 for YOLOv5-L). In addition, they can be more accurately identified on the deeper network of TE-YOLOv5-L with a confidence of 0.7. In the second image (second row), YOLOv5-S incorrectly identifies the edge artifacts in the DWI image as lesions; however, in the networks of the present application (TE-YOLOv5-S and TE-YOLOv5-L), these artifacts can be effectively removed.
[0079] Figure 11 A typical example of the lesion detection result of the TE-YOLOv5-S model of the present application. The main function of the AP module is to fuse the features of each layer, making the features more discriminative. On the other hand, the main function of the RA module is to make the boundary clearer through the reverse attention mechanism. The combination of the AP module and the RA module is mutually promoting, so that the micro-lesions with blurred boundaries can be better detected. Figure 11 The detection results of the TE-YOLOv5-S model of the present application on small and blurred lesions in different regions are shown. For example, the first column / second row shows that small-sized lesions (both the relative width and height are less than 5% of the entire image size) can be correctly detected. The first column / third row shows that lesions with blurred boundaries can also be correctly detected.
[0080] Figure 12 Comparison between manually labeled targets and automatically detected targets. This paper aims to develop an automatic detection method for cerebral apoplexy. The factors affecting the accuracy of automatic detection of cerebral apoplexy can be divided into two categories. The first category is the annotation standard and annotation accuracy of the data set, and the second category is the advantages and disadvantages of the designed algorithm. A series of comparisons are made between the original artificial annotation facts and the detection results of the network of the present application in terms of annotation standards. Figure 12 The left column shows that the radiologist considers the lesion area in the red region as a whole, while the algorithm of the present application considers it as two different lesion areas, which affects the accuracy of automatic detection. However, we found that some suspicious lesion areas were detected by the automatic detection method of the present application, but were not labeled by the radiologist on the DWI image, such as Figure 12 The middle and right columns.
[0081] Figure 13 A medical image abnormal signal intensity detection device as described above, characterized in that it comprises:
[0082] An acquisition unit for acquiring a digital medical image and operating on the cross-section of the digital medical image to obtain a plurality of cross-sectional feature maps;
[0083] An updating unit for performing multi-scale feature extraction operations on the plurality of cross-sectional feature maps to obtain multi-scale feature maps, and fusing the multi-scale feature maps and the global feature map to be updated to obtain an updated global feature map;
[0084] A boundary determination unit for performing an inversion operation on the updated global feature map after each multi-scale feature extraction operation to obtain background information corresponding to the multi-scale abnormal signal intensity features; and fusing the background information corresponding to the multi-scale abnormal signal intensity features and the multi-scale feature maps to obtain multi-scale boundary feature maps.
[0085] detecting, according to the multi-scale feature map and the boundary feature map, an abnormal signal intensity region of the digital medical image.
[0086] In the present application and other possible embodiments, the detection of the abnormal signal intensity is the detection of the abnormal signal intensity of ischemic stroke. The execution subject of the abnormal signal intensity detection method can be an abnormal signal intensity detection device of ischemic stroke, for example, the abnormal signal intensity detection method can be executed by a terminal device or a server or other processing device, wherein the terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the abnormal signal intensity detection method can be realized by a processor calling computer readable instructions stored in a memory.
[0087] Those skilled in the art can understand that, in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process, and the specific execution order of each step should be determined by its function and possible internal logic.
[0088] In some embodiments, the device provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above abnormal signal intensity detection method embodiments. For brevity, they will not be repeated here.
[0089] The embodiments of the present disclosure also propose a computer readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the above abnormal signal intensity detection method. The computer readable storage medium can be a non-volatile computer readable storage medium.
[0090] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor is configured to implement the above abnormal signal intensity detection method. The electronic device can be provided as a terminal, a server or other forms of devices.
[0091] Figure 13 The block diagram of the abnormal signal intensity detection unit of the present application is as follows, Figure 3As shown, the abnormal signal intensity detection device comprises: an acquisition unit 101, configured to acquire a digital medical image, and perform cross-sectional operation on the digital medical image to obtain a plurality of cross-sectional feature maps; an updating unit 102, configured to perform multi-scale feature extraction operation on the plurality of cross-sectional feature maps to obtain multi-scale feature maps; and perform fusion on the multi-scale feature maps and a global feature map to be updated to obtain an updated global feature map; a boundary determination unit 103, configured to perform inversion operation on the updated global feature map after each multi-scale feature extraction operation to obtain background information corresponding to multi-scale abnormal signal intensity feature; and perform fusion on the background information corresponding to the multi-scale abnormal signal intensity feature and the multi-scale feature maps to obtain multi-scale boundary feature maps; and a detection unit 104, configured to detect an abnormal signal intensity region of the digital medical image according to the multi-scale feature maps and the boundary feature maps.
[0092] Figure 14 is a block diagram of an electronic device 800 according to an exemplary embodiment. The electronic device 800 can be, for example, a terminal such as a mobile phone, a computer, a digital broadcasting terminal, a message communication device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0093] Referring to Figure 14 , the electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0094] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the above methods. In addition, the processing component 802 can include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0095] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or nonvolatile memory, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.
[0096] The power supply component 806 supplies power for various components of the electronic device 800. The power supply component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0097] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. The front camera and / or the back camera can receive external multimedia data when the electronic device 800 is in an operation mode, such as a photographing mode or a video mode. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0098] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.
[0099] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, etc. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0100] The sensor component 814 includes one or more sensors for providing status assessments for various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components of the electronic device 800, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration / g-force and temperature changes of the electronic device 800. The sensor component 814 can include an optical sensor for detecting ambient light, a proximity sensor configured to detect proximity of an object, a motion sensor, a temperature sensor, a magnetic sensor, an acceleration sensor, a gyroscope sensor, or a pressure sensor.
[0101] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a corresponding communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0102] In an example embodiment, the electronic device 800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements, to perform the above-described methods.
[0103] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 804 including computer program instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to perform the above-described methods.
[0104] Figure 15 FIG. 19 is a block diagram of an electronic device 1900 according to an example embodiment of the present disclosure. The electronic device 1900 can be provided as a server, for example. The electronic device 1900 can include a bus 1901, a processor 1902, a memory 1903, a storage 1904, an input / output (I / O) interface 1905, a display 1906, and a communication interface 1907. Figure 15The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.
[0105] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0106] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0107] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0108] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0109] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0110] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing / processing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0111] The computer readable program instructions can also be loaded onto a computing / processing device, other programmable data processing apparatus, or other device to cause a series of operations to be performed on the computing / processing device, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computing / processing device, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0112] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, to cause a series of operational elements or steps to be performed on the computer to produce a computer implemented process. Such instructions can also be stored and / or executed by other computer-readable media. Computer-readable media storing the computer readable instructions can include computers, processors, or other programmable data processing apparatuses capable of receiving, storing, and / or executing instructions.
[0113] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational elements or steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable elements or other apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0114] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational elements or steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable elements or other apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0115] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive of the disclosed embodiments. Many modifications and variations of the disclosed embodiments are possible in light of the above teachings without departing from the scope and spirit of the described embodiments. It is, therefore, to be understood that there has been described the preferred embodiments of the disclosure.
[0116] The new architecture of the TE-YOLOv5 proposed in the present application can aggregate multi-layer high-level features and simultaneously track the edge relationship of the features using the RA module, so that the brain stroke lesions in the DWI image can be automatically detected. A large number of experimental studies show that the brain stroke lesions, including the tiny lesions and the fuzzy boundary problems, can be effectively detected. Meanwhile, the method can also be used for lesion detection of other diseases. In summary, the detection method of the present application can effectively help the radiologists to reduce the work intensity and reduce the possibility of misdiagnosis or missed diagnosis caused by visual fatigue, which has great help and error correction potential for actual clinical work and treatment, greatly saves medical resources and medical manpower allocation, and reduces medical errors. This efficient and fast process will have great social and economic benefits.
Claims
1. A method of detecting abnormal signal intensity in a medical image, characterized by, The method comprises the following steps: acquire a diffusion weighted (DWI) digital medical image, and operate on the cross section of the digital medical image to obtain a plurality of cross section feature maps; perform multi-scale feature extraction on the plurality of cross section feature maps based on a YOLOv5 model to obtain multi-scale feature maps; and fuse the multi-scale feature maps and a global feature map to be updated to obtain an updated global feature map; perform an inversion operation on the updated global feature map after each multi-scale feature extraction to obtain background information corresponding to multi-scale abnormal signal intensity features; fuse the background information corresponding to the multi-scale abnormal signal intensity features and the multi-scale feature maps to obtain multi-scale boundary feature maps; detect an abnormal signal intensity region of the digital medical image according to the multi-scale feature maps and the boundary feature maps; The method for fusing the multi-scale feature maps and the global feature map to be updated to obtain an updated global feature map comprises the following steps: perform convolution on the multi-scale feature maps to obtain feature maps to be fused; scale the feature maps to be fused or the global feature map to be updated to the same size; add the scaled feature maps to be fused or the global feature map to be updated to obtain an updated global feature map; Before performing the inversion operation on the updated global feature map to obtain background information corresponding to multi-scale abnormal signal intensity features, the method further comprises the following steps: performing a spatial attention extraction operation on the updated global feature map to obtain spatial attention features; and performing an inversion operation on the spatial attention features to obtain background information corresponding to multi-scale abnormal signal intensity features; The method for detecting an abnormal signal intensity region of the digital medical image according to the multi-scale feature maps and the boundary feature maps comprises the following steps: fuse the multi-scale feature maps and the boundary feature maps to obtain a feature bounding box with the highest confidence, and use the feature bounding box to detect the abnormal signal intensity region of the digital medical image.
2. The method for detecting abnormal signal intensity of medical images according to claim 1, characterized in that: The method for performing multi-scale feature extraction on the plurality of cross section feature maps to obtain multi-scale feature maps comprises the following step: performing convolution on the plurality of cross section feature maps to obtain multi-scale feature maps.
3. The method of claim 1, wherein the step of detecting abnormal signal intensity in the medical image is characterized by, The method for fusing the background information corresponding to the multi-scale abnormal signal intensity features and the multi-scale feature maps to obtain multi-scale boundary feature maps comprises the following step: multiplying the background information corresponding to the multi-scale abnormal signal intensity features and the multi-scale feature maps to obtain multi-scale boundary feature maps.
4. The method of claim 1, wherein the step of detecting abnormal signal intensity in the medical image is characterized by, The method for fusing the multi-scale feature maps and the boundary feature maps to obtain a feature bounding box with the highest confidence comprises the following steps: splicing the multi-scale feature maps and the boundary feature maps to obtain a plurality of output feature maps consistent in number with the multi-scale feature maps; determine the maximum size of the plurality of output feature maps; up-sample the output feature maps of other sizes to the maximum size, and then splice the plurality of output feature maps to obtain a feature bounding box with the highest confidence.
5. The method of claim 1, wherein the step of detecting abnormal signal intensity in a medical image is characterized by, Before fusing the multi-scale feature map and the global feature map to be updated to obtain an updated global feature map, a multi-scale feature map to be fused is determined, and a determination method thereof comprises the following steps of: acquiring a set size; if the size of the multi-scale feature map is less than or equal to the set size, the multi-scale feature map is determined to be fused.
Citation Information
Patent Citations
Medical image recognition method, medical image recognition device, terminal equipment and medium
CN110647889A
Feature extraction method and device
CN110781923A