Foreign Object Detection Method for Chest X - ray Images Based on Attention Mechanism Convolutional Neural Network

Through a convolutional neural network based on attention mechanism, the problem of insufficient accuracy of foreign object detection on chest X-rays is solved, more efficient foreign object classification and positioning is achieved, and the reliability of chest X-ray diagnosis is improved.

CN115272215BActive Publication Date: 2025-07-22ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210863784.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-07-22
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

The prior art when detecting foreign objects on chest X-rays, it faces small sample/zero sample learning problems, resulting in insufficient accuracy and reliability of foreign object detection, which easily masks pathological results and increases the probability of incorrect diagnosis.

Method used

A convolutional neural network based on attention mechanism is adopted, and a convolutional neural network with bottom-up and top-down structures is constructed by pre-processing through the restricted contrast adaptive histogram equalization algorithm. The channel attention module is used to perform multi-scale feature fusion to enhance the accuracy of foreign object detection.

Benefits of technology

It improves the accuracy of foreign body detection, can better classify and locate foreign body locations, reduces the risk of misdiagnosis, and improves the reliability of chest X-ray diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272215B_ABST
    Figure CN115272215B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting foreign objects in chest X-ray images based on a convolutional neural network with an attention mechanism. First, the original image is preprocessed using the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm. Secondly, a convolutional neural network based on channel attention is constructed, where the backbone network extracts features to generate multi-scale feature maps, the feature pyramid structure performs multi-scale feature fusion, and the attention module enhances the feature fusion of the multi-scale feature maps in the channel dimension and introduces the feature maps output by the feature pyramid structure into the head for foreign object detection. Then, the neural network based on channel attention is trained with the training set data to obtain a trained network model. Finally, the test set data is input into the trained network model to regress the coordinates of the foreign object. The present invention performs multi-scale feature fusion in the channel dimension, fully utilizes the effective information of each channel, accurately classifies foreign objects, and locates the positions of foreign objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for detecting foreign objects in chest X-ray images based on an attention mechanism convolutional neural network. Background Art

[0002] Medical image processing plays an important role in medical research. Analyzing chest X-ray images is a common clinical method for diagnosing lung and heart diseases. However, foreign objects may occasionally appear on chest X-ray images. These foreign objects may obscure the pathology, thereby increasing the probability of misdiagnosis. They may also confuse doctors about the true pathological results. For example, buttons on chest X-ray images visually resemble nodules, thus increasing the false positive judgment.

[0003] Detecting foreign objects on chest X-ray images helps doctors diagnose lung and heart diseases more accurately. However, it is challenging for neural networks to detect foreign objects because specific types of foreign objects that appear in the test set may be rare or never seen in the training set, and the foreign object images of chest X-ray images are scarce in themselves, thus causing small sample / zero sample learning problems. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method for detecting foreign objects in chest X-ray images based on an attention mechanism convolutional neural network, which can improve the contrast of images and perform fusion of multi-scale features in the channel dimension.

[0005] The technical solution adopted by the present invention to solve the technical problems is as follows:

[0006] The specific steps are as follows:

[0007] Step 1, perform data preprocessing, and preprocess the original image with the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm.

[0008] Step 2, construct a convolutional neural network based on a bottom-up and top-down structure with channel attention, which includes a bottom-up structure backbone network and a top-down feature pyramid structure. The backbone network extracts features to generate multi-scale feature maps, and the top-down feature pyramid structure performs fusion of multi-scale feature maps.

[0009] The attention module strengthens the feature fusion of multi-scale feature maps in the channel dimension. The attention module introduces the feature map output by the top-down feature pyramid structure into the head of foreign object detection.

[0010] Further, the bottom-up structure backbone network of the first part is a ResNet50 network, and five downsampling operations are performed on the preprocessed image to obtain five layers of feature maps of the first part in sequence. The top-down feature pyramid structure of the second part is horizontally connected to the backbone network of the first part. First, the fifth-layer feature map of the first part is obtained after convolution operation to get the fifth-layer feature map of the second part. After the fifth-layer feature map of the second part undergoes an upsampling operation, the resulting feature map is concatenated with the feature map obtained after the convolution operation of the fourth-layer feature map of the first part, and then a first convolution block operation is performed to obtain the fourth-layer feature map of the second part; the feature map obtained after the fourth-layer feature map of the second part undergoes an upsampling operation is concatenated with the feature map obtained after the convolution operation of the third-layer feature map of the first part, and then a first convolution block operation is performed to obtain the third-layer feature map of the second part; the sixth-layer feature map of the second part is obtained by performing a convolution operation on the fifth-layer feature map of the second part, and the seventh-layer feature map of the second part is obtained by performing a convolution and ReLU function operation on the sixth-layer feature map of the second part.

[0011] For the attention module, the upsampled fifth-layer feature map of the second part obtained above and the fourth-layer feature map of the second part are first subjected to element-wise addition, and then successively passed through a binary adaptive average pooling layer, a fully connected layer, a ReLU function, a fully connected layer, and a sigmoid activation function operation to obtain channel weights. The channel weights are then multiplied with the fifth-layer feature map of the second part and the fourth-layer feature map of the second part respectively at the channel level to obtain new feature maps. Finally, they are respectively jump-connected with the fifth-layer feature map of the second part and the fourth-layer feature map of the second part for element-wise addition to obtain the fifth-layer feature map of the third part and the first component of the fourth-layer feature map of the third part; at the same time, the fourth-layer feature map of the second part is upsampled, and then it and the third-layer feature map of the second part first undergo element-wise addition, and then successively pass through a binary adaptive average pooling layer, a fully connected layer, a ReLU function, a fully connected layer, and a sigmoid activation function operation to obtain channel weights. The channel weights are then multiplied with the fourth-layer feature map of the second part and the third-layer feature map of the second part respectively at the channel level to obtain new feature maps. Finally, they are respectively jump-connected with the fourth-layer feature map of the second part and the third-layer feature map of the second part for element-wise addition to obtain the second component of the fourth-layer feature map of the third part and the third-layer feature map of the third part. The first component and the second component of the fourth-layer feature map of the third part are added element-wise to obtain the complete fourth-layer feature map of the third part.

[0012] Step 3: Input the training set data into the convolutional neural network based on channel attention for training to obtain a trained network model.

[0013] Step 4: Input the test set data into the trained network model to obtain the extracted feature maps, and then perform prediction and regression to obtain the coordinates of foreign objects in the X-ray film.

[0014] Further, the first convolutional block in the feature pyramid of the top-down structure includes: convolution, batch normalization, and activation function.

[0015] Further, the head of the foreign object detection includes:

[0016] (1) The foreign object classification branch: The sixth and seventh layer feature maps of the second part and the third to fifth layer feature maps of the third part first undergo operations of 4 second convolutional blocks, and then the classification result is obtained through convolution operations, where the second convolutional block includes convolution, group normalization, and activation function;

[0017] (2) The foreign object coordinate regression branch: The sixth and seventh layer feature maps of the second part and the third to fifth layer feature maps of the third part first undergo operations of 4 second convolutional blocks, and then convolution and activation function operations are performed, and finally the foreign object coordinates are regressed.

[0018] Advantages of the present invention: The present invention provides a channel attention module in the channel dimension, which can more effectively perform multi-scale feature fusion, make full use of the effective information of each channel, accurately classify foreign objects in X-ray films, and locate the positions of foreign objects. Description of the Drawings

[0019] Figure 1 is the flowchart of foreign object detection in medical image chest X-ray films of the present invention;

[0020] Figure 2 is the structure diagram of the convolutional neural network based on channel attention of the present invention;

[0021] Figure 3 is the structure diagram of the attention module of the present invention;

[0022] Figure 4 is the schematic diagram of the foreign object detection result of the present invention. Detailed Embodiments

[0023] The following further describes the present invention in conjunction with the drawings and embodiments.

[0024] In this embodiment, the dataset provided by Zhejiang Provincial People's Hospital is selected to train the convolutional neural network model and evaluate its performance. The training set and test set of the dataset are two batches of different random data, and new types of foreign objects may appear in the test set. There are 2663 images with foreign objects in the training images of the dataset, and 712 test images. In order to reduce the number of parameters, all the pictures in the dataset need to be re-clipped to 512*800.

[0025] As shown in Figure 1 Figure Figure 1 , a foreign object detection method for chest X-ray images based on an attention mechanism in a convolutional neural network includes the following steps:

[0026] Step 1, perform data preprocessing. Use the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm to preprocess the original image and enhance the image contrast.

[0027] Step 2, construct a convolutional neural network with a bottom-up and top-down structure based on channel attention, including a backbone network with a bottom-up structure and a top-down feature pyramid structure. The backbone network extracts features to generate multi-scale feature maps, and the feature pyramid performs multi-scale feature map fusion.

[0028] The attention module strengthens the feature fusion of multi-scale feature maps in the channel dimension. The attention module introduces the feature maps of the feature pyramid structure into the head of foreign object detection.

[0029] As shown in Figure 2 Figure Figure 2 , the backbone network of the bottom-up structure in the first part is the ResNet50 network. Five downsampling operations are performed on the preprocessed image to obtain five layers of feature maps in the first part, which are the first-layer feature map C1, the second-layer feature map C2, the third-layer feature map C3, the fourth-layer feature map C4, and the fifth-layer feature map C5 from bottom to top. Since the semantic information levels of the feature maps C1 and C2 are too low, they do not participate in subsequent foreign object predictions.

[0030] The top-down feature pyramid structure in the second part is horizontally connected to the backbone network in the first part. First, the fifth-layer feature map C5 in the first part passes through a 1×1 convolution to obtain the fifth-layer feature map P5 in the second part. Then, after an upsampling operation, the resulting feature map and the feature map obtained by passing the fourth-layer feature map C4 in the first part through a 1×1 convolution are subjected to a channel concatenation operation, and then a first convolution block operation is performed to obtain the fourth-layer feature map P4 in the second part; the feature map obtained after the upsampling operation of P4 and the feature map obtained by passing the third-layer feature map C3 in the first part through a 1×1 convolution are subjected to a channel concatenation operation, and then after a first convolution block operation, the third-layer feature map P3 in the second part is obtained; the sixth-layer feature map P6 in the second part is the feature map obtained by performing a 3×3 convolution operation on the fifth-layer feature map P5, and the seventh-layer feature map P7 in the second part is the feature map obtained by performing a 3×3 convolution and ReLU function operation on the sixth-layer feature map P6.

[0031] As shown in Figure 3As shown in the figure, the attention module first performs element-wise addition on the upsampled feature map P5 of the fifth layer of the second part and the feature map P4 of the fourth layer, and then successively performs operations including binary adaptive average pooling layer, fully connected layer, ReLU function, fully connected layer, and sigmoid activation function to obtain channel weights. The channel weights are then multiplied with the feature map P5 of the fifth layer of the second part and the feature map P4 of the fourth layer at the channel level to obtain new feature maps. Finally, they are respectively element-wise added to the skip connections of the feature map P5 of the fifth layer of the second part and the feature map P4 of the fourth layer to obtain the feature map M5 of the fifth layer of the third part and the first component of the feature map of the fourth layer of the third part.

[0032] At the same time, after upsampling the feature map P4 of the fourth layer of the second part and performing element-wise addition on the feature map P3 of the third layer of the second part, operations including binary adaptive average pooling layer, fully connected layer, ReLU function, fully connected layer, and sigmoid activation function are successively performed to obtain channel weights. The channel weights are then multiplied with the feature map P4 of the fourth layer of the second part and the feature map P3 of the third layer at the channel level to obtain new feature maps. Finally, they are respectively element-wise added to the skip connections of the feature map P4 of the fourth layer of the second part and the feature map P3 of the third layer to obtain the second component of the feature map of the fourth layer of the third part and the feature map M3 of the third layer of the third part. The first component and the second component of the feature map of the fourth layer of the third part are element-wise added to obtain the complete feature map M4 of the fourth layer of the third part.

[0033] Step 3: Set the network training learning rate to 0.000001 and the optimizer to the Adam optimizer. Input the training set data into the convolutional neural network based on channel attention for training to obtain a trained network model.

[0034] Step 4: Input the test set data into the trained network model to obtain the extracted feature maps, and then perform prediction to regress the coordinates of the foreign object, as Figure 4 shown.

[0035] The first convolutional block in the feature pyramid of the top-down structure includes: 3x3 convolution, batch normalization, and ReLU activation function.

[0036] The head of the foreign object detection includes:

[0037] (1) The foreign object classification branch: After the feature maps M3, M4, M5, P6, and P7 pass through 4 second convolutional blocks, they are then passed through a 3x3 convolution to obtain the classification result. The second convolutional block includes 3x3 convolution, group normalization, and ReLU activation function.

[0038] (2) The foreign object coordinate regression branch is obtained by performing operations on the feature maps M3, M4, M5, P6, and P7 through the 4th second convolutional block, followed by 3x3 convolution and ReLU activation function operations, and finally regressing to obtain the foreign object coordinates.

[0039] On the test set provided by the Zhejiang Provincial People's Hospital, the mean average precision mAP reached 0.75, while the mAP of the fully convolutional one-stage method was 0.71, an increase of 4%. The calculation of mAP is as follows:

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] In formulas (1) and (2), TP (true positive) indicates that the prediction is correct and the actual is also a positive example, FP (false positive) indicates that the prediction is incorrect and the actual is a negative example, and FN (false negative) indicates that the prediction is incorrect and the actual is a positive example. The calculation method of mAP is to first calculate the precision and recall using formulas (1) and (2), and the AP calculation is defined as the area enclosed by the interpolated precision-recall curve and the X-axis. The interpolation process is that for a given recall value r, as shown in formula (3), P interp is the maximum precision value between the next recall value r' and the current r value. In formula (4), r1, r2,..., r n is the recall value corresponding to the first interpolation point of the precision interpolation segment arranged in ascending order. Calculate the AP for all K categories, and then take the average to calculate mAP, as shown in formula (5).

Claims

1. A method for detecting foreign objects in chest X-ray images based on an attention mechanism convolutional neural network, characterized in that It includes the following steps: Step 1, preprocess the original chest X-ray image using the Limited Contrast Adaptive Histogram Equalization algorithm; Step 2, construct a convolutional neural network based on the channel attention bottom-up and top-down structures, including the backbone network of the bottom-up structure and the top-down feature pyramid structure; The attention module strengthens the feature fusion of multi-scale feature maps in the channel dimension; The backbone network extracts features to generate multi-scale feature maps; The top-down feature pyramid structure performs multi-scale feature map fusion; The attention module introduces the feature map output by the top-down feature pyramid structure into the head of foreign object detection; Step 3, input the training set data into the convolutional neural network based on channel attention for training to obtain a trained network model; Step 4, input the test set data into the trained network model to obtain the extracted feature map, and then perform prediction to regress the coordinates of foreign objects in the X-ray image.

2. The method for detecting foreign objects in chest X-ray images based on an attention mechanism convolutional neural network according to claim 1, wherein: In Step 2, the backbone network of the bottom-up structure in the first part is the resnet50 network, and five downsampling operations are performed on the preprocessed image to obtain five layers of feature maps in the first part in sequence; The top-down feature pyramid structure in the second part is horizontally connected to the backbone network in the first part. Specifically: The fifth layer feature map in the first part undergoes a convolutional operation to obtain the fifth layer feature map in the second part; The feature map obtained after the fifth layer feature map in the second part undergoes an upsampling operation and the feature map obtained after the fourth layer feature map in the first part undergoes a convolutional operation are subjected to a channel concatenation operation; then a first convolutional block operation is performed to obtain the fourth layer feature map in the second part; The feature map obtained after the fourth layer feature map in the second part undergoes an upsampling operation and the feature map obtained after the third layer feature map in the first part undergoes a convolutional operation are subjected to a channel concatenation operation, and then a first convolutional block operation is performed to obtain the third layer feature map in the second part; The sixth layer feature map in the second part is obtained by performing a convolutional operation on the fifth layer feature map in the second part; The seventh layer feature map in the second part is obtained by performing a convolutional operation and an activation function operation on the sixth layer feature map in the second part; For the said attention module, the upsampled fifth layer feature map in the second part and the fourth layer feature map in the second part are first subjected to element-wise addition, and then successively pass through a binary adaptive average pooling layer, a fully connected layer, a ReLU function, a fully connected layer, and a sigmoid activation function operation to obtain the weights of each channel. The weights of each channel are then respectively subjected to channel-wise multiplication with the fifth layer feature map and the fourth layer feature map in the second part to obtain new feature maps. Finally, they are respectively jump-connected to the fifth layer feature map and the fourth layer feature map in the second part for element-wise addition to obtain the fifth layer feature map in the third part and the first component of the fourth layer feature map in the third part; After upsampling the feature map of the fourth layer of the second part and the feature map of the third layer of the second part, first perform element-wise addition, and then successively perform binary adaptive average pooling layer, fully connected layer, ReLU function, fully connected layer, and sigmoid activation function operations to obtain the weights of each channel. Then, the weights of each channel are respectively multiplied with the feature maps of the fourth layer and the third layer of the second part at the channel level to obtain new feature maps. Finally, perform element-wise addition with the feature maps of the fourth layer and the third layer of the second part through skip connections respectively to obtain the second component of the fourth layer feature map of the third part and the third layer feature map of the third part; The first component of the fourth layer feature map of the third part and the second component of the fourth layer feature map of the third part are added element-wise to obtain the complete fourth layer feature map of the third part.

3. The method for detecting foreign objects in chest X-ray images based on the convolutional neural network with attention mechanism according to claim 2, wherein: The first convolutional block includes convolution, batch normalization, and activation function.

4. The method for detecting foreign objects in chest X-ray images based on an attention mechanism convolutional neural network according to claim 2, wherein: The head of the foreign object detection described in step 2 includes a foreign object classification branch and a foreign object coordinate regression branch; For the foreign object classification branch, the feature maps of the sixth layer of the second part, the seventh layer of the second part, and the third to fifth layers of the third part first go through four second convolutional block operations, and then through a convolution operation to obtain the classification result; For the foreign object coordinate regression branch, the feature maps of the sixth layer of the second part, the seventh layer of the second part, and the third to fifth layers of the third part first go through four second convolutional block operations, and then through convolution and activation function operations to regress the foreign object coordinates.

5. The method for detecting foreign objects in chest X-ray images based on the convolutional neural network with attention mechanism according to claim 4, wherein: The second convolutional block includes convolution, group normalization, and activation function.