Robust x-ray image target detection method based on same kind target fusion data augmentation
Patent Information
- Application Number
- CN202410996894.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-07-24
AI Technical Summary
但是以上这些方法大多都依赖于一个大规模的数据集来进行模型训练,但是由于X光图片中存在着严重的物体重叠遮挡的情况,这导致数据集的标注相比其它任务更为困难
[0047]采用上述的技术方案,本发明与现有技术相比,其具有的有益效果为:本方案方法可以有效地消除标签噪声对X光安检图像目标检测器训练产生的影响,并且通过图像融合很好的模拟了X光图像中重叠遮挡的情况,让模型可以更好的学习到X光图片的固有特征。该方案不仅在多个公开地数据集上都取得了良好的性能,同时相比于传统的噪声标签学习方法,是一种更加灵活,且贴近实际需求地解决噪声标签情况下X光安检图像目标检测的方案。除此之外,通过实验结果发现,本方案方法不仅在X光数据集中有效,在一些通用目标检测数据集中也有较好的效果。
Smart Images

Figure CN118968075B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a robust X-ray image target detection method based on augmentation of similar target fusion data. Background Technology
[0002] Object detection is a crucial problem in computer vision, referring to the identification of target objects within an image and the determination of their location, size, and category. Object detection has wide applications in computer vision, such as autonomous vehicles, drones, and security monitoring. In security monitoring applications, one type of image is obtained through X-ray imaging, such as the images output by security screening equipment in airports, train stations, and subway stations. This necessitates object detection in X-ray images to address the automated detection of dangerous goods during security checks.
[0003] In recent years, with the development of deep learning, more and more computer vision applications have begun to use deep learning-based methods, and X-ray security image detection has also entered the era of deep learning. Early X-ray image target detection used methods based on handcrafted features, and these methods were not very effective. Because statistical deep learning requires a large amount of data for training, but there were no large datasets specifically designed for X-ray security detection at the time, convolutional neural networks were initially not well applied to this field. However, the subsequent emergence of transfer learning methods changed this situation. In 2016, Akcay et al. ( S, Kundegorski ME, Devereux M, et al. Transfer learning using convolutional neural networks for object classification within x-ray baggage security imagery[C] / / 2016 IEEE International Conference on Image Processing (ICIP). IEEE, 2016: 1057-1061.) To overcome the problem of insufficient data, transfer learning was used for the first time in the field of X-ray security image detection. Based on AlexNet, the parameters of convolutional and fully connected layers were fine-tuned. This method achieved better results than previous methods, which was the first study of deep convolutional neural networks in X-ray security image detection. Since then, more and more deep learning-based methods have been applied in this field. Miao et al. (Miao C, Xie L, Wan F, et al. Sixray: A large-scale security inspection x-ray benchmark for prohibited item discovery in overlapping images[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2019:2119-2128.) proposed a large-scale X-ray security inspection dataset, SIXray. They also used a class-balanced hierarchical refinement method, and experiments showed that this method can effectively solve the problem of imbalanced positive and negative samples. They also designed a class-balanced loss function to mitigate the noise introduced by negatively oriented samples. Wei (Wei Y, Tao R, Wu Z, et al. Occluded Prohibited Items Detection: An X-ray Security Inspection Benchmark and De-occlusion Attention Module[C]. ACM Multimedia 2020.) et al. proposed a high-quality dataset, OPIXray, designed for target detection tasks in security inspection scenarios. All images in this dataset contain dangerous items that were manually labeled by professional security inspectors at an airport.Furthermore, this paper proposes a method for detecting obscured contraband in security check scenarios. The core of this method utilizes an attention mechanism to enhance the edge and material information of contraband in X-ray images. However, most of these methods rely on a large-scale dataset for model training. Due to severe object overlap and occlusion in X-ray images, labeling the dataset is more difficult than for other tasks. Therefore, obtaining a large-scale X-ray image dataset with completely accurate labels is extremely challenging. In fact, most current X-ray datasets contain noise to varying degrees, including inaccurate bounding boxes and incorrect object classifications. When a model is trained on a noisy dataset, its training accuracy can be severely affected by this noise. Therefore, designing a noise-robust X-ray contraband detection method is essential. Summary of the Invention
[0004] In view of this, the purpose of this invention is to propose a robust X-ray image target detection method based on the fusion data of similar targets, which is reliable, versatile, and has good noise resistance.
[0005] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows:
[0006] A robust X-ray image target detection method based on augmentation of similar target fusion data includes the following steps:
[0007] A. Prepare a dataset of X-ray security inspection images, and divide it into a training set and a test set. The image data in both the training set and the test set have category labels and bounding box labels pointing to the target objects.
[0008] B. Add label noise to the image data in the training set to simulate noise conditions. The noise can be added manually and includes category label noise and target bounding box label noise.
[0009] C. For each target object T in each image data in the training set, find K-1 images with the same category label as the target object T from the image data in the dataset. Then, based on the target bounding box label corresponding to the image data, crop out the target object and convert the resolution of the cropped image to the same resolution as the bounding box of the target object T.
[0010] D. Use the edge smoothing mask algorithm to re-paste the cropped target object image back into the original position of the original image data;
[0011] E. Input the image data of the training set after step D into a deep learning-based object detector for model training.
[0012] F. In the loss function of the target detector model, a pre-set large loss suppression mechanism is used to suppress the large loss caused by label errors in image data, thereby mitigating the negative impact of label noise on model training.
[0013] G. Use the backpropagation algorithm to train the model end-to-end, and obtain the trained model after the model passes the test set.
[0014] H. Use the trained model to perform target detection in X-ray machine images and analyze the detection results output by the model.
[0015] As one possible implementation, further, in step A of this scheme, the prepared X-ray security inspection image dataset includes OPIXray, PIDray, and MS-COCO.
[0016] The dataset OPIXray used in this scheme is an X-ray image dataset containing folding knife, straight knife, scissor, utility knife and multi-tool knife, consisting of 8885 X-ray images, of which 7109 images were assigned to the training set and the remaining 1776 images were assigned to the test set.
[0017] The PIDray dataset is a large X-ray dataset containing 47,677 images across 12 different categories. 29,457 images were used for training, and 18,220 images were used for testing.
[0018] The MS-COCO dataset is a large, publicly available, general-purpose object detection dataset containing 135,000 training images and 5,000 test images, across 80 different categories. This approach uses the OPIXray and PIDray X-ray datasets to test the performance of our method on X-ray contraband detection with noisy labels, and uses MS-COCO to verify that our method also generalizes well to other tasks.
[0019] Since it is difficult to accurately determine the noise level of a dataset, and it is also difficult to find a dataset with a specific noise rate, in order to test the performance of this scheme under different noise rates, as a preferred implementation method, step B of this scheme, which involves adding label noise to the image data of the training set to simulate noise conditions, includes adding label noise to the category label of the target object and the target bounding box label respectively;
[0020] The method for adding label noise to the category labels is as follows: the original category labels of the image data are replaced with labels of any other category in the dataset with a preset probability. This preset probability is the category noise rate of the dataset.
[0021] The method for adding label noise to the target bounding box labels is as follows: The original bounding boxes of the image data are offset and scaled with a preset probability; this preset probability is the target bounding box noise rate of the dataset. The coordinates (x, y, w, h) are set as the bounding box of the image data, and the following operations are performed with the preset probability:
[0022]
[0023] Where, Δ x Δ y Δ w Δ h It is obtained by random sampling from a uniform distribution U(-δ,δ), where δ represents the noise perturbation level, and x, y, w, h are the bounding box coordinates pointing to the target object in the image data. The bounding box coordinates after adding noise; as an example, δ = 0.3 can be set.
[0024] As a preferred implementation method, step C of this scheme preferably includes: for each target object T in each image data in the training set, the label corresponding to the target object T contains a target bounding box label pointing to the location information of the target object and a category label indicating the category information of the target object; based on the category information provided by the label, find K-1 images of target objects with the same category information from the original dataset, and crop out the target objects in the image data according to their corresponding target bounding box labels, and put them into the result set;
[0025] In this process, an image is randomly selected from the dataset each time. If the image contains a target object with the same category information as the target object T in its label, the target is cropped out according to its corresponding target bounding box label. The image resolution of the cropped target object is then changed to the same resolution as the target object T and added to the result set. Otherwise, an image is randomly selected from the dataset again until the result set contains K-1 targets with the same resolution.
[0026] As a preferred implementation method, step D of this solution preferably includes:
[0027] For all target objects in the result set and their corresponding original target objects T, they are fused into the same image to increase the probability of target objects with the corresponding category labels appearing in the target bounding box, thereby reducing the probability of target objects with the corresponding category labels not appearing in the target bounding box due to noise. To make the fusion smoother and the final result look more natural, an edge smoothing mask algorithm is used to fuse the image data. The corresponding calculation formula is as follows:
[0028]
[0029] Where, d i,j It is the distance between the pixel at position (i,j) in the target bounding box and the nearest edge of the bounding box, and its value is min(i,j,h). o -i,w o -j), h o ,w o Here, λ represents the height and width of the cropped target, respectively; β is a threshold that controls the range of the smooth boundary; and λ is a number randomly sampled from the Beta distribution.
[0030] The images in the result set are fused using an edge-smoothing mask. The fusion formula is as follows:
[0031]
[0032] Where α is the edge smoothing mask, B a B is the image patch corresponding to the original target object T. n For each image patch in the result set, K represents the total number of images to be merged. Element-wise multiplication;
[0033] Finally, the merged image is pasted back into the original image data at the location of the target object T, completing the fusion process.
[0034] The image fusion process described above can significantly reduce the noise rate of both category noise and bounding box noise. Suppose a dataset with a noise rate of P contains a label of category C. The probability that the bounding box actually contains a target of category C is 1-P. When K image patches of the same category are fused, the probability that the fused bounding box contains the target becomes 1-P. K Since P is a number between 0 and 1, the probability of the target being contained in the bounding box will be greatly increased.
[0035] As a preferred implementation method, in step E of this scheme, the deep learning-based object detector uses Faster R-CNN as the basic object detection network. Its core components include a basic feature extraction network, a region proposal network, and a region classification and boundary regression network, which are used for feature extraction, selecting candidate regions from the feature map, and classifying and regressing bounding boxes for each candidate region, respectively. Faster R-CNN achieves efficient object detection through end-to-end training, significantly improving both speed and accuracy compared to traditional two-stage methods. Although this invention primarily uses Faster R-CNN for experiments, since this method mainly modifies images, it can also be applied to other object detection networks depending on the task requirements.
[0036] After fusing multiple targets in step E, this method may encounter multiple targets of different categories at a single location due to noise, even though they share a single label. Even if the model predicts all targets in the image, other target categories appearing due to noise, lacking correct labels, will be suppressed as incorrect predictions, leading to a decline in model performance. To mitigate the impact of these "false errors" on model training, as a preferred implementation method, step F of this scheme preferably includes:
[0037] The predictions of the target detector model are divided into four categories:
[0038] (1)PB neg The model predicts that the IoU between the target bounding box and the label bounding box is less than a threshold.
[0039] (2)PB pp The model predicts that the IoU between the target bounding box and the label bounding box is greater than the threshold, but the predicted category is different from the label category and is not the background.
[0040] (3)PB fb The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, and the predicted category is background.
[0041] (4)PB pos The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, and the predicted category is the same as the label.
[0042] Among them, "false error" prediction mainly refers to PB. pp The included prediction results are discarded from the loss function calculation, and only the classification losses of the other three parts are calculated. Finally, the total loss function becomes:
[0043] L = L bbox +Lclsneg +L clspos +L clsfb
[0044] Among them, L bbox This represents the regression loss, which is left unchanged. clsneg L clspos L clsfb Representing the above PB respectively neg PB pos ,PB fb The corresponding classification loss.
[0045] As a preferred implementation method, step G of this scheme further includes data preprocessing, which includes: using random flipping data augmentation to normalize the data of all training and test sets; the basic feature extraction network is initialized and trained on ResNet pre-trained on ImageNet, with stochastic gradient descent as the optimizer, wherein the initial learning rate is 0.0025, and the learning rate is reduced by a factor of 10 in the 17th and 21st rounds of training, respectively, the weight decay parameter is 0.0001, and the momentum parameter is 0.9; the mini-batch size for each iteration is set to 2, and the entire network is trained for 24 rounds; during the training process, the preset large loss suppression mechanism is used to modify the loss function, and after calculating the loss, the gradient descent algorithm is used for backpropagation to update the model parameters.
[0046] As a preferred implementation method, step H of this scheme includes: transforming the resolution of the input image to be detected to 1333*800, then inputting it into a trained model for target detection of X-ray security inspection images, using the model to input the extracted features into a regression network to obtain the coordinate values of the detection bounding box, and then inputting them into a classification network to predict its corresponding category, thereby generating a detection result.
[0047] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: The proposed method can effectively eliminate the impact of label noise on the training of X-ray security image target detectors, and through image fusion, it effectively simulates the overlapping and occlusion situation in X-ray images, allowing the model to better learn the inherent features of X-ray images. This solution not only achieves good performance on multiple publicly available datasets, but also, compared with traditional noisy label learning methods, is a more flexible and practical solution for X-ray security image target detection under noisy label conditions. Furthermore, experimental results show that the proposed method is not only effective on X-ray datasets, but also performs well on some general target detection datasets. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a flowchart illustrating the entire implementation process of an embodiment of the present invention.
[0050] Figure 2 This is a diagram of the overall network structure according to an embodiment of the present invention.
[0051] Figure 3 Some test results on the OPIXray dataset before and after using the proposed method for the same model (Faster RCNN). Both models were trained at the same noise rate (60% bounding box noise and class noise). Figure 3 (a) Detection results of the original Faster RCNN model. The blue box represents the true class label and bounding box, the red box represents the predicted label and bounding box, and the parentheses represent the prediction confidence. Figure 3 (b) shows the detection results of the Faster RCNN model after using this method. It can be seen that the quality of the predicted bounding boxes and the accuracy of the predicted categories have been greatly improved after using this method.
[0052] Figure 4 This is a simplified illustration of the average accuracy measured by the area enclosed by the PR curve. Detailed Implementation
[0053] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] See Figure 1 As shown, the solution in this embodiment is a robust X-ray image target detection method based on similar target fusion data augmentation, which includes the following steps:
[0055] A. Prepare a dataset of X-ray security inspection images, dividing it into training and testing sets. Both sets contain category labels and bounding box labels pointing to the target objects. The training sets consist of two commonly used datasets for X-ray security inspection image target detection and one general object detection dataset. The X-ray security inspection image target detection datasets are OPIXray and PIDray, and the general object detection dataset is MS-COCO. OPIXray is an X-ray image dataset containing folding knives, straight knives, scissors, utility knives, and multi-tool knives, consisting of 8885 X-ray images. 7109 images are used in the training set, and the remaining 1776 images are used in the testing set. PIDray is a larger X-ray dataset containing 47677 images across 12 different categories. 29457 images are used for training, and 18220 images are used for testing. MS-COCO is a publicly available, large-scale, general-purpose object detection dataset containing 135,000 training images and 5,000 test images, across 80 different categories. This approach uses the OPIXray and PIDray X-ray datasets to test the performance of its method on X-ray contraband detection tasks with noisy labels, and uses MS-COCO as the test set to verify that the invention also exhibits good generalization performance on other tasks.
[0056] B. To simulate the presence of noise, some noise is manually added to the training data labels in the dataset.
[0057] Because it is difficult to accurately determine the noise level of a dataset, and also difficult to find datasets with a specific noise rate, this solution directly adds noise of a specific noise rate to the original labels of the dataset to test its performance under different noise rates. Since the labels for object detection tasks consist of two parts: the bounding box and the object category, noise needs to be added to both parts separately. Specifically, for category noise, this solution replaces the original category labels in the dataset with labels of any other category in the dataset with a certain probability; this probability is the category noise rate of the dataset. For bounding box noise, this solution offsets and scales the original bounding boxes in the dataset with a certain probability; this probability is the target bounding box noise rate of the dataset. Specifically, for a bounding box with coordinates (x, y, w, h), this solution performs the following operations with a certain probability:
[0058]
[0059] Where, Δ x Δy Δ w Δ h It is obtained by random sampling from a uniform distribution U(-δ,δ), where δ represents the disturbance level of noise. As an example, this scheme is set to δ = 0.3 in all experiments.
[0060] C. For each target T in each image in the training dataset, find K-1 targets in the dataset that have the same class label as the target T. Crop out the target based on the bounding box provided by the label, and convert the cropped image to the same resolution as the bounding box of the target T.
[0061] For each target T in the training images, the corresponding label contains the target's location and category information. This scheme uses the category information provided by the label to find K-1 images with targets of the same category from the original dataset. The targets are then cropped from these images based on their corresponding bounding box labels and added to the result set. Specifically, this scheme randomly selects an image from the dataset each time. If the image contains a target of the same category as in label A, the target is cropped based on its corresponding bounding box label. The cropped image is then updated to the same resolution as target T and added to the result set. Otherwise, another image is randomly selected from the dataset until the result set contains K-1 targets of the same resolution.
[0062] D. Combination Figure 2 As shown, an edge-smoothing masking algorithm is used to merge the above images and paste them back into their original positions in the original image.
[0063] For all targets in the result set and the original target T, this scheme fuses them into the same image to increase the probability of corresponding category targets appearing in the target bounding box and reduce the probability of corresponding category targets not appearing in the target bounding box due to noise. To make the fusion smoother and the final result look more natural, this method proposes an edge smoothing mask algorithm to fuse the images. The edge smoothing mask calculation formula is as follows:
[0064]
[0065] Where d i,j It is the distance between the pixel at position (i,j) in the target bounding box and the nearest edge of the bounding box, and its value is min(i,j,h). o -i,w o -j), h o ,w o λ represents the height and width of the cropped target, respectively; β is a threshold that controls the range of the smooth boundary; and λ is a number randomly sampled from the Beta distribution.
[0066] Based on the above, this scheme uses an edge smoothing mask to fuse the images in the result set. The fusion formula is as follows:
[0067]
[0068] Where α is the edge smoothing mask, B a For the image patch corresponding to the original target T, B n For each image patch in the result set, K represents the total number of images to be merged. This involves element-wise multiplication. Finally, the merged image is pasted back into the original image at the location of the target T, thus completing the fusion process.
[0069] The image fusion process described above can significantly reduce the noise rate of both category noise and bounding box noise. Suppose a dataset with a noise rate of P contains a label of category C. The probability that the bounding box actually contains a target of category C is 1-P. When K image patches of the same category are fused, the probability of the fused bounding box containing the target becomes 1-P. K Since P is a number between 0 and 1, the probability of the target being contained in the bounding box will be greatly increased.
[0070] E. The processed image is then fed into a deep learning-based object detector.
[0071] This method uses Faster R-CNN as the base object detection network. Faster R-CNN is a classic two-stage object detection method, whose core components include a basic feature extraction network, a region proposal network, and a region classification and boundary regression network. These networks are used for feature extraction, selecting candidate regions from the feature map, and classifying and regressing bounding boxes for each candidate region. Faster R-CNN achieves efficient object detection through end-to-end training, showing significant improvements in both speed and accuracy compared to traditional two-stage methods. Although this invention primarily uses Faster R-CNN for experiments, since the method mainly modifies the image level, it can also be applied to other object detection networks depending on the task requirements.
[0072] F. Combination Figure 2 As shown, a large loss suppression mechanism is used in the loss function to suppress some large losses that may be caused by label errors, thereby mitigating the negative impact of label noise on model training.
[0073] In the result step E, after fusing multiple targets, this method encounters a problem: due to noise, multiple targets of different categories may appear at a single location, all sharing the same label. Even if the model predicts all targets in the image, those other target classes appearing due to noise, lacking correct labels, will be suppressed as incorrect predictions, leading to a degraded model performance. To mitigate the impact of these "false errors" on model training, this method designs a large loss suppression mechanism to modify the loss function. Specifically, this method categorizes all model predictions into four classes: PB. neg The model predicts that the IoU between the target bounding box and the label bounding box is less than a threshold; PB pp The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, but the predicted category is different from the label category and is not the background; PB fb The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, and the predicted category is background; PB pos The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, and the predicted category is the same as the label; the aforementioned "false error" predictions mainly refer to PB (Physical Boron). pp The included prediction results. Therefore, this method discards these prediction results from the loss function calculation, only calculating the classification loss of the other three parts. The final total loss function becomes:
[0074] L = L bbox +L clsneg +L clspos +L clsfb
[0075] Among them, L bbox This represents the regression loss, which will not be changed in this scheme. clsneg L clspos L clsfb Representing the above PB respectively neg PB pos ,PB fb The corresponding classification loss.
[0076] G. Use the backpropagation algorithm to perform end-to-end training to obtain a trained model.
[0077] During training, in addition to the methods mentioned above, this scheme uses random flipping data augmentation for data preprocessing and normalizes all training and testing data. The feature extraction network is initialized using a ResNet pre-trained on ImageNet, with stochastic gradient descent as the optimizer. The initial learning rate is 0.0025. In epochs 17 and 21, the learning rate is reduced by a factor of 10, the weight decay parameter is 0.0001, and the momentum parameter is 0.9. The mini-batch size is set to 2 for each iteration, and the entire network is trained for 24 epochs. During training, this scheme uses the large loss suppression method mentioned above to modify the loss function. After calculating the loss, gradient descent is used for backpropagation to update the model parameters.
[0078] To test the detection method of this scheme, this embodiment compares the detection method of this scheme with the traditional Faster R-CNN detection model. The OPIXray dataset is used as the test set for the comparison. Both models are trained at the same noise rate (60% bounding box noise and class noise). The results are as follows: Figure 3 As shown, it illustrates the detection results of the proposed model and the traditional Faster R-CNN model on X-ray images. It can be seen that... Figure 3 In the image, (a) shows the detection results of the original Faster R-CNN model. The blue boxes represent the true class labels and bounding boxes, the red boxes represent the predicted labels and bounding boxes, and the numbers in parentheses represent the prediction confidence. Figure 3 In the middle, (b) shows the detection results of the Faster RCNN model after using this method. It can be seen that the quality of the predicted bounding box and the accuracy of the predicted category have been greatly improved after using this method.
[0079] In addition, this study compares the performance of our proposed method with several existing X-ray image target detection methods on the OP IXray dataset, which includes 60% class label noise and 60% target bounding box offset noise. The results are shown in Table 1.
[0080] Faster RCNN 56.7 4h 27min 18.7 LIM 72.7 19h 18min 7.3 SDANet 52.4 6h 20min 16.2 SCE 48.6 4h 28min 19.3 LNCIS 65.9 4h 30min 19.3 OA-MIL 56.4 4h 33min 19.4 This embodiment's solution 81.8 4h 30min 19.0
[0081] Remark:
[0082] Faster R-CNN corresponds to the method proposed by S. Ren et al. (S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Adv. Neural Inform. Process. Syst., 2015.)
[0083] LIM corresponds to the method proposed by R.Tao et al. (R.Tao, Y.Wei,
[0084] SDANet corresponds to the method proposed by L. Zhang et al. (L. Zhang, L. Jiang, R. Ji, and H. Fan, “PIDray: A large-scale X-ray benchmark for real-world prohibited item detection,” in Proc. Int. J. Comput. Vis., 2023.)
[0085] SCE corresponds to the method proposed by Y. Wang et al. (Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, and J. Bailey, “Symmetric cross entropy for robust learning with noisy labels,” in Proc. Int. Conf. Comput. Vis., 2019.)
[0086] LNCIS corresponds to the method proposed by L. Yang et al. (L. Yang, F. Meng, H. Li, Q. Wu, and Q. Cheng, “Learning with noisy class labels for instance segmentation,” in Proc. Eur. Conf. Comput. Vis., 2020.)
[0087] OA-MIL corresponds to the method proposed by C. Liu et al. (C. Liu, K. Wang, H. Lu, Z. Cao, and Z. Zhang, “Robust object detection with inaccurate bounding boxes,” in Proc. Eur. Conf. Comput. Vis., 2022.)
[0088] The training time refers to the time required to train the model on the OP IXray dataset for 24 epochs. Model speed is measured in frames per second (fps), which represents the number of images the model can process per second; a higher value indicates faster speed. Precision is measured in mAP@0.5 (mean average precision), which represents the average precision at an IoU threshold of 0.5. This metric combines precision and recall in IoU (Intersection over Union) calculations.
[0089] The IoU value between two bounding boxes is calculated by dividing the area of their intersection by the area of their union, and is used to measure the quality of the predicted bounding box. The calculation formula is:
[0090]
[0091] The methods for calculating precision and recall are as follows:
[0092] First, all predicted boxes are divided into True Positive (TP): the number of predicted boxes whose IoU with the ground truth is greater than a threshold (0.5) and whose prediction score is greater than the threshold. Each Ground Truth label is counted only once. When multiple predicted boxes have IoU with the Ground Truth greater than the threshold, the value with the largest prediction score is taken.
[0093] False Positive (FP): The number of predicted bounding boxes with an IoU less than the threshold (0.5) plus the number of redundant detection boxes that detect the same Ground Truth.
[0094] False Negative (FN): The number of Ground Truths that were not detected.
[0095] True Negative (TN, not included in the calculation): The number of bounding boxes with IoU less than the threshold and predicted labels as background. Since negative classes will not be labeled in the experiment, this is not important and can be ignored.
[0096]
[0097] Precision, used to determine whether the model's predictions are accurate (precision rate), and recall, used to determine whether the model can find all objects (recall rate). Neither metric alone is sufficient to judge model performance. For example, if the model predicts no objects, the precision is 1; similarly, if the model predicts a sufficient number of bounding boxes, even if the majority are incorrect, the recall can still be 1. Therefore, this approach uses Mean Precision (AP) to test model performance, combining precision and recall.
[0098] Average accuracy via PR curve (reference) Figure 4 The area enclosed by the confidence threshold and the coordinate axis is used as the metric. Multiple recall and precision values are obtained by continuously adjusting the confidence threshold, and then these points are connected by a straight line to obtain the PR curve. The area enclosed by the PR curve and the coordinate axis is the AP value. In actual calculation, the jagged curve is usually smoothed first, replacing the precision corresponding to each recall with the maximum precision to its right. Therefore, the final smoothed curve will monotonically decrease. The average precision (AP) calculated in this way will be less sensitive to small changes. The following figure shows an example of a PR curve. The green curve is the smoothed curve, and the orange curve is the unsmoothed curve. This scheme divides the recall value from 0 to 1 into 11 points (0, 0.1, 0.2, ..., 1.0), and then calculates the average of the precision (the corresponding value on the smoothed curve) for these points as the AP: (AP r (0) represents the precision value when the recall rate is 0.
[0099]
[0100] mAP is calculated by taking the AP value for each category and then averaging them. By measuring mAP, both recall and precision of the model are considered, thus providing a better reflection of the model's performance.
[0101] This approach primarily addresses the issue of noisy labels on the working dataset, where model training performance is severely impacted by noise. To mitigate the effects of noisy labels on model training, this approach proposes a novel data augmentation technique. This technique enables X-ray security image target detection to maintain good performance even under significant noise, greatly improving the model's robustness without modifying its complexity. Specifically, due to noisy labels, the target location in the image corresponding to the label may not contain the correct target. To alleviate this problem, this approach fuses multiple target image patches of the same category during the training phase, significantly increasing the probability of the target object appearing at the labeled location. Furthermore, this image fusion operation can simulate the object occlusion features of X-ray security images, allowing the model to better learn these features. Simultaneously, this approach proposes a large loss suppression strategy. During the fusion process, due to category noise, the fused result may contain other targets. Since these targets lack corresponding correct labels, even if the model correctly predicts them, they will be considered incorrect predictions and suppressed during the loss calculation phase, negatively impacting model training. Therefore, this scheme mitigates the negative impact of these predictions on model training by removing the loss calculation for these predictions during the loss calculation process. The noise-resistant X-ray security image target detection method proposed in this invention achieves excellent results on noisy X-ray security image datasets. Furthermore, this method also achieves good results on general target detection datasets, demonstrating its superiority and versatility.
[0102] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A robust X-ray image target detection method based on augmented data from fusion of similar targets, characterized in that, It includes the following steps: A. Prepare a dataset of X-ray security inspection images, and divide it into a training set and a test set. The image data in both the training set and the test set have category labels and bounding box labels pointing to the target objects. B. Add labeled noise to the image data in the training set to simulate noise conditions; C. For each target object T in each image data in the training set, find K-1 images with the same category label as the target object T from the image data in the dataset. Then, based on the target bounding box label corresponding to the image data, crop out the target object and convert the resolution of the cropped image to the same resolution as the bounding box of the target object T. D. Use the edge smoothing mask algorithm to re-paste the cropped target object image back into the original position of the original image data; E. Input the image data of the training set after step D into a deep learning-based object detector for model training. F. In the loss function of the target detector model, a pre-set large loss suppression mechanism is used to suppress the large loss caused by label errors in image data, thereby mitigating the negative impact of label noise on model training. G. Use the backpropagation algorithm to train the model end-to-end, and obtain the trained model after the model passes the test set. H. Use the trained model to perform target detection in X-ray machine images and analyze the detection results output by the model.
2. The robust X-ray image target detection method based on similar target fusion data augmentation as described in claim 1, characterized in that, In step A, the dataset of X-ray security inspection images prepared includes OPIXray, PIDray, and MS-COCO.
3. The robust X-ray image target detection method based on similar target fusion data augmentation as described in claim 1 or 2, characterized in that, In step B, adding label noise to the image data in the training set to simulate noise includes adding label noise to the category label of the target object and the bounding box label of the target object respectively; One method for adding label noise to the category labels is to replace the original category labels of the image data with labels of any other category in the dataset based on the category noise rate of the dataset. The method for adding label noise to the target bounding box labels is as follows: Offset and scale the original bounding boxes of the image data based on the target bounding box noise rate in the dataset; then adjust the coordinates... Set the bounding box for the image data and perform the following operations: in, From a uniform distribution Obtained by random sampling. Indicates the level of noise disturbance. The coordinates of the bounding box pointing to the target object in the image data. These are the coordinates of the bounding box after noise has been added.
4. The robust X-ray image target detection method based on fusion data augmentation of similar targets as described in claim 3, characterized in that, Step C includes: for each target object T in each image data in the training set, the label corresponding to the target object T contains a target bounding box label pointing to the location information of the target object and a category label indicating the category information of the target object; based on the category information provided by the label, find K-1 images of target objects with the same category information from the original dataset, and crop out the target objects in the image data according to their corresponding target bounding box labels and put them into the result set; In this process, an image is randomly selected from the dataset each time. If the image contains a target object with the same category information as the target object T in its label, the target is cropped out according to its corresponding target bounding box label. The image resolution of the cropped target object is then changed to the same resolution as the target object T and added to the result set. Otherwise, an image is randomly selected from the dataset again until the result set contains K-1 targets with the same resolution.
5. The robust X-ray image target detection method based on similar target fusion data augmentation as described in claim 4, characterized in that, Step D includes: For all target objects in the result set and their corresponding original target objects T, they are fused into the same image to increase the probability of target objects with the corresponding category labels appearing in the target bounding box, thereby reducing the probability of target objects with the corresponding category labels not appearing in the target bounding box due to noise. To make the fusion smoother and the final result look more natural, an edge smoothing mask algorithm is used to fuse the image data. The corresponding calculation formula is as follows: in, The location within the target bounding box The distance between the pixel and the nearest edge of the bounding box, its value is... , These are the height and width of the cropped target, respectively. It is a threshold that controls the range of smooth boundaries. It is a number randomly sampled from the Beta distribution; The images in the result set are fused using an edge-smoothing mask. The fusion formula is as follows: in, For edge smoothing mask, This refers to the image patch corresponding to the original target object T. For each image patch in the result set, K represents the total number of images to be merged. Element-wise multiplication; Finally, the merged image is pasted back into the original image data at the location of the target object T, completing the fusion process.
6. The robust X-ray image target detection method based on fusion data augmentation of similar targets as described in claim 5, characterized in that, In step E, the deep learning-based object detector uses Faster RCNN as the basic object detection network. Its core components include a basic feature extraction network, a region proposal network, and a region classification and boundary regression network, which are used for feature extraction, selecting candidate regions from the feature map, and classifying and regressing bounding boxes for each candidate region, respectively.
7. The robust X-ray image target detection method based on fusion data augmentation of similar targets as described in claim 6, characterized in that, Step F includes: The predictions of the target detector model are divided into four categories: (1) The model predicts that the IoU between the target bounding box and the label bounding box is less than a threshold. (2) The model predicts that the IoU between the target bounding box and the label bounding box is greater than the threshold, but the predicted category is different from the label category and is not the background. (3) The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, and the predicted category is background. (4) The model predicts that the IoU between the target bounding box and the label bounding box is greater than a threshold, and the predicted category is the same as the label. Among them, "false error" prediction is The included prediction results are discarded from the loss function calculation, and only the classification losses of the other three parts are calculated. Finally, the total loss function becomes: in, This represents the regression loss; we will not change it. The above are respectively represented The corresponding classification loss.
8. The robust X-ray image target detection method based on similar target fusion data augmentation as described in claim 7, characterized in that, Step G also includes data preprocessing, which includes: using random flipping data augmentation to normalize the data in all training and test sets; the basic feature extraction network is initialized and trained on ResNet pre-trained on ImageNet, with stochastic gradient descent as the optimizer, wherein the initial learning rate is 0.0025, and the learning rate is reduced by a factor of 10 in the 17th and 21st epochs of training, respectively, with a weight decay parameter of 0.0001 and a momentum parameter of 0.9; the mini-batch size for each iteration is set to 2, and the entire network is trained for 24 epochs; during training, the preset large loss suppression mechanism is used to modify the loss function, and after calculating the loss, the gradient descent algorithm is used for backpropagation to update the model parameters.
9. The robust X-ray image target detection method based on fusion data augmentation of similar targets as described in claim 8, characterized in that, Step H includes: transforming the resolution of the input image to be detected to 1333*800, then inputting it into a trained model for target detection in X-ray security inspection images, using the model to input the extracted features into a regression network to obtain the coordinate values of the detection bounding box, and then inputting them into a classification network to predict the corresponding category and generate the detection result.