Improved YOLOv8 security check image contraband detection method and computer equipment
By introducing the SimIRB attention mechanism module and C2f-DWR module into the backbone and neck network of YOLOv8, and using Dynamic Head and Powerful-IoU loss functions, the problem of insufficient target detection accuracy and low recall in X-ray security image contraband detection is solved, and high-precision and fast contraband detection is achieved.
Patent Information
- Application Number
- CN202510133511.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the target detection accuracy of X-ray security image contraband detection does not meet the standards and the model recall rate is low.
Add the SimIRB attention mechanism module between the C2f module and the SPPF module located in the backbone network of YOLOv8, and use the C2f-DWR module to replace the original C2f module in the neck network, use the Dynamic Head to detect the head, and use Powerful-IoU to lose the loss function.
The recall rate and detection accuracy of the model are improved, and the rapid, accurate and automatic judgment of whether there are contraband in the luggage is achieved, which reduces the false alarm rate and missed detection rate of manual security inspection, alleviates the working pressure of security inspectors and improves work efficiency.
Smart Images

Figure CN120031844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning image detection, and in particular relates to a method and computer equipment for detecting contraband in security inspection images based on an improved YOLOv8. Background Art
[0002] X-ray contraband detection has important value in the field of security, effectively preventing prohibited items from causing danger or interference to transportation. Traditional X-ray security inspection is to let luggage items pass through the X-ray security inspection machine to generate X-ray images, and the on-site manual operator uses the naked eye to identify and judge whether there are prohibited items, so the manual security inspection of the manual operator plays a vital role. Security inspectors must improve their visual inspection expertise through long-term practice and training to cope with it, which results in high personnel and time costs. Although experience and knowledge are important elements for effective inspections, security inspectors will face environmental, physical, and mental influences, and cannot guarantee the stability and accuracy of the inspection work, which poses a safety hazard. Therefore, it is of great significance to study how to apply computer technology to X-ray security inspection images to achieve automatic, fast, and accurate detection and identification of prohibited items.
[0003] With the rapid development of hardware level and detection technology, target detection based on deep learning has become a hot topic for many scholars to study and discuss. Traditional detection methods such as gradient histogram or support vector machine mainly use sliding window to detect targets, which is time-consuming and the manually designed features are not robust. The emergence of deep learning has shifted computer vision research from traditional feature extraction and shallow models to deep network models. Deep learning in convolutional neural networks has high real-time and accuracy, and because it completes target detection through feature learning of a large number of samples, it has good robustness when facing complex image recognition problems. Therefore, target detection algorithms based on deep learning have become the mainstream method in the current field of machine vision.
[0004] The X-ray security inspection image contraband detection method based on deep learning generally adopts the basic idea and framework of general target detection. For example, the YOLO series model is one of the most widely used models, and more and more scholars are considering improving on the basis of existing ideas and frameworks. For example, the Chinese invention patent with application publication number CN117218583A and application publication date 2023.12.12 discloses a security inspection image contraband image detection method based on the YOLOv5-Mobilenet network model. In this method, the YOLOv5 backbone network is built through the inverted residual structure of the MobilenetV2 network and the deep separable convolution of the MobilenetV1 network to obtain the YOLOv5-Mobilenet network model to reduce the amount of model calculation; and the SE attention mechanism is introduced in the inverted residual structure of the MobilenetV2 network of the YOLOv5-Mobilenet network model to improve the detection accuracy of small target contraband, so that the network is more in line with the requirements of lightweight and can improve the detection accuracy of the model for small targets.
[0005] However, due to the unique properties of X-ray security inspection images that are different from natural images, such as cluttered backgrounds, large variations in the shapes and scales of contraband, and severe overlapping and occlusion, the current detection accuracy still cannot meet the needs of practical applications. Summary of the invention
[0006] The purpose of the present invention is to provide a security inspection image contraband detection method and computer equipment that improves YOLOv8, so as to solve the problems of substandard target detection accuracy and low model recall rate in the prior art.
[0007] To solve the above technical problems, the present invention provides a method for detecting contraband in security inspection images based on an improved YOLOv8, the method comprising: obtaining a security inspection image and inputting it into a trained contraband detection model to identify the location and type of contraband in the security inspection image; wherein the contraband detection model is improved on YOLOv8, and the improvement to YOLOv8 comprises: adding a SimIRB attention mechanism module between the C2f module and the SPPF module located at the deepest layer of the network in the backbone network of YOLOv8, and the SimIRB attention mechanism module is a module obtained by replacing the multi-head attention mechanism in the iRMB module with the SimAM attention mechanism.
[0008] Furthermore, the improvement to YOLOv8 also includes: the C2f module in the neck network of YOLOv8 is replaced by a C2f-DWR module, and the C2f-DWR module is improved on the C2f module, and the improvement to the C2f module is to replace the second convolutional layer in the Bottleneck of the C2f module with DWR_Conv, and DWR_Conv includes a convolutional layer, a DWR, a batch normalization layer and an activation function layer connected in sequence.
[0009] Furthermore, the activation function layer is a GeLU activation function layer.
[0010] Furthermore, the improvements to YOLOv8 include: YOLOv8's detection head uses Dynamic Head.
[0011] Furthermore, the loss function used in training the contraband detection model is the Powerful-IoU loss function.
[0012] Further, the types of prohibited items include knives and liquid containers.
[0013] To solve the above technical problems, the present invention also provides a computer device, including a processor, wherein the processor is used to execute a computer program to implement the steps of the above-mentioned improved YOLOv8 security image contraband detection method.
[0014] The present invention is an improved invention creation, and its beneficial effects are as follows: the present invention improves YOLOv8 for contraband detection, specifically, a SimIRB attention mechanism module is added between the C2f module and the SPPF module located in the deepest layer of the network in the backbone network, and the SimIRB attention mechanism module is a module after the multi-head attention mechanism in the iRMB module is replaced with the SimAM attention mechanism, so as to simultaneously capture long-distance dependencies and global dependencies between features to enhance the feature extraction capability of objects, and the two complement each other to achieve excellent feature extraction effects, improve the recall rate of the model, and improve the detection accuracy of the contraband detection model, so as to realize fast, accurate and automatic judgment of whether there are contraband in the luggage. Moreover, the specific setting position of the SimIRB attention mechanism module is between the deepest C2f module and the SPPF module, and the feature integration capability of the SimIRB attention mechanism module is utilized to retain important features, eliminate redundant or unimportant features to avoid the influence of these features on the model detection accuracy, and improve the model detection accuracy. Overall, the present invention significantly improves the accuracy and reliability of contraband detection, greatly reduces the false alarm rate and missed detection rate of manual security checks, greatly alleviates the work pressure of security inspectors and improves work efficiency, and provides technical support for ensuring public transportation safety and passenger personal safety as well as the construction of a public intelligent transportation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a network architecture diagram of the improved YOLOv8 of the present invention;
[0016] Figure 2(a) is a structural diagram of the 1D attention mechanism;
[0017] Figure 2(b) is the structural diagram of the 2D attention mechanism;
[0018] FIG2( c ) is a structural diagram of the 3D-SimAM attention mechanism used in the present invention;
[0019] Figure 3 is a structural diagram of SimIRB of the present invention;
[0020] Figure 4 It is the structural diagram of the dilated residual module;
[0021] Figure 5 is a structural diagram of the C2f-DWR module of the present invention;
[0022] Figure 6 It is a schematic diagram of the specific theoretical implementation of Dynamic Head used in the present invention;
[0023] Figure 7 Detailed structural diagram of the Dynamic Head used in the present invention. DETAILED DESCRIPTION
[0024] The core improvement of the present invention is to improve YOLOv8 for security inspection image contraband detection, and add a SimIRB attention mechanism module between the C2f module and the SPPF module in the deepest layer of the network in the backbone network of YOLOv8. The SimIRB attention mechanism module is a module after the multi-head attention mechanism in the iRMB module is replaced with the SimAM attention mechanism, so as to simultaneously capture long-distance dependencies and global dependencies between features to enhance the feature extraction capability of objects and improve the reliability of model detection. Based on this, an improved YOLOv8 security inspection image contraband detection method and a computer device of the present invention can be realized. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0025] An improved YOLOv8 security image contraband detection method embodiment:
[0026] The present invention proposes an improved YOLOv8 network model, and uses the improved network model to detect prohibited items in X-ray security inspection images. The improved network model is called a prohibited item detection model. After the training of the prohibited item monitoring model is completed, the entire method flow implemented based on the trained model is: obtain the X-ray security inspection image, and input it into the trained contraband detection model, and use the model to identify the location and type of contraband in the X-ray security inspection image. The types specifically include 5 types of knives (blades, daggers, knives, scissors, Swiss Army knives) and 7 types of liquid containers (cans, carton beverages, glass bottles, plastic bottles, vacuum cups, spray cans, tin cans).
[0027] The YOLOv8 network model includes the backbone network Backbone, the neck network Neck and the detection head. Figure 1 As shown, the improvement of the YOLOv8 network model in this embodiment involves four aspects: First, a SimIRB attention mechanism module is added between the C2f module and SPPF (Spatial Pyramid Pooling with Pooling, a pooling structure used in deep learning) in the backbone network Backbone, and the SimIRB attention mechanism module is used to capture long-distance dependencies and global dependencies between features at the same time to enhance the feature extraction ability of objects. Second, the C2f-DWR module based on the dilation-wise residual block (DWR) is used in the neck network Neck to replace the original C2f module, which can more efficiently obtain multi-scale context information and enhance the model's understanding of diverse objects. Third, Dynamic Head is used in the detection head to effectively combine scale, space, and task awareness to improve the model's responsiveness to complex environments and different shapes and sizes of contraband. Fourth, the Powerful-IoU loss function is used to accelerate the convergence speed and obtain accurate bounding box regression. Figure 1 This is the network architecture diagram of the improved model. The following is a detailed introduction to these four aspects.
[0028] 1. SimIRB attention mechanism.
[0029] 1.SimAM.
[0030] In recent years, in the field of computer vision, the effectiveness of attention mechanisms in generating clearer and more refined feature representations has been well demonstrated. Common attention mechanisms include spatial attention mechanisms, channel attention mechanisms, and hybrid attention mechanisms. Channel and spatial attention mechanisms are 1D and 2D attention mechanisms, which either only treat channels differently or only treat spatial positions differently, and both have certain limitations. The hybrid attention mechanism is also a combination of 1D and 2D, but they do not work together.
[0031] The SimAM attention mechanism is a parameter-free 3D attention mechanism, which means that features can be weighted in multiple dimensions, not just channel or spatial dimensions, and the weight of each position can be calculated, thereby improving the interpretability of the model, allowing the model to more clearly understand the importance of each position in the feature map, and improving the performance of the model without increasing network parameters. The basic structures of the existing 1D and 2D attention mechanisms and the SimAM attention mechanism based on the 3D weight method are shown in Figure 2(a), Figure 2(b), and Figure 2(c), respectively.
[0032] The working principle of the SimAM attention mechanism is realized through the energy function, and the definition of the energy function is related to the phenomenon of spatial inhibition in the field of neuroscience. That is, active neurons will inhibit the performance of surrounding information-deficient neurons. Therefore, the importance of neurons can be expressed by the energy function. The expression of the energy function is shown in the following formula:
[0033]
[0034] In the formula, e t represents the energy function; w t and b t are the corresponding weights and biases respectively; y is the label for calculating the discrimination between the target neuron and other neurons; t, x i are the target neuron and other neurons in the input feature map respectively; i is the index in the spatial dimension; M is the number of all neurons in the channel; λ is the regularization coefficient; w is obtained by calculation t and b t Theoretically, each channel has M energy functions. Based on the above formula, we can get t and b t The fast closed-form solution of is:
[0035]
[0036] In the formula, and is to calculate the mean and variance of all neurons in this channel except t. Therefore, the minimum energy for:
[0037]
[0038] In the formula, From formula (4), we can see that the lower the energy, the greater the difference between the t neuron and other neurons, and the more important it is. The SimAM module is finally optimized as:
[0039]
[0040] Where E is The sum over all channels and spatial dimensions; ⊙ represents an element-by-element multiplication operation; X represents the input feature tensor; Sigmoid represents the Sigmoid function.
[0041] 2. SimIRB attention mechanism module.
[0042] The Inverted Residual Block (IRB) is a lightweight CNN infrastructure. The main idea of SimIRB is to innovatively combine the CNN architecture with the SimAM attention mechanism to create a more efficient network. As a hybrid network module, SimIRB integrates the depthwise separable convolution (3x3 DW-Conv) and the SimAM attention mechanism. The structure of the SimIRB attention mechanism is as follows Figure 3 As shown in , it is obtained by improving the iRMB module. Specifically, the multi-head attention mechanism in the iRMB module is replaced by the SimAM attention mechanism, thus obtaining the following Figure 3 The structure shown in Figure 1. Among them, 1x1 Conv plays a key role in the channel dimension. By effectively compressing and expanding the number of channels, it greatly optimizes the computing efficiency, allowing faster calculations and reducing resource consumption when processing X-ray image data. The depthwise separable convolution (3×3DW-Conv) focuses on capturing the spatial features in X-ray images and accurately locating key information such as the contour, shape, and position relationship of objects in the image. The SimAM attention mechanism is committed to capturing the global dependencies between features. It can comprehensively consider the associations between features in various regions of the image, such as identifying the potential connection between contraband and the surrounding environment or other items.
[0043] Moreover, it should be noted that the SimIRB attention mechanism module is set between the deepest C2f module and SPPF to utilize the integration effect of the SimIRB attention mechanism module to retain important features, remove redundant features and unimportant features, and improve positioning accuracy and efficiency.
[0044] 2. C2f-DWR module.
[0045] 1. Dilated Residual (DWR) module.
[0046] Given the significant differences in the scale and shape of contraband, it is particularly important to further improve the network's multi-scale feature extraction and fusion capabilities. The dilated residual module (DWR) structure is introduced, which significantly improves the efficiency of capturing multi-scale information. The DWR structure is an efficient two-step residual feature extraction method, which is divided into two steps: Region Residualization (RR) and Semantic Residualization (SR). Inside the residual, a two-step method is used to efficiently draw multi-scale contextual information, and then the feature maps generated by the multi-scale receptive fields are fused. The basic structure of the DWR structure is as follows: Figure 4 As shown in the figure. The purpose of RR is to generate concise regional feature maps of different sizes from the input features. This step consists of a normal 3×3 convolution layer, a batch normalization (BN) layer, and a ReLU (activation function) layer. The SR step introduces three 3×3 convolution layers with different expansion rates. These layers can effectively expand the receptive field of the convolution kernel and capture contextual information in a wider range, thereby achieving multi-scale feature extraction.
[0047] Through the above two steps, the role of multi-rate deep dilated convolution changes from laboriously acquiring as much complex semantic information as possible to simply performing morphological filtering with a required receptive field on each concisely expressed feature map. Finally, the dimension is adjusted through a 1×1 convolution layer and added to the original input, and a residual connection mechanism is introduced, which helps to alleviate the gradient vanishing problem, allowing the network to train more stably and better learn the residual information between input and output. At the same time, it also enables the model to fuse features extracted at different scales and output the final feature representation.
[0048] 2. C2f-DWR module.
[0049] In the X-ray contraband detection task, in order to build a network architecture that better meets the detection requirements, the present invention innovatively combines the DWR module with the C2f module to form a C2f-DWR module. The structure of the C2f-DWR module is as follows: Figure 5 As shown in Figure 2, the C2f-DWR module is improved on the basis of the C2f module. The second convolutional layer in the Bottleneck of the C2f module is replaced with DWR_Conv. Both the add-False and add-True branches are replaced. The replacement results are shown in Figure 2. Figure 5 The structure of DWR_Conv is shown in the yellow box on the left and the green box in the middle. Figure 5 As shown in the purple box on the right side of the figure, DWR_Conv adds an additional 1×1 convolution layer and batch normalization layer ( Figure 5 BatchNorm2d in) and the activation function layer (specifically, the GeLU activation function can be selected, that is, Figure 5 GELU), which specifically includes 1×1 convolutional layers connected in sequence (i.e. Figure 5 1×1Conv), DWR, batch normalization layer and activation function layer in .
[0050] In the X-ray contraband detection scenario, the batch normalization layer can normalize the data according to the distribution characteristics of the X-ray image data, so that the input data distribution of each layer of the model remains relatively stable during the training process, thereby effectively improving the convergence speed of the model and allowing the model to learn the characteristic patterns of contraband in X-ray images more quickly. The activation function layer introduces nonlinear factors into the model, enhancing the model's ability to fit the distribution of complex X-ray image data, enabling it to accurately capture the characteristic performance of contraband in different forms, densities, etc., while reducing the risk of overfitting caused by complex and changeable data, ensuring that the model can maintain high accuracy and stability when facing a variety of X-ray contraband detection tasks. Through this combination, efficient reuse and significant enhancement of features are achieved, which has effectively promoted the improvement of model training efficiency and the expansion of generalization capabilities.
[0051] 3. Dynamic Head.
[0052] In the object detection task, the original detection head of YOLOv8 has some limitations, including the following aspects: First, when dealing with targets of different sizes, it is impossible to effectively balance the detection effects of small targets and large targets, resulting in poor detection effects on targets in certain scale ranges. Secondly, when dealing with complex backgrounds or overlapping targets, the shapes, positions, and inclinations of the targets may vary, which makes it difficult to accurately locate these targets. In addition, target detection tasks may require different feature representations, and the original detection head cannot flexibly adapt to the needs of different tasks. The present invention uses Dynamic Head to solve the above problems and improve the model's responsiveness to complex environments and situations where contraband has different shapes and sizes.
[0053] Dynamic Head contains three different attention mechanisms: scale-aware attention, space-aware attention, and task-aware attention. Each mechanism focuses on a different perspective. The specific theoretical implementation of Dynamic Head is as follows: Figure 6 shown.
[0054] Given a feature layer F, if self-attention is applied to the feature, there is a formula:
[0055] W(F)=π(F)·F (6)
[0056] In the formula, F represents the input feature, π(·) is the attention function, and the naive attention function solution is implemented through a fully connected layer. Due to the high dimension of the tensor, directly learning the attention function of all dimensions is not only computationally expensive, but also practically unbearable. Therefore, the attention is divided into three dimensions, and each attention only focuses on one perspective. The resulting formula is:
[0057] W(F)=π C (π S (π L (F)·F)·F)·F (7)
[0058] In the formula, π L , π S and π C They represent the attention functions applied to three different dimensions: L, S, and C. Each dimension is introduced in turn below:
[0059] Scale-aware attentionπ L :Give different weights to feature layers at different levels, so that the model can adaptively fuse according to the importance of features at that level:
[0060]
[0061] Where f(·) is a linear approximation implemented by a 1×1 convolutional layer; It is a hard-sigmoid function, where S and C represent the scale of the feature and the quantity related to the channel dimension, respectively.
[0062] Spatially-aware attentionπ S : Use deformable convolution for sparse sampling and then aggregate features of each level at the same spatial position:
[0063]
[0064] Where L represents the number of feature levels; K is the number of sparse sampling locations; ω l,k represents the weight coefficient; l represents the level of the feature; c represents the channel dimension of the feature; p k +Δp k is the position after displacement; Δm k is the position p k importance.
[0065] Task-aware attentionπ C: To achieve joint learning and generalize different task representations of the target, task-aware attention is deployed at the end. It dynamically switches the on and off channels of features to support different tasks.
[0066] π C (F)·F=max(α 1 (F)·F c +β 1 (F),α 2 (F)·F c +β 2 (F)) (10)
[0067] In the formula, F c is the slice of the feature in the cth channel; [α 1 ,α 2 ,β 1 ,β 2 ] = θ(·) is a number that is learned to control the activation threshold; θ(·) first performs global average pooling on L×S dimensions to reduce the dimensionality, then uses two fully connected layers and a normalization layer, and finally applies a shifted sigmoid function to normalize the output to [-1, 1].
[0068] In X-ray contraband detection, the scale-aware attention of Dynamic Head can accurately learn the features of contraband of different sizes, reducing misjudgments and omissions caused by scale differences. Spatial-aware attention can enhance the distinction of spatial features of contraband and reduce confusion with complex backgrounds. Task-aware attention allows the model to flexibly focus on feature channels according to different detection tasks, maintaining a balance between positioning and recognition. The three work together to effectively improve the accuracy and robustness of detection. Finally, since the above three attention mechanisms are applied sequentially, π L , π S and π C The detailed configuration of Dynamic head is as follows Figure 7 shown.
[0069] 4. Powerful-IoU.
[0070] Object localization is a key task in object detection and relies heavily on the evaluation and optimization of the bounding box regression (BBR) loss function. Therefore, the bounding box regression loss function plays an important role in object detection. The original YOLOv8 model uses CIoU as the initial loss function. CIoU is an improved loss function based on IoU. Not only the overlapping area of the predicted box and the true box is considered, but also their center distance and geometric factors such as aspect ratio. The calculation formula of the CIoU loss function is as follows:
[0071]
[0072]
[0073] Where CIOU represents the CIOU loss value; IOU represents the IOU loss value; ρ(b,b gt ) is the center point b of the predicted box and the center point b of the real box gt The Euclidean distance between the two boxes; c is the diagonal length of the rectangle circumscribing the two boxes; w and h are the width and height of the prediction box respectively; w gt 、h gt are the width and height of the real frame respectively.
[0074] However, the existing IoU-based loss function still has unreasonable penalty factors, which leads to the expansion of the anchor box and slow convergence during the regression process. In addition, in the X-ray contraband detection task, low-quality samples are often caused by factors such as environmental complexity. The shape of the target may vary greatly, resulting in unstable aspect ratio. Using aspect ratio as a metric may impose too much penalty on these low-quality samples, thereby reducing the generalization ability of the model. In order to overcome these limitations, the present invention uses Powerful-IoU (PIoU) to replace the CIOU in the original YOLOv8, which combines a penalty factor with the target anchor size as the denominator and a function that adapts to the quality of the anchor. This method can guide the anchor box to effectively regress along a more direct path, thereby accelerating the convergence speed and improving the accuracy, so that it can better adapt to changes in the contraband detection scenario.
[0075] PIoU proposes a penalty factor P that adapts to the target size. The penalty factor is defined as follows:
[0076]
[0077] Where, dw 1 、dw 2 dh 1 dh 2 They represent the absolute value of the distance between the corresponding edges of the prediction box and the target box; w gt 、h gt Represent the width and height of the target box respectively. The penalty factor P is only associated with the size of the target box, and has nothing to do with the size of the anchor box and the minimum external box of the target box. Even if the anchor box is enlarged, it will not be affected, which improves the adaptability to the target size.
[0078] To further improve the loss function, PIoU proposed a penalty function adapted to the quality of the anchor box after referring to the characteristics of the ideal loss function listed by EIoU:
[0079]
[0080] When the penalty factor P is large, it indicates that there is a significant difference between the anchor box and the true box, and the f(P) value is small, which suppresses harmful gradients from low-quality anchor boxes. When P is near 1, it indicates that the anchor box is near the true box, and a larger f(P) will accelerate regression. When P approaches 0, f(P) gradually decreases as the quality of the anchor box increases, and the anchor box is continuously optimized to align with the true box.
[0081] The penalty function makes the medium-quality anchor boxes have the largest gradient, allowing these anchor boxes to quickly regress to the vicinity of the target box to become high-quality anchor boxes, so that the target detector can focus on the anchor boxes with moderate quality. Finally, the PIoU and loss calculation method can be obtained, namely:
[0082] PIoU=IoU-f(P),-1≤PIoU≤1 (16)
[0083] L PIoU =1-PIoU=L IoU +f(P),0≤L PIoU ≤2 (17)
[0084] So far, the whole method flow has been clearly introduced.
[0085] The above method is applied to specific practical contraband detection to help the modern security inspection industry to quickly and accurately identify contraband, which is introduced in detail below.
[0086] 1) Data preparation.
[0087] SIXray: Built by the Pattern Recognition and Intelligent System Development Laboratory of the University of Chinese Academy of Sciences, this dataset is a subway security inspection dataset with image-level category annotations provided by security inspectors, suitable for real-time classification, detection, and segmentation applications. In real-world detection scenarios, contraband often appears at extremely low frequencies. This dataset has a total of 1,059,231 images, and the positive samples in the database account for 0.84% of the total samples, which can truly reflect the security inspection scene. Among them, there are 8,929 positive samples containing contraband, and there are 5 categories of contraband that can be used for target detection, namely guns, knives, wrenches, pliers, and scissors. The experimental data is divided into training and validation sets in a ratio of 8:2.
[0088] CLCXray: Jointly constructed by Tongji University, Beijing University of Posts and Telecommunications, and University of Chinese Academy of Sciences, the dataset contains 9565 X-ray images, of which 4543 X-ray images (real data) are from real subway scenes, and 5022 X-ray images (simulated data) are from manually designed luggage scans. All images are acquired using a similar X-ray scanner (TECHIK, model THXS6550). All labels are marked by 8 junior employees (less than 5 years of work experience and students) and reviewed by 2 senior employees (more than 5 years of work experience). The dataset has 12 categories, including 5 types of knives (blades, daggers, knives, scissors, Swiss Army knives) and 7 types of liquid containers (cans, carton beverages, glass bottles, plastic bottles, vacuum cups, spray cans, tin cans), where the dataset is divided into training and test sets, with a ratio of about 8:2.
[0089] 2) Model training.
[0090] This study conducted experiments on the Windows 10 operating system. The GPU used for training was NVIDIA GeForceRTX 4060Ti, the CPU was a 12th Gen Intel(R)Core(TM)i5-12600KF 3.70GHz processor, and the training environment was configured with PyTorch 2.3.1, Python 3.8.19, and CUDA 12.1. During the training of the X-ray contraband detection model, the parameter settings are shown in Table 1.
[0091] The dataset is converted into a txt label format suitable for the YOLO model; combined with YOLOV8, the SimIRB attention mechanism is introduced at the Backbone end, the C2f-DWR module is used to replace the original C2f module in Neck, the detection head is changed to DynamicHead, and the loss function is improved to Powerful-IoU; the improved YOLOv8 model is trained with the training set to obtain the X-ray security image detection model; the X-ray security image detection model is tested with the test set.
[0092] Table 1
[0093]
[0094] 3) Model evaluation.
[0095] In order to evaluate the performance of the model, Precision (P), Recall (R), mean average precision (mAP) and F1score are selected as standard evaluation indicators. The prediction results are divided into positive samples and negative samples. The instance is a positive sample and the prediction is a positive sample is defined as TP, the instance is a negative sample and the prediction is a positive sample is defined as FP, and the instance is a negative sample and the prediction is a negative sample is defined as FN. The definitions of these indicators are as follows:
[0096]
[0097] The accuracy rate is based on the prediction results and indicates the proportion of correct predictions in the positive samples. The recall rate is based on the actual results and indicates the proportion of correct predictions in the actual positive samples. The average accuracy (AP) is equal to the area under the accuracy-recall curve. The mean average precision (mAP) is the result of the weighted average of the AP values of all sample categories. It is used to measure the detection performance of the model in all categories. The higher the mAP, the better the overall performance of the model. AP(i) in formula (21) represents the AP value of the category index value i, and n represents the number of categories of samples in the training data set. The F1 value is an indicator that comprehensively measures the Precision (P) and Recall (R) of the model. The higher its value, the better the balance between the precision and recall of the model. These indicators comprehensively evaluate the algorithm and are important evaluation indicators for improving the performance of X-ray contraband detection.
[0098] 4) Ablation experiment.
[0099] In order to prove the effectiveness of the proposed module in the model, this section conducts an ablation study on the SIXray dataset. The experimental strategy of each group of experiments remains consistent. When each improvement strategy is applied to the baseline model, it improves the detection performance to varying degrees. √ indicates that this improvement strategy is used. The experimental results are shown in Table 2.
[0100] Table 2
[0101]
[0102] The results of ablation experiments show that each improvement strategy improves the detection performance to varying degrees when applied to the baseline model. After replacing the original CIoU with Powerful-IoU, the penalty factor and the function adapted to the anchor quality guide the anchor box to effectively regress along a more direct path, which speeds up the convergence speed and improves the accuracy. The mAP50 is increased from 90.3% to 91.1%, and the accuracy is improved by 2.2%. The SimIRB attention mechanism module is introduced into the backbone network. SimIRB improves the focus on key information in the feature map, thereby achieving a certain improvement in mAP50 and accuracy. In the neck network, the C2f-DWR module is used to replace the original C2f module, which significantly improves the efficiency of capturing multi-scale information, and the accuracy and mAP50 are increased by 3.1% and 1.1% respectively. After using Dynamic Head, the scale, space, and task awareness are effectively combined, and all indicators are greatly improved, especially the recall rate and mAP50, which are increased by 3.2% and 3% respectively.
[0103] As the modules are gradually superimposed on the model, all indicators are rising. There is a synergistic effect between the improved modules, resulting in better performance. When all four modules are used together in the network, the model achieves the best performance, with a 3.8% increase in mAP50, a 3.7% increase in precision, a 4.6% increase in recall, and a 4.1% increase in F1 value compared to the baseline model. These improvements enable the model to detect contraband more accurately and have high reliability and usability in practical applications.
[0104] 5) Comparative experiment.
[0105] In order to prove the superiority of the model, comparative experiments were conducted with several mainstream target detection models, including Faster R-CNN, SSD, YOLOv5, YOLOv6, YOLOv10, YOLO11 and the original model YOLOv8. The results are shown in Table 3. The indicators of the proposed algorithm far exceed the one-stage classic model. Compared with the mainstream YOLO series, this model achieves the best results. The mAP50, precision, recall and F1 of the latest model YOLO11 are 4.2%, 3.9%, 5.7% and 4.9% higher.
[0106] Table 3
[0107]
[0108] A computer device embodiment:
[0109] A computer device embodiment of the present invention includes a memory, a processor, an internal bus and a computer program stored in the memory, and the processor and the memory communicate and exchange data with each other through the internal bus. The processor executes the computer program to implement the steps of the method described in the embodiment of the security inspection image contraband detection method of improving YOLOv8 of the present invention. Among them, the processor can be a processing device such as a microprocessor MCU, a programmable logic device FPGA, etc.; the memory can be various memories that use electrical energy to store information, such as RAM, ROM, etc., or it can be a memory using other methods.
[0110] In summary, the present invention aims at the problems of large object scale variation, random placement and overlap in X-ray security inspection images, and proposes an X-ray contraband detection method based on the improved YOLOv8 architecture to improve the detection accuracy of contraband and reduce the probability of missed detection. This model has strong robustness and generalization, and can quickly, accurately and automatically determine whether there are contraband in backpack luggage, and determine its corresponding position and type, which greatly reduces the false alarm rate and missed detection rate of manual security inspection, and also greatly relieves the work pressure of security inspectors and improves work efficiency, and provides technical support for ensuring public transportation safety and passenger personal safety and the construction of public intelligent transportation system, and has extremely important research and application value. The present invention provides an efficient and accurate detection tool for the security inspection field, which can meet the needs of high precision and high recall in X-ray contraband detection tasks, provide support and help in the security inspection industry, meet the strict requirements of real-world security applications, and promote the progress and development of the security inspection field.
[0111] Specific implementation methods are given above, but the present invention is not limited to the described implementation methods. The basic idea of the present invention lies in the above basic scheme. For ordinary technicians in this field, it does not take creative work to design various deformed models, formulas, and parameters according to the teachings of the present invention. Changes, modifications, substitutions, and variations of the implementation methods without departing from the principles and spirit of the present invention still fall within the scope of protection of the present invention.
Claims
1. An improved YOLOv8 security image contraband detection method, characterized in that: The method comprises: acquiring a security inspection image and inputting it into a trained contraband detection model to identify the location and type of contraband in the security inspection image; Among them, the contraband detection model is improved on YOLOv8. The improvements to YOLOv8 include: adding a SimIRB attention mechanism module between the C2f module and the SPPF module located in the deepest layer of the network in the backbone network of YOLOv8. The SimIRB attention mechanism module is a module that replaces the multi-head attention mechanism in the iRMB module with the SimAM attention mechanism.
2. The method for detecting contraband in security inspection images based on improved YOLOv8 according to claim 1, characterized in that: The improvements to YOLOv8 also include: the C2f module in the neck network of YOLOv8 is replaced by the C2f-DWR module, and the C2f-DWR module is improved on the C2f module, and the improvement to the C2f module is to replace the second convolutional layer in the Bottleneck of the C2f module with DWR_Conv, and DWR_Conv includes a convolutional layer, DWR, a batch normalization layer, and an activation function layer connected in sequence.
3. The method for detecting contraband in security inspection images based on improved YOLOv8 according to claim 2, characterized in that: The activation function layer is a GeLU activation function layer.
4. The method for detecting contraband in security inspection images based on improved YOLOv8 according to claim 1, characterized in that: Improvements to YOLOv8 also include: YOLOv8's detection head uses Dynamic Head.
5. The method for detecting contraband in security inspection images based on improved YOLOv8 according to claim 1, characterized in that: The loss function used to train the contraband detection model is the Powerful-IoU loss function.
6. The method for detecting contraband in security inspection images based on the improved YOLOv8 according to any one of claims 1 to 5, characterized in that: Types of prohibited items include knives and liquid containers.
7. A computer device comprising a processor, characterized in that: The processor is used to execute a computer program to implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Security check picture contraband image detection method and device based on YOLOv5-Mobilenet network model and computer storage medium
CN117218583A
Cited By
Unmanned missile loading vehicle road defect detection method and related device
CN121032954A