Improved X-ray security check image dangerous goods high-precision detection system combined with YOLOv8-seg

By improving the SPPF-LSKA and MC-Conv modules of the YOLOv8-seg model, combined with data enhancement strategies, the problem of contraband detection in X-ray security inspection is solved, and high-precision and efficient automated security inspection is achieved.

CN120339581APending Publication Date: 2025-07-18HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510420809.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-06
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When existing X-ray security inspection technology deals with complex scenarios and blocked items, it is difficult to achieve high-precision detection of contraband, and the manual identification method is inefficient, which can easily lead to fatigue and false detection.

Method used

Combining the YOLOv8-seg model, by improving the SPPF-LSKA module and MC-Conv module, the C2f-MSC module is built to enhance the multi-scale feature extraction and global correlation capabilities, and combined with data enhancement strategies, the generalization ability and robustness of the model are improved.

Benefits of technology

It significantly improves the detection accuracy and robustness of occluded hazardous items in X-ray security images, and can efficiently identify multiple categories of contraband in complex scenarios, reduce calculation complexity, adapt to heterogeneous materials and multi-angle interference, and achieve high-precision automated security inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339581A_ABST
    Figure CN120339581A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of X-ray hazardous article detection, and discloses an X-ray hazardous article detection algorithm based on improved YOLOv8-seg, which can detect the types and specific areas of hazardous articles in disordered security check articles. The method comprises the steps that a multi-channel interactive convolution module is designed, and based on a channel decoupling and recombination mechanism and a heterogeneous convolution kernel collaborative optimization strategy, the multi-scale feature discrimination capability is enhanced while calculation redundancy is reduced; the method comprises the following steps: constructing an SPPF-LSKA (Space Purpose Pooling Function-Learning SKA) composite module, and effectively improving the boundary retention degree of an occluded target through deep fusion of a large-kernel space attention mechanism and adaptive pyramid pooling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and object detection. Background Art

[0002] With the continuous improvement of the living standards of the general public, passengers' requirements for travel safety, service quality and efficiency are increasing day by day. As an important link in ensuring passenger safety, using X-ray security inspection equipment to check whether there are prohibited items in passengers' luggage has become an effective security inspection method. However, the traditional method of identifying prohibited items relies on manual judgment, and there are two main problems with manual identification: on the one hand, the objects in X-ray images are often blocked by other items, making it difficult to identify prohibited items; on the other hand, most X-ray images do not contain prohibited items, which can cause fatigue and decreased attention of security inspectors after long hours of work, thus affecting the accuracy and efficiency of identification. Therefore, it is particularly important to establish a real-time and accurate X-ray luggage security inspection system.

[0003] In terms of complex scene processing, The BoVW (Bag of Visual Words, BoVW) detection framework constructed by [authors' names] et al. includes three key steps: First, the DoG (Difference of Gaussians, DoG) Gaussian difference detector is used to locate key points and extract scale-invariant feature transform (SIFT); Subsequently, a visual dictionary is constructed through K-means clustering, and the matching histogram between the feature descriptor and the dictionary is calculated; Finally, an SVM (Support Vector Machine, SVM) classifier is trained to complete the image classification and retrieval task. For multi-view sequence analysis, Mery et al. developed a dynamic recognition system based on structure from motion, established the correspondence between different views through a multi-view geometric model, and performed joint segmentation and feature description on potential targets in sequence images, effectively solving the problem of incomplete single-view information. However, in contrast, deep learning-based algorithms can automatically learn and extract multi-layer semantic information in security inspection images, can express image content more completely, and improve the recognition accuracy. Liu et al. proposed a few-shot learning-based contraband detection model based on edge detection and reverse verification. This model pays attention to the edge information of occluded objects, designs an edge detection module, fuses the features output by the edge detection module with the high-level semantic information generated by the Transformer encoder, and proposes a reverse verification strategy to assist training, thereby improving the accuracy of the final output. Mademlis et al. conducted a comparative experiment on three common object detection models. The experimental results show the superiority of the DINO (DETR with Improved deNoising anchor boxes, DINO) method with CSPDarkNet-53 as the backbone on the SIXray dataset. However, the authors only considered the mAP metric and did not consider other performance parameter metrics, which only shows that the DINO method has a certain advantage in terms of accuracy. In addition, they integrated SimAM attention and deformable convolution into the downsampling stage of the neck to improve key feature capture. However, these above detection algorithms can only perform preliminary classification and rough localization of contraband items in security inspection X-ray images and cannot obtain the shape information of contraband items. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention proposes a high-precision detection system for dangerous goods in X-ray security inspection images improved by combining YOLOv8-seg. The traditional Spatial Pyramid Pooling Fusion (SPPF) module of the YOLOv8-seg model is combined with Large Separable Kernel Attention (LSKA) and improved into a SPPF-LSKA (Spatial Pyramid Pooling Fusion-Large Separable Kernel Attention) module fusion. By integrating the multi-scale local feature extraction ability of SPPF and the axial decomposition large kernel global perception advantage of LSKA, long-range spatial dependence relationships are constructed while retaining the details of small targets. Moreover, a novel Multiple Channel Convolution (MC-Conv) module is proposed and combined with the C2f module to be improved into a C2f-MSC module. By separating the original decay features and multi-scale context information through a channel decoupling strategy, local details and global semantics are extracted in parallel by heterogeneous convolution, and a dynamic weight adaptive fusion mechanism is used to significantly improve the feature discrimination ability of dangerous goods with complex materials in X-rays.

[0005] A high-precision detection system for dangerous goods in X-ray security inspection images improved by combining YOLOv8-seg according to the present invention includes:

[0006] Step 1: Perform data preprocessing and enhancement on the X-ray security inspection image dataset, which is beneficial to improving the quality and diversity of the dataset, reducing bias, enhancing the generalization ability of the model, covering more scenarios, and significantly improving the detection accuracy;

[0007] Step 11: The present invention uses the PIDray X-ray security inspection image dataset, which contains 12 types of typical dangerous goods (such as knives, guns, explosives, flammable liquids, etc.) in complex stacking and semi-occluded scenarios. The original annotation format of the PIDray dataset is converted from JSON to the txt format required for YOLO training. The bounding box coordinates are normalized to the range [0-1] and converted into the standard format of "category index x1,y1,x2,y2,...,x n ,y n " to ensure compatibility with the YOLOv8-seg object detection framework;

[0008] Steps 1-2: Divide the PIDray dataset in the ratio of 8:2. The training set contains 50,472 images and the test set contains 12,618 images. The 8:2 division strategy aims to ensure that there are sufficient samples in the training phase to learn the feature expressions of complex scenarios, while ensuring that the test set is large enough to comprehensively evaluate the generalization performance of the model;

[0009] Steps 1-3: Perform data augmentation using horizontal flipping and set the parameter (flip) to 0.5 to force the model to learn the orientation invariance features of the image content, significantly enhancing the model's robustness to object orientation changes;

[0010] Steps 1-4: Perform data augmentation using Scale Transformation and set the parameter (scale) to 0.5 to enable the model to adapt to significant changes in the target size, which is of great significance for detecting the same object at different shooting distances;

[0011] Step 2: Combine the SPPF module and the LSKA module of the YOLOv8-seg model and improve it to the SPPF-LSKA module fusion; propose the MC-Conv module, combine it with the C2f module, and improve it to the C2f-MSC module to build an improved YOLOv8-seg network;

[0012] Steps 2-1: Combine the traditional SPPF module of the YOLOv8-seg model with the LSKA and improve it to the SPPF-LSKA module fusion. The SPPF of YOLOv8 adds a feature fusion mechanism on the basis of SPP, which improves the perception and detection performance to a certain extent, but still lacks attention to the global context, especially in the case of multi-object overlap and redundant local details common in X-ray images.

[0013] LSKA first uses a depth convolutional layer to obtain local spatial features, reduces parameter coupling through channel-independent calculations; then introduces an atrous convolutional layer to expand the feature capture range, and uses an interval sampling mechanism to break through the perception limit of conventional convolution; a special axial decomposition design decouples the traditional two-dimensional convolutional kernel into cascaded horizontal and vertical one-dimensional convolutional kernels, significantly reducing the computational dimension while maintaining the completeness of feature extraction. Through the above three-stage decomposition strategy, efficient feature extraction is achieved.

[0014] By fusing the multi-scale local feature extraction ability of SPPF and the axial decomposition large-kernel global perception advantage of LSKA, long-range spatial dependence relationships are constructed while retaining small target details, significantly enhancing the cross-region correlation feature modeling ability of occluded dangerous goods in X-ray images, and maintaining a low computational complexity;

[0015] Step 22: Propose the MC-Conv module, which is used to achieve multi-scale interaction in the cross-channel dimension while reducing the computational complexity. Combine it with the C2f module to construct an adaptive feature extraction architecture for X-ray security inspection imaging characteristics - the C2f-MSC module.

[0016] Separate the original attenuation features and multi-scale context information through the channel decoupling strategy, combine heterogeneous convolutions to extract local details and global semantics in parallel, and use the dynamic weight adaptive fusion mechanism to effectively decouple the material boundary fusion phenomenon caused by the energy spectrum hardening effect, significantly improving the feature discrimination ability of X-ray complex material dangerous goods, and providing a more robust solution for multi-category contraband detection in the security inspection scenario;

[0017] Step 3: Use the PIDray training set to systematically train the improved YOLOv8-seg model to obtain a model that can efficiently and accurately detect X-ray security inspection images;

[0018] Step 4: Use the PIDray test set to comprehensively test the trained improved YOLOv8-seg model, test the generalization ability and robustness of the model, and accurately measure its performance in actual application scenarios.

[0019] Furthermore, in the present invention, in Step 1, an X-ray security inspection image dataset is selected for the training and testing of the model. Through this deep utilization and data augmentation processing of the public dataset, the training data can be greatly enriched, reducing the model performance bottleneck caused by data deviation, thereby significantly enhancing the generalization ability of the model, enabling it to exhibit excellent detection accuracy and stability when facing various complex and changing actual security inspection scenarios.

[0020] Furthermore, in the present invention, in Step 11, first convert the original annotation format of the PIDray dataset from JSON to the TXT format required for YOLO training, convert the original annotation format of the PIDray dataset from JSON to the txt format required for YOLO training, normalize the bounding box coordinates to the range of [0-1] and convert them to the standard format of "category index x1,y1,x2,y2,...,x n ,y n " to ensure compatibility with the YOLOv8-seg object detection framework; the standardized coordinate system enhances the robustness of the model to image size changes, and the fine annotation system for dangerous goods categories provides a semantically complete supervision signal for multi-scale object recognition in complex stacking scenarios, ensuring the cross-device migration ability of the algorithm in actual security inspection scenarios and laying a reliable data foundation for subsequent model architecture innovation.

[0021] Furthermore, in the present invention, in steps one and two, the PIDray dataset is divided in a ratio of 8:2. Among them, the training set contains 50,472 images, and the test set contains 12,618 images. This dataset division strategy, through a scientific training-test ratio configuration, ensures that the model fully learns the feature expressions of complexly stacked and semi-occluded objects in X-ray images, while guaranteeing the generalization verification of unknown scenarios. Strict sample isolation avoids the risk of data leakage, provides a reliable evaluation benchmark for multi-class dangerous goods detection algorithms, supports the robustness optimization of the model in real scenarios such as heterogeneous materials and multi-angle interference, and helps the smooth transition of the high-precision detection system from theoretical verification to industrial-grade security applications.

[0022] Table 1 PIDray Dataset Structure Division Table

[0023]

[0024] Furthermore, in the present invention, in step thirteen, a horizontal flipping operation based on probability distribution is implemented. By setting a flipping probability of 0.5, a bidirectional feature mapping space is constructed, forcing the model to simultaneously learn the feature expressions of the original and mirrored images during training, effectively suppressing the weight allocation of direction-sensitive features, and strengthening the ability to extract essential features such as object shape and material. This bidirectional feature learning mechanism forms a collaborative optimization with the multi-scale attention module, enabling the attention mechanism to preferentially focus on key regions with direction invariance when dynamically allocating weights on feature maps at different levels, thereby enhancing the robustness of feature selection.

[0025] Furthermore, in the present invention, in step fourteen, a Scale Transformation with a parameter of 0.5 is adopted to randomly scale the image size, introducing the diversity of target sizes in the training data and forcing the model to learn the cross-scale feature expression ability. This scale enhancement strategy is deeply coupled with the FPN structure of YOLOv8, generating multi-scale samples on feature maps of different resolutions, and prompting the attention mechanism to effectively capture the scale change laws of targets at all levels of the feature pyramid. Through this dual enhancement strategy, the system can construct a more generalizable feature representation space, enabling the model to maintain a stable feature extraction ability when facing complex situations such as random placement of items and significant size differences in X-ray security inspection scenarios.

[0026] Furthermore, in the present invention, in Step 2, the SPPF module of the YOLOv8-seg model is combined with the LSKA module and improved into a fused SPPF-LSKA module; the MC-Conv module is proposed and combined with the C2f module to be improved into a C2f-MSC module, and an improved YOLOv8-seg network is constructed. The former overcomes the problem of cross-region semantic association of complex occluded targets through the collaborative optimization of multi-scale local perception and axial large-kernel global modeling; the latter, based on the channel decoupling and dynamic heterogeneous convolution fusion mechanism, effectively suppresses the material feature coupling interference caused by spectral hardening. The two work together to achieve multi-level feature expression from pixel-level edge fidelity to target-level semantic decoupling, providing a lightweight solution for high-precision real-time detection of multi-category and heterogeneous material dangerous goods in complex stacking scenarios, and promoting the intelligent security inspection system to cross over to industrial-level reliability;

[0027] Furthermore, in the present invention, in Step 21, the traditional SPPF module of the YOLOv8-seg model is combined with LSKA and improved into a fused SPPF-LSKA module. Although the traditional SPPF module enhances the local feature expression ability through multi-scale pooling fusion, its fixed-scale pyramid structure is difficult to model the cross-region semantic relevance, especially in X-ray images with dense target stacking and blurred edges, which is prone to feature breakage and semantic fragmentation. The LSKA module, through the cascaded design of axial decomposition of large kernels, disassembles the traditional two-dimensional large-kernel convolution into a sequence of horizontal and vertical one-dimensional convolution operations. While significantly reducing the number of parameters, it constructs a long-range dependence network covering a larger spatial range. Its dilated convolution and depthwise separable strategy further balance the contradiction between receptive field expansion and computational efficiency. The innovative fusion of the two enables the SPPF-LSKA module to embed the global attention mechanism of axial large kernels in the multi-scale local feature pyramid, forming a two-way enhanced link of "microscopic detail capture - macroscopic semantic association": on the one hand, the multi-level pooling structure of SPPF finely preserves the edge texture of small targets and the gradient information of heterogeneous materials, avoiding the feature annihilation of tiny dangerous goods caused by downsampling; on the other hand, the axial large kernel of LSKA reconstructs the topological continuity of occluded targets through cross-channel spatial interaction, effectively suppressing the interference of local noise on global semantics. This multi-granularity coupling mechanism of local and global features not only solves the contradiction that it is difficult to achieve both detail fidelity and semantic coherence in traditional single-path feature extraction, but also provides an extensible lightweight solution for multi-scale target detection in complex security inspection scenarios under industrial computing resource constraints, promoting the collaborative optimization of the model in three dimensions: the recognition accuracy of occluded targets, the discrimination of heterogeneous materials, and the real-time inference efficiency, and building a core technical barrier for the precise interception of high-risk contraband.

[0028] Furthermore, in the present invention, in step 22, an MC-Conv module is proposed and combined with the C2f module to construct a C2f-MSC module. The C2f-MSC module proposed in the present invention constructs an adaptive feature decoupling system for complex physical imaging characteristics in response to the feature coupling problem caused by the energy spectrum hardening effect of multi-material targets in the X-ray security inspection scenario. Traditional convolution operations are vulnerable to the interference of differences in X-ray attenuation coefficients and energy spectrum drift in heterogeneous material representations, resulting in artifact fusion of the boundary features of metals, liquids, and composite dangerous goods, seriously weakening the semantic discrimination ability of the model. The C2f-MSC module separates the original attenuation signal from the multi-scale context information through a channel decoupling strategy, breaking through the limitation of a single feature stream in representing complex physical effects - the heterogeneous convolution branches capture local micro-textures and global topological structures in parallel with different receptive fields to form a complementary feature space; the dynamic weight fusion mechanism then realizes the adaptive calibration of cross-scale features through learnable channel attention, suppressing the blurring of material boundaries caused by energy spectrum hardening at the pixel level. This collaborative architecture of "physical perception feature decoupling - cross-scale semantic enhancement" has achieved three major breakthroughs: First, by establishing a mapping relationship between the material attenuation characteristics and the feature expression space, the edge sharpness of heterogeneous targets in stacked packages and the internal structure identification accuracy are significantly improved; second, the dynamic weight mechanism endows the model with an adaptive compensation ability for X-ray energy spectrum drift, effectively alleviating the feature distribution shift caused by equipment differences or scanning parameter fluctuations; third, the lightweight heterogeneous convolution design avoids the computational redundancy of traditional large-kernel convolutions while ensuring the completeness of multi-granularity feature extraction. This innovation not only provides a cross-modal feature decoupling paradigm for the accurate detection of high-risk prohibited items such as metal knives and liquid explosives in the security inspection scenario, but also promotes the embedded deployment of high-robustness algorithms in edge X-ray imaging devices through industrial-level computational efficiency and hardware compatibility optimization, laying the core technical foundation for building an all-weather, fully automatic intelligent security inspection ecosystem;

[0029] Furthermore, in the present invention, in step three, after completing the relevant preliminary preparation work, it enters the crucial model training stage. The purpose of this stage is to carry out systematic training on the improved YOLOv8-seg model using the divided training set, so as to obtain a high-quality model that can efficiently and accurately detect X-ray security inspection images. First, modify the cfg file of YOLOv8, including modifying the classes in the yaml file of the data to the number of classes labeled in the dataset. Set the hyperparameters of the network model, including the input image size when training the dataset and testing the model performance, the input data volume batch per batch, the number of training epochs, the learning rate lr0, and lrf. Then, input the training set data into the improved YOLOv8-seg model, and strictly follow the established training strategy to let the model gradually learn the feature information of various targets in the X-ray security inspection images, including the unique manifestations of the shapes, contours, and materials of the items under the X-ray images. At the same time, closely monitor various performance indicators of the model during the training process, such as the change of the loss function, accuracy, recall rate, etc. By continuously adjusting the parameters of the model, the model can better fit the training data. Through this systematic training, we can more deeply understand the impact of different training strategies on the model performance, so as to find the optimal training plan and obtain a model that can perform excellently in the X-ray security inspection image detection task, achieving efficient and accurate detection effects.

[0030] Furthermore, in the present invention, in step four, use the divided test set to conduct a comprehensive and in-depth multi-dimensional verification on the improved YOLOv8-seg model to verify the generalization ability of the model. Calculate core indicators such as the mean average precision (mAP) and frames per second (FPS), and combine the confusion matrix and heat map visualization to carefully analyze the localization and classification performance of the model in the object detection task, providing a scientific basis for model iteration and actual application deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flowchart of the method for detecting dangerous goods in X-ray security inspection images of the improved YOLOv8-seg in the method of the present invention.

[0032] Figure 2 is a schematic diagram of the network structure of the improved YOLOv8-seg for detecting dangerous goods in X-ray security inspection images in the method of the present invention.

[0033] Figure 3 is a schematic diagram of the network structure of the SPPF module of the improved YOLOv8-seg in the method of the present invention.

[0034] Figure 4 is a schematic diagram of the network structure of the SPPF-LSKA module of the improved YOLOv8-seg in the method of the present invention.

[0035] Figure 5 It is a schematic diagram of the MC-Conv module network structure of the improved YOLOv8-seg in the method of the present invention.

[0036] Figure 6 It is an example diagram of the redundant feature map of the improved YOLOv8-seg in the method of the present invention.

[0037] Figure 7 It is a visualization result diagram of the ablation experiment training process of the improved YOLOv8-seg in the method of the present invention.

[0038] Figure 8 It is a visualization diagram of the detection results of the improved YOLOv8-seg and other models in the method of the present invention.

[0039] Figure 9 It is a comparison diagram of the heat maps of the improved YOLOv8-seg and YOLOv8n in the method of the present invention. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0041] The high-precision dangerous goods detection method for X-ray security inspection images improved in combination with YOLOv8-seg under this specific implementation manner has a flowchart as Figure 1 shown, and includes the following steps:

[0042] Step a, adopt the PIDray X-ray security inspection image dataset and perform data preprocessing and enhancement.

[0043] Step a1, in this specific implementation manner, adopt the publicly available dataset PIDray dataset, convert the original annotation format of the PIDray dataset from JSON to the TXT format required for YOLO training, normalize the bounding box coordinates to the range of [0-1] and convert them to the standard format of "category index x1,y1,x2,y2,...,x n ,y n " to ensure compatibility with the YOLOv8-seg object detection framework;

[0044] Step a2, in this specific implementation manner, divide the dataset into a training set and a test set. The training set uses train, and the test set uses test. The data distribution of the PIDray dataset is shown in Table 2.

[0045] PIDray Dataset: The PIDray X-ray security inspection image dataset covers 12 types of typical dangerous goods (batons, pliers, hammers, power banks, scissors, wrenches, guns, bullets, sprayers, handcuffs, knives, lighters) in multi-angle scanning, complex stacking, and semi-occlusion scenarios, containing 63,090 images. It is used to train and evaluate the automatic segmentation and recognition capabilities of computer vision models for contraband in complex scenarios. Its annotation information covers the location, category, and occlusion status of objects, supporting the performance verification and optimization of algorithms in security inspection scenarios. Among them, the training set contains 50,472 images (80%), and the test set contains 12,618 images (20%).

[0046] Table 2 Data Distribution of the PIDray Dataset

[0047]

[0048] Step a3: In this specific embodiment, horizontal flipping is used for data augmentation, and the parameter (flip) is set to 0.5 to force the model to learn the orientation invariance features of the image content, significantly enhancing the model's robustness to object orientation changes;

[0049] Step a4: In this specific embodiment, Scale Transformation is used for data augmentation, and the parameter (scale) is set to 0.5, enabling the model to adapt to significant changes in the target size, which is of great significance for detecting the same object at different shooting distances;

[0050] Step b: Combine the SPPF module and the LSKA module of the YOLOv8-seg model and improve it to the SPPF-LSKA module fusion; propose the MC-Conv module, combine it with the C2f module, and improve it to the C2f-MSC module to construct an improved YOLOv8-seg network. The structure diagram of the improved YOLOv8-seg network is as Figure 2 shown.

[0051] Step b1: In this specific embodiment, in the X-ray detection scenario, in order to handle complex object overlaps and severe occlusions, it is usually necessary to pay attention to both local details and global information at the same time. Traditional SPP modules obtain multi-scale feature information through multi-scale pooling. However, when dealing with small objects or occluded areas, these modules still face challenges in fully extracting global information. The SPPF of YOLOv8 adds a feature fusion mechanism on the basis of SPP, which improves the perception and detection performance to a certain extent, but still lacks attention to the global context, especially in the case of multi-object overlaps and redundant local details common in X-ray images. The SPPF network structure is as Figure 3As shown. LSKA first uses a deep convolutional layer to obtain local spatial features, reducing parameter coupling through channel-independent calculations; then introduces a dilated convolutional layer to expand the feature capture range, and uses an interval sampling mechanism to break through the perception limit of conventional convolutions; a special axial decomposition design decouples the traditional two-dimensional convolutional kernel into cascaded horizontal and vertical one-dimensional convolutional kernels, significantly reducing the computational dimension while maintaining the completeness of feature extraction. Through the above three-stage decomposition strategy, efficient feature extraction is achieved. To address the problem of long-range dependence modeling caused by severe occlusion in X-ray security inspection images, the traditional SPPF module of the YOLOv8-seg model is combined with LSKA and improved into a SPPF-LSKA module fusion. The SPPF-LSKA network structure is as Figure 4 shown. By integrating the multi-scale local feature extraction ability of SPPF and the axial decomposition large-kernel global perception advantage of LSKA, long-range spatial dependence relationships are constructed while retaining the details of small targets, significantly enhancing the cross-region correlation feature modeling ability of occluded dangerous goods in X-ray images, and maintaining a low computational complexity;

[0052] Step b2: To address the severe occlusion challenges in X-ray security inspection images, we propose MC-Conv. X-ray images usually contain complex texture patterns and stacked objects, which makes object detection particularly difficult. Our method aims to reduce the redundancy of feature maps, reduce computational complexity and the number of parameters, while enhancing the network's ability to capture multi-scale information, which is crucial for detecting occluded objects. As Figure 5 shows the structural design of the MC-Conv layer.

[0053] In object detection tasks, feature maps after convolutional operations usually contain redundant information. These feature maps include both high-frequency edge information and low-frequency overall contour information. By reasonably performing cross-channel segmentation and interaction on these features, the computational amount can be reduced while improving the recognition ability of occluded targets. As Figure 6 shown in the example of redundant feature maps, in the YOLOv8 network, it can be observed from the feature maps extracted from the first convolutional layer that there are multiple similar feature maps. We call this pair of similar feature maps "phantom feature maps" because one feature map can be generated by a linear transformation of the other feature map.

[0054] Therefore, we propose an MC-Conv module to achieve multi-scale interaction in the cross-channel dimension while reducing computational complexity. The specific principle is as follows:

[0055] Let the input feature map be first evenly divided into two parts along the channel dimension:

[0056]

[0057] where, X cheapThe original features are retained and only passed through the residual connection, with a computational complexity of 0; X complex It is further split into two sub-channel groups:

[0058]

[0059] Among them, X3 and X5 are processed using 3×3 and 5×5 convolutions respectively to form Y3 and Y5:

[0060] Y3 = Conv 3×3 (X3) (4)

[0061] Y5 = Conv 5×5 (X5) (5)

[0062] Finally, X cheap , Y3, and Y5 are concatenated along the channel dimension and fused through 1×1 convolution:

[0063] Y = Conv 1×1 ([X cheap ; Y3; Y5]) (6)

[0064] In the YOLOv8 object detection framework, we construct an adaptive feature extraction architecture for X-ray security inspection imaging characteristics through the MC-Conv and C2f modules. Aiming at the core problems such as the superposition of multi-material attenuation projections, Compton scattering noise interference, and target edge blurring commonly existing in X-ray penetration imaging, this design adopts a channel decoupling strategy to divide the input features into a native path and a multi-scale analysis path: the former retains the original attenuation coefficient distribution features to maintain the information integrity in a low signal-to-noise ratio environment, and the latter parallelly extracts local detail gradients and wide-domain spatial correlations through heterogeneous convolution groups (3×3 and 5×5 kernels), effectively decoupling the material boundary fusion phenomenon caused by the energy spectrum hardening effect. The dynamic feature fusion stage realizes feature fusion through 1×1 convolution, adaptively balancing the high-frequency edge response of metal products and the contribution degree of the continuous attenuation features in the organic matter region.

[0065] To verify the effect of its complexity reduction, we will now conduct a comparative analysis of the computational complexity between the MC-Conv model and the traditional convolution model.

[0066] Assume that the dimension of the input feature map is C×H×W. The computational complexity of performing 3×3 and 5×5 convolutions on a quarter of the channels is:

[0067]

[0068] The complexity of the traditional full-channel multi-scale convolution method is:

[0069] o direct =(9 + 25)C·H·W (9)

[0070] The total complexity of MC-Conv is as follows:

[0071]

[0072] The second term corresponds to the 1×1 convolution operation. It can be seen that while maintaining feature diversity, MC-Conv significantly reduces the computational cost.

[0073] Step c1: In this specific embodiment, modify the cfg file of the improved YOLOv8-seg and change the 'classes' in the data yaml file to the number of classes in the dataset.

[0074] Step c2: In this specific embodiment, the hyperparameters of the network model, including the input image size of 640×640 for training the dataset and testing the model performance, the input data volume batch per batch is 32, the number of training epochs is 300, and the learning rate lr is 0.01.

[0075] Step c3: In this specific embodiment, the invention uses the evaluation system in the object detection method to evaluate the model, including the precision AP, recall, F1-score, accuracy (Precision), and mean average precision mAP@0.5 (mean average precision, IoU threshold is 0.5) for each class. The calculation formulas are as follows:

[0076]

[0077] TP represents the number of correctly identified positive samples, TN represents the number of correctly identified negative samples, FP is the number of negative samples misidentified as positive samples, and FN is the number of positive samples misidentified as negative samples. The F1 score can be further calculated using accuracy and recall. f TP refers to the correctly identified positive samples, f TN is the correctly identified negative samples, f TN is the misidentified negative samples. The average precision (Average Precision, AP) for each target class can be calculated through the area enclosed by the curve (P-R curve) formed by accuracy (Precision) and recall (Recall) and the coordinate axes.

[0078] For the instance segmentation task of prohibited items in X-ray security inspection images, this study verified the effectiveness of the algorithm improvement through systematic controlled variable experiments. The ablation experiment results are shown in Table 3.

[0079] Table 3 Ablation Experiment Results

[0080]

[0081] In the ablation experiment design, the baseline model (original YOLOv8) achieved an mAP50 index of 80.8% without introducing any improved modules. When the C2f-MSE module was integrated alone, the model achieved an mAP50 of 82.6% through a multi-scale feature fusion mechanism, an increase of 1.8 percentage points; the introduction of the SPPF-LSKA module increased the mAP50 by 1.1 percentage points to 81.9% through a hierarchical spatial kernel aggregation strategy. When the two types of improved modules work together, the model shows a significant performance synergy effect, with an mAP50 of 84.8%, an increase of 4 percentage points over the baseline model. At the same time, the simultaneous optimization of average precision (83.2%) and recall rate (80.3%) verifies the robustness of the algorithm in dense occlusion scenes.

[0082] Visual analysis of the training process further reveals the advantageous characteristics of the improved model. Figure 8 As shown in a), the mAP curve of the improved model enters the stable convergence stage after 150 epochs, 37% earlier than the baseline model in terms of training cycles (reduced from 240 epochs to 150 epochs), and the peak mAP50 reaches 84.8%, 5.2 percentage points higher than the original model. The loss curve is shown in Figure 8 b), the loss after improvement is lower than the original YOLOv8 model.

[0083] Fine-grained category evaluation shows that the improved algorithm exhibits differentiated detection performance for prohibited items of different forms. Batons with high contrast features achieve the best index (mAP50 = 91.2%, Dice coefficient = 0.89), and their regular shape stabilizes the contour prediction error within 3.2 pixels. However, due to the flat structure and frequent occlusion of knife-like items, the recall rate is only 78.4%, and 72% of its false detection cases are due to the overlap with items such as metal buttons. Handcuffs, with clear geometric features, break through 65% in the mAP50-95 index, and the Jaccard index of its spatial distribution prediction reaches 0.83, which is significantly better than other categories. The algorithm achieves an average accuracy of 84.3% and a mAP50-95 of 63.0% on 12 types of dangerous goods in the PIDray dataset, proving that it has the generalization ability to handle multi-scale and multi-form prohibited items, and provides a reliable technical solution for the intelligent upgrade of X-ray security inspection systems.

[0084] Table 4 Analysis of training results for each category on the PIDray dataset

[0085]

[0086] Step d: Use the test set to test the trained improved YOLOv8-seg model.

[0087] Step d1: In this specific embodiment, to verify the detection performance of the improved YOLOv8 algorithm on X-ray dangerous goods images, this study conducted a comparative experiment on the PIDray dataset by comparing the proposed algorithm with mainstream object detection algorithms such as YOLOv11, the original YOLOv8, and YOLACT. The experiment used Precision, Recall, mAP50, and mAP50-95 as evaluation metrics, and the results are shown in Table 5.

[0088] Table 5 Comparison of experimental results between improved YOLOv8 and mainstream segmentation algorithms

[0089]

[0090] As shown in Table 5, the improved YOLOv8 algorithm shows significant performance improvement on the PIDray dataset. Specifically, the algorithm outperforms YOLOv11 and the original YOLOv8 in all four core metrics: the Precision reaches 83.5%, which is 3.2 and 3.7 percentage points higher than YOLOv11 (80.3%) and YOLOv8 (79.8%) respectively; the Recall is increased to 80.8%, which is 1.3% (vs YOLOv11) and 2.9% (vs YOLOv8) higher than the comparison algorithms. In terms of localization accuracy, the improved model achieves excellent results of 84.7% and 63.1% in the mAP50 and mAP50-95 metrics respectively, showing obvious advantages over the baseline model (such as the corresponding metrics of YOLOv8 are 81.4% / 59.6%).

[0091] Step d1: In this specific embodiment, to systematically evaluate the model optimization effect, this study verified the effectiveness of the improvement strategy through multi-dimensional visualization comparison experiments. Figure 8 Shows the performance evolution process in scenarios of occluded object detection, single-object recognition, and multi-object mutual occlusion. The first column is the visualization image of the dataset annotation, the second column is the detection result of the baseline model, the third column is the detection result after adding C2f-MSC, the fourth column is the detection result after adding SPPF-LSKA, and the fifth column is the detection result of using both modules simultaneously. Figure 9 Then, the optimization mechanism of the algorithm's attention area is intuitively revealed through heatmap analysis.

[0092] 1. Performance analysis of single-object detection in complex scenarios

[0093] For the single-object scenario, such as Figure 9(As shown in (a), the confidence level achieved by the baseline model is 84.0%. The Cf-MSC module increased the confidence level to 86% through the boundary detail enhancement network; by integrating two improvement strategies, a confidence level of 89.0% was achieved, verifying the effectiveness of the local-global feature joint optimization mechanism.)

[0094] 2. Occluded Object Detection Performance Analysis

[0095] In the occluded scenario, as Figure 8 (shown in (b), due to insufficient feature extraction ability, the baseline model did not detect the occluded scissors. By introducing the C2f-MSC module, the model successfully detected the scissors, but the achieved confidence level was 29.0%, which was relatively low. Finally, through the synergistic effect of the two improvement points, the model successfully detected the scissors with a confidence level of 48.0%, which was 19% higher than the case with one improvement point.)

[0096] 3. Multi-scale Multi-object Detection Performance Analysis

[0097] In the multi-object occlusion experiment, Figure 8 (as shown in (c), the baseline model only detected one of the dangerous goods, Baton, and only achieved a confidence level of 36.0%. After introducing the C2f-MSC module, two dangerous goods were detected, namely Baton and Bullet, with confidence levels of 54.0% and 39.0% respectively. After introducing the SPPF-LSKA module, the confidence level of Baton increased to 62.0%; the final model successfully achieved full object recognition, fully verifying the key role of the C2f-MSC and SPPF-LSKA modules in complex scenarios and providing a better solution for occluded object detection.)

[0098] The heatmaps of the baton and bullet are as Figure 9 shown Figure 9 ((b) is the heatmap obtained by the baseline model, Figure 9 (c) is the heatmap obtained by the improved model. The baseline model has a significant feature dispersion problem, and the hot regions are relatively scattered, with multiple scattered small hot spots, and the algorithm's attention is relatively blurred, indicating relatively weak recognition ability. The improved model realizes cross-scale feature interaction through the heterogeneous kernel group parallel computing of the SPPF-LSKA module. The detected hot regions are significantly concentrated within the target object region, with the focus being precise and the range being clear, indicating that the model can effectively focus on the target features.)

[0099] The experimental results show that the collaborative optimization of local feature extraction and global dependency modeling can effectively improve the performance of occluded object detection; the multi-scale attention mechanism can significantly improve the feature focusing ability in complex scenarios; the improved model's comprehensive performance improvement in single-object, multi-object, and occluded scenarios verifies the universality of the algorithm design. Subsequent research will explore a dynamic weight allocation mechanism to further optimize the computational efficiency.

[0100] Although the present invention elaborates on the implementation manner of the technical solution through specific embodiments, it should be clear that these technical examples are only used to illustrate the core principles and potential application scenarios of the present invention. Based on this, any person skilled in the art can adjust the structural parameters, optimize the process, or reorganize the functional modules of the embodiments, and can also reconstruct other technical solutions, as long as the technical essence does not deviate from the protection scope defined by the claims. In particular, it should be noted that the combination methods of the dependent claims in the claim system are not limited to the logical relationships described in the specification, and the technical features disclosed in different embodiments can be mutually compatible and called. For example, the feature components recorded in a certain embodiment can form an innovative combination with the algorithm process in other embodiments, and this integration of technical elements across embodiments also belongs to the scope of innovative practices protected by this patent.

Claims

1. A high-precision detection system for dangerous goods in X-ray security inspection images improved by combining YOLOv8-seg, characterized in that, It includes the following steps: Step 1: Perform data preprocessing and enhancement on the X-ray security inspection image dataset, which is beneficial to improving the quality and diversity of the dataset, reducing bias, enhancing the generalization ability of the model, covering more scenarios, and significantly improving detection accuracy; Step 1: The present invention uses the PIDray X-ray security inspection image dataset, which contains 12 types of typical dangerous goods (such as knives, guns, explosives, flammable liquids, etc.) in complex stacking and semi-occluded scenarios. The original annotation format of the PIDray dataset is converted from JSON to the txt format required for YOLO training. The bounding box coordinates are normalized to the range of [0-1] and converted into the standard format of "class index x1,y1,x2,y2,...,x n ,y n " to ensure compatibility with the YOLOv8-seg object detection framework; Step 2: Divide the PIDray dataset in a ratio of 8:

2. The training set contains 50,472 images, and the test set contains 12,618 images. The 8:2 division strategy aims to ensure that there are sufficient samples in the training stage to learn the feature expressions of complex scenarios, while ensuring that the test set is large enough to comprehensively evaluate the generalization performance of the model. Step 3: Use horizontal flipping for data augmentation processing and set the parameter (flip) to 0.5 to force the model to learn the orientation invariance features of the image content, significantly enhancing the robustness of the model to object orientation changes; Step 4: Use Scale Transformation for data augmentation and set the parameter (scale) to 0.5 to enable the model to adapt to significant changes in the target size; Step 5: Combine the SPPF module of the YOLOv8-seg model with the LSKA (Large Separable Kernel Attention, LSKA) module and improve it to the SPPF-LSKA module fusion; propose the MC-Conv (Multiple Channel Convolution, MC-Conv) module, combine it with the C2f module, and improve it to the C2f-MSC module to construct an improved YOLOv8-seg network; Step 6: Combine the traditional SPPF module of the YOLOv8-seg model with LSKA and improve it to the SPPF-LSKA module fusion. The SPPF of YOLOv8 adds a feature fusion mechanism on the basis of SPP, which improves the perception and detection performance to a certain extent, but still lacks attention to the global context, especially in the case of multi-object overlap and redundant local details common in X-ray images. LSKA first uses a depth convolutional layer to obtain local spatial features, reduces parameter coupling through channel-independent calculation; then introduces a dilated convolutional layer to expand the feature capture range, and uses an interval sampling mechanism to break through the perception limit of conventional convolution; a special axial decomposition design decouples the traditional two-dimensional convolutional kernel into cascaded horizontal and vertical one-dimensional convolutional kernels, significantly reducing the computational dimension while maintaining the completeness of feature extraction. Through the above three-stage decomposition strategy, efficient feature extraction is achieved. By fusing the multi-scale local feature extraction ability of SPPF and the axial decomposition large-kernel global perception advantage of LSKA, long-range spatial dependence relationships are constructed while retaining small target details, significantly enhancing the cross-region correlation feature modeling ability of occluded dangerous goods in X-ray images, and maintaining a low computational complexity; Step 2. Propose the MC-Conv module, which is used to achieve multi-scale interaction in the cross-channel dimension while reducing the computational complexity. Combine it with the C2f module to construct an adaptive feature extraction architecture for X-ray security inspection imaging characteristics, namely the C2f-MSC module, and separate the original attenuation features and multi-scale context information through the channel decoupling strategy. Combine heterogeneous convolutions to parallelly extract local details and global semantics, and use the dynamic weight adaptive fusion mechanism to effectively decouple the material boundary fusion phenomenon caused by the energy spectrum hardening effect, significantly improving the feature discrimination ability of dangerous goods with complex materials in X-rays, and providing a more robust solution for multi-category contraband detection in security inspection scenarios; Step 3. Use the PIDray training set to systematically train the improved YOLOv8-seg model to obtain a model that can efficiently and accurately detect X-ray security inspection images. Step 4. Use the PIDray test set to comprehensively test the trained improved YOLOv8-seg model, verify the generalization ability and robustness of the model, and accurately measure its performance in actual application scenarios.

2. A high-precision detection system for dangerous goods in X-ray security inspection images improved by combining YOLOv8-seg, characterized in that , in the data preprocessing step: Divide the PIDray dataset according to 8:

2. The training set contains 50,472 images, and the test set contains 12,618 images. The 8:2 division strategy ensures that the model has a strong generalization evaluation ability for unknown samples while fully learning the complex features of X-ray images through a reasonable training-test ratio configuration. Its strict data isolation principle effectively avoids the risk of information leakage and provides a reliable verification benchmark for the optimization direction of multi-scale object detection models in security inspection scenarios.

3. A high-precision dangerous goods detection system for X-ray security inspection images improved by combining YOLOv8-seg, characterized in that, In the model improvement step: Combine the traditional SPPF module of the YOLOv8-seg model with LSKA and improve it to the SPPF-LSKA module fusion. By fusing the multi-scale local feature extraction ability of SPPF and the axial decomposition large kernel global perception advantage of LSKA, this design is based on a lightweight architecture, deeply integrating local detail perception and global spatial reasoning, which not only enhances the edge information fidelity of tiny dangerous goods but also reconstructs the target semantic coherence through cross-region attention interaction, providing an efficient computational paradigm for multi-scale feature decoupling in complex occlusion scenarios. Propose the MC-Conv module and combine it with the C2f module to construct the C2f-MSC module. Through the dynamic feature decoupling and cross-scale semantic calibration mechanism of heterogeneous convolutions, the problem of feature confusion at the boundaries of multi-material objects caused by the energy spectrum hardening effect in X-ray security inspection is overcome. This design reconstructs the feature representation space through channel decoupling, and effectively suppresses the artifact coupling interference between heterogeneous materials through the complementarity of local texture focusing and global semantic perception, establishing a multi-granularity feature expression system for the accurate identification of metals, liquids, and composite contraband in stacked packages.

4. A high-precision dangerous goods detection system for X-ray security inspection images improved by combining YOLOv8-seg, characterized in that, In Step 3, adjust the model configuration file and set hyperparameters. The specific steps are as follows: The image input size is 640×640, the input data volume per batch (batchsize) is 32, the learning rate is 0.01, and the number of training epochs is 300. The present invention uses evaluation metrics widely recognized in the field of object detection to measure the performance of the model, specifically including recall, precision, F1-score, and mean average precision mAP@0.5 (mean average precision, IoU threshold is set to 0.5).

5. A high-precision dangerous goods detection system for X-ray security inspection images improved by combining YOLOv8-seg, characterized in that, In step four, the trained model is evaluated using the test set. The specific steps are as follows: The present invention applies the improved YOLOv8-seg model, as well as mainstream object detection models such as YOLOv11, the original YOLOv8, and YOLACT, to the PIDray dataset for hazardous material detection. The experimental results show that the improved detection model of the present invention has achieved varying degrees of improvement in detection accuracy and speed compared to other mainstream methods.

Citation Information

Cited By

  • Underwater biological target detection method based on deep learning

    CN121366343A

  • Crushed material and waste material sorting method and system based on improved YOLOv8

    CN122265663A