A security check contraband detection method based on feature fusion and semantic interaction
By employing a method that integrates full-process feature fusion and semantic-dependent interaction, the accuracy and speed issues of contraband detection in X-ray images with complex backgrounds and severe occlusions were resolved, achieving efficient and accurate security inspection of contraband.
Patent Information
- Application Number
- CN202211556622.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-12-06
AI Technical Summary
Existing methods for detecting prohibited items in security checks are insufficient in terms of accuracy and speed when faced with complex backgrounds and severely obscured X-ray images, especially in public places where they are inefficient and expensive.
A method for detecting contraband in X-ray images for security checks is adopted based on full-process feature fusion and semantic dependency interaction. By constructing a coordinated fusion backbone network, an adaptive fine-tuning fusion pyramid network, and a selective dense feature interaction network, multi-scale and multi-level features are fully integrated. Semantic interaction is achieved by combining local convolution and visual Transformer, and deep-level contextual interaction is carried out.
It significantly improves the accuracy and speed of detecting prohibited items in security checks. The algorithm achieves an accuracy of 82.4% on security X-ray image datasets, meeting the requirements for real-time detection and demonstrating strong generalization and robustness.
Smart Images

Figure CN115937774B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a security inspection method for detecting prohibited items based on deep learning and convolutional neural networks. Background Technology
[0002] Public security checks are the process of assessing the access and control of visitors according to established standards or guidelines. They play a crucial role in maintaining public safety by promptly identifying potential risks in public places such as airports and train stations. For passenger baggage inspections, professional security personnel typically identify each piece of baggage in X-ray images scanned by security scanners to detect prohibited items such as controlled knives and firearms. However, this manual inspection method is often inefficient, time-consuming, and expensive due to factors such as the subjective role of security personnel and the quality of security equipment. With the development of artificial intelligence and deep learning, deep learning-based methods for detecting prohibited items in X-ray images are increasingly being applied to security systems, making it possible to build automated, accurate, and efficient security inspection methods and establish intelligent and information-based transportation systems.
[0003] Most research works regard the problem of identifying prohibited items in security checks as a target detection task in the field of computer vision, and propose various deep learning-based methods for detecting prohibited items, which greatly improves the detection efficiency of prohibited items in security checks. Many researchers use deep convolutional neural networks to learn and extract features of prohibited items in X-ray images, and try to analyze and understand the semantic information contained in the features, so as to effectively identify prohibited items with different postures, types and so on. Reference [1] proposes an asymmetric convolutional multi-view detection network, using a small convolutional asymmetric network and a multi-scale feature map fusion strategy to improve the accuracy of dangerous item identification in X-ray security check images under background interference. Reference [2] proposes a single-stage composite network detection algorithm, which combines a composite backbone network and introduces a feature enhancement module, achieving a balance between accuracy and detection speed on a specific dataset. Reference [3] constructs a large-scale dataset with millions of images and proposes a class-balanced hierarchical refinement model component to solve the problem of extreme imbalance between positive and negative sample classes in prohibited item images. Reference [4] published a large-scale security inspection image dataset for real-world scenarios and proposed a selective dense attention network module to focus on detecting deliberately hidden prohibited items.
[0004] The above research has made great progress in the task of detecting prohibited items in security checks, but it is only for specific security check image datasets. Based on the general object detector, it proposes plug-and-play modules or attention enhancement components to solve some of the prohibited item detection problems. Existing object detection algorithms have greatly improved the visual representation ability of images and have made continuous progress in detection accuracy and speed. References [5] and [6] redesigned a simpler and more elegant network structure, introduced structural reparameterization and other training techniques, and further improved the detection accuracy of the model while reducing the model complexity. Reference [7] proposed a cyclic feature pyramid and optional dilated convolution to promote the fusion between multi-scale features and improve the accuracy of object detection. Reference [8] divided the image feature map into non-overlapping rectangular windows to perform self-attention and obtained the global receptive field by moving the window to perform cross-window interaction. Reference [9] developed a cross-shaped window self-attention mechanism to calculate the self-attention of parallel horizontal and vertical stripes in the cross-shaped feature window and improve the image feature extraction ability.
[0005] References
[0006] [1] Zhang Youkang, Su Zhigang, Zhang Haigang, Yang Jinfeng. Multiscale detection of contraband in X-ray security inspection images [J]. Signal Processing, 2020, 36(07):1096-1106.
[0007] [2] Guo Shouxiang, Zhang Liang. Yolo-C: Contraband Detection in X-ray Images Based on Single-Stage Networks [J]. Progress in Laser & Optoelectronics, 2021, 58(08):75-84.
[0008] [3] C.Miao, L.Xie, F.Wan, C.Su, H.Liu, J.Jiao, and Q.Ye, "SIXray: A large-scale security inspection X-ray benchmark for prohibited item discovery inoverlapping images," in IEEE Conf.Comput.Vis.Pattern Recog., 2019.
[0009] [4] B.Wang, L.Zhang, L.Wen, X.Liu, and Y.Wu, "Towards real-world prohibited item detection: Alarge-scale X-ray benchmark," in Int.Conf.Comput.Vis., 2021
[0010] [5]C Y.Wang,A.Bochkovskiy,and HYM.Liao,“YOLOv7:Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,”in arXivpreprint arXiv:2207.02696,2022.
[0011] [6]C.Li,L.Li,H.Jiang,K.Weng,Y.Geng,L.Li,Z.Ke,Q.Li,M.Cheng,W.Nie,Y.Li,B.Zhang,Y.Liang,L.Zhou,X.Xu,X.Chu,X.Wei,and X.Wei,“YOLOv6:ASingle-StageObject Detection Framework for Industrial Applications,”in arXiv preprintarXiv:2209.02976,2022.
[0012] [7]S.Qiao,L.-C.Chen,and A.Yuille,“Detectors:Detecting objects withrecursive feature pyramid and switchable atrous convolution,”in IEEEConf.Comput.Vis.Pattern Recognit.,2021.
[0013] [8]Z.Liu,Y.Lin,Y.Cao,H.Hu,Y.Wei,Z.Zhang,S.Lin,and B.Guo,“Swintransformer:Hierarchical vision transformer using shifted windows,”inInt.Conf.Comput.Vis.,2021.
[0014] [9]Dong X,Bao J,Chen D,W.Zhang,N.Yu,L.Yuan,D.Chen,and B.Guo,“Cswintransformer:Ageneral vision transformer backbone with cross-shaped windows,”in IEEE Conf.Comput.Vis.Pattern Recog.,2022.
[0015]
[10] J.Glenn.,yolov5,2022.[Online].Available:https: / / github.com / ultralytics / yolov5.
[0016]
[11] A.Dosovitskiy,L.Beyer,A.Kolesnikov,et al,“An image is worth 16x16words:Transformers for image recognition at scale,”in arXiv preprint arXiv:2010.11929,2020.
[0017]
[12] Y.Cao,J.Xu,S.Lin,F.Wei,and H.Hu,“GCNet:Non-local networks meetsqueeze-excitation networks and beyond,”in Int.Conf.Comput.Vis.Workshop.,2019.
[0018]
[13] Liu S,Huang D,Wang Y.Learning spatial fusion for single-shotobject detection[J].arXiv preprint arXiv:1911.09516,2019.
[0019]
[14] Liu S,Qi L,Qin H,et al.Path aggregation network for instancesegmentation[C] / / Proceedings of the IEEE conference on computer vision andpattern recognition.2018:8759-8768. Summary of the Invention
[0020] This invention addresses the problems of complex backgrounds and severe obscuring of target objects in X-ray images. It proposes a method for detecting contraband in security X-ray images based on full-process feature fusion and semantically dependent interaction, employing the following technical solution and including the following steps:
[0021] A method for detecting prohibited items in security checks based on feature fusion and semantic interaction includes the following steps:
[0022] Step 1: Collect raw security X-ray images using a security inspection machine, and perform ground truth annotation on the raw images of prohibited items, including the location and category information of the prohibited items, to create a security X-ray image dataset;
[0023] Step 2: Enhance the image data;
[0024] Step 3: Construct the main architecture of the network model, including: constructing a coordinated fusion backbone network for X-ray image feature extraction; constructing an adaptive refinement fusion pyramid network for multi-scale feature fusion; constructing a selective dense feature interaction network for deep feature interaction; and constructing a feature regression and classification network, using the following methods:
[0025] Step 3.1: Construct a coordinated fusion backbone network for X-ray image feature extraction. This backbone network consists of two branches: a feature extraction propagation backbone and a feature fusion detection backbone. Both branches are structurally identical, each consisting of four module groups (stage_1, stage_2, stage_3, stage_4) connected in series. Each module group contains a 3x3 convolutional layer with a stride of 2 and either a C3 module or a ConvFormer composite module from the YOLOv5 object detection model. In both branches, the C3 modules are placed in the first three module groups, with stacking numbers of 3, 6, and 9 respectively. The ConvFormer composite module is placed in the last module group, with a stacking number of 6. The ConvFormer composite module is composed of three components in sequence: a local position-aware unit, a spatial position-aware self-attention, and a lightweight feedforward network. The local position-aware unit is a 3x3 depthwise separable convolution with skip connections. The spatial position-aware self-attention is achieved by embedding a cosine function (cos) to obtain scaled cosine attention during the self-attention calculation between feature sequences, and adding a learnable positional bias (r). ij v pos , where v pos Represents learnable positional embedding, r ijThe relative distance between feature pixels i and j is given; then, the scaled attention map and positional bias information are normalized using softmax respectively; the lightweight feedforward network is mainly composed of three 1*1 convolutions and one 3*3 depthwise separable convolution. The 3*3 convolution is placed between the two 1*1 convolutions to form a branch, while the other 1*1 convolution connects the input and output of the current branch.
[0026] Input image X C×H×W The image features are loaded into the coordinated fusion backbone network and extracted via the feature extraction and transmission backbone branch. These features are then processed by modules (stage_1, stage_2, stage_3, stage_4) within this branch to form a pyramid-structured feature group {C2, C3, C4, C5}. This pyramid-structured feature group is simultaneously input into the modules within the feature fusion and detection backbone branch. In the feature fusion and detection backbone branch, feature C2 is fused with the original input features as input to module stage_1, and then outputs feature P2. Similarly, feature C3 is fused with feature P2 as input to module stage_2, and then outputs feature P3. This process continues to obtain the output features P4 and P5 of the other two module groups, ultimately forming the fused pyramid-structured feature group {P2, P3, P4, P5}. Within a single ConvFormer composite module, the input feature X of the current module... in A feature map Y with location information is generated by a local location-aware unit. Y is then passed through a spatial location-aware self-attention unit to obtain a location self-attention map Z. Z is then passed through a lightweight feedforward network to obtain a mapping-enhanced feature X. out Finally, after calculation by 6 ConvFormer composite modules, features C5 and P5 are generated in the two branches respectively;
[0027] Step 3.2: Construct an adaptive refinement fusion pyramid network for multi-scale feature fusion, using the following method:
[0028] The improved path aggregation feature pyramid PANet introduces a non-local feature attention module GCNet, an adaptive feature fusion module, and a C3 module; and utilizes learnable weight parameters. Different weight coefficients are assigned to the multiple features to be fused, and these coefficients are continuously updated during network training. The values are dynamically adjusted to adjust the weights of the features to be fused, thereby achieving the purpose of adaptively fusing feature maps; the GCNet module is used to further refine global features and dependencies; in terms of structural position, the GCNet module, the adaptive feature fusion module and the C3 module are arranged in series to form a combination, which are then embedded into two feature upsampling and two feature downsampling in PANet, respectively, and finally form an adaptive refinement fusion pyramid network.
[0029] The pyramid feature group {P3, P4, P5} generated in step 3.1 is input in parallel into the adaptive refinement and fusion pyramid network for multi-scale, multi-level feature fusion to deeply mine fine-grained contextual clues. First, the P5 feature is reduced in channel size by a 1*1 convolution to generate the M5 feature map. The M5 feature map is then upsampled by a factor of 2 using a transposed convolution. At the first adaptive feature fusion point, P4 and M5 are weighted and adaptively fused. The fused features are then processed by the GCNet and C3 modules to generate the M4 feature map. At the second adaptive feature fusion point, P3 and the upsampled M4 features are weighted and adaptively fused. The N3 feature map is then generated sequentially through the GCNet and C3 modules. At the third adaptive feature fusion stage, P4, M4, and the N3 features downsampled by 3x3 convolutions are weighted and adaptively fused, and then the N4 feature map is generated sequentially through the GCNet and C3 modules. At the fourth adaptive feature fusion stage, P5, M5, and the N4 features downsampled by 3x3 convolutions are weighted and fused, and then the N5 feature map is generated sequentially through the GCNet and C3 modules. The four adaptive feature fusions use different learnable weight parameters, ultimately generating a refined and polished pyramid fused feature group {N3, N4, N5}.
[0030] Step 3.3: Construct a selective dense feature interaction network for deep feature interaction. For the pyramid feature group {N3, N4, N5} generated in Step 3.2, learn the spatial fusion weights for each scale feature map through identity scaling and adaptive weighted fusion, and then perform adaptive spatial feature fusion; the method is as follows: Generate F a Features: First, N4 and N5 features are identically scaled to the size of N3 feature using 3x3 convolution or max pooling downsampling with a stride of 2. Then, three features of the same size are assigned different learnable weight matrix parameters. Finally, the weighted features are added together and fused to obtain the deep fusion feature F in this module. a ; Generate F b Features: The N3 and N5 features are identically scaled to the N4 feature scale using transposed convolution upsampling and max pooling downsampling, respectively. Then, the three features are adaptively fused to obtain the deep fused feature F. b ; Generate F c Features: The N3 and N4 features are upsampled using transposed convolution to be identically scaled to the N5 feature scale. Then, the three features are adaptively fused to obtain the deep fused feature F. c Finally, a deep fusion feature group {F} is generated. a ,F b ,F c};
[0031] Step 3.4: Construct a regression classification network, using the deep fusion feature group {F} generated in step 3.3.a ,F b ,F c In this process, feature regression and classification of the target of interest are performed to obtain the regression and classification values of the target.
[0032] Step 4: Model training and testing to obtain the final optimal network model;
[0033] Step 5: Model Inference: Obtain the security X-ray image from the real scene and input it into the optimal network model to detect whether there are contraband items in the image; if there are contraband items in the image, the model generates detection result information, including the relative position coordinates (x,y) of the contraband target in the original image, the target width and height (w,h), and the contraband category c.
[0034] This patent proposes a method for detecting contraband in X-ray images during security checks based on full-process feature fusion and semantic dependency interaction. Compared with existing methods, it has the following advantages:
[0035] 1) Integrating multi-scale and multi-level features to improve the model's feature learning ability. This method fully integrates image features from different stages, scales, and levels throughout the network model, effectively extracting the visual representation of contraband in images. Features at different scales have different spatial detail information and semantic contextual meanings, while features at different stages and levels have different abstract expressive capabilities. Fully integrating these features greatly improves the model's feature learning ability.
[0036] 2) Modeling semantic dependencies and interactions enhances the model's feature representation capabilities. This method organically combines local convolution and visual Transformer to effectively model long-term dependencies and perform local-global semantic relationship interactions. Local convolution extracts local spatial information of the region of interest, while visual Transformer models long-range semantic dependencies. By fully combining the two, it effectively extracts identifiable and discriminative feature cues, thereby enabling efficient visual detection representation.
[0037] 3) Significantly improved performance in detecting prohibited items during security checks. This method fully integrates multi-scale and multi-level features from the network flow for effective feature learning. Simultaneously, it models long-term semantic dependencies between features and performs deep contextual interactions, thereby effectively increasing model capacity, enhancing feature representation and generalization ability, and greatly improving the performance in detecting prohibited items during security checks. Trained and tested on the PIDray security X-ray image dataset, the algorithm model achieves an average accuracy of up to 82.4%, and the detection speed meets the real-time requirement of 30 frames per second. It exhibits strong generalization and robustness, meeting the needs of practical applications. Attached Figure Description
[0038] Figure 1 This is a diagram showing the overall structure of the method model;
[0039] Figure 2 To coordinate and integrate the backbone network structure diagram;
[0040] Figure 3 The diagram shows the structure of the ConvFormer composite module (left) and the C3 module (right).
[0041] Figure 4 To adaptively refine the fusion pyramid network structure diagram;
[0042] Figure 5 A diagram of a selectively densely interactive network structure;
[0043] Figure 6 Visualize the results of X-ray images used in security checks. Detailed Implementation
[0044] Classical convolutional neural network-based object detection methods have achieved good detection results on various benchmark datasets and in real-world scenarios, but few methods have achieved good performance in contraband detection. Due to the complex background of security X-ray images, severe occlusion and stacking between objects, and a large amount of redundant noise information, general object detection methods cannot achieve satisfactory accuracy and speed in detecting contraband. Therefore, this patent improves the original method to enhance the overall performance of contraband detection. This method performs full-process feature fusion throughout the detection network, fully integrating high-level semantic features and low-level spatial geometric features to deeply mine fine-grained contextual information clues. At the same time, it combines local convolution and visual Transformer to model global long-term dependencies and perform local-global semantic dependency interactions.
[0045] To address the challenges of complex backgrounds and severe obscuring of target objects in X-ray images, a method for detecting contraband in security X-ray images based on full-process feature fusion and semantically dependent interaction is proposed. The method employs the following technical solution, including the following steps:
[0046] Step 1: Collect raw security X-ray images using a security scanner. Utilize specialized image annotation tools to perform ground truth annotation on the raw contraband images according to a common benchmark data format (including contraband location and category information), creating a security X-ray image dataset. The raw security X-ray images are shown below. Figure 6 The “truth chart” shows that prohibited items include, but are not limited to, handguns, pliers, wrenches, scissors, lighters, and knives.
[0047] Step 2: Divide the dataset obtained in Step 1 into a training set and a test set in a 4:1 ratio, and perform enhancement processing on the image data (including image translation, rotation, scaling, cropping, Gaussian noise, etc.). Store the training set and the test set in different folders.
[0048] Step 3: Construct the main architecture of the network model, including: constructing a coordinated fusion backbone network for X-ray image feature extraction; constructing an adaptive refinement fusion pyramid network for multi-scale feature fusion; constructing a selective dense feature interaction network for deep feature interaction; and constructing a feature regression and classification network, using the following methods:
[0049] Step 3.1: Construct a coordinated fusion backbone network for X-ray image feature extraction. This backbone network consists of two branches: a feature extraction propagation backbone and a feature fusion detection backbone, with the following structure: Figure 2 As shown. The two branches are structurally identical, each consisting of four module groups stage_1, stage_2, stage_3, and stage_4 connected in series. Each module group contains a 3*3 convolutional layer with a stride of 2 and a C3 module or ConvFormer composite module of the object detection model YOLOv5
[10] . In the two branches, the C3 module is set in the first three module groups, with stacking numbers of 3, 6, and 9 respectively; the ConvFormer composite module is set in the last module group, with a stacking number of 6; the ConvFormer composite module is composed of three components in sequence: a local position-aware unit, a spatial position-aware self-attention, and a lightweight feedforward network. The local position-aware unit is a 3*3 depthwise separable convolution with skip connections. The spatial position-aware self-attention is: when performing self-attention
[11] calculation between feature sequences, a scaled cosine attention is obtained by embedding a cosine function cos, and a learnable position bias r is added. ij v pos , where v pos Represents learnable positional embedding, r ij Let i be the relative distance between feature pixels i and j; then, softmax normalization is performed on the scaled attention map and the positional bias information respectively. The lightweight feedforward network mainly consists of three 1*1 convolutions and one 3*3 depthwise separable convolution. The 3*3 convolution is placed between the two 1*1 convolutions to form a branch, while the other 1*1 convolution connects the input and output of the current branch.
[0050] Input image X C×H×WThe image features are loaded into the coordinated fusion backbone network and extracted via the feature extraction and transmission backbone branch. These features are then processed by modules (stage_1, stage_2, stage_3, stage_4) within this branch to form a pyramid-structured feature group {C2, C3, C4, C5}. This pyramid-structured feature group is simultaneously input into the modules within the feature fusion and detection backbone branch. In the feature fusion and detection backbone branch, feature C2 is fused with the original input features as input to module stage_1, outputting feature P2; feature C3 is fused with feature P2 as input to module stage_2, outputting feature P3, and so on, obtaining the output features P4 and P5 of the other two module groups, ultimately forming the fused pyramid-structured feature group {P2, P3, P4, P5}. Within a single ConvFormer composite module, the input feature X of the current module... in A feature map Y with location information is generated by a local location-aware unit. Y is then passed through a spatial location-aware self-attention unit to obtain a location self-attention map Z. Z is then passed through a lightweight feedforward network to obtain a mapping-enhanced feature X. out Finally, after calculations by 6 ConvFormer composite modules, features C5 and P5 are generated in the two branches respectively.
[0051] Step 3.2: Construct an adaptive refinement fusion pyramid network for multi-scale feature fusion: Improve the path aggregation feature pyramid PANet
[14] by introducing the non-local feature attention GCNet
[12] module, the adaptive feature fusion module
[13] , and the C3 module. Learnable weight parameters are used to refine the fusion pyramid network. Different weight coefficients are assigned to the multiple features to be fused, and these coefficients are continuously updated during network training. The values are dynamically adjusted to change the weights of the features to be fused, thus achieving adaptive feature map fusion. The GCNet module is used for further global feature and dependency refinement. Structurally, the GCNet module, adaptive feature fusion module, and C3 module are arranged in series to form a combination. After upsampling two features and downsampling two features in PANet, they finally form an adaptive refinement fusion pyramid network.
[0052] The pyramid feature groups {P3, P4, P5} generated in step 3.1 are input in parallel into the adaptive refinement and fusion pyramid network for multi-scale, multi-level feature fusion to deeply mine fine-grained contextual clues. First, the P5 feature is channel-reduced by a 1*1 convolution to generate the M5 feature map. The M5 feature map is then upsampled by a factor of 2 using a transposed convolution. At the first adaptive feature fusion point, P4 and M5 are weighted adaptively fused, and the fused features are then passed through the GCNet and C3 modules to generate the M4 feature map. At the second adaptive feature fusion point, P3 and the upsampled M4 features are weighted adaptively fused, and then passed through the GCNet and C3 modules to generate the N3 feature map. At the third adaptive feature fusion point, P4, M4, and the N3 features downsampled by a 3*3 convolution are weighted adaptively fused, and then passed through the GCNet and C3 modules to generate the N4 feature map. At the fourth adaptive feature fusion stage, P5, M5, and N4 features downsampled by 3*3 convolution are weighted and fused, and then the N5 feature map is generated sequentially through GCNet and C3 modules. The four adaptive feature fusions use different learnable weight parameters, which ultimately generate a refined pyramid fused feature group {N3, N4, N5}.
[0053] Step 3.3: Construct a selective dense feature interaction network for deep feature interaction. For the pyramid feature group {N3, N4, N5} generated in Step 3.2, learn the spatial fusion weights for each scale feature map through identity scaling and adaptive weighted fusion, and then perform adaptive spatial feature fusion. The method is as follows: Generate F... a Features: First, N4 and N5 features are identically scaled to the size of N3 feature using 3x3 convolution or max pooling downsampling with a stride of 2. Then, three features of the same size are assigned different learnable weight matrix parameters. Finally, the weighted features are added together and fused to obtain the deep fusion feature F in this module. a Generate F b Features: The N3 and N5 features are identically scaled to the N4 feature scale using transposed convolution upsampling and max pooling downsampling, respectively. Then, the three features are adaptively fused to obtain the deep fused feature F. b Generate F c Features: The N3 and N4 features are upsampled using transposed convolution to be identically scaled to the N5 feature scale. Then, the three features are adaptively fused to obtain the deep fused feature F. c Finally, a deep fusion feature group {F} is generated. a ,F b ,F c}
[0054] Step 3.4: Construct a regression classification network, using the deep fusion feature group {F} generated in step 3.3.a ,F b ,F c In this process, feature regression and classification of the target of interest are performed to obtain the regression and classification values of the target.
[0055] Step 4: Model Training. The network model implemented in Step 4 is repeatedly trained iteratively on the training set partitioned in Step 2. The network model parameters are progressively updated and optimized using stochastic gradient descent. Once the loss function reaches its minimum and stabilizes, the model structure and parameters are saved. The model loss function consists of classification loss, bounding box regression loss, and confidence loss. The classification and confidence losses are implemented using the cross-entropy function, and the bounding box regression loss is implemented using the CIoU loss function.
[0056] Step 5: Model Testing and Inference. Using the optimal model parameters and structure saved in Step 5, perform multiple model tests on the test set defined in Step 2, calculating the average accuracy detection index for contraband. If the average accuracy reaches the predicted standard, the final optimal network model is obtained; if it does not, Step 5 needs to be repeated to adjust the model's hyperparameters, optimize model training, and finally execute this step to obtain the optimal model. Obtain real-world security X-ray images and input them into the optimal network model to detect the presence of contraband in the images. If contraband is present in the image, the model generates detection result information, including the relative position coordinates (x, y) of the contraband target in the original image, the target's width and height (w, h), the contraband category c, and the confidence score s, and outputs the result. Figure 6 The image visualization results of the "detection results".
[0057] Thus far, steps 1 to 5 have completed the method for detecting contraband in X-ray images for security checks based on feature fusion and semantic interaction. First, a coordinated fusion backbone network extracts and initially fuses image features, constructing a pyramid feature set rich in spatial detail and semantic context. Then, the ConvFormer composite module further models local-global semantic dependency interactions, fully mining deep semantic dependencies and detailed information. Next, the adaptive fine-tuning fusion pyramid receives the output features from the coordinated fusion backbone as input, refining and fusing the pyramid features to eliminate information redundancy between multi-scale features and bridge semantic inconsistencies between different levels of features. Then, a selective dense feature interaction network module further fuses the contextual information of deep features, preserving meaningful semantic and spatial information. Finally, a regression classification detection network performs regression and classification of the contraband targets on the final deep fused pyramid features, outputting the contraband detection results and completing the detection operation for contraband items in security checks.
Claims
1. A method for detecting prohibited items in security checks based on feature fusion and semantic interaction, comprising the following steps: Step 1: Collect raw security X-ray images using a security inspection machine, and perform ground truth annotation on the raw images of prohibited items, including the location and category information of the prohibited items, to create a security X-ray image dataset; Step 2: Enhance the image data; Step 3: Construct the main architecture of the network model, including: A coordinated fusion backbone network for X-ray image feature extraction is constructed; an adaptive refinement fusion pyramid network for multi-scale feature fusion is constructed; a selective dense feature interaction network for deep feature interaction is constructed; and a feature regression classification network is constructed, using the following methods: Step 3.1: Construct a coordinated fusion backbone network for X-ray image feature extraction. This backbone network consists of two branches: a feature extraction propagation backbone and a feature fusion detection backbone. Both branches are structurally identical, each consisting of four module groups (stage_1, stage_2, stage_3, stage_4) connected in series. Each module group contains a 3x3 convolutional layer with a stride of 2 and either a C3 module or a ConvFormer composite module from the YOLOv5 object detection model. In both branches, the C3 modules are placed in the first three module groups, with stacking numbers of 3, 6, and 9 respectively. The ConvFormer composite module is placed in the last module group, with a stacking number of 6. The ConvFormer composite module is composed of three components in sequence: a local position-aware unit, a spatial position-aware self-attention, and a lightweight feedforward network. The local position-aware unit is a 3x3 depthwise separable convolution with skip connections. The spatial position-aware self-attention is achieved by embedding a cosine function (cos) to obtain scaled cosine attention during the self-attention calculation between feature sequences, and adding a learnable positional bias (r). ij v pos , where v pos Represents learnable positional embedding, r ij The relative distance between feature pixels i and j is given; then, the scaled attention map and positional bias information are normalized using softmax respectively; the lightweight feedforward network is mainly composed of three 1*1 convolutions and one 3*3 depthwise separable convolution. The 3*3 convolution is placed between the two 1*1 convolutions to form a branch, while the other 1*1 convolution connects the input and output of the current branch. Input image X C×H×W The image features are loaded into the coordinated fusion backbone network and extracted via the feature extraction and transmission backbone branch. These features are then processed by modules stage_1, stage_2, stage_3, and stage_4 within this branch, forming a pyramid-structured feature group {C2, C3, C4, C5}. Simultaneously, this pyramid-structured feature group is input in parallel to modules stage_1, stage_2, stage_3, and stage_4 within the feature fusion and detection backbone branch. Within the feature fusion and detection backbone branch, feature C2 is fused with the original input features as input to module stage_1, outputting feature P2; feature C3 is fused with feature P2 as input to module stage_2, outputting feature P3, and so on, yielding the output features P4 and P5 of the other two module groups, ultimately forming the fused pyramid-structured feature group {P2, P3, P4, P5}. Within a single ConvFormer composite module, the input feature X of the current module... in A feature map Y with location information is generated by a local location-aware unit. Y is then passed through a spatial location-aware self-attention unit to obtain a location self-attention map Z. Z is then passed through a lightweight feedforward network to obtain a mapping-enhanced feature X. out Finally, after calculation by 6 ConvFormer composite modules, features C5 and P5 are generated in the two branches respectively; Step 3.2: Construct an adaptive refinement fusion pyramid network for multi-scale feature fusion, using the following method: The improved path aggregation feature pyramid PANet introduces a non-local feature attention module GCNet, an adaptive feature fusion module, and a C3 module; and utilizes learnable weight parameters. Different weight coefficients are assigned to the multiple features to be fused, and these coefficients are continuously updated during network training. The values are dynamically adjusted to adjust the weights of the features to be fused, thereby achieving the purpose of adaptively fusing feature maps; the GCNet module is used to further refine global features and dependencies; in terms of structural position, the GCNet module, the adaptive feature fusion module and the C3 module are arranged in series to form a combination, which are then embedded into two feature upsampling and two feature downsampling in PANet, respectively, and finally form an adaptive refinement fusion pyramid network. The pyramid feature group {P3, P4, P5} generated in step 3.1 is input in parallel into the adaptive refinement and fusion pyramid network for multi-scale, multi-level feature fusion to deeply mine fine-grained contextual clues. First, the P5 feature is reduced in channel size by a 1*1 convolution to generate the M5 feature map. The M5 feature map is then upsampled by a factor of 2 using a transposed convolution. At the first adaptive feature fusion point, P4 and M5 are weighted and adaptively fused. The fused features are then processed by the GCNet and C3 modules to generate the M4 feature map. At the second adaptive feature fusion point, P3 and the upsampled M4 features are weighted and adaptively fused. The N3 feature map is then generated sequentially through the GCNet and C3 modules. At the third adaptive feature fusion stage, P4, M4, and the N3 features downsampled by 3x3 convolutions are weighted and adaptively fused, and then the N4 feature map is generated sequentially through the GCNet and C3 modules. At the fourth adaptive feature fusion stage, P5, M5, and the N4 features downsampled by 3x3 convolutions are weighted and fused, and then the N5 feature map is generated sequentially through the GCNet and C3 modules. The four adaptive feature fusions use different learnable weight parameters, ultimately generating a refined and polished pyramid fused feature group {N3, N4, N5}. Step 3.3: Construct a selective dense feature interaction network for deep feature interaction. For the pyramid feature group {N3, N4, N5} generated in Step 3.2, learn the spatial fusion weights for each scale feature map through identity scaling and adaptive weighted fusion, and then perform adaptive spatial feature fusion; the method is as follows: Generate F a Features: First, N4 and N5 features are identically scaled to the size of N3 feature using 3x3 convolution or max pooling downsampling with a stride of 2. Then, three features of the same size are assigned different learnable weight matrix parameters. Finally, the weighted features are added together and fused to obtain the deep fusion feature F in this module. a ; Generate F b Features: The N3 and N5 features are identically scaled to the N4 feature scale using transposed convolution upsampling and max pooling downsampling, respectively. Then, the three features are adaptively fused to obtain the deep fused feature F. b ; Generate F c Features: The N3 and N4 features are upsampled using transposed convolution to be identically scaled to the N5 feature scale. Then, the three features are adaptively fused to obtain the deep fused feature F. c Finally, a deep fusion feature group {F} is generated. a ,F b ,F c }; Step 3.4: Construct a regression classification network, using the deep fusion feature group {F} generated in step 3.
3. a ,F b ,F c In this process, feature regression and classification of the target of interest are performed to obtain the regression and classification values of the target. Step 4: Model training and testing to obtain the final optimal network model; Step 5: Model Inference: Obtain the security X-ray image from the real scene and input it into the optimal network model to detect whether there are contraband items in the image; if there are contraband items in the image, the model generates detection result information, including the relative position coordinates (x,y) of the contraband target in the original image, the target width and height (w,h), and the contraband category c.
Citation Information
Patent Citations
Multi-scale contraband detection system and method based on X-ray image
CN111951226A
X-ray security check image prohibited article detection method based on style feature channel attention
CN113657493A