X-ray security check image contraband detection method based on large kernel feature injection mechanism

The three-tier progressive feature optimization architecture for YOLOv7 enhances X-ray image detection by improving feature extraction and label allocation, addressing the challenges of overlapping and hidden objects in X-ray images, resulting in higher precision and speed.

CN120318805AActive Publication Date: 2025-07-15HANGZHOU DIANZI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510771943.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-15
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing single-stage object detection algorithms perform poorly in X-ray security images, making it difficult to effectively detect overlapping and hidden contrabands, and the detection accuracy and speed are difficult to take into account.

Method used

Using a method based on the large-kernel feature injection mechanism, feature extraction is enhanced through incremental large-kernel attention-efficient layer aggregation network (LKNA), neck feature fusion module (IFFBN) and task alignment are designed to simplify the optimal transmission allocation strategy (TATA), optimize the feature extraction and label allocation process, and improve detection accuracy and speed.

Benefits of technology

It significantly improves the detection accuracy of contraband in X-ray security images, while maintaining a high detection speed, solving the detection problem of overlapping and hiding contraband.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318805A_ABST
    Figure CN120318805A_ABST
Patent Text Reader

Abstract

The invention discloses an X-ray security check image contraband detection method based on a big kernel feature injection mechanism, and the method comprises the steps: firstly obtaining an X-ray security check image data set containing a target bounding box and a category label, and carrying out the preprocessing of a security check image; secondly, improvement is carried out based on a YoV7 network, incremental convolution feature extraction is carried out on a trunk part for the preprocessed security check image, feature fusion is carried out through an alignment injection fusion mechanism, and a refined feature map is obtained; and finally, carrying out classification and regression calculation on the refined feature map to obtain a classification and regression calculation value of each detection anchor frame, outputting an X-ray security check image contraband detection result, and carrying out training and testing. According to the invention, high precision is obtained in X-ray security check image set detection, the real-time requirement is met, and high detection speed is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and object detection, and specifically, a single-stage X-ray security inspection image contraband object detection method based on a large kernel feature injection mechanism is proposed. Background Art

[0002] With the continuous increase in the passenger flow at public transportation stations, public safety guarantee has become increasingly important. In the security inspection process, X-ray scanners are widely used for luggage fluoroscopy inspection to generate X-ray security inspection images. Security inspectors identify potential dangerous items in luggage by observing the differences in the absorption and scattering abilities of different substances to X-rays and combining visual features such as the color, contour, and shape of the items. However, compared with natural light images, X-ray security inspection images have the following characteristics: (1) Overlap: Overlapping phenomena will occur in the image for items stacked together, resulting in blurred surface textures of the items; (2) Concealment: Contraband items are difficult to be detected by the detection system due to their small size and overlap with other items.

[0003] In recent years, the development of deep learning object detection technology has promoted the automation of X-ray contraband detection. Common object detection frameworks include single-stage, two-stage, and sparsity object detection frameworks. Single-stage object detection algorithms are widely used due to their excellent real-time performance, scalability, and accuracy. However, most single-stage object detection algorithms are designed for natural light images, and their performance in X-ray contraband detection is not ideal. Summary of the Invention

[0004] The present invention proposes an X-ray security inspection image contraband detection method based on a large kernel feature injection mechanism. It optimizes the single-stage object detection algorithm (taking YOLOv7 as an example) as a whole, and proposes a three-level progressive feature optimization architecture to break through the bottleneck of the existing technology, that is, the adaptive improvement of the feature extraction backbone based on the large kernel spatial perception module, the neck alignment fusion injection mechanism, and the global task alignment label assignment strategy, so as to significantly improve the detection accuracy of contraband in X-ray security inspection images and maintain a high detection speed. The present invention improves the detection accuracy by optimizing and improving the backbone, neck, and label assignment process of the single-stage object detection framework. Specifically, first, through an incremental large kernel structure, an LKNA backbone network is designed to enhance the model's ability to extract global context features of contraband. Then, by aligning and fusing some features in the backbone and injecting them into the neck network, a neck feature fusion module IFFBN is designed to make up for the loss of small target information caused by frequent upsampling and downsampling operations. Finally, from a global perspective, a consistency metric factor is introduced in the label assignment process to intelligently match positive and negative samples, and a TATA label assignment strategy is designed, so as to effectively combine the advantages of global optimization and task consistency and further improve the performance of the detection framework. This method is based on an incremental large kernel structure, a neck alignment fusion injection mechanism, and a global consistency label assignment strategy, aiming to improve the performance and accuracy of contraband detection in the security inspection field.

[0005] The X-ray security inspection image contraband detection method based on the large kernel feature injection mechanism includes the following steps:

[0006] Step (1) First, obtain an X-ray security inspection image dataset containing target bounding boxes and class annotations, divide the dataset into a training set and a validation set, and preprocess the security inspection images.

[0007] In the backbone part of the YoloV7 network, replace the efficient layer aggregation network (ELAN) with an incremental large kernel hierarchical aggregation network LKNA, add an incremental convolutional kernel structure, and add an attention mechanism inside LKNA to extract spatial feature information in the image. Through the B2, B3, B4, and B5 layers of the LKNA backbone network, a set of feature maps with different resolution sizes are obtained respectively . The Calculate again in the attention module to obtain a further refined feature map and .

[0008] In the neck network fusion stage of the YoloV7 network, use the alignment injection fusion mechanism IFFBN proposed in this paper. Through alignment, fusion, and injection operations, inject the feature information into the neck network to generate corresponding , and features. Then, these features are used to perform a fusion interaction operation with the feature maps and to obtain a refined feature map .

[0009] In step (4), the lightweight network RepVGG is used to perform classification and regression calculations on the refined feature map to obtain the classification and regression calculation values for each detection anchor box. In the label assignment stage, the task-aligned simplified optimal transport assignment strategy TATA is used to replace the original assignment strategy SimOTA. The optimal anchor box that matches the ground truth GT is found within the global scope as the positive sample, and the remaining anchor boxes are used as negative samples for assignment. During the matching process, the system measures the consistency between the classification sub-task and the regression sub-task through the classification and regression calculation values, alleviating the misalignment problem of ambiguous anchor boxes.

[0010] In step (5), the training parameters are set, and the training dataset obtained in step (1) is input into the model constructed in steps (2)-(5) for iterative training, and verified on the validation set to obtain the optimal parameter model, and the prohibited item detection effect diagram is output.

[0011] Furthermore, step (2) is specifically as follows:

[0012] (2-1) In the backbone part, the original ELAN is improved to obtain LKNA. The images in the training dataset are input into LKNA to extract multi-level features, and the output feature map sets of the last four stages B2, B3, B4, and B5 layers of the backbone network are obtained .

[0013] Specifically, in LKNA, a progressive large kernel design is adopted, that is, large core convolution kernels of different sizes are used in the {B3, B4, B5} layers of the backbone network: . Specifically, LKNA divides the feature map into two parts through two convolution operations, that is, , where respectively represent the weights of the two convolution kernels. does not participate in the calculation. And the other part of the feature map , while being truncated by the transition layer in ELAN, passes through two large convolution kernels to learn the detailed information in the feature map, that is, . After passing through the CA attention module, the spatial position information of the items in the feature map is aggregated, that is, 。This spatial information aggregation operation is very important for X-ray security inspection images lacking surface texture features. Finally, LKNA will After splicing, it is transported to the transition layer to participate in subsequent calculations. In ELAN, the gradient flow is divided into four paths, and the shortest gradient path is 1.

[0014] (2-2) Based on the multi-level feature maps , two GAM attention modules are used for processing respectively to obtain further refined feature maps and , that is, , . And include the extracted small target spatial feature information.

[0015] Furthermore, step (3) is specifically:

[0016] (3-1) Between the neck and the trunk, an alignment injection fusion mechanism IFFBN is added. IFFBN includes alignment, fusion, and injection operations performed sequentially.

[0017] (3-1-1) Alignment operation. Align the multi-level features in the backbone, pool and splice the multi-level features, and then obtain the aligned features through partial convolutional cross-stage operations.

[0018] (3-1-2) Fusion operation. IFFBN adaptively fuses the multi-level features. For the two fused features X and Y, after adding the features X and Y element-wise, they are branched for feature extraction. One branch passes through a filter block to obtain local features, and the other branch passes through global pooling and then obtains global features by the filter block; after adding the global features and the local features element-wise, a weight factor m is obtained, and the weighted addition of features X and Y is performed to obtain the fused features, where the weight of feature X is m and the weight of feature Y is 1 - m; the filter block is composed of convolution, BN, and activation function SiLU connected in sequence.

[0019] (3-1-3) Injection operation. Inject the aligned features into the fused features. After respectively performing convolution extraction on the fused features and the aligned features, injection local features and injection global features are generated; after average pooling the injection global features, it is used as a weight to perform a weighted operation on the injection local features and add them element-wise to the injection global features to obtain the final injection features; different multi-level features generate different injection features, that is, .

[0020] (3-2) Optimize the feature map set obtained in step (2) , and three multi-level features obtained from (3-1) are input into the improved neck network to generate refined features .

[0021] (3-2-1) is processed by a convolutional network to reduce the number of channels and extract features, obtaining a middle feature map . Then after being processed by three-way max pooling MaxPooling, the results are concatenated to integrate information at different scales. Finally, the obtained feature map is further passed through a convolutional layer, a convolutional layer, and then concatenated with the feature map processed by the convolutional layer to form the final feature map .

[0022] (3-2-2) performs bilinear interpolation upsampling operation on the optimized feature map set , to obtain upsampled features , and , is injected into to obtain refined features . After performing upsampling operation on and performing convolutional operation through the ELAN module, the obtained feature map is fused with , and after passing through the ELAN module and downsampling operation, the feature map is obtained. Injecting into obtains refined features . Injecting through the ELAN module and downsampling operation, the downsampled feature map is obtained. Injecting and into obtains refined features .

[0023] Furthermore, step (4) is specifically:

[0024] (4-1) divides the image into a grid array, and each grid point is responsible for detecting whether the center of the object to be detected is included in the grid. In this process, the output of each grid is a prediction vector, denoted as , where respectively represent the center point coordinates and width and height of the candidate bounding box, represents the confidence that the center point of the target falls within the current bounding box, and Represents the predicted probability of the object category within the current bounding box.

[0025] (4-2) During the label assignment stage, use the task-aligned simplified optimal transport assignment strategy TATA assignment strategy to replace the original transport assignment SimOTA strategy.

[0026] (4-2-1) First, through prediction and processing operations, that is, perform data dimension conversion preprocessing operations on the input image representation, prior anchor box set, and ground truth annotation box GT to generate predicted bounding boxes , predicted class scores . The target bounding boxes obtained after preprocessing GT , and perform one-hot encoding on the GT label categories .

[0027] (4-2-2) Calculate the IoU value between the target bounding box and the predicted bounding box to generate the corresponding value. According to the value and the dynamic K-matching algorithm, calculate the number of positive samples required for each GT .

[0028] (4-2-3) According to the following formula, calculate the consistency metric , and multiply it by , which is equivalent to using the consistency metric value to weight , and finally obtain the predicted label value .

[0029] (4-2-4) By calculating the binary cross-entropy between and the true label category , obtain the classification loss ; calculate the regression loss through the formula. Through , calculate the cost matrix of the anchor boxes , and sort it.

[0030] (4-2-5) According to the dynamic matching count , select the top anchor boxes with the lowest cost as positive samples. When a fuzzy anchor box matches multiple GTs, select the GT with the minimum cost as the GT matched by this anchor box. Finally, in the non-maximum suppression NMS stage, delete the duplicate anchor boxes.

[0031] Compared with the prior art, the beneficial effects of the present invention are:

[0032] Different from the previous method of adding separately designed specific modules to the existing detection framework to process features such as shapes and materials in X-ray images, the present invention adopts an overall design strategy to improve the detection performance by enhancing the model's ability to process the spatial information of contraband. Specifically, first, the present invention uses a Progressive Large Kernel Attention Efficient Layer Aggregation Network (LKNA) to enhance the ability to capture global dependencies during the feature extraction process. Then, a feature integration and fusion module from backbone to neck (IFFBN) is designed to compensate for the loss of small object feature information due to frequent upsampling and downsampling operations. Finally, a Task Alignment Simplified Optimal Transport Assignment (TATA) strategy is used, which combines the advantages of local optimization and task consistency, optimizes the assignment of fuzzy anchor boxes, and improves the detection accuracy. Compared with the original single-stage object detection algorithm, a higher accuracy is obtained in the detection of X-ray security inspection image sets, and a high detection speed is maintained to meet the real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is the processing flow chart of the present invention;

[0034] Figure 2 is the overall structure diagram of the model of the present invention;

[0035] Figure 3 is the structure diagram of LKNA in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In order to better understand the purpose, structure and function of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings.

[0037] This example proposes a method for detecting contraband in X-ray security inspection images based on single-stage object detection. The overall process is as Figure 1 shown. The overall structure of the model is as Figure 2 shown. First, the present invention designs a LKNA backbone network, as Figure 3 shown. It enhances the model's ability to extract global context features of contraband through an incremental large kernel structure. Then, the present invention redesigns a neck feature alignment injection and fusion module (IFFBN), as Figure 2 shown. Through alignment, fusion, and injection operations, IFFBN not only compensates for the loss of small target information due to frequent upsampling and downsampling operations, but also reduces the computational cost while maintaining high computational efficiency through lightweight design. Finally, the present invention introduces a TATA label assignment strategy. It starts from a global perspective and effectively combines the advantages of global optimization and task consistency by introducing a consistency metric factor to intelligently match positive and negative samples, further improving the performance of the detection framework.

[0038] The method of the present invention includes the following steps:

[0039] Step (1): Download the X-ray security inspection image dataset, divide the two X-ray security inspection image datasets into a training set and a validation set, and perform preprocessing operations on the pictures. Specifically:

[0040] (1-1) Divide the image set into two main parts: one part is used for training the model, that is ; the other part is used to verify the accuracy of the model, that is . Where R is the real number field, represents the number of image samples in the training set, represents the i-th training image sample, represents the number of image samples in the validation set, represents the j-th validation image sample, H represents the image height, W represents the image width, and 3 represents the number of RGB channels;

[0041] (1-2) The corresponding label of each training sample ; where represents the number of targets contained in the image sample represents the image sample The number of targets in, represents in the -th true category of the target, where C represents the total number of categories in the dataset, represents in the -th bounding box of the target, which is composed of the abscissa x of the center point, the ordinate y of the center point, the width w of the target, and the height h.

[0042] (1-3) Before formal training, perform a series of data augmentation operations on the images. This includes randomly selecting four images from the training set, performing a series of transformation operations on them, such as flipping, scaling, and adjusting the hue, and then splicing these images into a new sample while retaining their original label information. In addition, the images are also resized, uniformly scaled to 640×640 pixels, and the pixel values are normalized to ensure the consistency and effectiveness of the model input.

[0043] Step (2): On the network backbone, improve the original ELAN to obtain LKNA, add an incremental large kernel structure, and add an attention mechanism between the large kernel structures to extract the spatial feature information in the image to obtain a set of feature maps . The LKNA network structure is as Figure 3 shown. The multi-level feature maps are processed again in the attention module to obtain further refined feature maps and . Specifically:

[0044] (2-1) In the LKNA of the backbone part, a progressive large kernel design is adopted, that is, large core convolutional kernels of different sizes are adopted in the {B3, B4, B5} layers of the backbone network: . The corresponding convolution operation is. Specifically, LKNA divides the feature map into two parts through two convolution operations, that is , where respectively represent the weights of the two convolutional kernels. does not participate in the calculation and is directly truncated by the transition layer. And the other part of the feature map , while being truncated by the transition layer, passes through two large convolutional kernels to learn the detailed information in the feature map, that is . After passing through the CA attention module ( Figure 2 at position ①), the spatial position information of the objects in the feature map is aggregated, that is . This spatial information aggregation operation is very important for X-ray security inspection images lacking surface texture features. Finally, LKNA will splice and send it to the transition layer to participate in subsequent calculations. There are four gradient flows in LKNA, and the shortest gradient path is 1.

[0045] (2-2) The images in the training dataset are input into LKNA to extract multi-level features, obtaining the output feature map sets of the B3, B4, and B5 layers of the backbone network ;

[0046] (2-3) Based on the multi-level feature maps , two GAM attention modules are respectively used for processing ( Figure 2 at positions ② and ③), obtaining further refined feature maps and , that is , . . and include the extracted small target spatial feature information.

[0047] In step (3), in the neck network fusion stage, the IFFBN is used to inject the feature information in the backbone into the neck network through alignment, fusion, and injection operations, respectively generating the corresponding , and features. Then, these features are used together with the feature maps and Perform a fusion interaction operation to obtain a refined feature map . Specifically:

[0048] (3-1) Add an alignment fusion injection mechanism IFFBN between the neck and the backbone. In IFFBN, the designs of the backbone and neck structures are balanced, ensuring that while the backbone network can effectively extract deep features of the image, the neck structure can further optimize and utilize these features to improve the performance of object detection. IFFBN is divided into alignment, fusion, and injection operations that are executed sequentially.

[0049] (3-1-1) Alignment operation. IFFBN aligns multi-level features in the backbone to obtain the aligned features , that is . Among them, represents the concatenation operation, represents the global pooling operation, the output resolution of is the same as the resolution of refers to the CSP cross-stage operation, that is, by dividing the basic layer feature map into two parts, one part of the feature map does not participate in the operation, and the other part is directly merged with the feature that does not participate in the operation in the transition layer after calculation through the dense block.

[0050] (3-1-2) Fusion operation. In the fusion stage, IFFBN does not simply concatenate multi-level features, but performs adaptive fusion. The fusion process can be expressed as:

[0051]

[0052]

[0053]

[0054]

[0055] Among them, the block consists of convolution, BN, and the activation function SiLU. represents the global pooling operation. X and Y represent two features to be fused. During the fusion process, first, the X and Y features are combined, that is, after the "+" operation, they are sent to the filtering module to extract local features and obtain . Similarly, after the average pooling operation, the global feature is obtained. Then, and are fused to calculate the weight factor m of the X feature. Finally, the weight factor is used to calculate the final fused feature This fusion strategy effectively solves the problem of inconsistency in semantics and scale among features of different layers or branches. It not only improves the overall quality of the fused features but also significantly enhances the fusion accuracy of the features of small objects that are easily overlooked under the action of the global channel attention mechanism.

[0056] (3 - 1 - 3) Injection operation. IFFBN injects into to obtain the final injected feature . The injection process is expressed as:

[0057]

[0058]

[0059]

[0060]

[0061] Among them, A and B respectively represent two features in the injection process: the global feature and the local feature, that is, and , represents the average pooling operation. In the injection process, after convolutional extraction of feature A and feature B respectively, local feature and global feature are generated. After is averaged and pooled, it is used as a weight to perform a weighted operation on to obtain the final injected feature . Different multi - level features will generate different injected features, that is, .

[0062] (3 - 2) Input the feature , the optimized feature map set , and the three multi - level features obtained from (3 - 1) into the improved neck network.

[0063] (3 - 2 - 1) passes through a convolutional layer, a convolutional layer, and a convolutional layer for processing to reduce the number of channels and extract features, obtaining the middle feature map . Then passes through three - way max - pooling MaxPooling processing, and the results are concatenated to integrate information of different scales. Finally, the obtained feature map passes through another convolutional layer, a After the convolutional layer, it is concatenated with After passing through the feature map processed by the convolutional layer to form the final feature map .

[0064] (3-2-2) Upsample the optimized feature map set , using bilinear interpolation to obtain the upsampled feature , and , inject it into to obtain the refined feature . After is upsampled and then undergoes convolutional operations through the ELAN module, the resulting feature map is fused with , and after passing through the ELAN module and downsampling operation, the feature map is obtained. Inject into to obtain the refined feature . After undergoes convolutional operations through the ELAN module and downsampling operation, the downsampled feature map is obtained. Inject and into to obtain the refined feature .

[0065] In step (4) during the label assignment stage, the task-aligned simplified optimal transport assignment strategy TATA assignment strategy is used to replace the original transport assignment SimOTA strategy. By globally searching for the optimal anchor box that matches the GT as the positive sample, and the rest as the negative sample for assignment. And during the assignment process, the consistency between the classification sub-task and the regression sub-task is measured to alleviate the misalignment problem of ambiguous anchor boxes. Specifically:

[0066] (4-1) Divide the image into grid arrays, and each grid point is responsible for detecting whether the center of the object to be detected is contained within the grid. During this process, the output of each grid is a prediction vector, denoted as , where respectively represent the center point coordinates and width and height of the candidate bounding box, represents the confidence that the target center point falls within the current bounding box, while represents the predicted probability of the object category within the current bounding box.

[0067] (4-2) During the label assignment stage, use the new TATA assignment strategy to replace the original SimOTA strategy.

[0068] (4-2-1) First, through prediction and processing operations, that is, performing data dimension conversion preprocessing operations on the input image representation, the prior anchor box set, and the true annotation box GT, predicted bounding boxes are generated. , predicted class scores . The target bounding boxes obtained after preprocessing the GT , and the GT label classes are one-hot encoded. .

[0069] (4-2-2) Calculate the IoU value between the target bounding box and the predicted bounding box to generate the value. According to the value and the dynamic K-matching algorithm, calculate the number of positive samples required for each GT .

[0070] (4-2-3) According to the following formula, calculate the consistency metric , and multiply it by , which is equivalent to using the consistency metric value to weight , and finally obtain the predicted label value .

[0071]

[0072] Among them, and represent the classification score and the IoU value between the current anchor box and the GT, and are hyperparameters that control the contribution degrees of the cls and reg tasks to the task consistency, and the default values are 1.0 and 6.0 respectively. represents the maximum value acquisition function.

[0073] (4-2-4) By calculating the binary cross-entropy between and the true label class , the classification loss is obtained; the regression loss is calculated through the formula. Through , calculate the cost matrix of the anchor boxes and sort it.

[0074] (4-2-5) According to the dynamic matching count , select the top anchor boxes with the lowest cost As positive samples. When a fuzzy anchor box matches multiple GTs, the GT with the minimum cost will be selected as the GT matched by this anchor box. Finally, in the non-maximum suppression (NMS) stage, duplicate anchor boxes are removed.

[0075] Step (5) sets the training parameters, inputs the training dataset obtained in step (1) into the model constructed by steps (2)-(4) for iterative training, validates on the validation set, obtains the optimal parameter model, and outputs the effect diagram of contraband detection. The specific operations are as follows:

[0076] (5-1) In the initial stage of model training, a set of key hyperparameters need to be set first, such as the initial learning rate, momentum factor, decay rate of the learning rate, batch size, number of GPUs used, selected optimization algorithm, and total number of training epochs. Subsequently, the prepared dataset is input into the model framework constructed through the previous steps and undergoes multiple rounds of training. During the training process, the mAP50 metric is used as a measurement metric to continuously monitor the performance of the model on the validation set.

[0077] Whenever the best mAP50 value is obtained during the training process, the corresponding model configuration is recorded and these configurations are saved as the optimal parameters. This approach helps to find the best combination of parameters during the training process, thereby improving the prediction accuracy of the model.

[0078] (5-2) After the training is completed, the validation set is used to validate the optimal model obtained in step (5-1), obtain the metric parameters of the final model for detecting the dataset, and mark the categories and confidence levels of the detected contraband on the detection results.

[0079] The above embodiments elaborate in detail the objectives, technical solutions, and their advantages of the present invention. It should be clear that these embodiments only represent some preferred implementation approaches of the present invention and do not limit its scope. Any form of modification, equivalent replacement, or improvement made under the premise of following the core idea and principles of the present invention should be regarded as being included within the protection scope of the present invention.

[0080] Embodiment:

[0081] In this experiment, two types of X-ray security inspection datasets were used, namely OPIXray and CLCXray. The OPIXray dataset poses significant challenges to the detection model due to its inclusion of small and possibly overlapping contraband images. The CLCXray dataset, on the other hand, focuses on the distinction between contraband and similar backgrounds, especially when the contraband overlaps with the background. The detailed data distributions of contraband in the two datasets are shown in detail in Figure 1.

[0082] Table 1 OPIXray and CLCXray Datasets

[0083]

[0084] Among them, the OPIXray dataset contains a total of 8,885 images and contraband of 5 knife types, namely: Folding knife, Straight knife, Scissor, Utility knife, and Multi-tool knife. The OPIXray focuses on occluded contraband knives, and these knives are usually small in size and easily blocked by other objects, posing a new challenge to X-ray security inspection. The CLCXray dataset contains a total of 9,565 images, of which 4,543 are from real subway station security inspection scenarios and 5,022 are from artificially intervened design scenarios. The biggest feature of the CLCXray dataset is that it contains multiple contraband items that are similar to the background and overlap with each other, and it contains a total of 12 categories, namely: Blade, Dagger, Knife, Scissors, Tin, Cans, Carton Drinks, Glass Bottle, Plastic Bottle, Vacuum Cup, Spray Cans, Swiss Army Knife.

[0085] Experimental comparison description:

[0086] Table 2 Detection comparison results of the present invention and other SOTA methods on OPIXray

[0087]

[0088] Table 2 lists the detection results of X-Safe and several other state-of-the-art (SOTA) algorithms on the OPIXray dataset. Among them, X-Safe achieved a high accuracy of 93.1%, significantly outperforming other comparison algorithms. Specifically, X-Safe performed better than or close to the existing best methods in each category. For example, compared with YOLOv7 using the SimOTA label assignment strategy, the mAP50 of X-Safe increased by 9.4%; compared with YOLOv7u6(TAL) using the TAL label assignment strategy, the mAP50 of X-Safe increased by 1.1%. In addition, compared with the baseline algorithm SSD+DOAM on the OPIXray dataset, the accuracy of X-Safe increased by 19.1%. These results demonstrate that X-Safe has significant advantages in the contraband detection task. In addition, while maintaining high detection accuracy for easily recognizable categories such as Scissor (98.4%) and Multi-toolKnife (96.4%), X-Safe effectively improved the detection effect of targets with complex surface texture features such as StraightKnife, proving its better feature representation ability. It achieved a breakthrough detection accuracy in the UtilityKnife category, an 8.9% improvement compared to the sub-optimal method POD-Y, indicating a significant enhancement in the model's ability to capture local spatial features of knives. These results show that X-Safe has significant advantages in dealing with complex scenarios and diverse targets, and can effectively improve the accuracy and reliability of contraband detection.

[0089] Table 3 Comparison results of the present invention and other SOTA methods in CLCXray detection

[0090]

[0091] Table 3 shows the performance comparison of different YOLO series models on the CLCXray dataset. The results show that there are significant differences in performance among the models in terms of the key indicator mAP50. YOLOv7 (SimOTA) and YOLOv9-m performed relatively prominently, with mAP50 reaching 83.6% and 84.7% respectively. YOLOv10-m achieved an mAP50 of 79.7% by optimizing the inference speed, but its accuracy was slightly lower than that of YOLOv9-m and YOLOv11-m (84.6%). The X-Safe model proposed in the present invention further improved the detection accuracy by comprehensively optimizing the feature extraction and training strategies, with mAP50 reaching 85.1%, becoming the model with the highest current performance.

[0092] The above embodiments have elaborated in detail the objectives, technical solutions and their advantages of the present invention. It should be clear that these embodiments only represent some preferred implementation approaches of the present invention and do not limit its scope. Any form of modification, equivalent replacement or improvement should be regarded as being included within the protection scope of the present invention on the premise of following the core idea and principles of the present invention.

Claims

1. An X-ray security inspection image contraband detection method based on a large kernel feature injection mechanism, characterized in that It includes the following steps: Step 1: Obtain an X-ray security inspection image dataset containing target bounding boxes and class annotations, divide the dataset into a training set and a validation set, and preprocess the security inspection images; Step 2: Based on the improvement of the YoloV7 network, the preprocessed security inspection images extract features through incremental convolution in the backbone part, and perform feature fusion through an alignment injection fusion mechanism to obtain a refined feature map; Step 3: Perform classification and regression calculations on the refined feature map to obtain the classification and regression calculation values of each detection anchor box, and output the detection results of contraband in the X-ray security inspection images; Step 4: Use the test set to iteratively train the model constructed in Steps 2 and 3, verify it on the validation set, obtain the optimal parameter model, and output the contraband detection effect diagram.

2. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 1, wherein, The specific implementation process of Step 2 is as follows: Step 2.1: In the backbone part of the YoloV7 network, improve the Efficient Layer Aggregation Network (ELAN) to obtain the Layered Incremental Aggregation Network (LKNA). Add an incremental convolutional kernel structure and an attention mechanism inside the LKNA to extract the spatial feature information in the image , and obtain a refined feature map; Step 2.

2. In the neck network fusion stage of the YoloV7 network, using the alignment injection fusion mechanism IFFBN, through alignment, fusion, and injection operations, inject the feature information into the neck network, and then perform a fusion interaction operation with the refined feature map to obtain a refined feature map.

3. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 2, wherein, The specific implementation process of Step 2.1 is as follows: Step 2.1.1: In the backbone part, improve the original ELAN to obtain LKNA. In LKNA, adopt a progressive kernel design, that is, use convolutional kernels with gradually increasing sizes in the backbone network: The images in the training dataset are input into LKNA to extract multi-level features, and the feature map sets are respectively output by the last four stages B2, B3, B4, and B5 layers of the backbone network ; Step 2.1.

2. Based on the multi-level feature maps , process them respectively using two GAM attention modules to obtain refined feature maps and .

4. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 3, characterized in that, The specific implementation of the incremental hierarchical aggregation network LKNA is as follows: the feature map is divided into two parts through two convolutional operations, namely , the feature map While being truncated by the transition layer in the ELAN block, it passes through two convolutional layers to learn the information in the feature map, obtaining the feature map , and then passes through the CA attention module to aggregate the spatial position information of the items in the feature map, obtaining the feature map ; After are concatenated, they are sent to the transition layer to participate in subsequent calculations.

5. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 4, wherein, The specific implementation process of Step 2.2 is as follows: Step 2.2.1: Add an alignment injection fusion mechanism IFFBN between the neck and the backbone. IFFBN includes alignment, fusion, and injection operations performed in sequence to obtain different multi-level features and generate different injection features; Step 2.2.2: Input the feature , the feature map set , and three multi-level features into the improved neck network to generate refined features.

6. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 5, wherein, The operations in the alignment injection fusion mechanism IFFBN are specifically as follows: Alignment operation: Align multi-level features in the backbone{ , pool and splice the multi-level features, and then obtain aligned features through partial convolution cross-stage operations; Fusion operation: IFFBN adaptively fuses multi-level features. For two features X and Y to be fused, after element-wise addition of features X and Y, feature extraction is performed in two branches. One branch passes through a filtering block to obtain local features, and the other branch passes through global pooling and then a filtering block to obtain global features; after element-wise addition of the global features and the local features, a weight factor m is obtained, and weighted addition of features X and Y is performed to obtain a fused feature. The weight of feature X is m, and the weight of feature Y is 1 - m; the filtering block is composed of convolution, BN, and the activation function SiLU connected in sequence; Injection operation: Inject the aligned features into the fused features. After respectively performing convolution extraction on the fused features and the aligned features, generate injection local features and injection global features; after average pooling the injection global features, use it as a weight to perform a weighted operation on the injection local features and add them element-wise to the injection global features to obtain the final injection features; Different multi-level features generate different injection features, i.e., .

7. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 6, wherein, The specific implementation process of Step 2.2.2 is as follows: It is processed through a convolutional network, and then after three-way max pooling, the results are concatenated to integrate information of different scales; finally, the obtained feature map is passed through two convolutional layers with increasing numbers of convolutional kernels, and then concatenated with the feature map processed by the convolutional layer to form the final feature map ; For the set of feature maps , perform bilinear interpolation upsampling to obtain the upsampled features , and inject , into to obtain the refined features ; For perform upsampling, and after performing convolutional operations through the ELAN module, fuse the resulting feature map with , and after passing through the ELAN module and downsampling operation, obtain the feature map ; After injecting into , the refined feature is obtained; After passes through the ELAN module and the downsampling operation, the downsampled feature map is obtained; After injecting and into , the refined feature is obtained.

8. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 7, wherein, The specific implementation process of Step 3 is as follows: Step 3.1: Divide the image into a grid array, and each grid point is responsible for detecting whether the center of the object to be detected is contained within the grid; in this process, the output of each grid is a prediction vector; Step 3.2: In the label assignment stage, use the task-aligned simplified optimal transport assignment strategy TATA assignment strategy to replace the original transport assignment SimOTA strategy.

9. The method for detecting contraband in X-ray security inspection images based on the large kernel feature injection mechanism according to claim 8, characterized in that, The specific implementation process of the task-aligned simplified optimal transport assignment strategy is as follows: Perform data dimension conversion preprocessing operations on the input image representation, prior anchor box set, and true annotation box GT to generate predicted bounding boxes and predicted class scores; for the target bounding boxes obtained after preprocessing GT, perform one-hot encoding on the GT label classes; Calculate the IoU value between the target bounding box and the predicted bounding box to generate the corresponding value. Through the dynamic K-matching algorithm, calculate the number of positive samples required for each GT ; Calculate the consistency metric index and multiply it by the predicted class scores, which is equivalent to using the consistency metric index value to perform a weighted operation on the predicted class scores, and finally obtain the predicted label values; The classification loss is obtained by calculating the binary cross-entropy between the predicted label value and the true label category; the regression loss is calculated through the formula, and the cost matrix of the anchor box is calculated through the weighted sum of the classification loss and the regression loss, and it is sorted. According to the dynamic matching count , select the top anchor boxes with the lowest cost as positive samples; when a fuzzy anchor box matches multiple GTs, select the GT with the minimum cost as the GT matched by the anchor box; finally, in the non-maximum suppression (NMS) stage, remove the duplicate anchor boxes.

Citation Information

Patent Citations

  • Flame detection method based on MPGD-YOLO network

    CN116758393A

  • X-ray image contraband detection method

    CN117058606A

  • Unmanned aerial vehicle-oriented power transmission line lightning stroke point and bolt defect detection method and system

    CN118014968A

  • X-ray image contraband detection method based on de-overlapping and associated attention mechanism

    CN118261853A

  • X-ray security check image forbidden article detection method based on improved YOLOv7

    CN118397303A