X-ray Image Contraband Detection Method Based on De-overlapping and Correlation Attention Mechanisms

By introducing top-down deoverlapping module TD-DOM, associated attention mechanism enhancement feature fusion module EFFM and detection head AF-Head without anchor frames in the YOLOv7 model, the problem of low detection accuracy of X-ray image contraband detection algorithm in scenarios with significant changes in within-class changes is solved, and higher detection accuracy and speed are achieved.

CN118261853BActive Publication Date: 2025-06-20HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311789211.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-20
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

The existing X-ray image contraband detection algorithm has low detection accuracy in scenarios where there is significant change in class, making it difficult to effectively deal with the problems of overlapping objects, large changes in target scales, and high proportion of small targets.

Method used

Add the top-down deoverlapping module TD-DOM in the YOLOv7 model, design the associated attention mechanism to enhance the feature fusion module EFFM, and add the detection head AF-Head without anchor boxes to the network head, and use the DFL loss function to improve the detection accuracy.

Benefits of technology

By removing the noise influence caused by object overlap, the model is improved to improve the robustness of the target scale changes, the number of positive samples for small targets is increased, and the accuracy and detection speed of X-ray image contraband detection are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118261853B_ABST
    Figure CN118261853B_ABST
Patent Text Reader

Abstract

The present invention discloses an X-ray image contraband detection method based on de-overlapping and associated attention mechanisms. First, an X-ray image data set containing target bounding boxes and class annotations is obtained. Secondly, a classifier is added at the end of the backbone part of the YOLOv7 network, and a top-down de-overlapping module is added between the backbone and the neck to output a set of feature maps. Then, the feature fusion module in the bidirectional fusion feature pyramid structure of the YOLOv7 neck is replaced with an enhanced feature fusion module to obtain a set of feature maps N. Finally, a detection head without anchor boxes is added to the head of the YOLOv7 network, the set of feature maps N is input, the target bounding boxes and classes predicted by the model are output, the training parameters are set for iterative training, and verification is performed on the validation set to output the contraband detection effect diagram. The present invention significantly improves the detection accuracy for X-ray image contraband detection and has a high detection speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and object detection, and specifically to an X-ray image contraband detection method based on de-overlapping and associated attention mechanisms. Background Art

[0002] With the increase in the population density of public transportation hubs, security detection plays an increasingly important role in ensuring public safety. In security inspections, X-ray machines are usually used to scan luggage to generate X-ray images. Security personnel can judge whether there are dangerous items in the luggage based on the different absorption and scattering attenuation degrees of X-rays by items of different materials, and the corresponding X-ray image colors generated by the items, combined with morphological features such as edges and shapes.

[0003] Compared with the pictures obtained in natural scenes, X-ray images have the following characteristics: (1) Object overlap: When using an X-ray machine to scan luggage to generate X-ray images, the objects in the luggage are stacked on top of each other, resulting in interference information of other objects in the target object in the generated X-ray image, interfering with the detection of the target object; (2) Large variation in target scale: Due to the arbitrary placement of objects in the luggage, the target scale changes drastically; (3) High proportion of small targets: Since most contraband items are small objects, security personnel need to maintain a highly concentrated state of attention, which will cause fatigue of security personnel over a long time and lead to missed detections.

[0004] With the development of convolutional neural networks, breakthrough progress has been made in the X-ray security inspection image contraband detection method based on deep learning. Zhao et al. designed a de-overlapping block DOB to solve the problem of severe overlap in X-ray images. It uses high-level semantics as a guiding signal to reverse-filter out irrelevant information in the intermediate features; Wei et al. emphasized the edge information and material information of contraband items and designed a de-occlusion attention module; Wu et al. introduced the concept of an anchor-free detector into the contraband detection task and proposed a contraband detection network SA-CenterNet based on scale-adaptive centers; Xu et al. proposed a new idea of full-process feature fusion and local-global semantic dependence interaction to improve the automatic detection of contraband items and designed a single-stage contraband detection network PIXDet. However, the above-mentioned existing contraband detection algorithms have low detection accuracy in scenarios with significant intra-class variations. Summary of the Invention

[0005] To address the above problems, the present invention proposes an X-ray image contraband detection method based on de-overlapping and correlation attention mechanisms. By adding a classifier CLS to the backbone network of YOLOv7 and designing a top-down de-overlapping module TD-DOM, TD-DOM is added between the backbone and neck of the YOLOv7 network to remove the noise impact caused by the overlap between complex backgrounds and target objects, thereby improving the accuracy of contraband detection in security inspection X-ray images. Secondly, the feature fusion module FFM in the neck is combined with the designed correlation attention mechanism to obtain an enhanced feature fusion module EFFM, and the FFM in the neck is replaced with EFFM to improve the robustness of the model to target scale changes and enrich the extracted features. Finally, an additional anchor-free detection head AF-Head is added to the head of the YOLOv7 network. Without the need to pre-set anchor boxes, the number of positive samples for small targets increases, thus improving the accuracy of small target recognition. Moreover, the DFL loss function used effectively addresses significant intra-class variations by flexibly modeling the bounding boxes, further improving the detection accuracy of the model.

[0006] The X-ray image contraband detection method based on de-overlapping and correlation attention mechanisms includes the following steps:

[0007] Step (1): First, obtain an X-ray image dataset with target bounding boxes and class annotations, divide the dataset into a training set and a validation set, and preprocess the images in the training dataset.

[0008] Step (2): Add a classifier CLS at the end of the backbone network of YOLOv7, use the training set as input to the backbone network for feature extraction to obtain a set of feature maps F = {F i |F i ∈R C×H′×W′ , i = 3, 4, 5}, and then add a top-down de-overlapping module TD-DOM between the backbone network and the neck to output the set of feature maps optimized by TD-DOM

[0009] Step (3): Replace the feature fusion module FFM in the bidirectional fusion feature pyramid PANet structure in the neck of YOLOv7 with an enhanced feature fusion module EFFM, and send the set of feature maps output by the top-down de-overlapping module TD-DOM into the modified bidirectional fusion feature pyramid. The bidirectional fusion feature pyramid fuses the input multi-level features in a top-down and bottom-up manner to obtain a set of fused feature maps N = {N i |N i ∈R C×H′×W′ , i = 3, 4, 5}.

[0010] Step (4) Add an additional anchor-free detection head AF-Head to the head of the YOLOv7 network. Input the set of feature maps N output from step (3) into the head of the YOLOv7 network, and output the predicted target bounding boxes and classes of the model. and Use the regression loss function to calculate the error between the predicted bounding box and the ground truth bounding box, and use the BCE loss function to calculate the error between the predicted class and the ground truth class.

[0011] Step (5) Set the training parameters, input the training dataset obtained in step (1) into the model constructed by steps (2)-(4), validate on the validation set, obtain the optimal parameter model, and output the detection effect diagram of contraband.

[0012] Furthermore, step (2) is specifically as follows:

[0013] (2-1) Add a classifier CLS to the end of the backbone part of the YOLOv7 network. The backbone part of the network is divided into P2, P3, P4, and P5 layers by stage.

[0014] (2-2) Input the images t in the training dataset I into the backbone part of the network to extract features, and obtain the set of output feature maps of P3, P4, and P5 layers and the classification feature map F cls ; where F i represents the output feature map of the P th layer of the training image sample image i , and C i , H i , W i represent the number of feature channels, height, and width of the P i th layer respectively, and F cls represents the output feature map after adding the classifier.

[0015] (2-3) Add three TD-DOM mechanisms between the backbone and the neck of the network to remove the noise effect caused by the overlap of complex backgrounds and target objects. Input the set of feature maps F obtained from the backbone part of the network into the TD-DOM mechanism. The TD-DOM mechanism includes a channel selection part and a spatial activation part, and the specific design is as follows:

[0016] (2-3-1) Channel selection part: The input part is the feature F of the i-th layer i , the feature of the upper layer processed by the top-down de-overlap module TD-DOM mechanism and the classification feature map F cls . After inputting into TD-DOM, first obtain the channel descriptors corresponding to the three feature maps Low = Sig(PWC(GAP(Fi ))), and High = Sig(PWC(F cls )); where GAP(·) is global average pooling, PWC(·) is pointwise convolution, and Sig(·) is the Sigmoid activation function. Then the channel attention weight C att = Sig(PWC(Concat(High; Mid; Low))); where Concat is concatenation along the channel dimension, and then the current layer feature F i is weighted by channels to obtain

[0017] (2 - 3 - 2) The feature map obtained from the channel selection part is input into the spatial activation part, and the corresponding spatial position weight is obtained after passing through the spatial activation part where DWdConv7(·) is a 7×7 depthwise convolution operation, and DWConv7(·) is a 7×7 depthwise dilated convolution with a dilation rate of 4. Then is spatially weighted to obtain the finally optimized output feature map The above operations are performed on the set of feature maps F of the backbone part to obtain the corresponding optimized set of feature maps

[0018] Furthermore, step (3) is specifically as follows:

[0019] (3 - 1) Replace the feature fusion module FFM in the bottom - up fusion path of PANet with an enhanced feature fusion module EFFM, and the enhanced feature fusion module EFFM is obtained by combining the feature fusion module FFM with an association attention mechanism.

[0020] (3 - 2) Input the optimized set of feature maps obtained in step (2) into the bidirectional fusion feature pyramid of the improved neck to obtain the set of output feature maps of the N3, N4, and N5 layers of the neck Specifically as follows:

[0021] (3 - 2 - 1) First, perform top - down fusion on the optimized set of feature maps Assume the input is the high - level feature and the low - level feature The superscript i represents the i - th layer. The high - level feature is upsampled to the same size by bilinear interpolation, and the upsampled high - level feature and the original low - level feature are input into the feature fusion module FFM to obtain the output feature map X i , for the optimized set of feature maps Performing the above operations yields a set of feature maps for the top - down fusion path in the neck

[0022] (3 - 2 - 2) Perform bottom - up fusion on the set of feature maps X of the top - down fusion path. Assume the input is the high - level feature X i+1 and the low - level feature X i . The low - level feature is downsampled to X i+1 in size by mean downsampling. Input the downsampled high - level feature and the original low - level feature into the enhanced feature fusion module EFFM: First, input them into the feature fusion module FFM to obtain sets of feature maps at different layers and the feature map of the final aggregation layer where the subscript j represents the j - th layer before the final aggregation layer in the multi - layer concatenated convolutional structure. Then, input and into the relational attention mechanism RA to obtain N i+1 . The relational attention RA consists of two parts: the bridging attention module BA and the bridging spatial attention module BSA. The specific design is as follows:

[0023] (3 - 2 - 2 - 1) The BA module bridges the feature map of the final aggregation layer in the feature fusion module FFM and the sets of feature maps at different layers in the multi - layer concatenated convolutional structure Input the sets of feature maps and into the BA module, so that the final aggregation layer can obtain the feature weights for different channels according to the contribution degree of the relationship between the channels of the feature maps at different layers where GAP(·) is global average pooling, retaining the original dimension, Sum(·) is element - wise addition, ReLU(·) is the ReLU activation function, is a 1×1 convolution, j represents the j - th convolution, e represents the convolution of the final aggregation layer, and Sig(·) is the Sigmoid activation function. Finally, perform attention enhancement on the input to obtain the required feature map

[0024] (3 - 2 - 2 - 2) Input the sets of feature maps and into the BSA module to obtain the relational spatial attention weights

[0025] where Ap(·) is the average pooling operation along the channel dimension, MP(·) is the max - pooling operation along the channel dimension, Conv7(·) is a 7×7 convolution, Conv1(·) is a 1×1 convolution, Conv21 (·) is a 21×21 convolution, and Sig(·) is the Sigmoid activation function. The output of (3-2-2-1) is enhanced to obtain the final output feature map of the module

[0026] (3-2-3) performs the operations in step (3-2-2) on the set of feature maps X to obtain the final set of feature maps at the neck

[0027] Furthermore, step (4) is specifically:

[0028] (4-1) Add an additional anchor-free detection head AF-Head to the head of the YOLOv7 network, and then input the output feature set N of the network neck into the two detection heads at the network head, namely the anchor-based detection AB-Head and the anchor-free detection AF-Head. The final set of predicted output feature maps is obtained and where the value of the output channel C is 3 * (total number of target categories + 5).

[0029] (4-2) Screen the feature map H obtained by AB-Head through the OTA label assignment strategy AB to obtain positive samples, and use the CIOU loss function to calculate the error between the predicted bounding box and the ground truth box in the positive samples Use the cross-entropy loss function to calculate the error between the predicted category and the ground truth category

[0030] (4-3) Screen the feature map H obtained by AF-Head through the TAL label assignment strategy AF to obtain positive samples, and use the CIOU loss function and the DFL loss function to calculate the error between the predicted bounding box and the ground truth box in the positive samples Use the BCE loss function to calculate the error between the predicted category and the ground truth category

[0031] Furthermore, during the training process of step (5), according to the calculation method in (4), calculate the target classification loss and the bounding box regression loss Iteratively train the model until convergence to obtain the optimal parameters of the X-ray security inspection image contraband detection model.

[0032] Compared with the prior art, the beneficial effects of the present invention are:

[0033] The present invention proposes an X-ray image contraband detection method based on de-overlapping and correlation attention mechanism, which improves the YOLOv7 model. 1) By using the designed top-down de-overlapping module TD-DOM to solve the problem of object overlap and improve the detection accuracy; 2) The feature fusion module FFM in the neck is combined with the designed correlation attention mechanism to obtain the enhanced feature fusion module EFFM, and the FFM in the neck is replaced with EFFM to improve the robustness of the model to target scale changes and enrich the extracted features; 3) A detection head AF-Head without anchor boxes is added to the head of the YOLOv7 network. The setting without anchor boxes increases the number of positive samples of small targets, thereby improving the accuracy of small target recognition. The DFL loss function used effectively deals with significant intra-class variations by flexibly modeling the bounding boxes, further improving the detection accuracy of the model. Compared with the existing X-ray image contraband detection models and the original YOLOv7 algorithm, the detection accuracy is significantly improved, and a high detection speed is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the processing flow chart of the present invention;

[0035] Figure 2 is the overall structure diagram of the network model of the present invention;

[0036] Figure 3 is the structure diagram of the top-down de-overlapping module TD-DOM in the present invention;

[0037] Figure 4 is the structure diagram of the enhanced feature fusion module EFFM in the present invention;

[0038] Figure 5 is the structure diagram of the correlation attention mechanism RA in the present invention;

[0039] Figure 6 is the comparison diagram of the detection effects of the original model and the improved model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] In order to better understand the purpose, structure and function of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the drawings and specific preferred embodiments.

[0041] This example presents an X-ray image contraband detection method based on de-overlapping and correlation attention mechanisms. First, a classifier CLS is added at the end of the network backbone, and then a top-down de-overlapping module TD-DOM is added between the backbone and the neck to solve the problem of object overlap. Second, the feature fusion module FFM in the neck is combined with the designed correlation attention mechanism to obtain an enhanced feature fusion module EFFM, and the FFM in the neck is replaced with EFFM, thereby improving the robustness of the model to target scale changes and enriching the extracted features. Finally, an additional anchor-free detection head AF-Head is added to the head of the YOLOv7 network. The anchor-free setting increases the number of positive samples of small targets, thereby improving the accuracy of small target recognition. The DFL loss function used effectively deals with significant intra-class variations by flexibly modeling the bounding box, further improving the detection accuracy of the model. The overall architecture of the finally obtained model is as Figure 2 shown. Compared with the existing X-ray image contraband detection algorithms and the original YOLOv7 algorithm, the accuracy of the improved YOLOv7 model is significantly improved, and it has a high detection speed.

[0042] As Figure 1 shown, the method of the present invention includes the following steps:

[0043] Step (1): Obtain the X-ray image dataset. Download the OPIXray dataset from the official website where OPIXray is located, divide the X-ray image dataset into a training set and a validation set, and perform preprocessing operations on the images. Specifically:

[0044] (1-1) Divide the obtained X-ray image dataset into a training set and a validation set where R is the real number field, N t represents the number of image samples in the training set, represents the i-th training image sample, N v represents the number of image samples in the validation set, represents the j-th validation image sample, H represents the image height, W represents the image width, and 3 represents the number of RGB channels.

[0045] (1-2) Each training image sample has a corresponding label where represents the number of targets included in the image sample , represents the true class index of the -th target in where C represents the total number of classes in the dataset, The bounding box of the th target is composed of the abscissa x of the center point, the ordinate y of the center point, the width w of the target, and the height h.

[0046] (1 - 3) Preprocess the image before the input network backbone part, and use mosaic data augmentation: randomly select four images from the data training set, and perform augmentation operations such as flipping, scaling, and color gamut transformation, and then perform random splicing of the images. The original labels are still retained during the splicing process to form a new sample image. The image pixels are adaptively scaled to 640 * 640, and the values of the pixel points are normalized.

[0047] Step (2) Construct an additional classifier CLS and a top - down de - overlapping module TD - DOM for solving the overlapping problem. TD - DOM is as Figure 3 shown. The input is the pre - processed training set image, and the output is a set of optimized feature maps filtered by the top - down de - overlapping module TD - DOM Specifically:

[0048] (2 - 1) Add a classifier CLS at the end of the YOLOv7 network backbone part. The backbone part of the network is divided into P2, P3, P4, and P5 layers by stage.

[0049] (2 - 2) Input the images t in the training data set I into the network backbone part to extract features, and obtain a set of output feature maps of the P3, P4, and P5 layers and the classification feature map F cls ; among them, F i represents the output feature map of the P i th layer of the training image sample image i , C i , H i , W i represent the number of feature channels, height, and width of the P cls th layer respectively. F

[0050] (2 - 3) As Figure 2 shown, add three TD - DOM mechanisms between the network backbone and the neck to remove the noise influence caused by the overlap between the complex background and the target object. Input the set of feature maps F obtained from the network backbone part into TD - DOM. The TD - DOM mechanism includes a channel selection part and a spatial activation part. As Figure 3 shown, the specific design is as follows:

[0051] (2-3-1) Channel selection part: By interacting with high-level semantic information, intermediate features, and low-level features, the information related to the target object in the low-level feature map is enhanced. The input part is the feature Fi of the i-th layer i , the features of the upper layer processed by the top-down de-overlapping module TD-DOM , and the classification feature map F cls . After inputting into TD-DOM, first obtain the channel descriptors corresponding to the three feature maps: Low = Sig(PWC(GAP(F i ))), and High = Sig(PWC(F cls )); where GAP(·) is global average pooling, PWC(·) is pointwise convolution, and Sig(·) is the Sigmoid activation function. Obtain the channel attention weight C att = Sig(PWC(Concat(High; Mid; Low))); where Concat is concatenation along the channel dimension. Then perform channel weighting on the current layer feature F i to obtain , thus achieving the purpose of removing the noise caused by the overlap of complex backgrounds and target objects.

[0052] (2-3-2) Input the feature map obtained from the channel selection part into the spatial activation part, and obtain the corresponding spatial position weight after passing through the spatial activation part . Among them, DWConv7(·) is a 7×7 depthwise convolution operation, and DWdConv7(·) is a 7×7 depthwise dilated convolution with a dilation rate of 4. Then perform spatial weighting on to obtain the final optimized output feature map Perform the above operations on the set of feature maps F of the backbone part to obtain the corresponding optimized set of feature maps

[0053] Step (3) Construct an enhanced feature fusion module EFFM as shown in Figure 4 , replace the feature fusion module FFM in the neck, and the input is the optimized set of feature maps in step (2) The output is the set of feature maps N after multi-scale feature fusion in the neck. Specifically:

[0054] (3-1) Modify the bidirectional fusion feature pyramid PANet structure of the YOLOv7 neck, as shown in Figure 4As shown within the dashed box, the Feature Fusion Module (FFM) consists of: sampling the input upper and lower layer feature maps to the same size, then concatenating these two layers of feature maps and inputting them into a multi-layer concatenated convolutional structure to obtain a set of feature maps at different layers, concatenating this set of feature maps, and aggregating features through a 1×1 convolution to obtain the final output feature map. Replace the Feature Fusion Module (FFM) in the bottom-up fusion path of PANet with the Enhanced Feature Fusion Module (EFFM). As Figure 4 shown, the Enhanced Feature Fusion Module (EFFM) is obtained by combining the Feature Fusion Module (FFM) with the Relevance Attention Mechanism (RA). The Relevance Attention Mechanism (RA) is as Figure 5 shown.

[0055] (3-2) Input the set of optimized feature maps obtained in step (2) into the bidirectional fusion feature pyramid of the improved neck to obtain the set of output feature maps of the N3, N4, and N5 layers of the neck Specifically as follows:

[0056] (3-2-1) First, perform top-down fusion on the set of optimized feature maps Assume the input is the high-level feature and the low-level feature The superscript i represents the i-th layer. The high-level feature is upsampled to the same size through bilinear interpolation. Input the upsampled high-level feature and the original low-level feature into the Feature Fusion Module (FFM) to obtain the output feature map X i , and perform the above operations on the set of optimized feature maps to obtain the set of feature maps in the top-down fusion path in the neck

[0057] (3-2-2) Perform bottom-up fusion on the set of feature maps X in the top-down fusion path. Assume the input is the high-level feature X i+1 and the low-level feature X i . The low-level feature is downsampled to the size of X i+1 through mean downsampling. Input the downsampled high-level feature and the original low-level feature into the Enhanced Feature Fusion Module (EFFM): First, input them into the Feature Fusion Module (FFM) to obtain a set of feature maps at different layers and the final aggregated feature map where the subscript j represents the j-th layer before the final aggregation layer in the multi-layer concatenated convolutional structure. Then input and into the Relevance Attention Mechanism (RA) to obtain N i+1。The related attention RA consists of two parts: the bridging attention module BA and the bridging spatial attention module BSA. The specific design is as follows:

[0058] (3-2-2-1) The BA module bridges the feature maps of the last aggregation layer in the feature fusion module FFM and the set of feature maps of different layers in the multi-layer cascaded convolution structure Input the set of feature maps and into the BA module, so that the final aggregation layer can obtain the feature weights for different channels according to the contribution degree of the relationship between the channels of different layers of feature maps to the current layer Among them, GAP(·) is global average pooling, retaining the original dimension, Sum(·) is element-wise addition, ReLU(·) is the ReLU activation function, is a 1×1 convolution, j represents the j-th layer of convolution, e represents the convolution of the last aggregation layer, Sig(·) is the Sigmoid activation function. Finally, the input is attentively enhanced to obtain the required feature maps

[0059] (3-2-2-2) Input the set of feature maps and into the BSA to obtain the related spatial attention weights

[0060] Among them, AP(·) is the average pooling operation along the channel dimension, MP(·) is the maximum pooling operation along the channel dimension, Conv7(·) is a 7×7 convolution, Conv1(·) is a 1×1 convolution, Conv 21 (·) is a 21×21 convolution, Sig(·) is the Sigmoid activation function. Then, enhance the output of (3-2-2-1) to obtain the final module output feature maps

[0061] (3-2-3) Perform the operations in step (3-2-2) on the set of feature maps X to obtain the final set of feature maps at the neck

[0062] Step (4) is as shown in the head part of Figure 2 . Add the AF-Head module to the head of the YOLOv7 network, use the detection mechanism of two detection heads, the input is the set of feature maps of the N3, N4, and N5 layers, and the output is the set of predicted feature maps of the two detection heads. Specifically:

[0063] (4-1) Add an additional anchor-free detection head, abbreviated as AF-Head, to the head of the YOLOv7 network. Then, input the output feature set N of the network neck into the two detection heads of the network head, namely the anchor-based head (AB-Head) based on anchor box detection and the AF-Head without anchor boxes. The final predicted output feature map set is obtained. and Among them, the value of the output channel C is 3 * (total number of target categories + 5).

[0064] (4-2) Screen the feature map H obtained by the AB-Head through the OTA label assignment strategy AB to obtain positive samples, and use the CIOU loss function to calculate the error between the predicted box and the ground truth box in the positive samples Use the cross-entropy loss function to calculate the error between the predicted category and the ground truth category

[0065] (4-3) Screen the feature map H obtained by the AF-Head through the TAL label assignment strategy AF to obtain positive samples, and use the CIOU loss function and the DFL loss function to calculate the error between the predicted box and the ground truth box in the positive samples Use the BCE loss function to calculate the error between the predicted category and the ground truth category

[0066] In step (5), set the training parameters, input the training set into the object detection network model constructed in steps (2)-(4), use the Adam optimizer for iterative training to obtain the optimal weight model, and then apply it to the validation set to output the final detection effect diagram. Specifically:

[0067] (5-1) Set the training parameters, including the initial learning rate, momentum parameter, decay coefficient, batch size, number of GPUs, optimizer, number of training times, etc., and input the preprocessed training set into the network model constructed in steps (2)-(4).

[0068] (5-2) Calculate the object classification loss and the bounding box regression loss Use the Adam optimization algorithm to train the model, observe the mAP of the model on the validation set, where the calculation method of mAP is the mean of the average precision AP of all categories, and iterate to train the model until convergence to obtain the optimal parameters of the X-ray security inspection image contraband detection model.

[0069] (5-3) Load the optimal model parameters obtained in step (5-2) into the model, select several pictures in the validation set as the pictures to be detected and input them into the model, and the model will mark the bounding boxes of the detected targets and the corresponding categories of the targets on the pictures. As Figure 6 shown, the improved YOLOv7 model can detect target objects that the original YOLOv7 model cannot detect.

[0070] The above examples further elaborate on the purpose, technical solutions and advantages of the present invention. It should be understood that the above examples are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An X-ray image contraband detection method based on a de-overlapping and correlation attention mechanism, characterized in that, It includes the following steps: Step 1: Obtain an X-ray image dataset containing target bounding boxes and class annotations, divide the dataset into training set I t and validation set I v , and preprocess the images in the training dataset; Step 2: Add a classifier CLS at the end of the backbone part of YOLOv7. Input the training set into the backbone part of the network to obtain a set of feature maps F. Then, add a top-down de-overlapping module TD-DOM between the backbone and the neck, and output the set of feature maps optimized by TD-DOM The specific process is as follows: 2-1. Add a classifier CLS at the end of the backbone part of the YOLOv7 network. The backbone part of the network is divided into P2, P3, P4, and P5 layers by stage; 2-2. Training dataset I t The images in are input into the backbone of the network to extract features, obtaining the output feature map sets of the P3, P4, and P5 layers and the classification feature map F output by the classifier cls ; among them, F i represents the output feature map of the P th layer of the training image sample image i , C i , H i , W i respectively represent the number of feature channels, height, and width of the P i th layer; 2-3. Add three TD-DOM mechanisms between the network backbone and the neck, input the feature map set F into the TD-DOM mechanism, and obtain the feature map set The TD-DOM mechanism includes a channel selection part and a spatial activation part, and the specific design is as follows: 2-3-1. Channel selection part: The input part is the feature F of the i-th layer i , the features of the upper layer processed by the top-down de-overlapping module TD-DOM mechanism and the classification feature map F cls . After inputting into TD-DOM, the channel descriptors corresponding to the three feature maps are obtained: Low = Sig(PWC(GAP(F i ))), and High = Sig(PWC(F cls )); where GAP(·) is global average pooling, PWC(·) is pointwise convolution, and Sig(·) is the Sigmoid activation function; the channel attention weight C att = Sig(PWC(Concat(High; Mid; Low))); where Concat is concatenation along the channel dimension; then the current layer feature F i is weighted by channels to obtain 2-3-2. Obtain the feature map by selecting the channel part Input it into the spatial activation part, and obtain the corresponding spatial position weights after passing through the spatial activation part Among them, DWdConv7(·) is a 7×7 depthwise convolution operation, and DWConv7(·) is a 7×7 depthwise dilated convolution with a dilation rate of 4; then Perform spatial weighting to obtain the optimized output feature map Perform the above operations on the set of feature maps F of the backbone part to obtain the corresponding optimized set of feature maps Step 3: Replace the feature fusion module FFM in the bidirectional fusion feature pyramid PANet structure of YOLOv7 neck with the enhanced feature fusion module EFFM, and send the feature map set to the modified bidirectional fusion feature pyramid to obtain the fused feature map set N. The specific process is as follows: 3-1. Replace the feature fusion module FFM in the bottom-up fusion path of PANet with an enhanced feature fusion module EFFM. The enhanced feature fusion module EFFM is obtained by combining the feature fusion module FFM with the correlation attention mechanism; 3-2. Input the optimized feature map set into the bidirectional fusion feature pyramid of the improved neck to obtain the output feature map sets of the N3, N4, and N5 layers of the neck The specific process of obtaining the output feature map set N is as follows: 3-2-1. For the optimized feature map set perform top-down fusion. Assume the input is the high-level feature and the low-level feature The superscript i represents the i-th layer; the high-level feature is upsampled to the same size by bilinear interpolation. The upsampled high-level feature and the original low-level feature are input into the feature fusion module FFM to obtain the output feature map X i For the optimized feature map set perform the above operations to obtain the feature map set of the top-down fusion path in the neck 3-2-2. Perform bottom-up fusion on the set of feature maps X in the top-down fusion path. Assume the input is the high-level feature X i +1 and the low-level feature X i . The low-level feature is downsampled to X i+1 size by mean. The downsampled high-level feature and the original low-level feature are input into the enhanced feature fusion module EFFM: First, they are input into the feature fusion module FFM to obtain the set of feature maps at different layers and the feature map of the last aggregation layer where the subscript j indicates the j-th layer before the final aggregation layer in the multi-layer concatenated convolutional structure; then and are input into the relational attention mechanism RA to obtain N i+1 ; The correlation attention RA consists of two parts: a bridging attention module BA and a bridging spatial attention module BSA. The specific design is as follows: 3-2-2-1. Feature maps of the last aggregation layer in the BA module bridging the Feature Fusion Module FFM And the set of feature maps of different layers in the multi-layer cascaded convolutional structure Input the set of feature maps and into the BA module, so that the aggregation layer can obtain the feature weights for different channels according to the contribution degree of the channels between the feature maps of different previous layers to the channels of the aggregation layer Among them, GAP(·) is global average pooling, retaining the original dimension, Sum(·) is element-wise addition, ReLU(·) is the ReLU activation function, is a 1×1 convolution, j represents the j-th layer of convolution, e is the convolution of the last aggregation layer, Sig(·) is the Sigmoid activation function; perform attention enhancement on the input to obtain the feature map 3-2-2-2. Input and into the BSA module to obtain the associated spatial attention weights: where AP(·) is the average pooling operation along the channel dimension, MP(·) is the max pooling operation along the channel dimension, Conv7(·) is the 7×7 convolution, Conv1(·) is the 1×1 convolution, and Conv 21 (·) is the 21×21 convolution, Sig(·) is the Sigmoid activation function; the output of (3-2-2-1) is enhanced to obtain the output feature map of the module 3-2-3. Perform the operation in step 3-2-2 on the set of feature maps X to obtain the final set of feature maps for the neck Step 4. Add an additional anchor-free detection head to the head of the YOLOv7 network. Input the feature map set N into the head of the YOLOv7 network, output the target bounding boxes and categories predicted by the model, and calculate the errors between the predicted values and the true values respectively; Step 5. Set the training parameters, input the preprocessed training data set into the model constructed in steps 2-4, perform iterative training, verify on the validation set, and output the detection effect diagram of contraband; 2. The X-ray image contraband detection method based on a de-overlapping and correlation attention mechanism according to claim 1, characterized in that, The specific process of Step 4 is as follows: 4-1. Add an additional anchor-free detection head AF-Head to the head of the YOLOv7 network, and then input the output feature set N of the network neck into the two detection heads of the network head, namely the anchor-based detection AB-Head and the anchor-free detection AF-Head, to obtain the predicted output feature map sets H AB and H AF ; 4-2. Screen the feature map H obtained by AB-Head through the OTA label assignment strategy to obtain positive samples, and use the CIOU loss function to calculate the error between the predicted box and the ground truth box in the positive samples AB Use the cross-entropy loss function to calculate the error between the predicted category and the ground truth category Use the cross-entropy loss function to calculate the error between the predicted category and the ground truth category 4-3. Screen the feature map H obtained by the AF-Head through the TAL label assignment strategy to obtain positive samples, and use the CIOU loss function and the DFL loss function to calculate the error between the predicted box and the ground truth box in the positive samples AF Use the BCE loss function to calculate the error between the predicted category and the ground truth category Use the BCE loss function to calculate the error between the predicted category and the ground truth category 3. The X-ray image contraband detection method based on a de-overlapping and correlation attention mechanism according to claim 2, characterized in that, During the training process described in step 5, the calculation of the loss is specifically as follows: calculate the target classification loss and the bounding box regression loss Iteratively train the model until convergence.

Citation Information

Patent Citations

  • Less-sample character style migration method based on associated attention

    CN114742014A

  • Small sample pipeline defect intelligent identification system and method

    CN115049600A