Contraband detection method and system based on mixed self-attention and implicit feature fusion
Through the contraband detection system that combines self-attention and implicit features, the problem of difficult extraction and high computational complexity of contraband characteristics in existing X-ray security machines is solved, and efficient and accurate contraband detection is achieved.
Patent Information
- Application Number
- CN202510444975.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult for existing X-ray security machines to effectively extract the characteristics of contraband, especially contrabands with severe overlap, lack of obvious characteristics and large size differences. The calculation complexity of self-attention mechanisms leads to low detection accuracy.
The contraband detection system based on hybrid self-attention and implicit feature fusion is adopted, including data preprocessing, complex feature parallel extraction module, implicit feature fusion module and detection head module. The resolution is reduced and global features are extracted through the hybrid self-attention mechanism. The high-resolution features are mapped using implicit fusion paths to solve the feature aliasing problem caused by upsampling in the feature pyramid of the neck network.
It significantly improves the accuracy and efficiency of contraband detection, reduces the computational complexity, maintains lightweight and has the performance of self-attention mechanism, solves the feature aliasing problem, and improves the accuracy of contraband detection.
Smart Images

Figure CN120339639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and object detection, and particularly to a contraband detection method and system based on hybrid self-attention and implicit feature fusion. Background Art
[0002] With the rapid development of social economy and technology, the flow of people is becoming more and more frequent. In order to maintain public safety and transportation safety, X-ray security inspection machines are widely used in crowded transportation hub areas. Modern dual-energy X-ray security inspection machines use two discrete energy levels to identify different types of materials. Since the internal structures of objects with different material compositions, different thicknesses, and different densities absorb X-ray energy to different degrees, X-rays have different penetrabilities for different objects. Based on this characteristic, dual-energy X-ray security inspection devices on the market emit X-ray beams to penetrate the items to be inspected on the conveyor belt and receive X-ray photons with different energies after penetration, obtaining original high-energy and low-energy X-ray fluoroscopic images. Then, combined with a pseudo-color mapping algorithm, different colors are assigned to the equivalent atomic numbers of different substances and finally displayed as a pseudo-color image for easy human observation and identification. Through a specific generation function, the identified information is fused into a pseudo-color X-ray image for easy interpretation of the contents of the luggage.
[0003] However, such X-ray security inspection machines still have the following deficiencies:
[0004] (1) Since X-rays are penetrative, in X-ray security inspection images, the features of contraband will be mixed with those of other non-contraband items, making it difficult for detectors to extract effective features. In addition, there are significant size differences among contraband items of the same category. The current backbone network based on Convolutional Neural Network (CNN) overly focuses on local detailed features, which may lead to false detections or missed detections. However, directly using the Self-Attention mechanism to extract global features will bring huge computational overhead and is not conducive to on-site deployment.
[0005] (2) The feature pyramid fusion strategy widely used in the neck network of existing detection methods will exacerbate the feature mixing of contraband. Current detectors need to obtain feature maps with different resolutions through downsampling and upsampling to achieve the fusion of different receptive fields. Currently, commonly used upsampling operators will cause feature mixing, making it more difficult for detectors to capture the weak features of contraband.
[0006] In addition, the current X-ray image inspection method still mainly relies on manual inspection. The efficiency of manual inspection is relatively low, and missed detections are likely to occur. Especially under high-intensity work, it will greatly affect the judgment of the inspector and pose a threat to public safety.
[0007] Therefore, there is an urgent need to propose a detection method that can automatically and efficiently identify contraband in X-ray images. Summary of the Invention
[0008] The purpose of the present invention is to address the weak ability of the backbone network of the existing detection model to extract global features and the high computational complexity of feature extraction through the self-attention mechanism; especially for contraband with severe overlap, unclear features, and large size differences, the existing models and methods have high complexity and low detection accuracy. A contraband detection method and system based on hybrid self-attention and implicit feature fusion are proposed, which can solve the feature aliasing problem caused by upsampling in the feature pyramid of the neck network.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] In the first aspect, the present invention provides a contraband detection system based on hybrid self-attention and implicit feature fusion, including a data preprocessing module, a complex feature parallel extraction module, an implicit feature fusion module, and a detection head module that are sequentially connected in series;
[0011] Among them, the data preprocessing module is used to preprocess the contraband detection image;
[0012] The complex feature parallel extraction module is used to reduce the resolution of the preprocessed contraband image, expand the number of channels, and extract global features based on the hybrid self-attention mechanism, and output three groups of feature maps P3, P4, and P5 with different resolutions;
[0013] The implicit feature fusion module is used to map the low-resolution features of the feature maps P3, P4, and P5 to high-resolution features based on the implicit fusion path, and output three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions; the implicit fusion path is implemented by the following method:
[0014] Construct a high-resolution feature coordinate network, encode the low-resolution features into the latent space, interpolate the latent space to the high-resolution feature coordinate network based on the point sampling function, obtain a high-resolution latent encoding map, and calculate the nearest latent encoding value of each coordinate in the high-resolution latent encoding map based on the neighboring latent encoding weighted fusion mechanism, and decode the nearest latent encoding value to obtain the high-resolution feature value;
[0015] The detection head module is used to detect the feature maps O3, O4, and O5 and output the detection results.
[0016] As a possible implementation, the complex feature parallel extraction module includes a first downsampling layer, a second downsampling layer, a first cross-stage local layer, a third downsampling layer, a second cross-stage local layer, a fourth downsampling layer, a first HyAtt-CNN layer, a fifth downsampling layer, a second HyAtt-CNN layer, and a spatial pyramid pooling sub-module that are sequentially connected in series.
[0017] As a possible implementation, the preprocessed contraband image is passed through the first downsampling layer to reduce the resolution and expand the number of channels to obtain a feature map P1; the feature map P1 is passed through the second downsampling layer to reduce the resolution and expand the number of channels, and then enters the first cross-stage local layer to extract local features, obtaining a feature map P2; the feature map P2 is passed through the third downsampling layer to reduce the resolution and expand the number of channels, and then enters the second cross-stage local layer to extract local features, obtaining a feature map P3; the feature map P3 is passed through the fourth downsampling layer to reduce the resolution and expand the number of channels, and then enters the first HyAtt-CNN layer to extract local and global features, obtaining a feature map P4; the feature map P4 is passed through the fifth downsampling layer to reduce the resolution and expand the number of channels, and then enters the second HyAtt-CNN layer to extract local and global features, and different-scale feature fusion is performed by the spatial pyramid pooling sub-module to output a feature map P5.
[0018] As a possible implementation, both the first HyAtt-CNN layer and the second HyAtt-CNN layer include a first convolutional path, a second convolutional path, and a third convolutional path. The first convolutional path includes a first convolutional block for retaining the features extracted in the previous stage of the feature map input to the HyAtt-CNN layer; the second convolutional path includes a second convolutional block, a fourth convolutional block, and a fifth convolutional block that are sequentially connected in series for extracting local features of the feature map input to the HyAtt-CNN layer; the third convolutional path includes a third convolutional block and a HyAtt block that are sequentially connected in series for extracting global features of the feature map input to the HyAtt-CNN layer; the outputs of the first convolutional path, the second convolutional path, and the third convolutional path are concatenated along the channel dimension and then output as the feature map P4 or P5.
[0019] As a possible implementation, the HyAtt block includes a first reshaping layer, a second reshaping layer, a first normalization layer, a second normalization layer, a third normalization layer, a global pooling layer, a first Softmax activation function layer, a second Softmax activation function layer, a first fully connected layer, a second fully connected layer, a third fully connected layer, and a fourth fully connected layer;
[0020] The input feature map of the HyAtt block is transformed in dimension by the first reshaping layer and then passed through the matrix W included in the first fully connected layer Q , the matrix W included in the second fully connected layer K and the matrix W included in the third fully connected layer V, three matrices Q, K, and V after dimensionality reduction are obtained; the first normalization layer normalizes matrix Q, the second normalization layer normalizes matrix K, and matrix V is multiplied by the normalized matrix K and then multiplied by the normalized matrix Q to obtain the first intermediate output feature;
[0021] The global pooling layer performs an aggregation operation on matrix Q and then multiplies it by matrix Q, and outputs after passing through the first Softmax activation function layer; the global pooling layer performs an aggregation operation on matrix Q and then multiplies it by matrix K, and outputs after passing through the second Softmax activation function layer; the output of the first Softmax activation function layer is multiplied by the output of the second Softmax activation function layer and then multiplied by matrix V to obtain the second intermediate output feature;
[0022] The first intermediate output feature and the second intermediate output feature are added together, normalized by the third normalization layer, and then enter the fourth fully connected layer. The output of the fourth fully connected layer is adjusted in dimension by the second reshaping layer to obtain the output of the HyAtt block.
[0023] As a possible implementation, the implicit feature fusion module includes a first implicit fusion path, a second implicit fusion path, a third implicit fusion path, a third cross-stage local layer, a fourth cross-stage local layer, a fifth cross-stage local layer, a sixth cross-stage local layer, a seventh cross-stage local layer, a sixth downsampling layer, a seventh downsampling layer, and an eighth downsampling layer;
[0024] The first implicit fusion path is used to increase the resolution of feature map P5, the second implicit fusion path is used to increase the resolution of feature map P4, and feature map P3 is concatenated with the upsampled P4 and the upsampled P5, and the number of channels is compressed by the third cross-stage local layer to output feature map O3;
[0025] Feature map O3 is processed by the sixth downsampling layer and concatenated with P4. After concatenation, the number of channels is compressed by the fourth cross-stage local layer to obtain the intermediate output M4; the intermediate output M4 is processed by the seventh downsampling layer and concatenated with P5. After concatenation, the number of channels is compressed by the fifth cross-stage local layer to obtain the intermediate output M5; the intermediate output M5 is upsampled by the third implicit fusion path and concatenated with the intermediate output M4. After concatenation, the number of channels is compressed by the sixth cross-stage local layer to output feature map O4;
[0026] Feature map O4 is processed by the eighth downsampling layer and concatenated with the intermediate output M5. After concatenation, the number of channels is compressed by the seventh cross-stage local layer to output feature map O5.
[0027] As a possible implementation, the detection head module includes three detection heads, which respectively detect feature maps O3, O4, and O5, and output the detection results after non-maximum suppression.
[0028] Second aspect, the present invention provides a contraband detection method based on hybrid self-attention and implicit feature fusion, which is implemented by using the contraband detection system based on hybrid self-attention and implicit feature fusion provided in the first aspect. The contraband detection method includes:
[0029] S1. Configure the contraband detection image and preprocess the contraband detection image;
[0030] S2. Input the preprocessed contraband detection image into the complex feature parallel extraction module to extract the global feature and local feature of the contraband, and obtain three groups of feature maps P3, P4, and P5 with different resolutions;
[0031] S3. Input the three groups of feature maps P3, P4, and P5 with different resolutions into the implicit feature fusion module, and obtain three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions through three implicit fusion paths;
[0032] S4. Use 3 detection heads to detect the three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions respectively, and output the detection results.
[0033] As a possible implementation, S2 includes:
[0034] S20. Perform two downsampling processes, one local feature extraction, another downsampling process, and another local feature extraction on the preprocessed contraband detection image in sequence to obtain the feature map P3;
[0035] S21. Perform one downsampling process and local-global feature extraction on the feature map P3 in sequence to obtain the feature map P4;
[0036] S22. Perform one downsampling process, local-global feature extraction, and spatial pyramid pooling on the feature map P4 in sequence to obtain the feature map P5.
[0037] As a possible implementation, S3 includes:
[0038] S30. Increase the resolutions of the feature map P4 and the feature map P5 respectively, cascade the feature map P3 with the P4 and P5 after increasing the resolutions, and perform channel number compression to obtain the feature map O3;
[0039] S31. Perform downsampling processing on the feature map O3 and then cascade it with the feature map P4. After cascading, perform channel number compression to obtain the intermediate output M4. Perform downsampling layer processing on the intermediate output M4 and then cascade it with the feature map P5. After cascading, perform channel number compression to obtain the intermediate output M5. Increase the resolution of the feature map M5 and then cascade it with the intermediate output M4. After cascading, perform channel number compression to obtain the feature map O4;
[0040] After performing downsampling layer processing on the feature map O4 and cascading it with the intermediate output M5, channel number compression is performed after cascading to obtain the feature map O5.
[0041] Advantageous Effects
[0042] A contraband detection method and system based on hybrid self-attention and implicit feature fusion proposed by the present invention has the following advantageous effects compared with the prior art:
[0043] 1. The contraband detection system based on hybrid self-attention and implicit feature fusion proposed by the present invention designs a lightweight hybrid self-attention mechanism (HyAtt), reducing the computational complexity of the self-attention mechanism from O(N 2 ) to O(N);
[0044] 2. The contraband detection system and method based on hybrid self-attention and implicit feature fusion proposed by the present invention combine the advantages of the self-attention mechanism at the same time, retain the Softmax activation function, and enable the hybrid self-attention (HyAtt) to have performance comparable to that of the self-attention mechanism while remaining lightweight.
[0045] 3. The contraband detection system based on hybrid self-attention and implicit feature fusion proposed by the present invention combines HyAtt and CNN, designs a HyAtt-CNN layer, and further forms a complex feature parallel extraction module, which can extract global features and local features simultaneously in the backbone network, significantly improving the ability of the system to extract fuzzy features of contraband;
[0046] 4. The contraband detection system based on hybrid self-attention and implicit feature fusion proposed by the present invention designs an implicit feature fusion module, replaces the upsampling operator commonly used in the existing method with an implicit function representation, and reconstructs the fusion path of high-resolution features, solving the feature aliasing problem in the existing method and further improving the accuracy of contraband detection. Description of the Drawings
[0047] The drawings described herein are used to provide a further understanding of the present invention, form a part of the present invention, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0048] Figure 1 It is a schematic structural diagram of the contraband detection system based on hybrid self-attention and implicit feature fusion in the embodiment of the present invention;
[0049] Figure 2 It is a schematic structural diagram of the complex feature parallel extraction module included in the contraband detection system in the embodiment of the present invention;
[0050] Figure 3Schematic diagram of the HyAtt-CNN layer structure included in the complex feature parallel extraction module in the embodiment of the present invention;
[0051] Figure 4 Schematic diagram of the HyAtt block structure included in the HyAtt-CNN layer in the embodiment of the present invention;
[0052] Figure 5 Schematic diagram of the implicit feature fusion module structure included in the contraband detection system in the embodiment of the present invention;
[0053] Figure 6 Schematic diagram of the implicit fusion path structure included in the implicit feature fusion module in the embodiment of the present invention;
[0054] Figure 7 Flowchart of the contraband detection method based on hybrid self-attention and implicit feature fusion in the embodiment of the present invention;
[0055] Figure 8 Detection effect diagram obtained by experiments using the detection system and detection method of the present invention in the embodiment of the present invention;
[0056] Reference numerals:
[0057] 10 - Data preprocessing module;
[0058] 20 - Complex feature parallel extraction module, 201 - First downsampling layer, 202 - Second downsampling layer, 203 - First cross-stage local layer, 204 - Third downsampling layer, 205 - Second cross-stage local layer, 206 - Fourth downsampling layer, 207 - First HyAtt-CNN layer, 2070 - First convolutional block, 2071 - Second convolutional block, 2072 - Fourth convolutional block, 2073 - Fifth convolutional block, 2074 - Third convolutional block, 2075 - HyAtt block, 20750 - First reshaping layer, 20751 - Second reshaping layer, 20752 - First normalization layer, 20753 - Second normalization layer, 20754 - Third normalization layer, 20755 - Global pooling layer, 20756 - First Softmax activation function layer, 20757 - Second Softmax activation function layer, 20758 - First fully connected layer, 20759 - Second fully connected layer, 20760 - Third fully connected layer, 20761 - Fourth fully connected layer, 208 - Fifth downsampling layer, 209 - Second HyAtt-CNN layer, 210 - Spatial pyramid pooling sub-module;
[0059] 30 - Implicit Feature Fusion Module, 300 - First Implicit Fusion Path, 301 - Second Implicit Fusion Path, 302 - Third Implicit Fusion Path, 303 - Third Cross - Stage Local Layer, 304 - Fourth Cross - Stage Local Layer, 305 - Fifth Cross - Stage Local Layer, 306 - Sixth Cross - Stage Local Layer, 307 - Seventh Cross - Stage Local Layer, 308 - Sixth Downsampling Layer, 309 - Seventh Downsampling Layer, 310 - Eighth Downsampling Layer;
[0060] 40 - Detection Head Module. Detailed Implementation Manner
[0061] For the convenience of clearly describing the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily mean different.
[0062] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0063] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. The following at least one (item) or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b or c can represent: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b and c can be single or multiple.
[0064] An embodiment of the present invention aims to provide a contraband detection method and system based on hybrid self-attention and implicit feature fusion, so as to solve the technical problems that the backbone network of the existing detection model has weak global feature extraction ability, the feature extraction by the self-attention mechanism has high computational complexity, and for contraband with serious overlap, unclear features and large size differences, the existing models and methods have high complexity and low detection accuracy.
[0065] In a first aspect, the present invention provides a contraband detection system based on hybrid self-attention and implicit feature fusion, see Figure 1 , which includes a data preprocessing module 10, a complex feature parallel extraction module 20, an implicit feature fusion module 30, and a detection head module 40 that are sequentially connected;
[0066] The data preprocessing module 10 is used to preprocess the contraband detection image. The purpose of the preprocessing is to increase the richness of the contraband detection image samples. Exemplarily, methods such as mirroring, flipping, and blurring are randomly combined to preprocess the contraband detection image.
[0067] The complex feature parallel extraction module 20 is used to reduce the resolution of the preprocessed contraband image, expand the number of channels, and extract global features based on the hybrid self-attention mechanism, and output three groups of feature maps P3, P4, and P5 with different resolutions.
[0068] See Figure 1 to Figure 2 , as a possible implementation, the complex feature parallel extraction module 20 includes a first downsampling layer 201, a second downsampling layer 202, a first cross-stage local layer 203, a third downsampling layer 204, a second cross-stage local layer 205, a fourth downsampling layer 206, a first HyAtt-CNN layer 207, a fifth downsampling layer 208, a second HyAtt-CNN layer 209, and a spatial pyramid pooling sub-module 210 that are sequentially connected.
[0069] See Figure 2 , as an example, the first downsampling layer 201, the second downsampling layer 202, the third downsampling layer 204, the fourth downsampling layer 206, and the fifth downsampling layer 208 are all composed of a convolutional layer with a convolution kernel of 3 and a stride of 2, a batch normalization layer, and a SiLU activation function. The first cross-stage local layer 203, the second cross-stage local layer 205, and the spatial pyramid pooling sub-module 210 are consistent with the settings of YOLOv8.
[0070] See Figure 2, As a possible implementation, the preprocessed contraband image is passed through the first downsampling layer 201 to reduce the resolution and expand the number of channels, obtaining the feature map P1; the feature map P1 is passed through the second downsampling layer 202 to reduce the resolution and expand the number of channels, and then enters the first cross-stage local layer 203 to extract local features, obtaining the feature map P2; the feature map P2 is passed through the third downsampling layer 204 to reduce the resolution and expand the number of channels, and then enters the second cross-stage local layer 205 to extract local features, obtaining the feature map P3; the feature map P3 is passed through the fourth downsampling layer 206 to reduce the resolution and expand the number of channels, and then enters the first HyAtt-CNN layer 207 to extract local and global features, obtaining the feature map P4; the feature map P4 is passed through the fifth downsampling layer 208 to reduce the resolution and expand the number of channels, and then enters the second HyAtt-CNN layer 209 to extract local and global features, and is subjected to different-scale feature fusion by the spatial pyramid pooling sub-module 210, outputting the feature map P5.
[0071] See Figure 2 to Figure 3 , As a possible implementation, both the first HyAtt-CNN layer 207 and the second HyAtt-CNN layer 209 include a first convolutional path, a second convolutional path, and a third convolutional path. The first convolutional path includes a first convolutional block 2070 for retaining the features extracted in the previous stage of the feature map input to this HyAtt-CNN layer; the second convolutional path includes a second convolutional block 2071, a fourth convolutional block 2072, and a fifth convolutional block 2073 connected in sequence for extracting local features of the feature map input to this HyAtt-CNN layer; the third convolutional path includes a third convolutional block 2074 and a HyAtt block 2075 connected in sequence for extracting global features of the feature map input to this HyAtt-CNN layer; the outputs of the first convolutional path, the second convolutional path, and the third convolutional path are concatenated along the channel dimension and then output the feature map P4 or P5.
[0072] See Figure 3 , As an example, the convolutional kernels of the first convolutional block 2070, the second convolutional block 2071, and the third convolutional block 2074 are all 1, the stride is 1, and the convolutional kernels of the fourth convolutional block 2072 and the fifth convolutional block 2073 are all 3, the stride is 1.
[0073] See Figure 3 , As an example, assuming that the number of channels of the feature map input to the HyAtt-CNN layer is C, the height and width are H and W respectively, the input feature map first passes through the first convolutional block 2070, the second convolutional block 2071, and the third convolutional block 2074 with a convolutional kernel of 1, a stride of 1, and a padding of 1 respectively, compressing the channel dimension to The first convolutional path obtains X in1 , The second convolutional path obtains X in2 , The third convolutional path obtains Xin3 , X in1 Keep the information from the previous stage without processing, X in2 Further extract local features through the fourth convolutional block 2072 and the fifth convolutional block 2073 with a convolutional kernel of 3, a stride of 1, and a padding of 0, X in3 After adding with the results of the previous group, it is sent to the HyAtt block 2075 to extract global features, and finally the three parts are added to obtain the output features. The specific calculation process is as follows:
[0074] X in1 = CBS_1(X in )
[0075] X in2 = CBS_2(X in )
[0076] X in3 = CBS_3(X in )
[0077] X out1 = X in1
[0078] X out2 = CBS_5(CBS_4(X in2 ))
[0079] X out3 = HyAtt(X in3 )
[0080] X out = X out1 + X out2 + X out3
[0081] Among them, X out1 represents the output of the first convolutional path, X out2 represents the output of the second convolutional path, X out3 represents the output of the third convolutional path.
[0082] See Figure 3 to Figure 4 , as a possible implementation, the HyAtt block 2075 includes a first reshaping layer 20750, a second reshaping layer 20751, a first normalization layer 20752, a second normalization layer 20753, a third normalization layer 20754, a global pooling layer 20755, a first Softmax activation function layer 20756, a second Softmax activation function layer 20757, a first fully connected layer 20758, a second fully connected layer 20759, a third fully connected layer 20760, and a fourth fully connected layer 20761;
[0083] The input feature map of the HyAtt block 2075 is transformed in dimension by the first reshaping layer 20750 and then passes through the matrix W included in the first fully connected layer 20758 Q , the matrix W included in the second fully connected layer 20759 K and the matrix W included in the third fully connected layer 20760 V , obtaining three matrices Q, K, V with reduced dimensions; the first normalization layer 20752 normalizes the matrix Q, the second normalization layer 20753 normalizes the matrix X, and after multiplying the matrix V by the normalized matrix K and then multiplying by the normalized matrix Q, the first intermediate output feature is obtained;
[0084] The global pooling layer 20755 performs an aggregation operation on the matrix Q and then multiplies it by the matrix Q, and outputs after passing through the first Softmax activation function layer 20756; the global pooling layer 20755 performs an aggregation operation on the matrix Q and then multiplies it by the matrix K, and outputs after passing through the second Softmax activation function layer 20757; the output of the first Softmax activation function layer 20756 is multiplied by the output of the second Softmax activation function layer 20757 and then multiplied by the matrix V, obtaining the second intermediate output feature;
[0085] The sum of the first intermediate output feature and the second intermediate output feature is normalized by the third normalization layer 20754 and then enters the fourth fully connected layer 20761. The output of the fourth fully connected layer 20761 is adjusted in dimension by the second reshaping layer 20751 to obtain the output of the HyAtt block 2075.
[0086] See Figure 4 , as an example, the input feature map of the HyAtt block 2075 is a feature map with the number of channels C, height H, and width W. First, the input feature map is transformed in dimension by the first reshaping layer 20750, changing from a three-dimensional feature to a two-dimensional feature to facilitate the calculation of global attention information. After the transformation, a two-dimensional feature of N×C is obtained, where N = H×W. Then the input feature passes through three fully connected layer matrices W Q , W K and W V to obtain three matrices Q, K, V with dimensions N×d, where d is less than N. This process can be expressed as:
[0087] Q = W Q (X in )
[0088] K = W K (X in )
[0089] V = W V (X in )
[0090] Among them, X in represents the input feature map. Next, based on the linear self-attention mechanism, Q and K are normalized. K T is first multiplied by V, and then Q is multiplied by the product of the two, that is:
[0091]
[0092] Among them, the resulting is a d×d matrix, and then multiplied by to obtain an N×d matrix. The computational complexity of this process is approximately linear complexity O(Nd 2 ), represents the normalization operation.
[0093] Since the normalization operation used in the linear self-attention mechanism does not include the Softmax non-linear activation function, the performance will be lost and it is unable to focus on key information. To supplement the missing partial information of the linear self-attention, in this embodiment, the global pooling layer 20755 is used to aggregate the matrix Q, retaining the Softmax normalization function and operation order of the original self-attention mechanism, thereby reducing the dimension of the operation matrix and achieving the purpose of reducing the computational complexity. The specific implementation is as follows:
[0094] q = Pooling(Q)
[0095] X out2 = softmax(Kq T )softmax(Kq T ) T V
[0096] Among them, Pooling represents the global pooling operation, which is equivalent to further aggregating the global information of the matrix Q, and the dimension of q is reduced to 3×d. Calculating X out2 has a computational complexity of approximately O(Ndn). Finally, the outputs of the two parts are added together, and through the fourth fully connected layer 20761, the second reshaping layer 20751 is used to adjust the dimension to the output dimension of the original features. Finally:
[0097] X out = Reshape(FullyConnectedLayer(X out1 + X out2 ))
[0098] The existing implementation of the self-attention mechanism calculates the global correlation by multiplying the matrix itself, and its computational complexity is O(N 2 ), and the specific implementation is as follows:
[0099] SelfAttention = Softmax(QK T )V
[0100] QK T What is obtained is an N*N matrix, and the computational complexity is O(N 2 ), which is not friendly to the operation speed of the algorithm. In the embodiments of the present invention, a hybrid self-attention mechanism HyAtt is introduced. Compared with the existing self-attention mechanism SelfAttention, HyAtt reduces the computational complexity of self-attention SelfAttention from O(N2) to O(N), and retains the non-linear Softmax activation function of self-attention, enabling the present application to focus on more important features.
[0101] As an example, the process of inputting the preprocessed contraband image into the complex feature parallel extraction module 20 and outputting three groups of feature maps P3, P4, and P5 with different resolutions after reducing the resolution, expanding the number of channels, and advancing the global features is as follows:
[0102] Assume that the width and height of the preprocessed contraband image are 640×640, and the number of channels is 3. The preprocessed contraband image first enters the first downsampling layer 201. The first downsampling layer 201 reduces the image resolution to 320×320 and expands the number of channels to 32 to obtain the feature map P1. The feature map P1 enters the second downsampling layer 202. The second downsampling layer 202 reduces the image resolution to 160×160 and expands the number of channels to 64, and then enters the first cross-stage local layer 203. The convolutional neural network in the first cross-stage local layer 203 is used to extract local features, that is, aggregate the pixel features within the range of 3×3 to obtain the feature map P2. The feature map P2 enters the third downsampling layer 204. The third downsampling layer 204 reduces the image resolution to 80×80 and expands the number of channels to 128, and then enters the second cross-stage local layer 205. The convolutional neural network in the second cross-stage local layer 205 is used to extract local features to obtain the feature map P3. The feature map P3 enters the fourth downsampling layer 206. The fourth downsampling layer 206 reduces the image resolution to 40×40 and expands the number of channels to 256, and then enters the first HyAtt-CNN layer 207 to extract local features and global features to obtain the feature map P4. The feature map P4 enters the fifth downsampling layer 208. The fifth downsampling layer 208 reduces the image resolution to 20×20 and expands the number of channels to 512, and then enters the second HyAtt-CNN layer 209 and the spatial pyramid pooling sub-module 210 in sequence, and the feature map P5. The feature dimensions after being processed by the complex feature parallel extraction module 20 are as follows:
[0103] P1: 320×320×32
[0104] P2: 160×160×64
[0105] P3: 80×80×128
[0106] P4: 40×40×256
[0107] P5: 20×20×512。
[0108] The implicit feature fusion module 30 is used to map the low-resolution features of the feature maps P3, P4, and P5 to high-resolution features based on the implicit fusion path, and output three groups of feature maps O3, O4, and O5 with different resolutions after implicit fusion;
[0109] See Figure 1 and 5 , the implicit feature fusion module 30 includes a first implicit fusion path 300, a second implicit fusion path 301, a third implicit fusion path 302, a third cross-stage local layer 303, a fourth cross-stage local layer 304, a fifth cross-stage local layer 305, a sixth cross-stage local layer 306, a seventh cross-stage local layer 307, a sixth downsampling layer 308, a seventh downsampling layer 309, and an eighth downsampling layer 310.
[0110] As an example, the sixth downsampling layer 308, the seventh downsampling layer 309, and the eighth downsampling layer 310 are all composed of a convolutional layer with a convolution kernel of 3 and a stride of 2, a batch normalization layer, and a SiLU activation function. The five cross-stage local layers are all consistent with the settings of YOLOv8.
[0111] The first implicit fusion path 300 is used to increase the resolution of the feature map P5, the second implicit fusion path 301 is used to increase the resolution of the feature map P4, the feature map P3 is concatenated with the upsampled P4 and the upsampled P5, and after the number of channels is compressed by the third cross-stage local layer 303, the feature map O3 is output.
[0112] The feature map O3 is processed by the sixth downsampling layer 308 and then concatenated with P4. After the number of channels is compressed by the fourth cross-stage local layer 304, the intermediate output M4 is obtained; the intermediate output M4 is processed by the seventh downsampling layer 309 and then concatenated with P5. After the number of channels is compressed by the fifth cross-stage local layer 305, the intermediate output M5 is obtained; the intermediate output M5 is upsampled by the third implicit fusion path 302 and then concatenated with the intermediate output M4. After the number of channels is compressed by the sixth cross-stage local layer 306, the feature map O4 is output.
[0113] The feature map O4 is processed by the eighth downsampling layer 310 and then concatenated with the intermediate output M5. After the number of channels is compressed by the seventh cross-stage local layer 307, the feature map O5 is output.
[0114] In this embodiment, the input of the implicit feature fusion module 30 is three sets of feature maps P3, P4, and P5 with different resolutions obtained by the complex feature parallel extraction module 20, and the output is three sets of feature maps O3, O4, and O5 with different resolutions after implicit feature fusion. Among them, the feature maps P3 and O3 are feature maps with 128 channels, and the height and width are 80 and 80 respectively; the feature maps P4, M4, and O4 are feature maps with 256 channels, and the height and width are 40 and 40 respectively; the feature maps P5, M5, and O5 are feature maps with 512 channels, and the height and width are 20 and 20 respectively. The feature map P3 has the highest resolution and contains a large amount of detailed texture information. The feature map P5 has the lowest resolution and the largest receptive field and contains highly extracted semantic information.
[0115] The design of the implicit feature fusion module 30 reconstructs the fusion path of the high-resolution feature maps O3 and O4. Using the implicit neural representation method, it maps the low-resolution features to high-resolution features, which can avoid feature aliasing caused by using upsampling.
[0116] See Figure 6 , as a possible implementation, the implicit fusion path is implemented by the following method: construct a high-resolution feature coordinate network, encode the low-resolution features into the latent space, interpolate the latent space to the high-resolution feature coordinate network based on the point sampling function to obtain a high-resolution latent encoding map, and calculate the nearest latent encoding value of each coordinate in the high-resolution latent encoding map based on the neighboring latent encoding weighted fusion mechanism, and decode the nearest latent encoding value to obtain the high-resolution feature value.
[0117] As an example, the implicit fusion path implicitly encodes the internal structure of the data through a layer of encoder to obtain a latent code. In the continuous two-dimensional latent space, the coordinates of the low-resolution features and the latent code learn a continuous mapping relationship through a neural network and map to the corresponding feature values under the high-resolution features, realizing the learning of low-resolution features to high-resolution features. The implicit function is usually continuously differentiable, supports sampling and interpolation at any resolution, and avoids the discretization defects of explicit representation, thereby reducing feature aliasing caused by upsampling. Its principle is:
[0118] H(x) = f θ (z * , x - x * )
[0119] where z * ∈R d is the nearest neighbor latent encoding of x, x * is its corresponding low-resolution coordinate, and f θ is a decoder with an input dimension of d + 2 (latent encoding + relative coordinates). The implementation process is divided into three steps: first, construct a high-resolution feature coordinate grid; second, use a 1×1 convolution to process the low-resolution feature F LIt is encoded into the latent space Z, and the point sampling function is used to interpolate it into a high-resolution feature coordinate grid to generate a high-resolution latent encoding map. Finally, for each coordinate x, its neighboring latent encodings are queried from the latent encoding map and input into the decoder to decode the feature values. To enhance feature continuity, a neighboring latent encoding weighted fusion mechanism is further introduced. For the coordinate x, the interpolation weights w of its four nearest neighboring latent encodings are calculated i (based on the relative distance and satisfying ∑w i = 1), and the final feature value is:
[0120]
[0121] See Figure 1 and Figure 5 . After construction, the feature maps P3, P4, and P5 output by the complex feature parallel extraction module 20 are input into the implicit feature fusion module 30, and the output feature map O3 is obtained in the following way:
[0122] O3 = CSP_3(Concat(First Implicit Fusion Path (P5), Second Implicit Fusion Path (P4)))
[0123] where CSP_3 represents compressing the number of channels by the third cross-stage partial layer 303, Concat represents the concatenation operation, the first implicit fusion path 300 and the second implicit fusion path 301 map the feature maps P5 and P4 to a resolution of 80×80 in width and height respectively, and CSP_3 compresses the number of channels to 128. The intermediate output M4 and the intermediate output M5 are obtained in the following way:
[0124] M4 = CSP_4(Concat(P4, Downsampling_6(O3)))
[0125] M5 = CSP_5(Concat(P5, Downsampling_7(M4)))
[0126] where CSP_4 represents compressing the number of channels by the fourth cross-stage partial layer 304, CSP_5 represents compressing the number of channels by the fifth cross-stage partial layer 305, Concat represents the concatenation operation, Downsampling_6 represents being processed by the sixth downsampling layer 308, and Downsampling_7 represents being processed by the seventh downsampling layer 309. The output feature map O4 is obtained in the following way:
[0127] O4 = CSP_6(Concat(Third Implicit Fusion Path (M5), M4))
[0128] Among them, CSP_6 represents the number of channels compressed by the sixth cross-stage local layer 306. The third implicit fusion path 302 maps M5 to a resolution of 40×40 in terms of width and height and then concatenates it with M4. The sixth cross-stage local layer 306 compresses the number of channels to 256 and then outputs the feature map O4. Finally, the feature map O4 undergoes one downsampling and is concatenated with M5, and then the seventh cross-stage local layer 307 adjusts the number of channels to 512 and outputs the feature map O5:
[0129] O5 = CSP_7(Concat(Downsampling_8(O4), M5))
[0130] Downsampling_8 represents being processed by the eighth downsampling layer 310.
[0131] The detection head module 40 is used to detect O3, O4, and O5 and output the detection results.
[0132] As a possible implementation, the detection head module 40 includes three detection heads, all of which are YOLOv8 detection heads. They respectively detect the feature maps O3, O4, and O5, and after non-maximum suppression, output the detection results.
[0133] In a second aspect, an embodiment of the present invention provides an illegal item detection method based on hybrid self-attention and implicit feature fusion, which is implemented by using the illegal item detection system based on hybrid self-attention and implicit feature fusion provided in the first aspect. Refer to Figure 7 , and the illegal item detection method includes:
[0134] S1. Configure the illegal item detection image and preprocess the illegal item detection image;
[0135] S2. Input the preprocessed illegal item detection image into the complex feature parallel extraction module to extract the global features and local features of the illegal items, and obtain three groups of feature maps P3, P4, and P5 with different resolutions;
[0136] As a possible implementation, S2 includes:
[0137] S20. The preprocessed illegal item detection image undergoes two downsampling processes, one local feature extraction, another downsampling process, and another local feature extraction in sequence to obtain the feature map P3;
[0138] S21. The feature map P3 undergoes one downsampling process and local-global feature extraction in sequence to obtain the feature map P4;
[0139] S22. The feature map P4 undergoes one downsampling process, local-global feature extraction, and spatial pyramid pooling in sequence to obtain the feature map P5.
[0140] S3. Input the three groups of feature maps P3, P4, and P5 with different resolutions into the implicit feature fusion module, and obtain three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions through three implicit fusion paths;
[0141] As a possible implementation, S3 includes:
[0142] S30. Respectively increase the resolutions of the feature maps P4 and P5, cascade the feature map P3 with the P4 and P5 after increasing the resolutions, and then perform channel number compression to obtain the feature map O3;
[0143] S31. After performing downsampling on the feature map O3, cascade it with the feature map P4, perform channel number compression after cascading to obtain the intermediate output M4, perform downsampling layer processing on the intermediate output M4 and then cascade it with the feature map P5, perform channel number compression after cascading to obtain the intermediate output M5, increase the resolution of the feature map M5 and then cascade it with the intermediate output M4, perform channel number compression after cascading to obtain the feature map O4;
[0144] S32. After performing downsampling layer processing on the feature map O4, cascade it with the intermediate output M5, perform channel number compression after cascading to obtain the feature map O5.
[0145] S4. Use 3 detection heads to detect the three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions respectively, and output the detection results.
[0146] Next, the detection effect of the contraband detection system of the present invention will be further elaborated in detail with specific experiments.
[0147] The experiments use the PIDray dataset (https: / / github.com / bywang2018 / security-dataset), the SIXray dataset (https: / / github.com / MeioJane / SIXray), and the OPIXray dataset (https: / / github.com / OPIXray-author / OPIXray).
[0148] Among them, the PIDray dataset contains 47,677 images, covering 12 types of contraband items. Among these, 29,457 images are used for training, and 18,220 images are used for testing. The test set is further divided into three subsets: Easy, Hard, and Hidden. The Easy subset contains 9,482 images with only a single contraband item, the Hard subset contains 3,733 images with multiple contraband items, and the Hidden subset contains 5,005 high-difficulty scenario images with contraband items deliberately hidden. The SIXray dataset contains 8,929 images with contraband items, labeled with 5 types of contraband items: guns, knives, wrenches, pliers, and scissors. The dataset is divided in an 8:2 ratio, where 80% (7,143 images) are used for training and 20% (1,786 images) are used for testing. The OPIXray dataset contains 8,885 images, labeled with 5 types of contraband items: folding knives, straight knives, scissors, utility knives, and multi-functional knives. The dataset is divided in an 8:2 ratio, where 7,109 images are used for training and 1,776 images are used for testing.
[0149] The computing platform for this experiment is a Linux server equipped with an NVIDIA GeForce RTX 3090 GPU, and the model is implemented using the PyTorch framework. In the experiment, the mean average precision (mAP) is used as an indicator to evaluate the performance of contraband detection.
[0150] Next, the specific experimental process is described. The experiment is divided into two stages: the training stage and the validation and testing stage. The training stage includes the following steps:
[0151] Step A.1: Construct the training sample set, validation sample set, and test sample set, specifically:
[0152] Download the PIDray, OPIXray, and SIXray datasets, and generate training, test, and validation set labels in the data format of YOLOV8. The PIDray dataset follows the default division method and contains 29,457 training images; the OPIXray dataset and the SIXray dataset are divided into a training set and a test set in an 8:2 ratio.
[0153] Step A.2: Use the training sample set in Step A.1 to train the contraband detection system proposed in this embodiment based on hybrid self-attention and implicit feature fusion. The training settings are: epoch is 300, batch is 32, the optimizer is SGD, the initial learning rate is 0.01, and it decays cosine to 0.0001; the loss function is set as:
[0154] Loss = L cls +L obj +L reg
[0155] Among them, L clsis the classification loss, L obj is the confidence loss, L reg is the regression loss, and the loss function is the same as the one used in the YOLOv8 method.
[0156] After training is completed and entering the verification and testing stage, the test sample set in step A.1 is used to evaluate the detection effect. The evaluation criteria include the mean average precision (mAP0.5:0.95) with the intersection over union (IoU) threshold ranging from 0.5 to 0.95, the mean precision (mAP0.5) with the IoU threshold of 0.5, the number of parameters, the computational cost, and the inference time.
[0157] The verification sample set in step A.1 is used to verify the detection effect. The detection effect is shown in Figure 8 , the first line is the true position of the contraband, the second line is the result detected by the comparison method YOLOv8, and the third line is the detection result of this application. From Figure 8 it can be seen that this application can accurately find the specific position of the contraband in scenarios where the contraband is severely occluded and there are large size differences, indicating that the application has excellent detection performance in solving scenarios where it is difficult to capture the characteristics of contraband and the background is complex.
[0158] Next, the system and method proposed in this application are used to conduct experimental simulations on three public datasets, PIDray, SIXray, and OPIXray, and are compared with other advanced methods. The detection results output on the PIDray dataset are shown in Table 1 below:
[0159] Table 1 Comparison of experimental results between this application and other models on the PIDray dataset
[0160]
[0161]
[0162] As can be seen from Table 1, the current mainstream self-attention architectures have more parameters and computational costs than this application, but this application still obtains excellent detection results.
[0163] Next, experiments are conducted on the other two large-scale security inspection datasets, SIXRay and OPIXRay. mAP50 and mAP50:95 are selected as comparison metrics. The experimental results are shown in Table 2:
[0164] Table 2 Comparison of experimental results between this application and other methods on the SIXRay and OPIXRay datasets
[0165]
[0166] As can be seen from Table 2, compared with the baseline YOLOv8-S, both the mAP50 and mAP50:95 of the present invention have been effectively improved. This further shows that this application is helpful for contraband detection and improves the detection performance of the general detection framework in the security inspection scenario.
[0167] Through experiments, it can be seen that the systems and methods provided by the present invention have obtained excellent detection results on all three test sets. At the same time, the model still remains lightweight. Compared with YOLOv8-S, this application has improved by more than 1% on the easy and difficult test sets, and has improved by more than 3% compared with the previous best X-ray security inspection method PIXDet-S. On the hidden test set, the mAP0.5:0.9 of this application has increased significantly. Compared with YOLOv8-S and PIXDet-S, the mAP0.5:0.9 of this application has increased by 3.6% and 1.9% respectively. At the same time, compared with PIXDet-S, this application has reduced the number of parameters by 18.7M and the amount of computation by 17G, achieving the best balance between performance and model lightweight.
[0168] Table 3 shows the gain comparison of the original network after adding HyAtt-CNN or the implicit feature fusion network. It can be seen that adding any one component can improve the detection accuracy of the original method, especially on the hidden test set, and does not bring significant consumption of computing resources. When adding both HyAtt-CNN and the implicit feature fusion network to the original model at the same time, the detection accuracy of this design method reaches the highest on all three subsets.
[0169] Table 3 Ablation experiments of this application
[0170]
[0171] Table 4 compares the detection performance of the HyAtt-CNN structure using HyAtt and standard self-attention respectively. It can be seen that HyAtt is lower than the standard self-attention mechanism in terms of the number of parameters and the amount of computation, and at the same time, the detection result on the hidden test set is better than that of the standard self-attention mechanism, indicating that HyAtt still has the feature extraction ability of the standard self-attention mechanism while maintaining lightweight.
[0172] Table 4 Comparison of the detection performance between HyAtt and the standard self-attention mechanism
[0173]
[0174] Table 5 shows the performance comparison between the implicit feature fusion network and other feature fusion networks. It can be seen that the implicit feature fusion network has achieved the best results on all three test sets while maintaining an approximate number of parameters, demonstrating the gain of the implicit feature fusion network for hidden contraband detection.
[0175] Table 5 Performance Comparison between the Implicit Feature Fusion Network and Other Feature Fusion Networks
[0176]
[0177] Although the present invention has been described in conjunction with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the description of the drawings, etc. In the specification, the term "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the specification. Certain measures are recited in mutually different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0178] Although the present invention has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, this specification and the drawings are merely exemplary descriptions of the present invention and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A contraband detection system based on hybrid self-attention and implicit feature fusion, characterized in that, It includes a data preprocessing module, a complex feature parallel extraction module, an implicit feature fusion module, and a detection head module that are sequentially connected: Among them, the data preprocessing module is used to preprocess the contraband detection image; The complex feature parallel extraction module is used to reduce the resolution of the preprocessed contraband image, expand the number of channels, and extract global features based on the hybrid self-attention mechanism, and output three groups of feature maps P3, P4, and P5 with different resolutions; The implicit feature fusion module is used to map the low-resolution features of the feature maps P3, P4, and P5 to high-resolution features based on the implicit fusion path, and output three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions; The implicit fusion path is implemented by the following method: Construct a high-resolution feature coordinate network, encode the low-resolution features into the latent space, interpolate the latent space to the high-resolution feature coordinate network based on the point sampling function, obtain a high-resolution latent encoding map, and calculate the nearest neighbor latent encoding value of each coordinate in the high-resolution latent encoding map based on the neighboring latent encoding weighted fusion mechanism, and decode the nearest neighbor latent encoding value to obtain high-resolution feature values; The detection head module is used to detect the feature maps O3, O4, and O5 and output the detection results.
2. The contraband detection system based on hybrid self-attention and implicit feature fusion according to claim 1, characterized in that The complex feature parallel extraction module includes a first downsampling layer, a second downsampling layer, a first cross-stage local layer, a third downsampling layer, a second cross-stage local layer, a fourth downsampling layer, a first HyAtt-CNN layer, a fifth downsampling layer, a second HyAtt-CNN layer, and a spatial pyramid pooling sub-module that are sequentially connected.
3. The contraband detection system based on hybrid self-attention and implicit feature fusion according to claim 2, wherein The preprocessed contraband image is reduced in resolution and the number of channels is expanded by the first downsampling layer to obtain a feature map P1; the feature map P1 enters the first cross-stage local layer to extract local features after being reduced in resolution and the number of channels is expanded by the second downsampling layer, obtaining a feature map P2; the feature map P2 enters the second cross-stage local layer to extract local features after being reduced in resolution and the number of channels is expanded by the third downsampling layer, obtaining a feature map P3; the feature map P3 enters the first HyAtt-CNN layer to extract local and global features after being reduced in resolution and the number of channels is expanded by the fourth downsampling layer, obtaining a feature map P4; the feature map P4 enters the second HyAtt-CNN layer to extract local and global features after being reduced in resolution and the number of channels is expanded by the fifth downsampling layer, and different-scale feature fusion is performed by the spatial pyramid pooling sub-module to output a feature map P5.
4. The contraband detection system based on hybrid self-attention and implicit feature fusion according to claim 3, wherein Both the first HyAtt-CNN layer and the second HyAtt-CNN layer include a first convolutional path, a second convolutional path, and a third convolutional path. The first convolutional path includes a first convolutional block for retaining the features extracted from the feature map input to the HyAtt-CNN layer in the previous stage; the second convolutional path includes a second convolutional block, a fourth convolutional block, and a fifth convolutional block connected in sequence for extracting local features of the feature map input to the HyAtt-CNN layer; the third convolutional path includes a third convolutional block and a HyAtt block connected in sequence for extracting global features of the feature map input to the HyAtt-CNN layer; the outputs of the first convolutional path, the second convolutional path, and the third convolutional path are concatenated along the channel dimension and then the feature map P4 or P5 is output.
5. The contraband detection system based on hybrid self-attention and implicit feature fusion according to claim 4, wherein The HyAtt block includes a first reshaping layer, a second reshaping layer, a first normalization layer, a second normalization layer, a third normalization layer, a global pooling layer, a first Softmax activation function layer, a second Softmax activation function layer, a first fully connected layer, a second fully connected layer, a third fully connected layer, and a fourth fully connected layer; The input feature map of the HyAtt block is transformed in dimension by the first reshaping layer and then passes through the matrix W included in the first fully connected layer Q , the matrix W included in the second fully connected layer K and the matrix W included in the third fully connected layer V , obtaining three matrices Q, K, and V with reduced dimensions; the first normalization layer normalizes the matrix Q, the second normalization layer normalizes the matrix K, and after multiplying the matrix V by the normalized matrix K and then multiplying by the normalized matrix Q, a first intermediate output feature is obtained; The global pooling layer performs an aggregation operation on the matrix Q and then multiplies it with the matrix Q, and outputs after passing through the first Softmax activation function layer; the global pooling layer performs an aggregation operation on the matrix Q and then multiplies it with the matrix K, and outputs after passing through the second Softmax activation function layer; the output of the first Softmax activation function layer is multiplied with the output of the second Softmax activation function layer and then multiplied with the matrix V to obtain a second intermediate output feature; The first intermediate output feature and the second intermediate output feature are added together, then normalized by the third normalization layer and enter the fourth fully connected layer. The output of the fourth fully connected layer is adjusted in dimension by the second reshaping layer to obtain the output of the HyAtt block.
6. The contraband detection system based on hybrid self-attention and implicit feature fusion according to claim 1, wherein The implicit feature fusion module includes a first implicit fusion path, a second implicit fusion path, a third implicit fusion path, a third cross-stage local layer, a fourth cross-stage local layer, a fifth cross-stage local layer, a sixth cross-stage local layer, a seventh cross-stage local layer, a sixth downsampling layer, a seventh downsampling layer, and an eighth downsampling layer; The first implicit fusion path is used to increase the resolution of the feature map P5, and the second implicit fusion path is used to increase the resolution of the feature map P4. The feature map P3 is concatenated with the upsampled P4 and the upsampled P5, and after compressing the number of channels by the third cross-stage local layer, the feature map O3 is output; The feature map O3 is processed by the sixth downsampling layer and then concatenated with P4. After compressing the number of channels by the fourth cross-stage local layer, the intermediate output M4 is obtained; the intermediate output M4 is processed by the seventh downsampling layer and then concatenated with P5. After compressing the number of channels by the fifth cross-stage local layer, the intermediate output M5 is obtained; the intermediate output M5 is upsampled by the third implicit fusion path and then concatenated with the intermediate output M4. After compressing the number of channels by the sixth cross-stage local layer, the feature map O4 is output; The feature map O4 is processed by the eighth downsampling layer and then concatenated with the intermediate output M5. After compressing the number of channels by the seventh cross-stage local layer, the feature map O5 is output.
7. The contraband detection system based on hybrid self-attention and implicit feature fusion according to claim 1, wherein The detection head module includes three detection heads, which respectively detect the feature maps O3, O4, and O5. After non-maximum suppression, the detection results are output.
8. A contraband detection method based on hybrid self-attention and implicit feature fusion, characterized in that, Implemented by using the contraband detection system based on hybrid self-attention and implicit feature fusion according to any one of claims 1 to 7, the contraband detection method includes: S1. Configure the contraband detection image and preprocess the contraband detection image. S2. Input the preprocessed contraband detection image into the complex feature parallel extraction module to extract the global features and local features of the contraband, and obtain three groups of feature maps P3, P4, and P5 with different resolutions. S3. Input the three groups of feature maps P3, P4, and P5 with different resolutions into the implicit feature fusion module, and obtain three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions through three implicit fusion paths. S4. Use 3 detection heads to respectively detect the three groups of implicitly fused feature maps O3, O4, and O5 with different resolutions, and output the detection results.
9. The contraband detection method based on hybrid self-attention and implicit feature fusion according to claim 8, wherein The said S2 includes: S20. Perform two downsampling processes, one local feature extraction, another downsampling process, and another local feature extraction on the preprocessed contraband detection image in sequence to obtain the feature map P3. S21. Perform one downsampling process and local-global feature extraction on the feature map P3 in sequence to obtain the feature map P4. S22. Perform one downsampling process, local-global feature extraction, and spatial pyramid pooling on the feature map P4 in sequence to obtain the feature map P5.
10. The contraband detection method based on hybrid self-attention and implicit feature fusion according to claim 8, wherein The said S3 includes: S30. Respectively increase the resolutions of the feature map P4 and the feature map P5, cascade the feature map P3 with the P4 and P5 after increasing the resolutions, and perform channel number compression to obtain the feature map O3. S31. Perform downsampling processing on the feature map O3 and then cascade it with the feature map P4. After cascading, perform channel number compression to obtain the intermediate output M4. Perform downsampling layer processing on the intermediate output M4 and then cascade it with the feature map P5. After cascading, perform channel number compression to obtain the intermediate output M5. Increase the resolution of the feature map M5 and then cascade it with the intermediate output M4. After cascading, perform channel number compression to obtain the feature map O4. S32. Perform downsampling layer processing on the feature map O4 and then cascade it with the intermediate output M5. After cascading, perform channel number compression to obtain the feature map O5.