Subway hazardous article terahertz detection method based on adaptive downsampling and multi-dimensional attention mechanism

The terahertz detection method for hazardous materials in subways, which uses adaptive downsampling and multi-dimensional attention mechanisms, solves the problems of weak feature extraction capabilities and low robustness in traditional detection methods, and achieves improved detection accuracy and robustness. It is suitable for hazardous materials detection in scenarios such as subways, rail transit, and aviation hubs.

CN120635557AActive Publication Date: 2025-09-12GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510727740.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Traditional terahertz detection methods have problems in dangerous goods detection, such as weak feature extraction capability, low robustness, and insufficient utilization of spatial information, resulting in reduced detection accuracy.

Method used

A terahertz detection method for subway hazardous materials adopts adaptive downsampling and multi-dimensional attention mechanism. It processes feature maps through parallel pooling and convolution operations, combines the Transformer encoder for global context modeling, and uses strip pooling layers to capture long-range context dependencies and enhance target edge features.

Benefits of technology

The detection accuracy and robustness have been improved, making it suitable for real-time dangerous goods detection in scenarios such as subways, rail transit, and aviation hubs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635557A_ABST
    Figure CN120635557A_ABST
Patent Text Reader

Abstract

The invention discloses a metro dangerous goods terahertz detection method based on adaptive downsampling and a multi-dimensional attention mechanism. The metro dangerous goods terahertz detection method comprises the following steps: acquiring a terahertz image data set; building a deep learning network model, wherein the deep learning network model comprises a backbone feature extraction network, a check feature extraction network, a down-sampling module, a fusion channel and space attention module, a space attention module and a YoloHead detection head which are connected in sequence; performing dual-path processing on the input feature map through parallel pooling and convolution operation; in the fusion channel and space attention module, global context modeling is carried out on the feature map; in the space attention module, long-range context dependence is captured through a strip-shaped pooling layer; inputting the multi-scale feature map output by the check network into a YoloHead detection head, and outputting a target detection frame; a redundant detection frame is removed through a non-maximum suppression algorithm, and a detection result with the highest confidence coefficient is reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dangerous goods detection, and in particular relates to a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism. Background Art

[0002] As an emerging nondestructive testing solution in the security field, terahertz wave detection technology offers significant advantages due to its non-ionizing radiation properties. This technology achieves penetrating identification by analyzing the unique molecular vibrational response of a substance. This technology can effectively overcome complex environmental interference and accurately locate potentially dangerous items hidden within clothing, packages, or containers. Its biocompatibility ensures that the detection process does not cause ionizing damage to human tissue or the environment. This technology is particularly well-suited for intelligent screening of hidden threats in high-security scenarios such as airports and rail transit hubs, providing key technical support for the development of a non-invasive security inspection system.

[0003] The rapid development of deep learning technology has provided powerful tools for image recognition and detection. However, traditional object detection tools still have many limitations for hazardous material detection in the terahertz field: weak feature extraction capabilities: traditional downsampling modules tend to lose high-frequency details, affecting the detection of small targets; low robustness: conventional attention mechanisms such as SE and CBAM modules struggle to effectively distinguish hazardous materials from background noise; and spatial information utilization without pooling operations tends to weaken target edge features, resulting in reduced detection accuracy. Therefore, it is urgent to propose a terahertz detection method for subway hazardous materials based on adaptive downsampling and multi-dimensional attention mechanisms. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, which improves the detection capability of subway dangerous goods.

[0005] To achieve the above objectives, the present invention provides a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, comprising:

[0006] Obtaining a terahertz image dataset, wherein the dataset includes an XML file marking the location of dangerous goods;

[0007] Build a deep learning network model, which includes a backbone feature extraction network, a neck feature extraction network, a downsampling module, a fusion channel and spatial attention module, a spatial attention module and a YoloHead detection head connected in sequence;

[0008] In the downsampling module, the input feature map is processed in a dual path through parallel pooling and convolution operations;

[0009] In the fusion channel and spatial attention module, the feature map is modeled with global context through the Transformer encoder;

[0010] In the spatial attention module, long-range context dependencies are captured through striped pooling layers;

[0011] Input the multi-scale feature map output by the neck network into the YoloHead detection head and output the target detection frame in the normalized coordinate system;

[0012] Redundant detection boxes are removed through the non-maximum suppression algorithm, and the detection results with the highest confidence are retained.

[0013] Optionally, dual-path processing of the input feature map through parallel pooling and convolution operations includes: performing average pooling preprocessing on the input feature map; dividing the preprocessed feature map into two paths along the channel dimension; the first path captures local details through 3×3 convolution downsampling; the second path fuses channel information through maximum pooling and 1×1 convolution; and splicing the two outputs along the channel dimension.

[0014] Optionally, global context modeling of feature maps through Transformer encoders includes: flattening the input feature maps into sequence form and adding two-dimensional sine-cosine position encoding; using a multi-head self-attention mechanism to capture long-range dependencies; and enhancing feature expression through a feedforward network.

[0015] Optionally, capturing long-range contextual dependencies through a strip pooling layer includes: performing horizontal and vertical strip pooling on the input feature map; modulating the pooling result through a 1×1 convolution; generating spatial attention weights and weighting them to the original feature map.

[0016] Optionally, the construction process of the downsampling module includes: in the dual-path processing, the step size of the first 3×3 convolution is set to 2, and the padding value is 1; the window size of the second maximum pooling is 3×3, the step size is 2, and the padding value is 1; the two output channels have the same number and feature fusion is achieved by element-by-element addition.

[0017] Optionally, the construction process of the fusion channel and spatial attention module includes: when the position code is generated, the temperature parameter temperature is set to 10000, and pos_dim is set to 128; the number of heads of the multi-head self-attention mechanism is 8, and the dimension of each head is 64.

[0018] Optionally, the construction process of the spatial attention module includes: the strip pooling layer includes two parallel branches in the horizontal direction and the vertical direction, each branch contains a 3×1 convolution kernel; when calculating the dynamic weight, the number of channels of the 1×1 convolution is reduced by a ratio of 4.

[0019] Optionally, the structure of the neck network includes: connecting an SPPF layer after the downsampling module, and the pyramid scale of the SPPF layer is set to 5, 9, and 13; the input of the fusion channel and the spatial attention module is the feature map of the first two scales of the neck network.

[0020] Technical effect of the invention: The present invention discloses a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism. By introducing adaptive downsampling and multi-dimensional attention, the high-frequency information loss in the downsampling process is reduced, global context modeling and local feature refinement are combined, and long-range dependencies are captured through strip pooling layers. The target edge and shape features are optimized, so that the method has significant improvements in feature extraction capability, detection accuracy and robustness compared with traditional target detection algorithms. It is suitable for real-time applications and various scenarios of terahertz image dangerous goods detection. It has broad application prospects in scenarios such as subway dangerous goods detection, rail transportation, aviation hubs and large-scale event security. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0022] Figure 1 Schematic diagram of computer detection of terahertz images in terahertz dangerous goods detection according to an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram of the overall architecture of the detection network according to an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of the network architecture of the downsampling module in the overall detection network architecture according to an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of the fusion channel and spatial attention module network in the overall detection network architecture of an embodiment of the present invention;

[0026] Figure 5 Schematic diagram of the network architecture of the spatial attention module in the overall detection network architecture of an embodiment of the present invention. DETAILED DESCRIPTION

[0027] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0028] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0029] like Figure 1 As shown, this embodiment provides a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, including:

[0030] By using a Yolo-based target detector and a network built with downsampling and multi-dimensional attention enhancement, the computer runs the network for target detection. The implementation steps are as follows:

[0031] Step 1: Download the Terahertz Dataset I from GitHub h =[I h1 , I h2 ,...I hK ], where dataset I h The total number of elements is K=2841, the image size is 3×640×640, and the image annotation information file format is xml.

[0032] Step 2.1: Build Figure 2 The network model shown, the deep learning network model includes a backbone feature extraction network, a neck feature extraction network, a downsampling module, a fusion channel and spatial attention, spatial attention and YoloHead detection head structure.

[0033] Use size n ch The RGB image of size h×w is input into the backbone of the network model, and passes through the convolution, C3k2 layer, downsampling module, SPPF layer and fusion channel and spatial attention module, and outputs three feature maps of different scales I b1 , I b2 and I b3 , the scales are 128×80×80, 256×40×40 and 512×20×20 respectively. The number of C3k2 layers is set to 2, 2, 2, 2. The fusion channel and spatial attention module are connected after the SPPF layer, and the 8-head windowed axial attention mechanism is used to b1 , I b2 Perform dynamic weight calibration across channels and spaces to achieve lossless feature transfer between backbone and neck.

[0034] In the overall structure of the neck, the downsampling module is used to replace the original Conv module to reduce the loss of high-frequency information in the downsampling process. The strip pooling of the introduced spatial attention module captures long-range context dependencies, and then the dynamic sparse attention matrix is ​​used to calculate the cross-feature map correlation weights. The neck outputs three feature maps of different scales I n1 , I n2 and I n3 The scales are 128×80×80, 256×40×40 and 512×20×20 respectively.

[0035] Step 2.2: Downsampling module such as Figure 3 As shown in the figure, the core idea is to utilize different pooling and convolution operations in parallel to reduce the amount of computation while retaining diverse feature information. By combining channel segmentation with dual-path processing, the spatial resolution of the input feature map is reduced and downsampled, and the number of channels is adjusted. First, the input is preprocessed by performing 2×2 average pooling on the input x with a stride of 1 and padding of 0. This initially reduces the spatial dimension and alleviates the computational pressure of subsequent operations. Formula:

[0036] X pool =AvgPool2d(x, kernel_size=2, stride=1);

[0037] In the above formula, x represents the input tensor, X pool It is the result of the two-dimensional average pooling operation on the input tensor x. AvgPool2d() represents the calculation of the local area average value for each channel of the input tensor. kernel_size represents the size of the pooling window, and stride represents the step size of each sliding of the pooling window.

[0038] Then perform channel segmentation and convert X pool The channel dimension is divided into two parts x1 and x2, with the purpose of separating feature channels and allowing the two paths to independently process different feature patterns. In path 1, 3×3 convolution is used for downsampling, with stride = 2 and padding = 1 to capture local details while compressing the spatial dimension to 1 / 2 of the original input. Formula: x1 out =Conv 3×3 (x1), x1 out Is the result of convolution of input x1, Conv 3×3 () represents a 3×3 convolution downsampling operation. In path 2, it first passes through 3×3 maximum pooling with stride = 2 and padding = 1, and then adjusts the channel through 1×1 convolution. The formula is:

[0039] x2 out =Conv 1×1 (MaxPool2d 3×3 (x2));

[0040] In the above formula, x2 out After the input x2 is subjected to maximum pooling to emphasize the significant features, the channel information is fused through convolution. 1×1 () represents 1×1 convolution downsampling operation, MaxPool2d 3×3 () indicates the use of a two-dimensional maximum pooling function with a pooling size of 3×3.

[0041] Finally, x1 out and x2 out Splicing along the channel dimension, the output shape is [b, c2, h / / 2, w / / 2], where c2 is the target number of output channels, fusing local details (path one) with salient features (path two) to enhance the representation ability after downsampling.

[0042] Step 2.3: Fusion channel and spatial attention module as Figure 4 As shown in Figure 1, this is a Transformer encoder-based visual feature enhancement module designed specifically for processing two-dimensional image feature maps. It captures global context through a self-attention mechanism and introduces a two-dimensional sine-cosine positional encoding to preserve spatial position information, making it suitable for tasks requiring long-range dependency modeling. First, the input feature map x of shape [b, c, h, w] is flattened into a sequence of form [b, h × w, c], adapting to the Transformer's sequence processing paradigm, treating spatial positions as sequence elements.

[0043] Then generate a position code that matches the size of the feature map, generate sine-cosine codes for the width w and height h respectively, and then splice them into complete position information. Taking width as an example, the formula is:

[0044]

[0045] PE(w)=[sin(w·ω), cos(w·ω)];

[0046] In the above two formulas, i is an index variable used to traverse each frequency component dimension of the position encoding, w i represents the frequency component and is the key parameter for controlling the wavelength and frequency distribution of position encoding. pso_dim represents the position in each direction. temperature is the temperature parameter that controls the attenuation rate of the frequency component. PE(w) is the position encoding component in the width direction. sin(w·ω) represents the sine encoding component of position w, and cos(w·ω) represents the cosine encoding component of position w.

[0047] The final positional encoding dimension shape is [1, H×w, c].

[0048] Multi-Head Attention and Feedforward Network (FFN) input sequence features [b, h×w, c] plus position encoding PE, and output enhanced sequence features [b, h×w, c]. The formula is:

[0049] y seq =TransformerEncoderLayer(x flat +PE);

[0050] In the above formula, y seq Represents the sequence after Transformer encoding. The vector of each position incorporates the global dependency. TransformerEncoderLayer() represents the standard Transformer encoding layer, which includes multi-head self-attention and feedforward networks. flat It represents the sequence of flattened input feature maps, and PE represents two-dimensional sine-cosine position encoding.

[0051] Finally, the output sequence [b, h×w, c] is restored to its original shape [b, c, h, w] to maintain compatibility with the downstream convolutional network. The formula is:

[0052] y=Reshape(y seq );

[0053] In the above formula, y represents the final output feature map, which has the same shape as the original input. Reshape() restores the sequence to a two-dimensional feature map and restores the spatial structure. seq Represents the sequence encoded by Transformer.

[0054] Step 2.4: Spatial attention module such as Figure 5 As shown in , it utilizes horizontal and vertical strip pooling operations to collect long-range context from different spatial dimensions. Let x∈R C×H×W is the input tensor, where C represents the number of channels, H and W are the spatial height and width respectively. First, X is input into two parallel paths, each path contains a horizontal or vertical strip pooling layer, followed by a one-dimensional convolution layer with a kernel size of 3, which is used to modulate the current position and its neighboring features. Define y h ∈R C×H and y v ∈R C×W In order to obtain an output z∈R that contains more useful global priors C×H×W , firstly, y h and y v Put together, as shown below, to produce y∈R C×H×W :

[0055]

[0056] Then, the output feature z is calculated as:

[0057] z = Scale(x,σ(f(y)));

[0058] In the above formula, Scale(,) refers to element-by-element multiplication, σ is the sigmoid function, which is used to adaptively weight multi-scale context features by generating spatial attention weights to enhance the feature response of important areas and suppress irrelevant background noise. f() is a 1×1 convolution.

[0059] Step 3: Specific training parameter configuration includes: base learning rate lr = 0.01, batch size batchsize = 16, training set and validation set split of 0.9:0.1, optimizer type select SGD, and total training cycle epochs = 300.

[0060] Step 4: Use the trained network to make predictions, input the test image, and output the terahertz image target prediction target box. First, the image to be tested I t Input to the network, the image size is n ch ×h×w, after network inference, the Yolohead output is obtained. The output feature maps are scaled 80×80, 40×40, and 20×20. The classification and regression prediction results are extracted from the feature maps of different scales and concatenated and dimensionally transformed. For ease of processing, the original channel dimension is permuted to the end, resulting in the shape of the class prediction branch and the bounding box prediction branch being (b,8400,80) and (b,8400,4), respectively. All objects are sorted in descending order by the confidence level of the object presence (conf=0.001). The inter-connection-of-union (IOU) with other predictions is then calculated from high to low, and predictions with an IOU greater than a certain threshold (iou=0.6) are discarded. Subsequently, according to the previous preprocessing process, the remaining detection boxes are restored to the original image scale before the network output and non-maximum suppression is performed to remove redundant detection boxes. The number of detection boxes output does not exceed the preset maximum number of detections (max_per_img=300).

[0061] The original coordinate parameters (x, y, w, h) of the bounding box of the detection target are converted to numerical values ​​in the normalized coordinate system (X, Y, W, H) and visually annotated on the test image. If the system identifies a valid detection frame in the image to be inspected, it is determined that the person being inspected is carrying dangerous items. If no detection frame is generated, the person being inspected is determined not to be carrying prohibited items.

[0062] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, characterized by: include: Obtaining a terahertz image dataset, wherein the dataset includes an XML file marking the location of dangerous goods; Build a deep learning network model, which includes a backbone feature extraction network, a neck feature extraction network, a downsampling module, a fusion channel and spatial attention module, a spatial attention module and a YoloHead detection head connected in sequence; In the downsampling module, the input feature map is processed in a dual path through parallel pooling and convolution operations; In the fusion channel and spatial attention module, the feature map is modeled with global context through the Transformer encoder; In the spatial attention module, long-range contextual dependencies are captured through striped pooling layers; Input the multi-scale feature map output by the neck network into the YoloHead detection head and output the target detection frame in the normalized coordinate system; Redundant detection boxes are removed through the non-maximum suppression algorithm, and the detection results with the highest confidence are retained.

2. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: The dual-path processing of the input feature map through parallel pooling and convolution operations includes: performing average pooling preprocessing on the input feature map; dividing the preprocessed feature map into two paths along the channel dimension; the first path captures local details through 3×3 convolution downsampling; the second path fuses channel information through maximum pooling and 1×1 convolution; and splicing the two outputs along the channel dimension.

3. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: Global context modeling of feature maps through Transformer encoder includes: flattening the input feature map into a sequence form and adding two-dimensional sine-cosine position encoding; using a multi-head self-attention mechanism to capture long-range dependencies; and enhancing feature expression through a feedforward network.

4. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism as claimed in claim 1 is characterized in that: Capturing long-range contextual dependencies through strip pooling layers includes: horizontally and vertically strip pooling the input feature map; modulating the pooling result through 1×1 convolution; generating spatial attention weights and weighting them to the original feature map.

5. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism as claimed in claim 1 is characterized in that: The construction process of the downsampling module includes: in the dual-path processing, the step size of the first 3×3 convolution is set to 2, and the padding value is 1; the window size of the second maximum pooling is 3×3, the step size is 2, and the padding value is 1; the number of output channels of the two paths is the same and feature fusion is achieved by element-by-element addition.

6. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: The construction process of the fusion channel and spatial attention module includes: when the position encoding is generated, the temperature parameter temperature is set to 10000 and pos_dim is set to 128; the number of heads of the multi-head self-attention mechanism is 8, and the dimension of each head is 64.

7. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: The construction process of the spatial attention module includes: the strip pooling layer includes two parallel branches in the horizontal direction and the vertical direction, each branch contains a 3×1 convolution kernel; when calculating the dynamic weight, the number of channels of the 1×1 convolution is reduced by a ratio of 4.

8. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism as claimed in claim 1, characterized in that: The structure of the neck network includes: connecting the SPPF layer after the downsampling module, and the pyramid scale of the SPPF layer is set to 5, 9, and 13; the input of the fusion channel and the spatial attention module is the feature map of the first two scales of the neck network.

Citation Information

Patent Citations

  • Hazardous chemical vehicle detection method based on improved YOLOv5

    CN116863227A

  • Terahertz image hazardous article detection method based on multi-scale decomposition convolution

    CN117095158A

  • Dangerous article detection method and device based on cross fusion attention mechanism

    CN117115583A

  • Absolute pose regression method based on cascade attention module

    CN118657831A

  • Terahertz target detection model training method using novel attention mechanism

    CN119027770A

Cited By

  • Retinal blood vessel image segmentation method, device, equipment and medium

    CN121053394A

  • Adaptive downsampling time sequence coding method and system

    CN121690216A