Subway hazardous article terahertz detection method based on adaptive downsampling and multi-dimensional attention mechanism
The terahertz detection method for hazardous materials in subways, which uses adaptive downsampling and multi-dimensional attention mechanisms, solves the problems of weak feature extraction capabilities and low robustness in traditional detection methods, and achieves improved detection accuracy and robustness. It is suitable for hazardous materials detection in scenarios such as subways, rail transit, and aviation hubs.
Patent Information
- Application Number
- CN202510727740.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Traditional terahertz detection methods have problems in dangerous goods detection, such as weak feature extraction capability, low robustness, and insufficient utilization of spatial information, resulting in reduced detection accuracy.
A terahertz detection method for subway hazardous materials adopts adaptive downsampling and multi-dimensional attention mechanism. It processes feature maps through parallel pooling and convolution operations, combines the Transformer encoder for global context modeling, and uses strip pooling layers to capture long-range context dependencies and enhance target edge features.
The detection accuracy and robustness have been improved, making it suitable for real-time dangerous goods detection in scenarios such as subways, rail transit, and aviation hubs.
Smart Images

Figure CN120635557A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of dangerous goods detection, and in particular relates to a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism. Background Art
[0002] As an emerging nondestructive testing solution in the security field, terahertz wave detection technology offers significant advantages due to its non-ionizing radiation properties. This technology achieves penetrating identification by analyzing the unique molecular vibrational response of a substance. This technology can effectively overcome complex environmental interference and accurately locate potentially dangerous items hidden within clothing, packages, or containers. Its biocompatibility ensures that the detection process does not cause ionizing damage to human tissue or the environment. This technology is particularly well-suited for intelligent screening of hidden threats in high-security scenarios such as airports and rail transit hubs, providing key technical support for the development of a non-invasive security inspection system.
[0003] The rapid development of deep learning technology has provided powerful tools for image recognition and detection. However, traditional object detection tools still have many limitations for hazardous material detection in the terahertz field: weak feature extraction capabilities: traditional downsampling modules tend to lose high-frequency details, affecting the detection of small targets; low robustness: conventional attention mechanisms such as SE and CBAM modules struggle to effectively distinguish hazardous materials from background noise; and spatial information utilization without pooling operations tends to weaken target edge features, resulting in reduced detection accuracy. Therefore, it is urgent to propose a terahertz detection method for subway hazardous materials based on adaptive downsampling and multi-dimensional attention mechanisms. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, which improves the detection capability of subway dangerous goods.
[0005] To achieve the above objectives, the present invention provides a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, comprising:
[0006] Obtaining a terahertz image dataset, wherein the dataset includes an XML file marking the location of dangerous goods;
[0007] Build a deep learning network model, which includes a backbone feature extraction network, a neck feature extraction network, a downsampling module, a fusion channel and spatial attention module, a spatial attention module and a YoloHead detection head connected in sequence;
[0008] In the downsampling module, the input feature map is processed in a dual path through parallel pooling and convolution operations;
[0009] In the fusion channel and spatial attention module, the feature map is modeled with global context through the Transformer encoder;
[0010] In the spatial attention module, long-range context dependencies are captured through striped pooling layers;
[0011] Input the multi-scale feature map output by the neck network into the YoloHead detection head and output the target detection frame in the normalized coordinate system;
[0012] Redundant detection boxes are removed through the non-maximum suppression algorithm, and the detection results with the highest confidence are retained.
[0013] Optionally, dual-path processing of the input feature map through parallel pooling and convolution operations includes: performing average pooling preprocessing on the input feature map; dividing the preprocessed feature map into two paths along the channel dimension; the first path captures local details through 3×3 convolution downsampling; the second path fuses channel information through maximum pooling and 1×1 convolution; and splicing the two outputs along the channel dimension.
[0014] Optionally, global context modeling of feature maps through Transformer encoders includes: flattening the input feature maps into sequence form and adding two-dimensional sine-cosine position encoding; using a multi-head self-attention mechanism to capture long-range dependencies; and enhancing feature expression through a feedforward network.
[0015] Optionally, capturing long-range contextual dependencies through a strip pooling layer includes: performing horizontal and vertical strip pooling on the input feature map; modulating the pooling result through a 1×1 convolution; generating spatial attention weights and weighting them to the original feature map.
[0016] Optionally, the construction process of the downsampling module includes: in the dual-path processing, the step size of the first 3×3 convolution is set to 2, and the padding value is 1; the window size of the second maximum pooling is 3×3, the step size is 2, and the padding value is 1; the two output channels have the same number and feature fusion is achieved by element-by-element addition.
[0017] Optionally, the construction process of the fusion channel and spatial attention module includes: when the position code is generated, the temperature parameter temperature is set to 10000, and pos_dim is set to 128; the number of heads of the multi-head self-attention mechanism is 8, and the dimension of each head is 64.
[0018] Optionally, the construction process of the spatial attention module includes: the strip pooling layer includes two parallel branches in the horizontal direction and the vertical direction, each branch contains a 3×1 convolution kernel; when calculating the dynamic weight, the number of channels of the 1×1 convolution is reduced by a ratio of 4.
[0019] Optionally, the structure of the neck network includes: connecting an SPPF layer after the downsampling module, and the pyramid scale of the SPPF layer is set to 5, 9, and 13; the input of the fusion channel and the spatial attention module is the feature map of the first two scales of the neck network.
[0020] Technical effect of the invention: The present invention discloses a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism. By introducing adaptive downsampling and multi-dimensional attention, the high-frequency information loss in the downsampling process is reduced, global context modeling and local feature refinement are combined, and long-range dependencies are captured through strip pooling layers. The target edge and shape features are optimized, so that the method has significant improvements in feature extraction capability, detection accuracy and robustness compared with traditional target detection algorithms. It is suitable for real-time applications and various scenarios of terahertz image dangerous goods detection. It has broad application prospects in scenarios such as subway dangerous goods detection, rail transportation, aviation hubs and large-scale event security. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0022] Figure 1 Schematic diagram of computer detection of terahertz images in terahertz dangerous goods detection according to an embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of the overall architecture of the detection network according to an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the network architecture of the downsampling module in the overall detection network architecture according to an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the fusion channel and spatial attention module network in the overall detection network architecture of an embodiment of the present invention;
[0026] Figure 5 Schematic diagram of the network architecture of the spatial attention module in the overall detection network architecture of an embodiment of the present invention. DETAILED DESCRIPTION
[0027] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0028] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0029] like Figure 1 As shown, this embodiment provides a terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, including:
[0030] By using a Yolo-based target detector and a network built with downsampling and multi-dimensional attention enhancement, the computer runs the network for target detection. The implementation steps are as follows:
[0031] Step 1: Download the Terahertz Dataset I from GitHub h =[I h1 , I h2 ,...I hK ], where dataset I h The total number of elements is K=2841, the image size is 3×640×640, and the image annotation information file format is xml.
[0032] Step 2.1: Build Figure 2 The network model shown, the deep learning network model includes a backbone feature extraction network, a neck feature extraction network, a downsampling module, a fusion channel and spatial attention, spatial attention and YoloHead detection head structure.
[0033] Use size n ch The RGB image of size h×w is input into the backbone of the network model, and passes through the convolution, C3k2 layer, downsampling module, SPPF layer and fusion channel and spatial attention module, and outputs three feature maps of different scales I b1 , I b2 and I b3 , the scales are 128×80×80, 256×40×40 and 512×20×20 respectively. The number of C3k2 layers is set to 2, 2, 2, 2. The fusion channel and spatial attention module are connected after the SPPF layer, and the 8-head windowed axial attention mechanism is used to b1 , I b2 Perform dynamic weight calibration across channels and spaces to achieve lossless feature transfer between backbone and neck.
[0034] In the overall structure of the neck, the downsampling module is used to replace the original Conv module to reduce the loss of high-frequency information in the downsampling process. The strip pooling of the introduced spatial attention module captures long-range context dependencies, and then the dynamic sparse attention matrix is used to calculate the cross-feature map correlation weights. The neck outputs three feature maps of different scales I n1 , I n2 and I n3 The scales are 128×80×80, 256×40×40 and 512×20×20 respectively.
[0035] Step 2.2: Downsampling module such as Figure 3 As shown in the figure, the core idea is to utilize different pooling and convolution operations in parallel to reduce the amount of computation while retaining diverse feature information. By combining channel segmentation with dual-path processing, the spatial resolution of the input feature map is reduced and downsampled, and the number of channels is adjusted. First, the input is preprocessed by performing 2×2 average pooling on the input x with a stride of 1 and padding of 0. This initially reduces the spatial dimension and alleviates the computational pressure of subsequent operations. Formula:
[0036] X pool =AvgPool2d(x, kernel_size=2, stride=1);
[0037] In the above formula, x represents the input tensor, X pool It is the result of the two-dimensional average pooling operation on the input tensor x. AvgPool2d() represents the calculation of the local area average value for each channel of the input tensor. kernel_size represents the size of the pooling window, and stride represents the step size of each sliding of the pooling window.
[0038] Then perform channel segmentation and convert X pool The channel dimension is divided into two parts x1 and x2, with the purpose of separating feature channels and allowing the two paths to independently process different feature patterns. In path 1, 3×3 convolution is used for downsampling, with stride = 2 and padding = 1 to capture local details while compressing the spatial dimension to 1 / 2 of the original input. Formula: x1 out =Conv 3×3 (x1), x1 out Is the result of convolution of input x1, Conv 3×3 () represents a 3×3 convolution downsampling operation. In path 2, it first passes through 3×3 maximum pooling with stride = 2 and padding = 1, and then adjusts the channel through 1×1 convolution. The formula is:
[0039] x2 out =Conv 1×1 (MaxPool2d 3×3 (x2));
[0040] In the above formula, x2 out After the input x2 is subjected to maximum pooling to emphasize the significant features, the channel information is fused through convolution. 1×1 () represents 1×1 convolution downsampling operation, MaxPool2d 3×3 () indicates the use of a two-dimensional maximum pooling function with a pooling size of 3×3.
[0041] Finally, x1 out and x2 out Splicing along the channel dimension, the output shape is [b, c2, h / / 2, w / / 2], where c2 is the target number of output channels, fusing local details (path one) with salient features (path two) to enhance the representation ability after downsampling.
[0042] Step 2.3: Fusion channel and spatial attention module as Figure 4 As shown in Figure 1, this is a Transformer encoder-based visual feature enhancement module designed specifically for processing two-dimensional image feature maps. It captures global context through a self-attention mechanism and introduces a two-dimensional sine-cosine positional encoding to preserve spatial position information, making it suitable for tasks requiring long-range dependency modeling. First, the input feature map x of shape [b, c, h, w] is flattened into a sequence of form [b, h × w, c], adapting to the Transformer's sequence processing paradigm, treating spatial positions as sequence elements.
[0043] Then generate a position code that matches the size of the feature map, generate sine-cosine codes for the width w and height h respectively, and then splice them into complete position information. Taking width as an example, the formula is:
[0044]
[0045] PE(w)=[sin(w·ω), cos(w·ω)];
[0046] In the above two formulas, i is an index variable used to traverse each frequency component dimension of the position encoding, w i represents the frequency component and is the key parameter for controlling the wavelength and frequency distribution of position encoding. pso_dim represents the position in each direction. temperature is the temperature parameter that controls the attenuation rate of the frequency component. PE(w) is the position encoding component in the width direction. sin(w·ω) represents the sine encoding component of position w, and cos(w·ω) represents the cosine encoding component of position w.
[0047] The final positional encoding dimension shape is [1, H×w, c].
[0048] Multi-Head Attention and Feedforward Network (FFN) input sequence features [b, h×w, c] plus position encoding PE, and output enhanced sequence features [b, h×w, c]. The formula is:
[0049] y seq =TransformerEncoderLayer(x flat +PE);
[0050] In the above formula, y seq Represents the sequence after Transformer encoding. The vector of each position incorporates the global dependency. TransformerEncoderLayer() represents the standard Transformer encoding layer, which includes multi-head self-attention and feedforward networks. flat It represents the sequence of flattened input feature maps, and PE represents two-dimensional sine-cosine position encoding.
[0051] Finally, the output sequence [b, h×w, c] is restored to its original shape [b, c, h, w] to maintain compatibility with the downstream convolutional network. The formula is:
[0052] y=Reshape(y seq );
[0053] In the above formula, y represents the final output feature map, which has the same shape as the original input. Reshape() restores the sequence to a two-dimensional feature map and restores the spatial structure. seq Represents the sequence encoded by Transformer.
[0054] Step 2.4: Spatial attention module such as Figure 5 As shown in , it utilizes horizontal and vertical strip pooling operations to collect long-range context from different spatial dimensions. Let x∈R C×H×W is the input tensor, where C represents the number of channels, H and W are the spatial height and width respectively. First, X is input into two parallel paths, each path contains a horizontal or vertical strip pooling layer, followed by a one-dimensional convolution layer with a kernel size of 3, which is used to modulate the current position and its neighboring features. Define y h ∈R C×H and y v ∈R C×W In order to obtain an output z∈R that contains more useful global priors C×H×W , firstly, y h and y v Put together, as shown below, to produce y∈R C×H×W :
[0055]
[0056] Then, the output feature z is calculated as:
[0057] z = Scale(x,σ(f(y)));
[0058] In the above formula, Scale(,) refers to element-by-element multiplication, σ is the sigmoid function, which is used to adaptively weight multi-scale context features by generating spatial attention weights to enhance the feature response of important areas and suppress irrelevant background noise. f() is a 1×1 convolution.
[0059] Step 3: Specific training parameter configuration includes: base learning rate lr = 0.01, batch size batchsize = 16, training set and validation set split of 0.9:0.1, optimizer type select SGD, and total training cycle epochs = 300.
[0060] Step 4: Use the trained network to make predictions, input the test image, and output the terahertz image target prediction target box. First, the image to be tested I t Input to the network, the image size is n ch ×h×w, after network inference, the Yolohead output is obtained. The output feature maps are scaled 80×80, 40×40, and 20×20. The classification and regression prediction results are extracted from the feature maps of different scales and concatenated and dimensionally transformed. For ease of processing, the original channel dimension is permuted to the end, resulting in the shape of the class prediction branch and the bounding box prediction branch being (b,8400,80) and (b,8400,4), respectively. All objects are sorted in descending order by the confidence level of the object presence (conf=0.001). The inter-connection-of-union (IOU) with other predictions is then calculated from high to low, and predictions with an IOU greater than a certain threshold (iou=0.6) are discarded. Subsequently, according to the previous preprocessing process, the remaining detection boxes are restored to the original image scale before the network output and non-maximum suppression is performed to remove redundant detection boxes. The number of detection boxes output does not exceed the preset maximum number of detections (max_per_img=300).
[0061] The original coordinate parameters (x, y, w, h) of the bounding box of the detection target are converted to numerical values in the normalized coordinate system (X, Y, W, H) and visually annotated on the test image. If the system identifies a valid detection frame in the image to be inspected, it is determined that the person being inspected is carrying dangerous items. If no detection frame is generated, the person being inspected is determined not to be carrying prohibited items.
[0062] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism, characterized by: include: Obtaining a terahertz image dataset, wherein the dataset includes an XML file marking the location of dangerous goods; Build a deep learning network model, which includes a backbone feature extraction network, a neck feature extraction network, a downsampling module, a fusion channel and spatial attention module, a spatial attention module and a YoloHead detection head connected in sequence; In the downsampling module, the input feature map is processed in a dual path through parallel pooling and convolution operations; In the fusion channel and spatial attention module, the feature map is modeled with global context through the Transformer encoder; In the spatial attention module, long-range contextual dependencies are captured through striped pooling layers; Input the multi-scale feature map output by the neck network into the YoloHead detection head and output the target detection frame in the normalized coordinate system; Redundant detection boxes are removed through the non-maximum suppression algorithm, and the detection results with the highest confidence are retained.
2. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: The dual-path processing of the input feature map through parallel pooling and convolution operations includes: performing average pooling preprocessing on the input feature map; dividing the preprocessed feature map into two paths along the channel dimension; the first path captures local details through 3×3 convolution downsampling; the second path fuses channel information through maximum pooling and 1×1 convolution; and splicing the two outputs along the channel dimension.
3. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: Global context modeling of feature maps through Transformer encoder includes: flattening the input feature map into a sequence form and adding two-dimensional sine-cosine position encoding; using a multi-head self-attention mechanism to capture long-range dependencies; and enhancing feature expression through a feedforward network.
4. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism as claimed in claim 1 is characterized in that: Capturing long-range contextual dependencies through strip pooling layers includes: horizontally and vertically strip pooling the input feature map; modulating the pooling result through 1×1 convolution; generating spatial attention weights and weighting them to the original feature map.
5. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism as claimed in claim 1 is characterized in that: The construction process of the downsampling module includes: in the dual-path processing, the step size of the first 3×3 convolution is set to 2, and the padding value is 1; the window size of the second maximum pooling is 3×3, the step size is 2, and the padding value is 1; the number of output channels of the two paths is the same and feature fusion is achieved by element-by-element addition.
6. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: The construction process of the fusion channel and spatial attention module includes: when the position encoding is generated, the temperature parameter temperature is set to 10000 and pos_dim is set to 128; the number of heads of the multi-head self-attention mechanism is 8, and the dimension of each head is 64.
7. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism according to claim 1 is characterized in that: The construction process of the spatial attention module includes: the strip pooling layer includes two parallel branches in the horizontal direction and the vertical direction, each branch contains a 3×1 convolution kernel; when calculating the dynamic weight, the number of channels of the 1×1 convolution is reduced by a ratio of 4.
8. The terahertz detection method for subway dangerous goods based on adaptive downsampling and multi-dimensional attention mechanism as claimed in claim 1, characterized in that: The structure of the neck network includes: connecting the SPPF layer after the downsampling module, and the pyramid scale of the SPPF layer is set to 5, 9, and 13; the input of the fusion channel and the spatial attention module is the feature map of the first two scales of the neck network.
Citation Information
Patent Citations
Hazardous chemical vehicle detection method based on improved YOLOv5
CN116863227A
Terahertz image hazardous article detection method based on multi-scale decomposition convolution
CN117095158A
Dangerous article detection method and device based on cross fusion attention mechanism
CN117115583A
Absolute pose regression method based on cascade attention module
CN118657831A
Terahertz target detection model training method using novel attention mechanism
CN119027770A
Cited By
Retinal blood vessel image segmentation method, device, equipment and medium
CN121053394A
Adaptive downsampling time sequence coding method and system
CN121690216A