Dangerous article detection method based on improved YOLOv8

By improving the YOLOv8 network structure, combining the subway station monitoring video data and public data sets, fasterNet Block and feature fusion module are used to solve the detection accuracy and real-time problems of YOLOv8 in complex environments, and more efficient detection of dangerous goods at subway stations is achieved.

CN120339940APending Publication Date: 2025-07-18UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510366892.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing YOLOv8 model is in complex environments, especially when there are dense crowds and large changes in light in subway stations, and the detection accuracy and real-time performance are insufficient, making it difficult to meet the needs of subway station safety inspection.

Method used

By improving the YOLOv8 network, fasterNet Block is used to replace the C2f module, feature alignment module FAM and information fusion module IFM are introduced, and the subway station monitoring video data and publicly disclosed hazardous goods detection data sets are combined to build a multi-scale feature fusion mechanism to optimize the model structure to improve detection accuracy and speed.

Benefits of technology

It significantly improves the detection accuracy and real-timeness of the model in complex environments, improves the ability to identify targets at different scales, and enhances technical support for subway station safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339940A_ABST
    Figure CN120339940A_ABST
Patent Text Reader

Abstract

The invention relates to a dangerous goods detection method based on improved YOLOv8, and the method comprises the steps: an improved YOLOv8 network comprises a backbone network backbone, a neck network check, a head network head, a FasterNet Block structure is adopted to replace a Bottleneck calculation unit, a C2f module in the backbone network backbone of the improved YOLOv8, a feature alignment module FAM and an information fusion module IFM are used in the neck network check to achieve the aggregation of different levels of features, and the feature alignment module FAM and the information fusion module IFM are used in the improved YOLOv8 backbone network backbone. And the fused information is distributed back to each level of the network through an information injection module Inject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of article detection, and particularly to a method for detecting dangerous goods based on improved YOLOv8. Background Art

[0002] In the field of object detection, the YOLO series of algorithms has become a benchmark in the industry due to its excellent performance and real-time nature. YOLOv8 is the result of innovative improvements by the Ultralytics team based on the YOLOv5 architecture, significantly enhancing the model's performance in tasks such as object detection, instance segmentation, and image classification. The model consists of a backbone network, a feature extraction network, and a detection head network, and particularly uses the C2f module to replace the C3 module, which not only retains the advantages of the CSP structure but also enhances the feature extraction ability through diverse gradient combinations. Although YOLOv8 has excellent performance, its complex network structure and large number of parameters also bring a high computational cost, especially in detection scenarios that require quick response, which is particularly obvious.

[0003] In the safety detection of public places such as subway stations, accurately and quickly identifying dangerous goods is crucial for ensuring public safety. Although existing detection algorithms based on the YOLO series can meet this requirement to a certain extent, in complex environments, especially in subway stations with dense crowds and large lighting variations, the accuracy and real-time nature of these algorithms still need to be improved. In addition, the combination of video data captured by subway station surveillance cameras and the dangerous goods detection dataset is crucial for constructing an effective training dataset, which has a decisive impact on improving the generalization ability and accuracy of the algorithm. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a method for detecting dangerous goods in subway stations based on improved YOLOv8. This method collects and annotates the video data captured by subway station surveillance cameras, and combines it with the publicly available dangerous goods detection dataset to establish a dangerous goods detection dataset for subway stations. On this basis, in response to the special challenges in the subway station scenario, the present invention can significantly improve the accuracy and effect of the YOLOv8 model in object detection tasks through a series of innovative optimization measures. By deeply analyzing the characteristics of the subway station environment, the present invention proposes targeted improvement solutions, which significantly enhance the detection accuracy and real-time nature of the model in complex environments while maintaining high performance.

[0005] Specifically, the method for detecting dangerous goods based on improved YOLOv8 of the present invention includes the following steps:

[0006] S1: Making a dangerous goods detection dataset;

[0007] S2: Constructing an improved YOLOv8 network;

[0008] S3: Train the improved YOLOv8 network using the detection dataset and save the trained model parameters;

[0009] S4: Obtain the dangerous goods images in the real scene, and use the trained model parameters to detect the dangerous goods in the dangerous goods images;

[0010] Among them, the improved YOLOv8 network includes: a backbone network backbone, a neck network neck, and a head network head. The fasterNet Block structure is used to replace the Bottleneck computing unit. In the neck network neck, a feature alignment module FAM and an information fusion module IFM are used to achieve the aggregation of features at different levels, and the fused information is distributed back to each level of the network through the information injection module Inject.

[0011] Further, the backbone network is sorted from low to high in resolution, and four feature maps B5, B4, B3, and B2 are output. These feature maps are processed through the low-order aggregation and distribution branch module and the high-order aggregation and distribution branch module. The low-order aggregation and distribution branch processes the low-resolution feature maps, and the high-order aggregation and distribution branch processes the high-resolution feature maps processed by the low-order aggregation and distribution branch. The processing results form three output feature maps N5, N4, and N3, and are transmitted to the head network head for final object detection.

[0012] Further, the low-order aggregation and distribution branch module includes a low-order alignment module and a low-order information fusion module, and the high-order aggregation and distribution branch module includes a high-order alignment module and a high-order information fusion module; the high-order alignment module and the low-order alignment module use the self-attention mechanism to align the feature maps at different levels, and the low-order information fusion module and the high-order information fusion module are used to fuse the aligned feature maps. The fused features inject global information into each level through the injection module inject;

[0013] The low-order aggregation and distribution branch module processes the four feature maps B5, B4, B3, and B2 to form the processed feature maps P5, P4, and P3 corresponding to the B5, B4, and B3 feature maps respectively. Then, the feature maps P5, P4, and P3 are processed by the high-order aggregation and distribution branch to form three output feature maps N5, N4, and N3.

[0014] Further, the low-order alignment module downsamples the input feature map through average pooling operation to unify the size; the low-order information fusion module contains multiple layers of reparameterized convolutional blocks and a segmentation operation for fusing the aligned features;

[0015] The high-order alignment module reduces the dimension of the input features through average pooling. The high-order information fusion module consists of multiple stacked transformer blocks, which include multi-head attention mechanisms and feed-forward networks for obtaining global information.

[0016] Further, the module such as inject combines the fused feature map with the original feature map through a series of convolutional operations and Sigmoid activation functions and injects it back into each layer of the network.

[0017] Further, the step S1 includes:

[0018] The network obtains pictures and annotation information of dangerous goods to obtain the public dataset a0;

[0019] Collect the monitoring videos of the actual scene, extract the video frame pictures containing dangerous goods, and annotate the video frame pictures to obtain the private dataset a1;

[0020] Fuse the public dataset a0 and the private dataset a1 and use them as the basis for Mosaic data augmentation for data expansion: adjust the images in the public dataset a0 and the private dataset a1 to the set size and adjust the bounding box coordinates accordingly; randomly select four image regions, splice them into a Mosaic pattern, randomly select a region in the Mosaic image for cutting to form a new sample containing different source image blocks, and obtain the augmented dataset a2;

[0021] Divide the augmented dataset a2 into a training set, a validation set and a test set according to the ratio of 8:1:1, so as to obtain the dangerous goods detection dataset.

[0022] Further, the detection method is used for detecting dangerous goods in subway stations.

[0023] In the present invention, the bottleneck computing unit in the C2f module of the backbone network of the traditional YOLOv8 is replaced by the fasterNet Block structure. This improvement utilizes the partial convolution (PConv) technology, which performs convolution operations only on a part of the input feature map while keeping the remaining channels unchanged, thereby reducing unnecessary calculations and memory accesses. The implementation of the PConv layer involves a masking operation that dynamically selects a part of the input feature map for convolution calculation while ignoring other regions, so as to achieve the purpose of accelerating the calculation.

[0024] Secondly, through the design of low-order (Low-GD) and high-order (High-GD) aggregation-distribution branches, multi-scale feature fusion is achieved. These two branches process feature maps of different sizes through a feature alignment module (FAM) and an information fusion module (IFM). Through this structural model, information from different depth levels of the network can be effectively fused, which is crucial for accurately detecting targets of different sizes.

[0025] Through these improvements, the present invention significantly enhances the feature extraction ability, detection accuracy, and speed of the model in the real scenario of dangerous item detection in subway stations, thus providing more reliable technical support for subway station safety management. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is the architecture diagram of the traditional YOLOv8 network;

[0027] Figure 2 is the architecture diagram of the backbone part of the traditional YOLOv8;

[0028] Figure 3 is the architecture diagram of the C2f module in the backbone part of the traditional YOLOv8;

[0029] Figure 4 is the architecture diagram of the Bottleneck in the C2f module of the traditional YOLOv8;

[0030] Figure 5 is the architecture diagram of the neck part of the traditional YOLOv8;

[0031] Figure 6 is the architecture diagram of the FasterNetBlock module of the present invention;

[0032] Figure 7 is the architecture diagram of the YOLOv8 network of the present invention;

[0033] Figure 8 is the network structure diagram of the Neck module of the YOLOv8 network of the present invention;

[0034] Figure 9 is the network structure diagram of the low-order aggregation and distribution (Low-GD) branch module of the present invention;

[0035] Figure 10 is the network structure diagram of the high-order aggregation and distribution (High-GD) branch module of the present invention;

[0036] Figure 11 is the network structure diagram of the information injection module (Inject) of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] SeeFigure 1 Figure 0 is the architecture diagram of traditional YOLOv8, which includes several parts: backbone, neck, and head. The backbone part is used to extract the features of the input image. Through multiple convolutional layers and pooling layers, it gradually reduces the spatial dimension of the image while increasing the number of channels, thereby obtaining a deep feature representation of the image. The Neck part is used for multi-scale feature fusion. It fuses the feature maps from different stages of the backbone to enhance the representation ability. It includes multiple convolutional layers and upsampling layers, which combine low-level detailed information and high-level semantic information to form a more abundant feature representation. The Head part is used for the specific implementation of the detection task. Based on the fused feature maps, it locates and classifies the targets, and outputs the bounding box coordinates and class probabilities of each target. In addition, there is a loss function part after the head, which is used to calculate the difference between the model prediction results and the true annotations, guiding the training and optimization process of the model. By minimizing the loss function, the detection performance and accuracy of the model are improved.

[0038] See Figure 2 Figure 1, the Backbone part is a key component in the YOLOv8 network responsible for feature extraction. It is implemented through multiple convolutional layers (Conv) and C2f modules. The role of the C2f module in the Backbone is to perform feature fusion to improve the performance of object detection. Specifically, the C2f module connects low-level feature maps and high-level feature maps, and adjusts their number of channels and sizes through appropriate convolutional operations, realizing information transfer and fusion between feature maps at different levels, thereby improving the detection ability of the object detection algorithm for objects of different scales.

[0039] See Figure 3 Figure 2, the C2f module includes a series of Conv layers and Bottleneck blocks, which are computing units that extract features at different scales and integrate information through channel concatenation (Concat). See Figure 4 Traditional Bottlenecks mainly rely on standard 3×3 convolutions for feature extraction, with a large amount of computation, high demand for computing resources during the inference process, and affecting the inference efficiency of the model.

[0040] See Figure 5, the Neck part of YOLOv8 is a crucial link in the network structure, responsible for multi-scale feature fusion. This part enhances the feature representation ability by fusing feature maps from different stages of the Backbone. The Neck part includes multiple convolutional layers that receive multi-scale feature maps output by the Backbone and further extract and integrate feature information through feature fusion techniques such as the Path Aggregation Network (PANet) and the Spatial Pyramid Pooling Fast Version (SPPF). During the feature fusion process, the Neck part first aggregates multi-layer features from the Backbone through a bottom-up path and then transfers high-level semantic information to the low-level feature maps through a top-down path. Such a design enables the model to more effectively utilize feature information at different scales and improve the detection ability for targets of different sizes. Finally, the feature maps processed by the Neck part will be passed to the Head part for the final object detection and classification tasks.

[0041] Based on the original YOLOv8 network, the present invention constructs the SDO-YOLO network model by introducing the fasterNet Block, the aggregation-distribution mechanism (GD), the feature alignment module (FAM), the information fusion module (IFM), and the information injection module (Inject).

[0042] The following details the improvements of the SDO-YOLO network model of the present invention compared to YOLOV8:

[0043] (1) Improve the C2f module of the Backbone using the fasterNet Block:

[0044] In the backbone network of the original YOLOv8 network, the C2f module is a key component that achieves multi-scale fusion of features through Cross Stage Partial. To further improve the efficiency and effect of feature extraction, the present invention uses the fasterNet Block structure to replace the Bottleneck calculation unit in the original C2f module.

[0045] Specifically, see Figure 6, the FasterNet Block replaces the standard 3×3 convolution calculation in the Bottleneck with partial convolution (PConv) and an MLP structure (1×1Conv + BN + ReLU + 1×1Conv), reducing computational redundancy while enhancing the effectiveness of feature extraction. Structurally, the FasterNet Block still maintains the overall framework of the C2f module, including parts such as the input Conv 1×1 for channel adjustment, Split for channel division, Concat for channel concatenation, and the output Conv 1×1 for feature fusion, but redesigned the Bottleneck computational unit. PConv only performs 3×3 convolution calculations on a part of the channels of the input feature map while keeping the remaining channels unchanged. This design reduces unnecessary calculations and memory accesses while retaining key feature information. The implementation of the PConv layer involves a masking operation that dynamically selects a part of the input feature map for convolution calculation while ignoring other areas, thereby achieving the purpose of accelerating the calculation. The MLP structure further optimizes the feature expression ability, enabling full interaction of channel information and improving the model's recognition ability for different targets.

[0046] Such improvements aim to reduce the computational amount, increase the running speed of the network, and maintain or improve the detection accuracy. Through this improvement, it is possible to reduce the computational burden of the model while maintaining the quality of feature extraction, and improve the performance of the model in the subway station dangerous goods detection task, especially the detection accuracy and speed when dealing with high-resolution images and complex scenes.

[0047] (2) Improve the feature fusion network of the original Neck module using the aggregation-distribution mechanism (GD):

[0048] The present invention improves the Neck module of YOLOv8 by introducing the aggregation-distribution mechanism (GD), which aggregates features at different levels through a feature alignment module (FAM) and an information fusion module (IFM), and distributes the fused information back to each level of the network through an information injection module (Inject).

[0049] In the Neck part of YOLOv8, there are usually three inputs and three outputs. These three input feature maps come from different levels of the Backbone, usually B5, B4, B3, which respectively correspond to the output feature maps of the 5th, 4th, and 3rd layers of the Backbone. See Figures 7 - 10 , in the present invention, the processing of the B2 feature map is integrated in the Neck part, making full use of the advantage of the high resolution of the B2 feature map, enhancing the model's detection ability for small-sized targets, and improving the overall detection performance through feature fusion and information injection.

[0050] These input feature maps are processed through the Low-GD (Low-order Aggregation Distribution) branch and the High-GD (High-order Aggregation Distribution) branch. The Low-GD branch processes the low-resolution feature maps, while the High-GD branch processes the high-resolution feature maps that have been processed by the Low-GD branch. The output feature maps N5, N4, and N3 are the results after being processed by the High-GD branch and will be passed to the Head part for final object detection.

[0051] See Figure 8 , where B2 - B5 represent the output feature maps of different levels of the Backbone part. B2 is the output of the second layer of the Backbone, with the highest resolution and containing the most detailed spatial information, suitable for detecting small objects. B5 corresponds to the output of the fifth layer, with the lowest resolution and containing the most extensive spatial information, suitable for detecting large objects. The feature map resolutions of B3 and B4 are between those of B2 and B5.

[0052] P3 - P5 represent the feature maps after being processed by the Low-GD branch. The Low-GD branch processes the feature maps of B2 - B5 through the Low-FAM (Low-order Feature Alignment Module), Low-IFM (Low-order Information Fusion Module), and Inject module to enhance the feature representation and prepare for passing to the High-GD branch. P3, P4, and P5 are these processed feature maps, which correspond to B3, B4, and B5 respectively.

[0053] N3 - N5 represent the feature maps after being processed by the High-GD branch. The High-GD branch further processes the P3 - P5 feature maps through the High-FAM, High-IFM, and Inject modules to improve the detection ability for large-sized objects. N3, N4, and N5 are these processed feature maps and will be passed to the Head part.

[0054] In the FAM, the self-attention mechanism is used to align the feature maps of different levels so that each feature map can obtain global context information. Specifically, the self-attention operation in the FAM calculates the correlation between each feature map and all other feature maps, and then adjusts the weights of the feature maps based on these correlations. The IFM is responsible for fusing the feature maps aligned by the FAM. This step involves weighted summation of the feature maps, where the weights are determined by the output of the FAM module. In this way, the IFM can effectively combine the feature maps of different levels to form a more rich feature representation.

[0055] Multi-scale feature fusion is a commonly used technique in object detection models, aiming to improve the model's detection ability for objects of different sizes. By combining features from different levels of the network, this technique can capture information at multiple scales from coarse to fine. Low-level features usually contain more details about small objects, while high-level features capture the semantic information of large objects. Multi-scale feature fusion enhances the model's representation ability by aggregating features from these levels, enabling the model to more accurately identify and locate objects of various sizes in the image.

[0056] The multi-scale feature fusion of the present invention is reflected in the design of the low-order (Low-GD) and high-order (High-GD) aggregation-distribution branches, which process feature maps of different sizes through a feature alignment module (FAM) and an information fusion module (IFM). Through this structural model, the network can effectively fuse information from different depth levels, which is crucial for accurately detecting objects of different sizes.

[0057] See Figures 9 - 10 , the low-order aggregation-distribution (Low-GD) branch includes a low-order feature alignment module (LoW-FAM) and a low-order information fusion module (Low-IFM), and the high-order aggregation-distribution (High-GD) branch contains a high-order feature alignment module (High-FAM) and a high-order information fusion module (High-IFM). These two branches are the key parts of the model for processing feature maps of different sizes and improving object detection performance. Through the feature alignment (FAM) and information fusion (IFM) modules at different scales, the model's ability to process features at different scales and improve object detection performance is enhanced.

[0058] The Low-GD branch is responsible for processing low-resolution feature maps such as B2, B3, B4, B5. It aligns features through the Low-FAM module, then fuses features through the Low-IFM module, and finally injects global information into each level through the Inject module. The High-GD branch processes the P3, P4, P5 feature maps output by the Low-GD branch. It aligns features through the High-FAM module, then fuses features through the High-IFM module, and finally injects global information into each level through the Inject module.

[0059] The low-FAM module downsamples the input feature map through an average pooling (AvgPool) operation to unify the size. The low-IFM module contains multiple layers of reparameterized convolutional blocks (RepBlock) and a segmentation operation for fusing the aligned features.

[0060] The high-FAM module reduces the dimension of the input features through average pooling (AvgPool). The high-IFM module consists of multiple stacked transformer blocks, which include multi-head attention mechanisms and feed-forward networks, to obtain global information.

[0061] See Figure 11 The role of the information injection module (Inject) is to distribute the information after fusing FAM and IFM back to each level of the network. This step is crucial for maintaining the multi-scale feature extraction ability of the network. The Inject module combines the fused feature map with the original feature map through a series of convolutional operations and Sigmoid activation functions. Specifically, the Inject module first uses 1x1 convolution to compress the number of channels of the fused feature map, and then generates a weight map through the Sigmoid activation function, which is used to control the fusion degree of the original feature map and the fused feature map. Finally, the Inject module adds the weighted fused feature map to the original feature map to inject the fused information back to each level of the network, forming the SDO-YOLO network.

[0062] The following takes the detection of dangerous goods in a subway station as an example to illustrate the present invention:

[0063] 1. Obtain the dataset required for the model, and the specific methods are as follows:

[0064] (1) Download the publicly available dangerous goods detection dataset from the roboflow website, with a total of 1000 pictures, including the annotation information of dangerous goods such as sticks and daggers, to obtain dataset a0;

[0065] (2) Obtain the surveillance video of the subway station in the real environment, select available segments, and extract video frame pictures;

[0066] (3) Refer to the publicly available dangerous goods detection dataset, use the annotation tool LabelImg to annotate the sticks and knives in the subway station video frame pictures, and automatically generate the corresponding TXT format annotation files, which contain the object names and the coordinate information of the bounding boxes. The categories are stick for clubs and knife for daggers, to obtain the private dataset a1 of 1000 pictures;

[0067] (4)Fuse the public dataset a0 and the private dataset a1, and use it as the basis for Mosaic data augmentation to expand the data. First, determine the target input shape, that is, set the size of the output image to 640×640, adjust the size of the images in each dataset to match the target shape, and correspondingly adjust the bounding box coordinates. Then, randomly select four image regions from the dataset and splice them into a Mosaic pattern. While splicing, merge and adjust the bounding boxes to ensure correct positions in the new image. Finally, randomly select a region in the Mosaic image for cutting to form a new sample containing different source image patches, and obtain the augmented dataset a2;

[0068] (5)Divide the dataset a2 into a training set, a validation set, and a test set according to the ratio of 8:1:1 to obtain the train.txt, val.txt, and test.txt files; obtain the final subway station dangerous goods detection dataset (SDO)

[0069] II. Set model parameters:

[0070] (1) Network input size setting: Define the input image size of the network model, and set both the width w and the height h to 640 pixels to match the resolution of the images in the dataset.

[0071] (2) Initialize the model structure: Construct the initial model structure model_body of the SDO-YOLO network. This model will include the improved C2f module and the feature alignment module, information fusion module, and information injection module required by the aggregation-distribution mechanism (GD).

[0072] (3) Configure training parameters: Set the number of training epochs epochs = 300, and iterate through the entire dataset 300 times. Set the initial learning rate lr = 0.001, and adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the model performance during training. Set the batch size batch size = 16, and process 16 images during each iteration update.

[0073] (4) Training and parameter update: Start training the SDO-YOLO network from the initial epoch = 0. At the end of each epoch, evaluate the model according to the performance on the validation set, and save the best model parameters best.pt. During the training process, use stochastic gradient descent (SGD) as the optimization algorithm to update the network parameters.

[0074] (5) Save the final model: When the training is completed for 300 epochs, save the final model weights last.pt, that is, the optimal model during the entire training process.

[0075] (6) Training completed: After all training steps are completed, the training process ends.

[0076] III. Use the saved optimal model for the dangerous item detection task in the real subway station scenario. The specific approach is as follows:

[0077] (1) Preprocess the input image: Obtain the real-time video stream from the subway station surveillance camera and perform necessary preprocessing on the acquired image, including adjusting the image size to match the model input requirements (640x640 pixels), and converting the image format to the format acceptable to the model.

[0078] (2) Image normalization processing: Perform normalization processing on the preprocessed image, that is, scale the image pixel values from [0, 255] to the range of [0, 1], which is usually achieved by dividing each pixel value by 255. Normalization helps to improve the convergence speed and prediction performance of the model.

[0079] (3) Load the optimal model: Load the weights last.pt of the trained and saved optimal model to ensure that the model is fully initialized and ready for inference.

[0080] (4) Model inference: Input the normalized image into the SDO - YOLO network and perform forward propagation for object detection. The model will output the detection results, including the category, confidence score, and bounding box coordinates of each detected object.

[0081] (5) Result parsing and display: Parse the detection results output by the model, extract information such as bounding box coordinates, category names, and confidence scores. Overlay the detection results on the original image to visually display the detected dangerous items, usually including drawing bounding boxes and displaying category labels on the image.

[0082] The following are the performance evaluation results:

[0083] (1) Improvement in recognition accuracy:

[0084] The SDO - YOLO model of the present invention has achieved a significant improvement in recognition accuracy on the SDO dataset. Specifically, the average precision (AP) of our model on the SDO dataset has reached 42.1%, which is a 2.4% increase compared to the 39.7% AP of YOLOv8.

[0085] (2) Detection of targets at different scales:

[0086] In terms of the detection ability for targets of different sizes, SDO-YOLO also demonstrates superiority. In the detection of small targets, the model AP of the present invention has increased from 19.8% of YOLOv8 to 22.5%. In the detection of medium-sized targets, the detection AP has increased from 38.2% to 41.3%, while in the detection of large targets, the detection AP has increased from 58.9% to 61.2%.

[0087] (3) In the actual subway station monitoring scenario, the SDO-YOLO model successfully identified 96% of the dangerous item carriers in a simulation test, while the recognition rate of YOLOv8 was 92%. This indicates that the model of the present invention has higher accuracy in complex environments.

[0088] To more intuitively demonstrate the improvement effect, the following is the performance comparison table of SDO-YOLO and YOLOv8 on the SDO dataset:

[0089] Performance Comparison of SDO-YOLO and YOLOv8 on the SDO Dataset

[0090]

[0091]

[0092] As can be seen from the table, SDO-YOLO is superior to YOLOv8 in all key performance indicators. Especially in the detection of small and large targets, the model of the present invention demonstrates better performance.

[0093] In summary, the SDO-YOLO model of the present invention not only improves the recognition accuracy but also enhances the processing speed in the subway station dangerous item detection task. These experimental results and evaluation data fully prove the effectiveness of the present invention in improving the detection performance.

[0094] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dangerous goods detection method based on improved YOLOv8, characterized in that, Including the following steps: S1: Produce a dangerous goods detection dataset; S2: Construct an improved YOLOv8 network; S3: Use the detection dataset to train the improved YOLOv8 network and save the trained model parameters; S4: Obtain dangerous goods images in real scenarios and use the trained model parameters to detect dangerous goods in the dangerous goods images; Among them, the improved YOLOv8 network includes: a backbone network backbone, a neck network neck, and a head network head. The fasterNet Block structure is used to replace the Bottleneck computing unit. In the neck network neck, a Feature Alignment Module (FAM) and an Information Fusion Module (IFM) are used to achieve the aggregation of features at different levels, and the fused information is distributed back to each level of the network through an Information Injection Module (Inject).

2. The method for detecting dangerous goods based on the improved YOLOv8 according to claim 1, wherein The backbone network sorts the feature maps in ascending order of resolution and outputs four feature maps B5, B4, B3, and B2. These feature maps are processed through a low-order aggregation and distribution branch module and a high-order aggregation and distribution branch module. The low-order aggregation and distribution branch processes the low-resolution feature maps, and the high-order aggregation and distribution branch processes the high-resolution feature maps processed by the low-order aggregation and distribution branch. The processing results form three output feature maps N5, N4, and N3 and are passed to the head network head for final object detection.

3. The method for detecting dangerous goods based on improved YOLOv8 according to claim 2, wherein, The low-order aggregation and distribution branch module includes a low-order alignment module and a low-order information fusion module. The high-order aggregation and distribution branch module includes a high-order alignment module and a high-order information fusion module. The high-order alignment module and the low-order alignment module use the self-attention mechanism to align the feature maps at different levels. The low-order information fusion module and the high-order information fusion module are used to fuse the aligned feature maps. The fused features inject global information into each level through an injection module inject; The low-order aggregation and distribution branch module processes the four feature maps B5, B4, B3, and B2 to form processed feature maps P5, P4, and P3 corresponding to the B5, B4, and B3 feature maps respectively. Then, the feature maps P5, P4, and P3 are processed by the high-order aggregation and distribution branch to form three output feature maps N5, N4, and N3.

4. The hazardous material detection method based on improved YOLOv8 according to claim 3, characterized in that, The low-order alignment module downsamples the input feature map through an average pooling operation to unify the size. The low-order information fusion module contains multiple layers of reparameterized convolutional blocks and a splitting operation for fusing the aligned features; The high-order alignment module reduces the dimension of the input feature through average pooling. The high-order information fusion module consists of multiple stacked transformer blocks, which include a multi-head attention mechanism and a feed-forward network for obtaining global information.

5. The method for detecting dangerous goods based on the improved YOLOv8 according to claim 4, wherein, The injection module inject combines the fused feature map with the original feature map through a series of convolutional operations and a Sigmoid activation function and injects it back to each level of the network.

6. The hazardous material detection method based on improved YOLOv8 according to claim 1, characterized in that, The step S1 includes: The network obtains pictures and annotation information of dangerous goods to obtain a public dataset a0; Collect the monitoring video of the actual scene, extract the video frame images containing dangerous goods, and annotate the video frame images to obtain the private dataset a1; Fuse the public dataset a0 and the private dataset a1, and use them as the basis for Mosaic data augmentation to expand the data: adjust the images in the public dataset a0 and the private dataset a1 to the set size and correspondingly adjust the bounding box coordinates; randomly select four image regions, splice them into a Mosaic pattern, randomly select a region in the Mosaic image for cutting, and form a new sample containing different source image blocks to obtain the augmented dataset a2; Divide the augmented dataset a2 into a training set, a validation set, and a test set according to the ratio of 8:1:1, so as to obtain the dangerous goods detection dataset.

7. The hazardous material detection method based on improved YOLOv8 according to any one of claims 1-6, characterized in that, The detection method is used for detecting dangerous goods in the subway station.