Rail transit foreign matter intrusion detection method and device, electronic equipment and storage medium
Through the improved YOLOv8 network, including a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network, the problem of low accuracy of small-scale object detection in the prior art is solved, and high-accuracy foreign object detection in complex environments is achieved.
Patent Information
- Application Number
- CN202510437531.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing deep learning-based rail transit foreign body intrusion detection methods have low accuracy in small-scale target detection in complex environments.
The improved YOLOv8 network is adopted, including a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network, for detecting foreign objects in rail transit.
It improves the detection accuracy of small-scale objects, is suitable for complex application scenarios, and realizes high-accuracy foreign object detection in complex environments.
Smart Images

Figure CN119964091A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a rail transit foreign body intrusion detection method, device, electronic equipment and storage medium. Background Art
[0002] When mountain rocks, idle personnel, etc. invade the rail transit boundary, it will cause serious harm to the safety of passengers' lives and property and the safety of rail infrastructure. At the same time, these incidents are sudden and need to be discovered in time. Foreign objects that invade the rail transit boundary area vary in size and are at different distances from the camera. Foreign objects with unclear image representation and small pixel area are easy to disappear in a deep network. At present, the traditional manual inspection method has a large workload, low efficiency, and poor real-time performance; the existing rail transit foreign object intrusion detection method based on deep learning has a low accuracy rate for small-scale target detection in complex environments. Summary of the invention
[0003] The present invention provides a rail transit foreign object intrusion detection method, device, electronic device and storage medium, which are used to solve the defect of the rail transit foreign object intrusion detection method based on deep learning in the prior art, which has low accuracy in detecting small-scale targets in complex environments.
[0004] The present invention provides a rail transit foreign body intrusion detection method, comprising: Acquire an original monitoring image of a target track, and preprocess the original monitoring image to obtain a monitoring image of the target track; Inputting the monitoring image of the target track into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model; The foreign body detection model is obtained by training based on sample monitoring images of the sample track and foreign body detection result labels of the sample track; Among them, the foreign body detection model is improved based on the YOLOv8 network, and the foreign body detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
[0005] In some embodiments, the lightweight backbone network is obtained by lightweight improvement based on the backbone network of YOLOv8, and the lightweight backbone network includes FasterNet, and the FasterNet includes multiple FasterNet Blocks, each FasterNet Block includes a partial convolution PConv layer and two point-by-point convolution PWConv layers.
[0006] In some embodiments, the feature-focused diffusion pyramid network is obtained by improving the neck network of YOLOv8, and the feature-focused diffusion pyramid network includes a feature focusing module, and the feature focusing module includes a local feature extraction layer, a local feature focusing layer, a contextual feature extraction layer and a feature fusion layer.
[0007] In some embodiments, the local feature extraction layer includes a downsampling layer, a convolution layer and an upsampling convolution layer; the context feature extraction layer includes multiple depth-separable convolution DWConv layers, and the feature fusion layer is used to fuse the focused local features and context features.
[0008] In some embodiments, the dynamic detection head network is improved based on the head network of YOLOv8, and the dynamic detection head network includes a dynamic detection head module based on an attention mechanism, and the dynamic detection head module includes a scale perception module, a space perception module and a task perception module.
[0009] In some embodiments, inputting the monitoring image of the target track into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model includes: Inputting the monitoring image of the target track into the lightweight backbone network to obtain a multi-scale feature map of the target track output by the lightweight backbone network; Inputting the multi-scale feature map of the target track into the feature-focused diffusion pyramid network to obtain a multi-scale fusion feature map of the target track output by the feature-focused diffusion pyramid network; The multi-scale fusion feature map of the target track is input into the dynamic detection head network to obtain the foreign object detection result of the target track output by the dynamic detection head network.
[0010] In some embodiments, the training process of the foreign body detection model includes: Acquire an original sample monitoring image of a sample track, and preprocess the original sample monitoring image to obtain a sample monitoring image of the sample track; Determining a foreign body detection result label of the sample track; Inputting the sample monitoring image of the sample track into the initial foreign object detection model to obtain the foreign object detection prediction result of the sample track output by the initial foreign object detection model; Calculate the Focal-EIoU loss function value based on the foreign body detection result label of the sample track and the foreign body detection prediction result of the sample track; According to the Focal-EIoU loss function value, the parameters of the initial foreign object detection model are iteratively optimized to obtain the foreign object detection model.
[0011] The present invention also provides a rail transit foreign body intrusion detection device, comprising: An acquisition unit, used for acquiring an original monitoring image of a target track, and preprocessing the original monitoring image to obtain a monitoring image of the target track; A detection unit, used for inputting the monitoring image of the target track into a pre-built foreign object detection model, and obtaining a foreign object detection result of the target track output by the foreign object detection model; The foreign body detection model is obtained by training based on sample monitoring images of the sample track and foreign body detection result labels of the sample track; Among them, the foreign body detection model is improved based on the YOLOv8 network, and the foreign body detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the rail transit foreign object intrusion detection method as described above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the rail transit foreign object intrusion detection methods described above.
[0014] The rail transit foreign object intrusion detection method, device, electronic device and storage medium provided by the present invention obtain the original monitoring image of the target track, pre-process the original monitoring image, obtain the monitoring image of the target track, input the monitoring image of the target track into a pre-built foreign object detection model, and obtain the foreign object detection result of the target track output by the foreign object detection model, wherein the foreign object detection model is improved based on the YOLOv8 network, including a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network, has high detection accuracy, can realize accurate detection of small-scale objects, and is suitable for complex application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 It is a schematic diagram of the structure of the YOLOv8 network provided by the prior art; Figure 2 It is a schematic diagram of the flow of a rail transit foreign body intrusion detection method provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the structure of a lightweight backbone network provided by an embodiment of the present invention; Figure 4 is a schematic diagram of the structure of a partial convolutional layer provided by an embodiment of the present invention; Figure 5 is a schematic diagram of the structure of a feature-focused diffusion pyramid network provided by an embodiment of the present invention; Figure 6 is a schematic diagram of the structure of a feature focusing module provided by an embodiment of the present invention; Figure 7 is a schematic structural diagram of a rail transit foreign object intrusion detection device provided by an embodiment of the present invention; Figure 8 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] It should be noted that the foreign body detection model in the embodiment of the present invention is improved based on the YOLOv8 network; YOLOv8 is one of the YOLO (You Only Look Once) series target detection models, and has higher detection accuracy and computational efficiency than the previous YOLO series models.
[0019] Figure 1 The schematic diagram of the structure of the YOLOv8 network provided by the prior art. Figure 1 As shown in the figure, the YOLOv8 network includes an input layer (Input), a backbone network (Backbone), a neck network (Neck) and a head network (Head).
[0020] It should be noted that in the YOLOv8 network, the input layer sends the image to the backbone network for image feature extraction; the neck uses an improved Feature Pyramid Network-Path Aggregation Network (PAN-FAN) structure to further fuse and enhance the extracted features, and reduces the model complexity by removing the convolution structure in the upsampling stage and using the feature fusion module; the head uses a decoupled-head structure to separate the target classification and bounding box regression tasks, and adopts the Anchor-Free idea, without pre-setting the anchor box, directly predicting the target and bounding box at each position of the image to generate the final detection result. The YOLOv8 network mainly includes convolutional layers, fusion layers, and pooling layers. Among them, the convolution layer can be a Convolution-Batch Normalization-SiLU activation function (Conv-BatchNorm2d-SiLu, CBS) layer. The CBS layer is composed of a series of convolution layers, batch normalization layers and SiLu activation function layers. The combination of these layers can effectively extract different features in the image; the fusion layer can be a Channel to Feature Map (C2f) layer. The C2f layer is composed of two convolution layers connected by a bottleneck residual block in the middle. The addition of skip-layer connections and additional split operations can enrich feature information. Because of its simple structure, it can reduce the complexity of the model; the pooling layer can be a fast spatial pyramid pooling (Spatial Pyramid Pooling-Fast, SPPF) layer. The SPPF layer defines pooling kernels of different scales, performs a maximum pooling operation on the input feature map, splices feature maps of different scales together, and improves the detection capability of targets of different sizes.
[0021] Figure 2 The following is a flow chart of a method for detecting foreign body intrusion in rail transit provided by an embodiment of the present invention. Figure 2 As shown, a rail transit foreign body intrusion detection method is provided, comprising the following steps: step 210 and step 220. The method flow steps are only a possible implementation of the present invention.
[0022] Step 210: acquiring an original monitoring image of the target track, and preprocessing the original monitoring image to obtain a monitoring image of the target track; Step 220: input the monitoring image of the target track into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model; Among them, the foreign body detection model is trained based on the sample monitoring images of the sample track and the foreign body detection result labels of the sample track; Among them, the foreign object detection model is improved based on the YOLOv8 network. The foreign object detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
[0023] Optionally, the original monitoring image of the target track is acquired through a camera, a drone or a satellite.
[0024] Optionally, the original monitoring image is preprocessed by denoising, image enhancement, geometric correction, etc.
[0025] Figure 3 The schematic diagram of the structure of the lightweight backbone network provided by the embodiment of the present invention is as follows. Figure 3 As shown, in some embodiments, the lightweight backbone network is obtained by lightweight improvement based on the backbone network of YOLOv8, and the lightweight backbone network includes a fast network FasterNet, FasterNet includes multiple fast network modules FasterNet Block, and each FasterNet Block includes a partial convolution (Partial Convolution, PConv) layer and two point-wise convolution (Point-wise Convolution, PWConv) layers.
[0026] It should be noted that FasterNet is divided into four hierarchical stages, each of which is preceded by an embedding layer (with a stride of 4 and a size of 4×4 convolution) or a merging layer (with a stride of 2 and a size of 2×2 convolution) for spatial downsampling and channel number expansion. Each stage consists of a FasterNet Block. The FasterNet Block in the last two stages consumes less access memory and has a higher FLOPs value; therefore, more computing resources can be allocated to the last two stages. Each FasterNet Block consists of a PConv layer followed by 2 PWConv layers or 2 1×1 size Conv convolution layers, which are combined together as a residual module, in which the middle layer has an expanded number of channels, and a shortcut connection is set to reuse the input features. The last three layers are a global average pooling layer, a 1×1 size convolution layer, and a fully connected layer, which are used together for feature conversion and classification.
[0027] It should be noted that the convolutional layer is a key submodule of many network models. ,use c Filters To calculate the output , the conventional convolution FLOPs is ,in h is the height of the input image, wFor width, k Indicates the size of the convolution kernel.
[0028] Figure 4 The schematic diagram of the structure of the partial convolutional layer provided by the embodiment of the present invention. Figure 4 As shown, the partial convolution Pconv layer works as follows: Partial convolutional layers apply regular convolution to a portion of the input channels for spatial feature extraction, while keeping the remaining channels unchanged; for continuous or regular memory accesses, the first or last continuous The channel represents the entire feature map for calculation. It is generally believed that the input and output feature maps have the same number of channels; therefore, the FLOPs value of PConv is only: , when taking a typical ratio value When , the FLOPs of PConv is only 1 / 16 of that of regular convolution; and the memory access of PConv is also smaller. When , this value is only 1 / 4 of the regular convolution, and the size is .
[0029] Figure 5 A schematic diagram of the structure of a feature-focused diffusion pyramid network provided by an embodiment of the present invention. Figure 5 As shown, in some embodiments, the feature-focused diffusion pyramid network is obtained by improving the neck network based on YOLOv8, and the feature-focused diffusion pyramid network includes a feature-focusing module (Feature Fusion Module, FFM).
[0030] It should be noted that the C2f module in the middle layer of the neck network of YOLOv8 is replaced by FFM, and the output layer (such as P3, P4, and P5 layers) in the lightweight backbone network is connected to the first FFM, and the output of the FFM is diffused, and the splicing operation is performed before the fusion layer (such as C2f) before the first and third detection heads, and the feature information of different scales is transferred to the detection heads of different scales. The features with rich contextual information are diffused to the detection heads of each scale through the diffusion mechanism. By adopting FFM and feature diffusion mechanism, the features of each scale can have detailed contextual information, which is more conducive to the subsequent detection and classification of objects of different scales.
[0031] Figure 6 Schematic diagram of the structure of the feature focusing module provided in the embodiment of the present invention. Figure 6As shown, in some embodiments, the feature focusing module includes a local feature extraction layer, a local feature focusing layer, a context feature extraction layer and a feature fusion layer; the local feature extraction layer includes a downsampling layer, a convolution layer and an upsampling convolution layer; the context feature extraction layer includes multiple depthwise separable convolution DWConv layers, and the feature fusion layer is used to fuse the focused local features and context features.
[0032] It should be noted that the input of FFM is composed of the focused output layers (P3, P4, and P5 layers) in the lightweight backbone network; P3 is used for downsampling operations, and the ADown downsampling layer in YOLOv9 is introduced. This module optimizes the number of parameters in the convolution layer and reduces the computational complexity; the P4 layer is only used for convolution operations to maintain the dimension of the features; the P5 layer is used for upsampling and convolution operations; the processed P3, P4, and P5 layers are spliced to focus on the local information of each layer, and multiple parallel depth-separable convolution layers are used to capture contextual information at multiple scales.
[0033] It is understandable that FFM can comprehensively consider features at different scales, obtain rich information features across multiple scales, and enhance the network's ability to detect targets.
[0034] Alternatively, the mathematical expression of FFM is as follows: ; ; in, X 1 is the input feature of P3, X 2 is the input feature of P4, X 3 is the input feature of P5, L 1 Yes X 1 , X 2 , X 3 Focusing on local features extracted, P m By m indivual k m × k m The contextual features extracted by the depthwise separable convolutional DWConv layer, k m =( m +1)×2+1.
[0035] Optionally, a 1×1 convolutional layer is used to fuse local features and context features to characterize the relationship between different feature layers, and finally the residual operation is used to further enhance the input features, so that features of different levels and scales are fused and focused. The mathematical expression of FFM is as follows: ; ; in, Y 1 is a multi-scale fusion feature. Z 1 is the final output feature.
[0036] It can be understood that 1×1 convolution and residual operation are used as channel fusion mechanisms to fuse features with different receptive fields. By adopting FFM, contextual information of different receptive fields can be learned without affecting the integrity of local texture features.
[0037] In some embodiments, the dynamic detection head network is improved based on the head network of YOLOv8, and the dynamic detection head network includes a dynamic detection head module based on an attention mechanism, and the dynamic detection head module includes a scale perception module, a space perception module and a task perception module.
[0038] It should be noted that the difference in target scale is related to the features of different levels; the geometric transformation of different target shapes is related to the features of different spatial positions; different target representations and tasks are related to the features of different channels; in order to better integrate the features of targets of different sizes, the dynamic detection head (Dynamic head, Dyhead) is introduced into the head network of YOLOv8. It is a dynamic target detection head based on the attention mechanism and integrates a scale perception module. π L , spatial perception module π S and task awareness module π C , its mathematical expression is: ; in, W For attention, π(∙) is the attention function, F is the input L×S×C three-dimensional tensor, L is the level of the feature map, S is the product of the width and height of the feature map, C is the number of channels of the feature map.
[0039] Among them, the scale perception module π L, which is used to dynamically fuse features of different scales based on semantic importance. Its attention function is: ; in, f It is a linear function approximated by a 1×1 convolutional layer. It is the hard-sigmoid function.
[0040] Among them, the spatial perception module based on fusion features π S Focus on discriminative regions that consistently coexist between spatial locations and feature hierarchies; consider S The high dimension of the module is decomposed into two steps: first, the attention learning is sparse by using deformable convolution, and then the features of different layers are aggregated at the same spatial position. The mathematical expression is: ; in, is the number of sparsely sampled locations, k Indicates k sparsely sampled locations, l Represents the feature map l layer, w l,k Represents the feature map l Layer k The attention of sparsely sampled positions, is the position after the move, is the self-learning space offset, It's location The self-learning importance scalar at .
[0041] Among them, the task perception module π C Different tasks are supported by dynamically switching the on / off state of the feature channel. The mathematical expression is as follows: ; in, F c It is c feature slices of channels, is a hyperparameter used to control the activation threshold.
[0042] Optionally, first L × S A global average pooling is performed on the dimension to reduce the dimensionality, followed by two fully connected layers and a normalization layer, and finally a shifted sigmoid function is applied to normalize the output to [-1, 1].
[0043] Optionally, you can nest scale-aware modules multiple times.π L , spatial perception module π S and task awareness module π C .
[0044] In some embodiments, step 220, inputting the monitoring image of the target track into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model, includes: Step 221: input the monitoring image of the target track into the lightweight backbone network to obtain a multi-scale feature map of the target track output by the lightweight backbone network; Step 222: input the multi-scale feature map of the target track into the feature-focused diffusion pyramid network to obtain a multi-scale fusion feature map of the target track output by the feature-focused diffusion pyramid network; Step 223: input the multi-scale fusion feature map of the target track into the dynamic detection head network to obtain the foreign body detection result of the target track output by the dynamic detection head network.
[0045] It can be understood that the embodiment of the present invention improves the YOLOv8 benchmark model through a lightweight backbone network structure, maintaining similar detection accuracy while reducing computational consumption; proposes a feature-focused diffusion pyramid network to improve the neck network structure of the benchmark model, strengthen the mutual exchange of features at different levels, and enhance the ability to recognize targets of different scales; and improves the basic model through a dynamic detection head to address the difficulties in target recognition of different scales, spaces, and tasks, thereby improving the loss of fine-grained feature information of targets in deep networks.
[0046] In an embodiment of the present invention, an original monitoring image of a target track is acquired and preprocessed to obtain a monitoring image of the target track. The monitoring image of the target track is input into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model. The foreign object detection model is improved based on the YOLOv8 network, including a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network. The model has high detection accuracy, can realize accurate detection of small-scale objects, and is suitable for complex application scenarios.
[0047] In some embodiments, the training process of the foreign body detection model includes: Acquire an original sample monitoring image of the sample track, and preprocess the original sample monitoring image to obtain a sample monitoring image of the sample track; Determine the foreign body detection result label of the sample track; Inputting the sample monitoring image of the sample track into the initial foreign object detection model to obtain the foreign object detection prediction result of the sample track output by the initial foreign object detection model; Based on the foreign body detection result label of the sample track and the foreign body detection prediction result of the sample track, calculate the Focal-EIoU loss function value; According to the Focal-EIoU loss function value, the parameters of the initial foreign body detection model are iteratively optimized to obtain the foreign body detection model.
[0048] Optionally, the initial foreign body detection model includes an initial lightweight backbone network, an initial feature-focused diffusion pyramid network and an initial dynamic detection head network, and the initial foreign body detection model is improved based on the YOLOv8 network.
[0049] Optionally, the sample monitoring image of the sample track is input into the initial foreign object detection model to obtain the foreign object detection prediction result of the sample track output by the initial foreign object detection model, including: Inputting the sample monitoring image of the sample track into the initial lightweight backbone network, obtaining a sample multi-scale feature map of the sample track output by the initial lightweight backbone network; Inputting the sample multi-scale feature map of the sample track into the initial feature focused diffusion pyramid network, obtaining the sample multi-scale fusion feature map of the sample track output by the initial feature focused diffusion pyramid network; The sample multi-scale fusion feature map of the sample track is input into the initial to dynamic detection head network to obtain the foreign object detection prediction result of the sample track output by the initial dynamic detection head network.
[0050] Among them, the calculation formula of Focal-EIoU loss function is as follows: ; ; ; In the formula, L Focal-EIoU is the Focal-EIoU loss, L EIoU is the EIoU loss, L IoU is the IoU loss, L Dis is the distance loss, L Asp is the direction loss, ρ () is the Euclidean distance between two points, b , ω , h are the parameters of the predicted bounding box, b gt , ω gt , h gt are the parameters of the true bounding box, w c and h c To predict the width and height of the minimum bounding rectangle of the bounding box and the true bounding box, the parameters α , β and γ It is a hyperparameter that controls the degree of suppression and is used to balance the impact of the gradient on the model. C is a constant value that keeps the function continuous.
[0051] It should be noted that in order to solve the problem that the length and width cannot increase or decrease synchronously, the loss function EIoU cancels the calculation of the aspect ratio of the bounding box and replaces it with the regression of the length and width values of the bounding box. At present, the types of foreign objects that invade the track limit are diverse and the difficulty levels vary. There is an imbalance problem between positive and negative samples, which leads to a decrease in the generalization ability of the model. The Focal Loss loss function balances the weights of samples of different categories and different qualities, thereby enhancing the focus on high-quality samples and improving the learning ability of difficult samples.
[0052] It should be noted that according to the deployment characteristics of surveillance cameras along the track, their monitoring range is generally within the railway limit area. Based on this, the collection content and method of the simulation data set of dangerous situations along the track are designed. The up and down rail areas in the track limit are divided into 1 track, between-track and 2 track areas. During the day and night under different weather conditions, three track surface areas of foreign objects and personnel intrusion are simulated respectively, and simulated collection of dangerous situation samples is carried out at different distances from the camera.
[0053] Table 1 is a statistical table of the foreign body sample collection scheme provided by an embodiment of the present invention. As shown in Table 1, dangerous foreign body samples include various types such as fallen rocks, branches, woven bags, and personnel, which are collected at different distances under different weather and time conditions.
[0054] Table 1 Statistics of foreign body sample collection plan
[0055] It should be noted that the simulation of track danger was conducted under different cameras along the track, combined with different time and weather conditions, and recorded to produce sample monitoring images. Finally, 36,000 sample monitoring images of sample tracks were annotated.
[0056] Table 2 is a data statistics table of sample monitoring images of sample tracks provided by an embodiment of the present invention. As shown in Table 2, the targets of the sample monitoring images are counted according to different sizes, large targets refer to targets with a size greater than 96 pixels × 96 pixels, medium targets refer to targets with a size between 32 pixels and 96 pixels, and small targets refer to targets with a size less than 32 pixels × 32 pixels.
[0057] Table 2 Statistics of sample monitoring images of sample tracks
[0058] Optionally, the collected, processed and labeled orbital hazard sample data are used as a data set and divided into a training set and a validation set in a ratio of 8:2; in the training process of the foreign object detection model, data enhancement technologies such as Mosaic and Mixup are used, the number of iterations is set to 300 rounds, the batchsize is 16, the number of working threads is 8, and the optimizer is selected as SGD.
[0059] Optionally, the detection accuracy (mean Average Precision, mAP) indicator commonly used in target detection algorithms is selected as the performance evaluation indicator of the foreign body detection model. This indicator is determined by the detection accuracy (Precision, P) and recall (Recall, R); at the same time, in practical applications, it is necessary to consider the detection speed (Frames per second, FPS) and the number of model parameters.
[0060] Optionally, according to the relationship between the model detection result and Ground-Truth, the model detection result is divided into true positive (TP), false positive (FP), false negative (FN), and true negative (TN), and the accuracy and recall of the model are defined accordingly. The accuracy is the ratio of true positives to detected intrusion cases, which characterizes the precision of the model. The recall is the ratio of true positives to all true intrusion cases, which characterizes the recall of the model. The calculation method is as follows: ; .
[0061] It should be noted that it is difficult to comprehensively evaluate the quality of the model based on the accuracy and recall rate alone. A high accuracy rate does not mean that as many true positive examples as possible can be detected, and a high recall rate may also mean that too many negative examples are detected as positive examples. Therefore, only by combining the two can the detection ability of the model be better evaluated. The accuracy P and recall rate R of each category are plotted into a PR curve, and the line is integrated to obtain the detection accuracy AP indicator.
[0062] Optionally, in order to verify the impact of the improved module FasterNet on the model, it replaces the YOLOv8-m backbone network and is compared with the native YOLOv8-m, YOLOv5-s, YOLOv7-tiny, and YOLOv8-s models. The comparative experimental results are shown in Table 3.
[0063] Table 3 shows the impact of the backbone network improvement module on the model. As shown in Table 3, the lighter s and tiny series models have lower recall, accuracy, mAP and other indicators than the m model, but higher FPS values; among them, YOLOv7-tiny achieved the highest FPS value of 289 with the lightest model size, but its mAP value was also the lowest value of 75.4; the mAP value of the largest model YOLOv8-m also achieved the highest value, and the mAP value of the improved model of the embodiment of the present invention ranked second, 1.1% lower than the highest value YOLOv8-m, and the accuracy was relatively close. Compared with the YOLOv8-s with a similar model size, the improved model of the embodiment of the present invention is 5.8% higher in accuracy, and the FPS value is 9% lower than the highest value YOLOv7-tiny, and the FPS value is 3 frames lower than the YOLOv8-s with a similar model size. It can be seen that the improved model of the embodiment of the present invention has achieved the best mAP value at the same model size. Among the models with similar mAP values, the algorithm has the highest FPS value. Therefore, the improved model of the embodiment of the present invention has a good effect in balancing the processing rate and recognition accuracy.
[0064] Table 3 The impact of backbone network improvement modules on the model
[0065] Optionally, in order to verify the influence of the feature-focused diffusion pyramid network module on the model recognition performance, the module is used to carry out an improved comparative experiment on the neck structure in the YOLOv8-m network. The experimental results are shown in Table 4.
[0066] Table 4 shows the influence of the feature focused diffusion pyramid network module on the model. As shown in Table 4, AP S ,AP M ,AP LThey represent the AP values of small, medium and large intrusion targets respectively. It can be seen that the AP of the optimized model is increased by 2.6%, and the recognition index AP for large, medium and small intrusion targets is increased by 0.5%, 1.3% and 6%. The enhancement of the index is gradually improved, and the improvement of small target recognition is the most significant. It can be seen that in the native YOLOv8 model, for small-scale targets, as the number of network layers increases, the fine-grained feature information of small scales will gradually disappear. Through the feature focusing module, the feature information of small and medium scales can be strengthened, and the diffusion operation can be continued in the neck structure to transfer the feature information to the deep network, thereby improving the recognition ability of small and medium-scale targets. Therefore, the fine-grained features of small and medium-sized targets of the model are retained by using the feature focusing diffusion pyramid network module, which can effectively improve the model detection capability.
[0067] Table 4 The impact of feature-focused diffusion pyramid network module on the model
[0068] Optionally, in order to verify the impact of the dynamic detection head module on the model recognition performance, the module is used to perform an improved comparative experiment on the detection head structure in the YOLOv8-m network. The experimental results are shown in Table 5.
[0069] Table 5 shows the impact of the dynamic detection head module on the model. As shown in Table 5, the AP of the optimized model has increased by 1.1%, and the recognition indexes of large, medium and small intrusion targets have increased by 0.4%, 1% and 1.9%, respectively. The recognition accuracy index will gradually improve according to the difficulty of the intrusion target, among which the recognition of small targets has improved significantly. The three improved structures of scale perception module, space perception module and task perception module strengthen the characteristics of the target at the detection head, better integrate the characteristics of different scales, and make it easier to distinguish the characteristics of targets of different scales and spatial positions during detection, thus improving the overall foreign body recognition capability.
[0070] Table 5 The impact of dynamic detection head module on the model
[0071] Optionally, in order to verify the impact of the Focal-EIoU loss function on the model recognition performance, the training process of the YOLOv8-m network was improved using this module, and a comparative experiment was conducted. The experimental results are shown in Table 6.
[0072] Table 6 shows the impact of the Focal-EIoU loss function on the model. As shown in Table 6, AP Person ,AP Stone ,AP Bough ,AP Package ,AP OtherThey represent the AP values of intrusion targets such as personnel, fallen rocks, branches, packages, and other intrusion targets respectively. It can be seen that the AP of the optimized model for different types of intrusion targets has been improved. Among them, personnel, fallen rocks, and packages have relatively uniform overall features and small differences in intra-class features, so the overall AP of the optimized model is relatively small, which are 0.3%, 0.5%, and 0.2% respectively. The native YOLOv8-m already has good recognition capabilities, but for intrusion targets such as branches and other (color-coated steel plates, light objects, etc.) with more complex intra-class feature differences, the AP of the optimized model has been greatly improved, which are 2% and 3.9% respectively. By setting different weights for samples of different categories and qualities through the loss function, the model's attention to such samples can be effectively enhanced, the learning ability of difficult samples can be improved, and the ability to recognize foreign objects can be improved.
[0073] Table 6 Impact of Focal-EIoU loss function on the model
[0074] It should be noted that in order to verify the influence of different improvement methods in the present invention on the model recognition performance, multiple groups of experiments were conducted and the improved modules were analyzed. All experiments set the same training parameters and were trained on the same data set. The detection results of different models are shown in Table 7.
[0075] Table 7 is the ablation experiment results. As shown in Table 7, "√" represents that the improved method of the embodiment of the present invention is used in the YOLOv8-m model, and "×" represents that the improved method is not used. It can be seen from Table 7 that on the data set of the embodiment of the present invention, the improved method of the embodiment of the present invention has achieved improvement in the detection accuracy index. Compared with the native model, the AP index is finally improved by 3.7%; among them, Improvement 1 reduces the size of the model, greatly improves the FPS value of the model, and keeps the AP value of the model close to that of the native YOLOv8-m. With the continuous superposition of the improved methods, the AP index has been continuously improved; Improvement 2 improves the AP value by 2.4% compared with the native model, especially the small target AP SThe improvement is the largest, reaching 5.9%, and the improvements for medium and large targets are 0.9% and 0.4% respectively. It can be seen that the feature focusing module can improve the model's detection performance for targets of different scales; Improvement 3 improves the AP value by 3.1% compared with the native model, and increases the AP value by 0.5% on the basis of Improvement 2. The improvements for small, medium and large targets are relatively balanced compared with Improvement 2, which are 0.9%, 0.8% and 0.4% respectively. It can be seen that the dynamic detection head can balance the features of targets of different scales in the final detection stage; Improvement 4 is the concentrated form of all improvements, with an AP value improvement of 3.6% compared with the native model, among which the maximum improvement for small targets is 7.4%, and the medium and large targets are improved by 2.3% and 1.1% respectively. All the improvement methods of the embodiments of the present invention can effectively improve the algorithm detection capability in different situations, so that the algorithm has better practical application effect.
[0076] Table 7 Ablation experiment results
[0077] Optionally, in order to further verify the effectiveness of the foreign body detection model provided by the embodiment of the present invention, mainstream detection models in the field of target detection are selected, such as Faster RCNN in the two-stage method, YOLOv5-m and YOLOv8-m in the one-stage detection method, and RT-DETR based on the Transformer architecture. These detection models are compared with the improved model proposed in the present invention. The experimental results are shown in Table 8.
[0078] Table 8 is a comparison table of the experimental results of this model and other models. As shown in Table 8, it can be seen that the AP value of the model of the present invention is greatly improved compared with other models. Its AP is 25.5%, 7%, 3.7% and 3.2% higher than Faster RCNN, YOLOv5-m, YOLOv8-m and RT-DETR, which proves that the overall detection capability of the foreign object detection model provided by the embodiment of the present invention is strong, especially for small-scale targets. The AP value is the only one among all methods that reaches more than 90%. It can be seen that in the case of targets of different sizes, the method of the embodiment of the present invention can better retain the detailed feature information of different scales, so that the model can detect intrusion targets of various shapes in a complex orbital environment.
[0079] Table 8 Comparison of experimental results of this model and other models
[0080] The rail transit foreign object intrusion detection device provided by an embodiment of the present invention is described below. The rail transit foreign object intrusion detection device described below and the rail transit foreign object intrusion detection method described above can be referenced to each other.
[0081] Figure 7A schematic diagram of the structure of a rail transit foreign body intrusion detection device provided by an embodiment of the present invention is shown in FIG. Figure 7 As shown, the rail transit foreign body intrusion detection device 700 includes: An acquisition unit 710 is used to acquire an original monitoring image of the target track, and pre-process the original monitoring image to obtain a monitoring image of the target track; The detection unit 720 is used to input the monitoring image of the target track into a pre-built foreign object detection model to obtain the foreign object detection result of the target track output by the foreign object detection model; Among them, the foreign body detection model is trained based on the sample monitoring images of the sample track and the foreign body detection result labels of the sample track; Among them, the foreign object detection model is improved based on the YOLOv8 network. The foreign object detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
[0082] Optionally, the lightweight backbone network is obtained by performing lightweight improvement on the backbone network based on YOLOv8, the lightweight backbone network includes FasterNet, FasterNet includes multiple FasterNet Blocks, and each FasterNet Block includes a partial convolution PConv layer and two point-by-point convolution PWConv layers.
[0083] Optionally, the feature-focused diffusion pyramid network is improved based on the neck network of YOLOv8, and the feature-focused diffusion pyramid network includes a feature focusing module, and the feature focusing module includes a local feature extraction layer, a local feature focusing layer, a context feature extraction layer and a feature fusion layer.
[0084] Optionally, the local feature extraction layer includes a downsampling layer, a convolution layer and an upsampling convolution layer; the context feature extraction layer includes multiple depth-separable convolution DWConv layers, and the feature fusion layer is used to fuse the focused local features and context features.
[0085] Optionally, the dynamic detection head network is improved based on the head network of YOLOv8, and the dynamic detection head network includes a dynamic detection head module based on an attention mechanism, and the dynamic detection head module includes a scale perception module, a space perception module and a task perception module.
[0086] Optionally, the monitoring image of the target track is input into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model, including: The monitoring image of the target track is input into the lightweight backbone network to obtain a multi-scale feature map of the target track output by the lightweight backbone network; Input the multi-scale feature map of the target track into the feature-focused diffusion pyramid network, and obtain the multi-scale fusion feature map of the target track output by the feature-focused diffusion pyramid network; The multi-scale fusion feature map of the target track is input into the dynamic detection head network, and the foreign object detection result of the target track is obtained as output by the dynamic detection head network.
[0087] Optionally, the training process of the foreign object detection model includes: Acquire an original sample monitoring image of the sample track, and preprocess the original sample monitoring image to obtain a sample monitoring image of the sample track; Determine the foreign body detection result label of the sample track; Inputting the sample monitoring image of the sample track into the initial foreign object detection model to obtain the foreign object detection prediction result of the sample track output by the initial foreign object detection model; Based on the foreign body detection result label of the sample track and the foreign body detection prediction result of the sample track, calculate the Focal-EIoU loss function value; According to the Focal-EIoU loss function value, the parameters of the initial foreign body detection model are iteratively optimized to obtain the foreign body detection model.
[0088] It should be noted here that the rail transit foreign object intrusion detection device provided in the embodiment of the present invention can implement all the method steps implemented in the above-mentioned rail transit foreign object intrusion detection method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0089] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the rail transit foreign body intrusion detection method, the method comprising: obtaining the original monitoring image of the target track, preprocessing the original monitoring image, and obtaining the monitoring image of the target track; inputting the monitoring image of the target track into a pre-built foreign body detection model, and obtaining the foreign body detection result of the target track output by the foreign body detection model; wherein the foreign body detection model is obtained by training based on the sample monitoring image of the sample track and the foreign body detection result label of the sample track; wherein the foreign body detection model is obtained by improving the YOLOv8 network, and the foreign body detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
[0090] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0091] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the rail transit foreign object intrusion detection method provided by the above-mentioned methods, the method comprising: acquiring an original monitoring image of a target track, preprocessing the original monitoring image, and obtaining a monitoring image of the target track; inputting the monitoring image of the target track into a pre-built foreign object detection model, and obtaining a foreign object detection result of the target track output by the foreign object detection model; wherein the foreign object detection model is trained based on a sample monitoring image of a sample track and a foreign object detection result label of the sample track; wherein the foreign object detection model is improved based on a YOLOv8 network, and the foreign object detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network, and a dynamic detection head network.
[0092] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A rail transit foreign body intrusion detection method, characterized in that: include: Acquire an original monitoring image of a target track, and preprocess the original monitoring image to obtain a monitoring image of the target track; Inputting the monitoring image of the target track into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model; The foreign body detection model is obtained by training based on sample monitoring images of the sample track and foreign body detection result labels of the sample track; Among them, the foreign body detection model is improved based on the YOLOv8 network, and the foreign body detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
2. The rail transit foreign body intrusion detection method according to claim 1, characterized in that: The lightweight backbone network is obtained by performing lightweight improvement on the backbone network based on YOLOv8. The lightweight backbone network includes FasterNet. The FasterNet includes multiple FasterNet Blocks. Each FasterNet Block includes a partial convolution PConv layer and two point-by-point convolution PWConv layers.
3. The rail transit foreign body intrusion detection method according to claim 1, characterized in that: The feature-focused diffusion pyramid network is obtained by improving the neck network of YOLOv8. The feature-focused diffusion pyramid network includes a feature focusing module, and the feature focusing module includes a local feature extraction layer, a local feature focusing layer, a context feature extraction layer and a feature fusion layer.
4. The rail transit foreign body intrusion detection method according to claim 3, characterized in that: The local feature extraction layer includes a downsampling layer, a convolution layer and an upsampling convolution layer; the context feature extraction layer includes a plurality of depth-separable convolution DWConv layers, and the feature fusion layer is used to fuse the focused local features and context features.
5. The rail transit foreign body intrusion detection method according to claim 1, characterized in that: The dynamic detection head network is obtained by improving the head network based on YOLOv8. The dynamic detection head network includes a dynamic detection head module based on an attention mechanism. The dynamic detection head module includes a scale perception module, a space perception module and a task perception module.
6. The rail transit foreign body intrusion detection method according to any one of claims 2 to 5, characterized in that: The step of inputting the monitoring image of the target track into a pre-built foreign object detection model to obtain a foreign object detection result of the target track output by the foreign object detection model includes: Inputting the monitoring image of the target track into the lightweight backbone network to obtain a multi-scale feature map of the target track output by the lightweight backbone network; Inputting the multi-scale feature map of the target track into the feature-focused diffusion pyramid network to obtain a multi-scale fusion feature map of the target track output by the feature-focused diffusion pyramid network; The multi-scale fusion feature map of the target track is input into the dynamic detection head network to obtain the foreign object detection result of the target track output by the dynamic detection head network.
7. The rail transit foreign body intrusion detection method according to claim 1, characterized in that: The training process of the foreign body detection model includes: Acquire an original sample monitoring image of a sample track, and preprocess the original sample monitoring image to obtain a sample monitoring image of the sample track; Determining a foreign body detection result label of the sample track; Inputting the sample monitoring image of the sample track into the initial foreign object detection model to obtain the foreign object detection prediction result of the sample track output by the initial foreign object detection model; Calculate the Focal-EIoU loss function value based on the foreign body detection result label of the sample track and the foreign body detection prediction result of the sample track; According to the Focal-EIoU loss function value, the parameters of the initial foreign object detection model are iteratively optimized to obtain the foreign object detection model.
8. A rail transit foreign body intrusion detection device, characterized in that: include: An acquisition unit, used for acquiring an original monitoring image of a target track, and preprocessing the original monitoring image to obtain a monitoring image of the target track; A detection unit, used for inputting the monitoring image of the target track into a pre-built foreign object detection model, and obtaining a foreign object detection result of the target track output by the foreign object detection model; The foreign body detection model is obtained by training based on sample monitoring images of the sample track and foreign body detection result labels of the sample track; Among them, the foreign body detection model is improved based on the YOLOv8 network, and the foreign body detection model includes a lightweight backbone network, a feature-focused diffusion pyramid network and a dynamic detection head network.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the rail transit foreign object intrusion detection method as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the rail transit foreign object intrusion detection method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Small-diameter pipe welding defect detection method based on strip shape perception and feature adjustment
CN117745679A
Train track obstacle detection method based on improved YOLOv8
CN118629002A
Abnormal behavior detection method and device based on improved YOLOv8 and storage medium
CN119152344A
Track foreign matter intrusion detection method based on improved YOLOv8 model
CN119339314A
Method and system for classifying images
GB202401994D0
Cited By
Target strike method and device based on unmanned aerial vehicle and computer equipment
CN120578195A
Train sensing system and method
CN121493051A