Escalator passenger abnormal behavior detection method and system based on YOLOv8
By making multi-scale improvements and lightweight improvements to the YOLOv8 algorithm, introducing attention modules and optimized loss functions, the problems of insufficient detection accuracy and insufficient deployment speed in the detection of abnormal behavior of escalator passengers is solved, and fast and accurate abnormal behavior recognition and real-time detection on edge computing devices are achieved.
Patent Information
- Application Number
- CN202510055809.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
The traditional YOLOv8 algorithm lacks attention to different scales and important features in the detection of abnormal behavior of escalator passengers, resulting in insufficient detection accuracy, high rates of missed and false detection, and insufficient speed when deployed on resource-constrained edge computing devices.
By making multi-scale improvements and lightweight improvements to the YOLOv8 algorithm, an attention module is introduced to enhance feature perception capabilities and optimize the loss function to improve detection accuracy and speed. The specific methods include using ShuffleNetV2 as the backbone network, integrating the ECA attention module and the C2f_DSConv module, and optimizing the loss function to WIoU.
It realizes rapid and accurate identification of abnormal behaviors of escalator passengers, reduces the rate of missed detection and false detection, improves the robustness and adaptability of detection methods, and facilitates real-time detection on edge computing devices.
Smart Images

Figure CN119942648A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, computer vision and target detection technology, and specifically, is a method and system for detecting abnormal behavior of escalator passengers based on YOLOv8. Background Art
[0002] Escalators are widely used in crowded places such as shopping malls, railway stations, and subway stations, bringing great convenience to passengers. However, due to passengers' irregular escalator riding, lack of safety awareness and other factors, safety accidents often occur. Therefore, by detecting abnormal behavior of escalator passengers, escalator safety accidents can be effectively prevented, which is conducive to ensuring passenger safety. With the rapid development and widespread application of deep learning in recent years, as well as the continuous improvement of the performance of various embedded devices and edge computing devices, more solutions have been provided to reduce or avoid the occurrence of escalator personal safety accidents.
[0003] Currently, common detection methods for abnormal behavior of escalator passengers include image processing-based and deep learning-based methods. The image processing-based method mainly relies on the designed feature extraction algorithm and classifier, but the effect is limited in complex scenarios, and there is a lot of room for performance improvement. The deep learning-based method relies on the extensive use of convolutional neural networks (CNNs), among which the YOLO series algorithm is a very common target detection algorithm, and the average performance in all aspects is relatively stable. Before the traditional YOLOv8 algorithm introduced the attention module, it lacked attention to different scales and important features. In the detection of abnormal behavior of escalator passengers, the characteristics of passengers in different postures are crucial to classification and positioning, and special attention needs to be paid to improve the accuracy of detection, reduce the missed detection and false detection rates, and improve the robustness of the detection method. In addition, in order to improve the detection speed of the algorithm, it is necessary to make lightweight improvements to the algorithm model so that it can be more appropriately deployed on resource-constrained edge computing devices. Summary of the invention
[0004] In order to solve the above technical problems, the present invention proposes a method and system for detecting abnormal behavior of escalator passengers based on YOLOv8. The YOLOv8 algorithm is used as the basic framework, and its network structure is improved at multiple scales. The detection speed is improved and the deployment of real-time detection is more convenient by improving the algorithm structure lightweight, and the algorithm accuracy is reduced due to the reduction of lightweight parameters by introducing an attention module. In this way, the information of passengers with abnormal behavior can be effectively and accurately identified, and it is convenient to deploy edge computing equipment for real-time detection.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] On the one hand, the present invention provides a method for detecting abnormal behavior of escalator passengers based on YOLOv8:
[0007] Step 1: Collect and preprocess the abnormal behavior data of escalator passengers, annotate the preprocessed image data, and construct an escalator passenger abnormal behavior detection image dataset;
[0008] Step 2: Using YOLOv8 as the baseline algorithm, introduce the lightweight feature extraction network ShuffleNetV2 and the designed C2f_DSConv module, integrate the ECA module and optimize the loss function to build the YOLOv8_SDE passenger abnormal behavior detection model;
[0009] Step 3: Use the escalator passenger abnormal behavior detection image dataset to train the YOLOv8_SDE passenger abnormal behavior detection model to obtain an optimized detection model;
[0010] Step 4: Use evaluation indicators to evaluate the performance of the optimized detection model, and deploy the optimized detection model that has passed the evaluation on the edge computing device to detect abnormal behavior of escalator passengers.
[0011] Furthermore, step 1 includes collecting abnormal behavior data of escalator passengers captured in real time and obtained through the network, performing data enhancement processing, marking the behavior postures of escalator pedestrians, including standing, sitting, bending, leaning, and falling, and constructing an escalator passenger abnormal behavior detection image dataset.
[0012] Furthermore, the construction of the YOLOv8_SDE passenger abnormal behavior detection model in step 2 includes:
[0013] The lightweight feature extraction network ShuffleNetV2 is used as the backbone network of YOLOv8 to extract features from images.
[0014] The neck-trunk network uses PAN-FPN to perform multi-scale feature fusion on the features extracted by the backbone network;
[0015] The detection head part includes a detection head and a classification head, which perform target detection and classification on the fused multi-scale features.
[0016] Furthermore, the neck-trunk network includes a C2f_DSConv module and a fused ECA attention module.
[0017] Furthermore, each C2f_DSConv module is connected to the fused ECA attention module and the loss function is optimized to WIoU.
[0018] Furthermore, the C2f_DSConv module is a new module formed by replacing the standard convolution block in C2f with a distribution shift convolution block.
[0019] Furthermore, the fused ECA attention module is a channel attention mechanism deep convolutional neural network, which uses a 1D convolutional layer to fuse information between channels.
[0020] Furthermore, the evaluation indicators in step 4 include precision, recall, accuracy, average accuracy value of all categories, model parameter quantity, and model weight size.
[0021] Furthermore, the optimized detection model is deployed on an edge computing device using a Rockchip RK3588 development board for inference.
[0022] On the other hand, the present invention also provides an escalator passenger abnormal behavior detection system based on YOLOv8, the system comprising:
[0023] A data set acquisition unit is used to collect and preprocess the abnormal behavior data of escalator passengers, annotate the preprocessed image data, and construct an escalator passenger abnormal behavior detection image data set;
[0024] The detection model building unit is used to use YOLOv8 as the baseline algorithm, introduce the lightweight feature extraction network ShuffleNetV2 and the designed C2f_DSConv module, integrate the ECA module and optimize the loss function to build the YOLOv8_SDE passenger abnormal behavior detection model;
[0025] An optimization model acquisition unit is used to train the YOLOv8_SDE passenger abnormal behavior detection model using an escalator passenger abnormal behavior detection image dataset to obtain an optimized detection model;
[0026] The model evaluation unit is used to evaluate the performance of the optimized detection model using evaluation indicators, and deploy the optimized detection model that has passed the evaluation on the edge computing device to detect abnormal behavior of escalator passengers.
[0027] The beneficial effects of the present invention are:
[0028] The present invention can quickly and accurately identify abnormal behaviors of passengers on escalators, and is convenient to be deployed on resource-constrained edge computing devices; the ECA module can enhance the model's ability to perceive the characteristics of passenger behavior, so that the model is further lightweight, with fewer parameters and improved efficiency, and can also automatically learn important features at different spatial positions; the loss function is optimized to WIoU, and the IoU is weighted, which solves the problem of border loss caused by deviation in the traditional IoU part; the data enhancement module can increase the diversity of data by performing various random transformations and expansions on the training data, and by randomly changing the size, rotation angle, brightness, etc. of the image during the training process, the model can have better robustness and generalization ability, helping the YOLOv8 model to better adapt to different passenger abnormal behavior detection scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the structure of the improved YOLOv8 model of the present invention;
[0030] Figure 2 It is a schematic diagram of the structure of the C2f_DSConv module of the present invention;
[0031] Figure 3 This is a schematic diagram of the ECA module structure of the present invention;
[0032] Figure 4 The present invention is a flowchart of an implementation method of a real-time detection method of abnormal behavior of escalator passengers based on improved YOLOv8. DETAILED DESCRIPTION
[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] The present invention discloses a method for detecting abnormal behavior of escalator passengers based on YOLOv8, the method comprising:
[0035] Step 1: Collect and preprocess the abnormal behavior data of escalator passengers, annotate the preprocessed image data, and construct an escalator passenger abnormal behavior detection image dataset;
[0036] Step 2: Using YOLOv8 as the baseline algorithm, introduce the lightweight feature extraction network ShuffleNetV2 and the designed C2f_DSConv module, integrate the ECA module and optimize the loss function to build the YOLOv8_SDE passenger abnormal behavior detection model;
[0037] Step 3: Use the escalator passenger abnormal behavior detection image dataset to train the YOLOv8_SDE passenger abnormal behavior detection model to obtain an optimized detection model;
[0038] Step 4: Use evaluation indicators to evaluate the performance of the optimized detection model, and deploy the optimized detection model that has passed the evaluation on the edge computing device to detect abnormal behavior of escalator passengers.
[0039] Specifically, we used a mobile phone camera to capture live video data, used Python scripts to extract frames from the video data, converted them into image data, and collected relevant resources on the Internet to create a data set. The data mainly collected five types of behavior postures of escalator pedestrians, including standing (Stand), sitting (Sit), bending (bow), leaning (lean), and falling (fall). The initial data set contains a total of 4,000 images. In order to increase the diversity of data and improve the generalization ability of the model, the initial data set is enhanced and combed. Specifically, Gaussian noise and salt and pepper noise are randomly added to each original data sample, and the brightness, exposure, hue and contrast are randomly adjusted to generate enhanced data samples; other methods can also be used to expand the data set, such as image mirroring, random flipping, translation, rotation, and blurring.
[0040] The enhanced data samples are mixed with the original data samples in the same quantity ratio to form a new data set. The embodiment of the present invention divides the training set, validation set and test set in a ratio of 7:2:1. Through data enhancement, the scale of the data set is effectively increased, the number of samples is greatly increased, the diversity of samples is enriched, and it is more conducive to training deep learning models. It helps the model to better capture different types of abnormal behavior postures of pedestrians during training, improve the robustness of the model, and make the model more reliable in actual scene applications.
[0041] like Figure 1 As shown, specifically, the improved YOLOv8 algorithm proposed in the embodiment of the present invention is mainly optimized by lightweighting the backbone network, integrating the ECA module and the attention module, designing a new C2f_DSConv module, and optimizing the loss function. The lightweight strategy of the backbone network is to replace the original backbone network with the ShuffleNetV2 network. The number of parameters of the replaced backbone network is greatly reduced compared to the original network, and the operating efficiency is significantly improved. The working principle of the ShuffleNetv2 module is to replace the standard convolution block in the module with a depth-separable convolution to reduce the computational complexity of the module, and to compensate for the disadvantage of reduced feature extraction capability caused by reduced information exchange between channels through the channel shuffling strategy.
[0042] The principle of the C2f_DSConv module designed in the embodiment of the present invention is to replace the standard convolution block in the original C2f module with a distribution shift convolution block. Among them, the DSConv distribution shift convolution is a deformation of the depthwise separable convolution, which is widely used in the field of computer vision. DSConv simulates the standard convolution layer by using quantization and distribution offset, and decomposes the traditional convolution kernel into two components: variable quantization kernel (VQK) and distribution offset. Lower memory usage and higher speed are achieved by storing only integer values in VQK; distribution shift mainly includes two parts: kernel distribution offset (KDS) and channel distribution offset (CDS). KDS is used to move the distribution of each block of VQK, and CDS is used to move the distribution in the channel. Compared with the standard convolution operation, it has higher efficiency and occupies less memory. It reduces the calculation amount of the model and improves the detection speed without affecting the accuracy of the model, and has a good lightweight effect. The structure of the C2f_DSConv module is as follows Figure 2 As shown in the figure, convolution is used to change the number of channels of the input features, and then the separation operation replaces the convolution layer segmentation feature. The receptive field and calculation speed are improved by superimposing multiple distribution shift convolutions. Some cross-level connections are used to reduce the amount of calculation and enrich the characteristic behaviors of passengers at different scales.
[0043] The ECA attention module is as follows Figure 3 As shown in the figure, it is an improved version of SENET, which can not only effectively avoid the side effects caused by the dimension reduction caused by the fully connected layer, but also adapt the convolution kernel size to improve the information exchange between channels, more efficiently capture the effective information in the image, and weaken the background information. In order to avoid the side effects of channel attention caused by dimension reduction, and there is no need to maintain the interdependence of all channels, it removes the fully connected layer in the original SENET and directly uses 1 after global average pooling. 1 convolutional layer effectively solves the dimension reduction problem and improves the ability to capture cross-channel interactions. The basic structure diagram of ECA is as follows Figure 4 As shown in the figure, H, W, and C are the height, width, and number of channels of the feature map, respectively. The feature map is adaptively combined with a one-dimensional convolutional layer of size K after global average pooling to determine the range of cross-channel interaction, and multiplied with the original feature map to complete the recalibration. The convolution kernel size changes through an adaptive function, which is related to the number of channels. The adaptive function is:
[0044] ,
[0045] Among them, C is the number of channels, b=1,γ=2, and k is the size of the one-dimensional convolution kernel. The ECA attention module only adds a very small number of additional parameters but significantly reduces the complexity of the model, enabling the model to achieve higher performance. It is a lightweight attention mechanism.
[0046] The bounding box loss function of the traditional YOLOv8 algorithm uses CIoU. The calculation formula of CIoU-Loss is as follows:
[0047] ,
[0048] ,
[0049] ,
[0050] Where IoU is the intersection over union ratio in the target detection task, which is calculated by dividing the intersection of the predicted box and the true box by the union of the former and the latter, reflecting the degree of overlap between the predicted box and the true box. is the centroid of the prediction box; w and h are the width and height of the prediction box; the corresponding gt is the parameter of the real box; v is the parameter to measure the consistency of the aspect ratio; the parameter α is the weight coefficient, and c is the diagonal length of the minimum circumscribed matrix that includes both the real box and the prediction box. When the aspect ratio of the predicted box and the real box is close, the penalty term of the loss function is prone to failure, which is an unreasonable situation. To address this problem, this paper uses the WIoU loss function instead of CIoU as the border regression loss function of the algorithm. The calculation formula of WIoU-Loss is as follows:
[0051] ,
[0052] In the formula, w i It represents the weight value of the i-th category, n is the number of prediction boxes, b i With g i is the coordinate of the i-th predicted box and the coordinate of the real labeled box. The WIoU loss function can more accurately measure the similarity between the predicted box and the real target by introducing the concept of weight, especially when multiple object parts are involved, which can improve the detection accuracy.
[0053] like Figure 4,The implementation steps of this method are as follows: first, collect escalator passenger behavior data and make a data set, divide the data set, use the improved network training data set to obtain the model, use the evaluation index to evaluate the model performance, export the model as an onnx model, and then convert the trained model into an rknn model through the conversion tool, deploy it to the rk3588 development board to accelerate reasoning and connect the camera for real-time detection. Among them, the model evaluation indicators for the improved algorithm training include: experimental model evaluation indicators select precision (Precision, P), recall (Recall, R), accuracy (Average Precision, AP), mean Average Precision (mAP), model parameter quantity (Parameters, Param), model weight size (Weights).
[0054] On the other hand, the present invention also provides an escalator passenger abnormal behavior detection system based on YOLOv8, the system comprising:
[0055] A data set acquisition unit is used to collect and preprocess the abnormal behavior data of escalator passengers, and to annotate the preprocessed image data to construct an escalator passenger abnormal behavior detection image data set;
[0056] The detection model building unit is used to use YOLOv8 as the baseline algorithm, introduce the lightweight feature extraction network ShuffleNetV2 and the designed C2f_DSConv module, integrate the ECA module and optimize the loss function to build the YOLOv8_SDE passenger abnormal behavior detection model;
[0057] An optimization model acquisition unit is used to train the YOLOv8_SDE passenger abnormal behavior detection model using an escalator passenger abnormal behavior detection image dataset to obtain an optimized detection model;
[0058] The model evaluation unit is used to evaluate the performance of the optimized detection model using evaluation indicators, and deploy the optimized detection model that has passed the evaluation on the edge computing device to detect abnormal behavior of escalator passengers.
[0059] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for detecting abnormal behavior of escalator passengers based on YOLOv8, characterized in that: The method comprises the following steps: Step 1: Collect and preprocess the abnormal behavior data of escalator passengers, annotate the preprocessed image data, and construct an escalator passenger abnormal behavior detection image dataset; Step 2: Using YOLOv8 as the baseline algorithm, introduce the lightweight feature extraction network ShuffleNetV2 and the designed C2f_DSConv module, integrate the ECA module and optimize the loss function to build the YOLOv8_SDE passenger abnormal behavior detection model; Step 3: Use the escalator passenger abnormal behavior detection image dataset to train the YOLOv8_SDE passenger abnormal behavior detection model to obtain an optimized detection model; Step 4: Use evaluation indicators to evaluate the performance of the optimized detection model, and deploy the optimized detection model that has passed the evaluation on the edge computing device to detect abnormal behavior of escalator passengers.
2. According to the YOLOv8-based escalator passenger abnormal behavior detection method of claim 1, it is characterized in that: The step 1 includes collecting the abnormal behavior data of escalator passengers captured in real time and obtained through the network, performing data enhancement processing, marking the behavior postures of escalator pedestrians, including standing, sitting, bending, leaning, and falling, and constructing an escalator passenger abnormal behavior detection image dataset.
3. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 1, characterized in that: The step 2 of constructing the YOLOv8_SDE passenger abnormal behavior detection model includes: The lightweight feature extraction network ShuffleNetV2 is used as the backbone network of YOLOv8 to extract features from images; The neck-trunk network uses PAN-FPN to perform multi-scale feature fusion on the features extracted by the backbone network; The detection head part includes a detection head and a classification head, which perform target detection and classification on the fused multi-scale features.
4. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 3 is characterized in that: The neck-trunk network includes a C2f_DSConv module and a fused ECA attention module.
5. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 4, characterized in that: Each C2f_DSConv module is connected to the fused ECA attention module and the loss function is optimized as WIoU.
6. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 5, characterized in that: The C2f_DSConv module is a new module formed by replacing the standard convolution block in C2f with a distributed shift convolution block.
7. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 5, characterized in that: The fused ECA attention module is a channel attention mechanism deep convolutional neural network, which uses a 1D convolutional layer to fuse information between channels.
8. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 1, characterized in that: The evaluation indicators in step 4 include precision, recall, accuracy, average accuracy of all categories, model parameters, and model weight.
9. The method for detecting abnormal behavior of escalator passengers based on YOLOv8 according to claim 1, characterized in that: The optimized detection model is deployed on the edge computing device and reasoned using the Rockchip RK3588 development board.
10. An escalator passenger abnormal behavior detection system based on YOLOv8, characterized in that: include: A data set acquisition unit is used to collect and preprocess the abnormal behavior data of escalator passengers, annotate the preprocessed image data, and construct an escalator passenger abnormal behavior detection image data set; The detection model building unit is used to use YOLOv8 as the baseline algorithm, introduce the lightweight feature extraction network ShuffleNetV2 and the designed C2f_DSConv module, integrate the ECA module and optimize the loss function to build the YOLOv8_SDE passenger abnormal behavior detection model; An optimization model acquisition unit is used to train the YOLOv8_SDE passenger abnormal behavior detection model using an escalator passenger abnormal behavior detection image dataset to obtain an optimized detection model; The model evaluation unit is used to evaluate the performance of the optimized detection model using evaluation indicators, and deploy the optimized detection model that has passed the evaluation on the edge computing device to detect abnormal behavior of escalator passengers.
Citation Information
Patent Citations
Lightweight pomegranate identification method based on improved YOLOv8s
CN116958961A
Yolo magnetic shoe surface defect detection method and system based on lightweight convolution
CN118196023A
Escalator pedestrian abnormal behavior recognition method, device and equipment and storage medium
CN118587760A
Abnormal behavior detection method and device based on improved YOLOv8 and storage medium
CN119152344A