Escalator pedestrian safety detection method based on improved YOLOv8

By introducing technologies such as attention mechanism, feature pyramid and fully convolutional mask autoencoder in YOLOv8, the pedestrian detection model is optimized, and the problems of insufficient pedestrian detection accuracy and low real-time performance in the escalator environment are solved, achieving more efficient and more accurate pedestrian detection.

CN120198853APending Publication Date: 2025-06-24NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510336583.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has problems with insufficient detection accuracy, collapse of features, low real-time performance, and repeated alarms and false alarms in pedestrian detection in escalator environments.

Method used

By introducing attention mechanism CA and feature pyramid CFP algorithm in YOLOv8, combining the fully convolution mask autoencoder framework FCMAE and the multi-layer perceptron MLP classifier, the pedestrian detection model is optimized.

Benefits of technology

It significantly improves the accuracy and robustness of pedestrian detection, reduces false alarms and repeated alarms, enhances the model's performance ability in complex environments, and meets the high real-time requirements in escalator environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BSA0000300035560000041
    Figure BSA0000300035560000041
  • Figure HSA0000300035570000011
    Figure HSA0000300035570000011
  • Figure HSA0000300035570000021
    Figure HSA0000300035570000021
Patent Text Reader

Abstract

The invention relates to the field of computer vision and deep learning, in particular to introduction of an attention mechanism CA in a Yov8 target detection algorithm in deep learning and introduction of a feature pyramid CFP algorithm in a feature fusion mode for enhancement. The invention also provides a related research and technical scheme for classifying human body key points detected on the escalator by using a full convolution mask auto-encoder frame (FCMAE) and an MLP classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and deep learning, and in particular to the introduction of the attention mechanism CA in the Yolov8 object detection algorithm in deep learning, the introduction of the feature pyramid CFP algorithm in the feature fusion method for enhancement, as well as the related research and technical solutions of the fully convolutional mask autoencoder framework FCMAE and the use of an MLP classifier to classify the human key points detected on the escalator. Background Art

[0002] With the acceleration of urbanization, people's dependence on public transportation facilities is increasing day by day. Among them, escalators, as important facilities connecting different floors, are widely used in public places such as shopping malls, subway stations, and airports. However, the high-traffic use of escalators and their special operating methods bring significant safety hazards. Especially during peak hours, the crowded situation on the escalator may lead to safety accidents such as falls and squeezes, posing a serious threat to pedestrian safety. Although traditional video surveillance systems can provide round-the-clock real-time monitoring, these systems still have significant limitations in automatically identifying specific behaviors, evaluating potential risks, and timely issuing warnings. Most existing surveillance systems rely on manual observation, which not only consumes resources but also has low efficiency, especially when dealing with a large amount of surveillance data.

[0003] In the field of object detection technology, although advanced algorithms such as YOLOv8 have been widely applied and demonstrated excellent capabilities in multi-object detection tasks, in specific escalator scenarios, these technologies still face many challenges. The uniqueness of the escalator environment, such as the frequent overlap between pedestrians, local occlusion, and the complex and changeable background and lighting conditions, greatly increases the difficulty of the detection task. In addition, real-time performance is extremely critical in escalator safety monitoring, but existing algorithms often struggle to meet the high real-time requirements when running on resource-constrained devices. Therefore, although existing technologies have basic object detection functions, their efficiency and accuracy in dealing with dynamic and complex escalator scenarios still need to be improved. This series of challenges prompts us to explore more optimized solutions to ensure more efficient and accurate pedestrian detection in critical safety areas such as escalator usage environments.

[0004] To solve the above problems and improve the accuracy and robustness of escalator pedestrian safety detection. By introducing the Channel Attention (CA) mechanism into YOLOv8, the learning ability of the model for important features can be effectively enhanced, thereby improving the accuracy of pedestrian detection; in addition, by introducing the Feature Pyramid (CFP) in the feature fusion method, features can be better fused and enhanced across different scales, thus enhancing the detection model's ability to recognize small and blurred targets; at the same time, the Fully Convolutional Mask Autoencoder Framework (FCMAE) can further improve the model's performance in dealing with occlusion and background complexity; finally, using a Multi-Layer Perceptron (MLP) classifier to classify the detected human key points not only improves the classification accuracy but also provides support for pedestrian pose analysis, which is crucial for predicting and preventing potential dangerous behaviors, further improving the accuracy and robustness of escalator pedestrian safety detection. Summary of the Invention

[0005] 1. The object of the present invention is to solve the problems of insufficient detection accuracy in existing escalator pedestrian safety detection, especially in the case of multiple people overlapping and complex environments, as well as problems such as feature collapse, low real-time performance, and repeated and false alarms.

[0006] To achieve the above object of the invention, the present invention proposes an escalator pedestrian safety detection method based on improved YOLOv8, aiming to improve the accuracy and robustness of escalator pedestrian safety detection. The method includes the following steps:

[0007] Step 1: Collect video data in the escalator environment, including different times, different lighting conditions, and different pedestrian flows. Preprocess the video data, including frame extraction, image cropping, and annotation of pedestrian positions and key points.

[0008] Step 2: Introduce the Channel Attention mechanism (CA) and the Feature Pyramid CFP algorithm into the backbone network of YOLOv8. Place these layers after the convolutional layer and before the activation layer, so that the features extracted by the convolutional layer can be directly adjusted. These improvements aim to enhance the model's local information processing ability and feature expressiveness. Perform average pooling operation on each feature map to compress the spatial information of each channel into a single scalar. Use one or more fully connected layers to learn the dependencies between different channels. This usually includes a fully connected layer for reducing the dimension and a fully connected layer for expanding the dimension, forming a "bottleneck" structure to reduce the number of parameters and increase the model complexity. Usually, the ReLU activation function is used between the fully connected layers to introduce non-linearity, and the Sigmoid function is used in the last layer to output the weight coefficients of each channel, with the range between 0 and 1. Through the previous structure, the CA layer will output a weight vector with the same number of channels as the feature map. This weight vector will be used to weight the importance of each channel in the original feature map, that is, apply the weight to the corresponding channel through element-wise multiplication (Hadamard Product). In this way, the model can enhance the response to important features in the image while suppressing unimportant information.

[0009] The integration of the attention mechanism is shown in Equation (1):

[0010] E cs = σ(W C *E C + W S *E S )(1)

[0011] where E cs is the final attention weight matrix, σ is the Sigmoid function, E C and E S represent the channel and spatial attention scores respectively, and W C and W S are the corresponding weights. Such a technical integration significantly improves the model's ability to focus on important features more precisely and reduces the impact of irrelevant or interfering information. This method can effectively reduce false alarms and repeated alarms by strengthening the model's learning and utilization of key features in pedestrian detection. By weighting the important feature channels, the model can more accurately distinguish individuals in complex escalator environments, such as when the crowd is crowded, thus reducing misidentifications.

[0012] Step 3: Introducing the Feature Pyramid CFP algorithm optimization in the feature fusion method is a crucial step. It greatly improves the model's ability to capture pedestrians of different sizes from near to far by effectively integrating features from different levels. This process first involves optimizing the basic feature extraction layer to ensure that the features extracted from each convolutional layer of the network can reflect information at different scales, covering the view from details to the global. When constructing the feature pyramid, we adopt multi-scale feature fusion technology. In specific operations, feature maps of different depths are hierarchically integrated according to their different spatial resolutions. The key to this step lies in adjusting the feature maps of different levels to a unified scale, usually achieved through upsampling and downsampling operations. Upsampling and downsampling techniques are used to adjust the spatial resolutions of feature maps at different levels so that they can be fused at the same scale. For deeper feature maps, upsampling (such as bilinear interpolation or transposed convolution) is used to increase their resolution to match that of shallower layers; conversely, for shallower feature maps, downsampling (such as max pooling or strided convolution) is used to reduce their resolution to adapt to deeper layers. In the feature fusion method, it is not just simple element-wise addition but can also include convolutional operations to further fuse features at different scales, ensuring that the fused feature map can take into account the advantages of features at each layer and further reducing false alarms and repeated alarms caused by changing conditions.

[0013] Step 4: For the accuracy and robustness of escalator pedestrian safety detection, the present invention uses an improved YOLOv8 algorithm for optimization. The FCMAE framework helps the model recover detailed information that may be lost in the input by encoding and decoding the features of local regions, especially in cases where pedestrians are partially occluded or overlapping.

[0014] Finally, the hybrid attention weight matrix E cs is multiplied element-wise with the input feature map E and added to the original input feature map to obtain the refined feature map E final , calculated as shown in formula (2):

[0015] E final = E + E ⊙ E cs (2)

[0016] where E is the initial feature map, ⊙ represents element-wise multiplication, and E cs is the weighted feature map obtained through the channel attention mechanism. This module emphasizes paying attention to meaningful features in both the channel and spatial dimensions, focusing on important features and suppressing invalid features, thereby improving the discriminability of features and the ability to capture key information. For pedestrian detection in complex environments such as occlusion or overlap, it helps to recover detailed information that may be lost in the image, significantly improving the accuracy and efficiency of the model in escalator pedestrian safety detection.

[0017] Step 5: Set the training hyperparameters. Based on the escalator pedestrian safety detection dataset established in Step 2, train the improved YOLOv8 network in Step 3. During the training phase, use a large amount of labeled data to train the MLP classifier and perform supervised learning using a large number of sample data with behavior labels. These data examples are labeled with different behavior patterns of pedestrians, such as normal walking, running fast, falling, etc. During the training process, by adjusting the weights and biases in the network, it learns how to correctly identify and classify the behavior patterns corresponding to the input data (such as human key points). During the training process, the MLP minimizes the prediction error by adjusting the weights and biases of its internal neural network to optimize the classification performance. After the improved YOLOv8 model detects a pedestrian and identifies its key points, this key point information is used as the input of the MLP classifier. The MLP classifier structure includes an input layer, several hidden layers, and an output layer, where the number of neurons and the number of layers in the hidden layer will be adjusted according to the complexity and characteristics of the training data. The key to using the MLP classifier lies in its ability to handle complex pattern recognition problems. By learning a large amount of pedestrian pose data, it can distinguish different behavior states and issue a timely warning when detecting that a pedestrian is in an abnormal state on the escalator.

[0018] After training, obtain the detection model and optimize the model parameters through the cross-validation method; input the video frame by frame into the escalator pedestrian safety detection model for detecting the human state, especially pay attention to the performance of the model in recognizing pedestrians of different sizes, obtain the action information and confidence of the pedestrians, and effectively handle the pedestrian detection task in dynamic environments such as escalators.

[0019] The beneficial effects of the present invention are as follows: Through the above steps, by improving Yolov8, the accuracy and efficiency of the YOLOv8 object detection algorithm in escalator pedestrian detection have been significantly improved, effectively solving problems such as multiple people overlapping, complex environments, and high real-time requirements, and ensuring the safety of pedestrians in the escalator environment. Description of the Drawings

[0020] Figure 1 It is a structural diagram of the improved version introducing the attention mechanism CA, enhancing in the feature fusion method by introducing the feature pyramid CFP algorithm, as well as the fully convolutional mask autoencoder framework FCMAE and using the MLP classifier. Figure 2 It is a detailed diagram of the improvement. Detailed Embodiment

[0021] The following combines the attached Figure 2 Describe the embodiments of the present invention in detail.

[0022] Step 1: Collect pedestrian image data in various escalator usage scenarios, paying particular attention to pedestrian overlap and occlusion. Use data augmentation techniques (such as image rotation, scaling, and color adjustment) to expand the dataset, and preprocess the images, including resizing and normalization, to optimize the subsequent model training effect. At the same time, adopt data augmentation methods such as data cutting and adding noise to improve robustness.

[0023] Step 2: Optimize using the improved YOLOv8 algorithm proposed in this paper. First, add a CA attention mechanism to the feature extraction network to enrich semantic information. It simultaneously realizes attention in the horizontal and vertical directions and retains the position information of the targets within the samples, effectively enhancing the model's ability to express features. First, perform global average pooling on each channel of the feature map to obtain the global spatial information of each channel, which will form a channel descriptor.

[0024] Assume the size of the input feature map X is C*H*W (where C is the number of channels, and H and W are the height and width of the feature map). The operation of global average pooling can be expressed as:

[0025]

[0026] where, F gap (c) is the average value of all pixel values in the c-th channel.

[0027] Then, use a small fully connected (FC) network to process these channel descriptors. This network usually contains two layers of FC. The first layer reduces the dimension (for example, the number of channels is divided by 16), and the second layer expands the dimension back to the original number of channels, aiming to learn the weights of different channels. The fully connected layer is usually used to process the channel descriptors obtained in Step 1. If there are C channels, then the input to the fully connected layer is a C-dimensional vector. First is a dimension reduction layer, and its output is:

[0028] F fc1 =ReLU(W1*F gap +b1) (4)

[0029] where, W1 is a weight matrix of dimension, b1 is the bias, and r is the dimension reduction ratio, usually 16.

[0030] Next is an expansion layer to expand the dimension back to C:

[0031] F fc2 =W2·F fc1 +b2 (5)

[0032] where, W2 is a weight matrix of dimension, b2 is the bias.

[0033] Pass the output of the FC network through the sigmoid activation function to obtain the weight factors for each channel.

[0034] F sigmoid = σ(F fc2 ) (6)

[0035] where σ represents the sigmoid activation function, which is used to normalize the weights to the range [0, 1].

[0036] Finally, these weight factors are used to weight each channel of the original feature map, that is, the weights are applied to the corresponding channels through element-wise multiplication (Hadamard product).

[0037] X out (c, i, j) = F sigmoid × X(c, i, j) (7)

[0038] Here, is the output feature map after channel attention weighting, c is the channel index, and i and j are the spatial position indices.

[0039] Step 3: In an object detection architecture like Yolov8, a Centralized Feature Pyramid (CFP) can be designed as a feature fusion mechanism to enhance the model's ability to integrate features from different network levels. This fusion pattern particularly focuses on effectively utilizing the unique information of each layer's features through centralized processing. First, select multiple feature layers with different depths from the base convolutional network (such as CSPNet or other backbone networks). Usually, deep (low resolution, high semantics), middle (medium resolution, medium semantics), and shallow (high resolution, low semantics) features are selected. Preprocess each selected feature layer, such as adjusting the number of channels using 1x1 convolutions to make the number of channels of each layer the same for subsequent processing. Adjust the spatial dimensions of the feature maps of each layer through upsampling or downsampling methods to make them reach the same resolution. Then use weighted summation, concatenation, or more complex fusion operations (such as further processing the fused features through convolutional layers) to combine the feature maps from different layers. Finally, integrate the CFP module into the feature extraction and prediction processes of Yolov8.

[0040] Step 4: Based on Step 2 and Step 3, at the key nodes, the features processed by FCMAE are re-injected into the main process. As an independent feature enhancement branch, it operates independently outside the main network architecture (such as the basic feature extraction and feature fusion stages). The fully convolutional mask autoencoder usually serves as an auxiliary module, working in parallel with the main detection process to improve the expressiveness of features and the robustness of the network. Among them, the encoder is usually composed of multiple convolutional layers, aiming to gradually reduce the spatial dimension of the feature map while possibly increasing the number of channels. For a given input feature map X, the encoder can be expressed as a series of convolutional operations:

[0041] X encoded = ReLU(W enc * X + b enc ) (8)

[0042] Where W enc and b enc are the weights and biases of the encoder convolutional layer respectively, * represents the convolutional operation, and ReLU is the activation function used to introduce non-linearity. This helps to capture more abstract feature representations. The encoder can directly receive the input from an intermediate feature layer of Yolov8. The decoder part uses transposed convolution or upsampling plus convolution to gradually restore the spatial size of the feature map while reducing the number of channels, aiming to reconstruct the high-resolution feature map of the original input, usually designed to be symmetric with the encoder:

[0043] X decoded = ReLU(W dec * X encoded + b dec ) (9)

[0044] Where W dec and b dec are the weights and biases of the decoder convolutional layer, used to reconstruct the feature map close to the original input. After obtaining the output X decoded processed by the encoder and decoder, this output is usually fused with the feature map X main of the main network flow:

[0045] X fused = λ * X decoded + (1 - λ) * X main ) (10)

[0046] Here, λ is the fusion coefficient, which may be fixed or learned through training, and is used to control the contribution degrees of the two feature sources. The output of FCMAE can be designed to fuse with the main network features. For example, the output of FCMAE can be merged with the main feature stream by element-wise addition or concatenation. In this way, FCMAE not only increases the network's depth of understanding of the input data but also may help improve the detection effect of occluded objects, especially in complex environments.

[0047] Step Five: After integrating advanced technologies such as CA (Channel Attention), CFP (Centralized Feature Pyramid), and FCMAE (Fully Convolutional Mask Autoencoder), fine-tuning of the last part of the framework is involved. First, before feeding the features into the MLP, it is usually necessary to preprocess the features. If the feature map size is large or the number of channels is large, global average pooling or other pooling techniques can be used first to reduce its dimension and convert it into a one-dimensional feature vector. The MLP usually contains multiple fully connected layers. The ReLU activation function can be used in the intermediate layers, and the softmax (single label) or sigmoid (multi-label) function can be used in the output layer according to whether it is multi-label classification. In the detection head, the MLP classifier receives the input features from the feature extraction network and performs the final classification task. Before this, the features have been processed by the CA module and enhanced by CFP and FCMAE. In the training stage, the MLP classifier is trained together with the entire network. The cross-entropy loss is used to optimize the classification task to ensure that the weights of the MLP are optimized through the backpropagation algorithm, while considering the performance and efficiency of the overall network. After integrated training, first input an image or video into the trained improved Yolov8 model for detection, identify the behavioral characteristics of people such as normal walking, running fast, falling, etc., and if it is found that a pedestrian is in a dangerous posture, an alarm will be given and the corresponding pictures and videos will be saved; the client is responsible for receiving the alarm information from the server and performing the corresponding viewing work.

[0048] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in the relevant technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A pedestrian safety detection method for escalators based on improved YOLOv8, characterized in that: The specific steps include: S1: Install multiple cameras on elevators in public areas to capture pedestrian activities in elevators from multiple angles and at different time periods. Perform data enhancement on the collected images and videos, including horizontal and vertical flipping, brightness and contrast adjustment, and random cropping of some areas to simulate various visual environments and increase the diversity and volume of the dataset; S2: Use YOLOv8 as the target detection framework and integrate the channel attention mechanism into it. This includes modifying the network structure of YOLOv8 and adding CA layers, which will learn to weight the importance of different channels to strengthen the model's response to discriminative features in the image; S3: Based on the improvement of YOLOv8, the feature pyramid CFP algorithm is integrated for feature fusion. The algorithm is designed to establish connections between features extracted at different levels, and enhance the adaptability to scale changes through the fusion of context information, especially the detection ability of small and partially occluded targets; S4: Implement the fully convolutional masked autoencoder framework (FCMAE), which segments and encodes human body regions more finely; further optimizes the model's handling of occlusion and background complexity; S5: Apply a multi-layer perceptron (MLP) classifier to accurately classify the detected human key points, support deeper behavior and posture analysis, and enhance the model's ability to predict and prevent potentially dangerous behaviors.

2. The escalator pedestrian safety detection method based on improved YOLOv8 according to claim 1 is characterized in that: The specific steps of step S2 are: S21: Introduce the channel attention mechanism (CA) in the backbone network of YOLOv8; place it after the convolution layer and before the activation layer, so that the features extracted by the convolution layer can be directly adjusted; S22: Perform an average pooling operation on each feature map to compress the spatial information of each channel into a single scalar; S23: Use one or more fully connected layers to learn dependencies between different channels. This usually includes a fully connected layer with reduced dimensions and a fully connected layer with expanded dimensions, forming a "bottleneck" structure to reduce the number of parameters and increase model complexity. S24: ReLU activation function is used between fully connected layers to introduce nonlinearity, and the last layer uses Sigmoid function to output the weight coefficient of each channel, ranging from 0 to 1.

3. The escalator pedestrian safety detection method based on improved YOLOv8 according to claim 1 is characterized in that: The specific steps of step S3 are: The specific steps of step S3 are: S31: Select key feature extraction layers in multiple layers of YOL0v8. These layers will include convolutional layers from different depths of the network, each of which represents feature information at different scales. After selecting these layers, design algorithms to establish effective connections between these layers. These connections not only transmit information vertically (in the depth direction), but also share features horizontally (between different modules at the same depth); S32: Use upsampling and downsampling techniques to adjust the spatial resolution of feature maps at different levels so that they can be fused at the same scale. For feature maps of deeper layers, upsampling (such as bilinear interpolation or transposed convolution) is used to increase their resolution to match that of shallower layers; conversely, for feature maps of shallower layers, downsampling (such as maximum pooling or stride convolution) is used to reduce their resolution to adapt to deeper features; S33: Fusion of contextual information is achieved by designing cross-layer direct connections and strengthening the semantic information expression in each layer.

4. The escalator pedestrian safety detection method based on improved YOLOv8 according to claim 1 is characterized in that: The specific steps of step S4 are: S41: Develop a fully convolutional network (FCN) as a base structure that can process input images of arbitrary size and output segmentation maps of the same size; S42: Add a mask generation layer to FCN, which is specifically used to predict whether each pixel belongs to the human body region. By using the sigmoid or softmax activation function, the output layer will generate a binary or multi-classification mask for each pixel, indicating the attributes of different regions. S43: Integrate autoencoders to improve feature encoding efficiency. Autoencoders compress the input into a low-dimensional representation through an encoder and reconstruct the output through a decoder, learning to capture the most critical information. S44: Use annotated human region data from real scenes to train FCMAE and optimize the network to ensure efficient segmentation and accurate occlusion handling. The training process may require adjusting the loss function to better handle occlusions and background noise.

5. The escalator pedestrian safety detection method based on improved YOLOv8 according to claim 1 is characterized in that: The specific steps of using a multi-layer perceptron (MLP) to classify the detected human key points in step S5 are: S51: Use an optimized target detection model (such as YOLOv8) to identify the human body in the image and locate key points. These key points include important parts such as the head, hands, and feet; S52: Design an MLP network that contains multiple hidden layers to process the input features obtained from the key point detection step. Each input feature represents the location information and possible context information of a key point; S53: Use the labeled posture data to train the MLP so that it can classify different behaviors or postures based on the configuration of key points. For example, distinguish between walking, sitting, falling, etc. S54: Optimize the classification accuracy and response speed of the MLP by adjusting network parameters, increasing the data set, or applying different training strategies; S55: Integrate the trained MLP classifier with the subject detection framework to ensure real-time analysis of behaviors and gestures in video streams and improve the ability to predict potentially dangerous behaviors; The beneficial effects of the present invention are as follows: through the above steps, the channel attention mechanism, the feature pyramid CFP algorithm, the fully convolutional masked autoencoder framework and the multi-layer perceptron classifier are combined to significantly improve the detection accuracy, the ability to handle complex backgrounds and occlusions, and the depth of behavior and posture analysis. The application of this integration and advanced algorithms enables the system to effectively identify and prevent potentially dangerous behaviors in a complex public escalator environment, thereby greatly improving the effectiveness of the public safety monitoring system and ensuring pedestrian safety.