Escalator safety detection method based on double-cascade YOLOv8 architecture

Through the escalator safety detection method based on the dual-cascade YOLOv8 architecture, the problems of background interference sensitive, small target miss detection and occlusion scene performance in escalator scenes are solved, and efficient and real-time passenger attitude detection is achieved, which is suitable for resource-constrained environments.

CN120039753APending Publication Date: 2025-05-27SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510341290.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

There are problems in escalator safety detection with sensitive background interference, high detection rate of small target leakage and degraded performance in occlusion scenes, and the existing technology is difficult to take into account both detection accuracy and efficiency.

Method used

The escalator safety detection method based on the dual-cascade YOLOv8 architecture is adopted, and by constructing the YOLOv8-Escalator-DET and YOLOv8-POSE-DET networks, it is used for escalator operation area positioning and passenger attitude detection respectively. This method combines multi-scale feature fusion and occlusion perception mechanism to optimize the network structure to improve detection accuracy and real-time.

Benefits of technology

It significantly improves the detection capabilities of small targets and occlusion targets, realizes efficient target detection, and meets the real-time computing needs in resource-constrained environments, providing an effective technical solution for the safety monitoring of escalators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120039753A_ABST
    Figure CN120039753A_ABST
Patent Text Reader

Abstract

The invention discloses an escalator safety detection method based on a double-cascade YOLOv8 architecture, which specifically comprises the following steps: adopting a staged task decoupling strategy: in the first stage, realizing positioning of an escalator running area (ROI) through a lightweight improved YOLOv8-Escalator-DET network, and effectively inhibiting interference of adjacent non-elevator areas (such as fixed stairs and floor transition zones); in the second stage, the ROI is input into a high-precision YOLOv8-POSE-DET network, and fine-grained analysis of the passenger posture is completed through multi-scale feature fusion and a shielding perception mechanism. According to the method, good balance between real-time performance and detection precision can be effectively realized, and the method is suitable for a real-time monitoring system with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to escalator safety detection technology, and in particular to an escalator safety detection method based on a double-cascade YOLOv8 architecture. Background Art

[0002] As the core transportation equipment of urban transportation hubs and commercial complexes, the safe operation of escalators is directly related to public safety and the safety of passengers' lives and property. However, with the surge in passenger flow and the complexity of passenger behavior, traditional detection methods based on manual monitoring and basic sensors have become difficult to cope with dynamic risks. Manual monitoring has the problem of missed detection due to distraction, while traditional technologies such as infrared sensors and pressure sensing devices can only detect anomalies in fixed areas (such as object obstruction), and are insufficient in identifying dynamic dangerous postures such as passengers falling or climbing handrails. For this reason, intelligent detection technology based on computer vision has become a research hotspot, among which target detection algorithms have become the preferred solution for the deployment of real-time monitoring systems due to their high efficiency and strong real-time performance.

[0003] Among the deep learning methods for escalator safety detection, the mainstream technologies can be divided into two categories: target detection and instance segmentation. Single-stage target detection models represented by the YOLO series can quickly locate the passenger position through bounding boxes, and can identify significant abnormal behaviors such as falls and reverse walking in real time (response within 30ms), meeting the stringent requirements of escalator scenarios for low latency. In contrast, although instance segmentation algorithms (such as Mask R-CNN) can output human contours through pixel-level classification, their computational complexity is high (it takes more than 150ms to process a single frame of 4K images), making it difficult to achieve real-time warnings in dense passenger flow scenarios, and the demand for hardware resources limits its deployment on edge devices. Although target detection has significant efficiency advantages, existing solutions still have three major bottlenecks:

[0004] 1) Sensitive to background interference: Dynamic targets in adjacent areas (such as pedestrian movement on non-automatic staircase floors and reflective objects on concrete stairs) are easily misdetected as passengers in the elevator, resulting in a significant increase in the false alarm rate.

[0005] 2) High missed detection rate for small targets: Small-scale targets such as long-distance passengers and personal belongings have insufficient feature representation capabilities under low-resolution input, and missed detection is a prominent problem during long-distance detection.

[0006] 3) Performance degradation in occluded scenes: Baggage occlusion or dense crowds reduce the visible area of ​​the target. Existing models lack occlusion modeling capabilities, resulting in a high misjudgment rate for partially occluded postures.

[0007] The current research mainly optimizes the object detection model through two paths: the first is the lightweight network transformation. For example, the team from Shenzhen University replaced the backbone network of YOLOv5 with GhostNet, reducing the number of parameters by 45% through feature channel compression, achieving 60FPS detection on the Jetson Xavier device, but the missed detection rate of small targets still reaches 32%. The team from Shanghai Jiao Tong University proposed the YOLOv7-Tiny scheme guided by spatial attention, reducing the false detection rate to 8.7%, but the detection accuracy for occluded targets is less than 65%. The second is the cascaded detection architecture. For example, the team from the Tokyo Institute of Technology designed a two-stage system: in the first stage, MobileNetv3 is used to segment the elevator area, and in the second stage, YOLOv8 is used to detect the passenger posture. The false detection rate of this scheme is reduced by 52%, but the end-to-end latency increases to 120ms, making it difficult to meet the real-time requirements. The MTR Corporation of Hong Kong, China, adopted the cascaded YOLOv5 model, reducing the computational amount through region screening, but still requires 8GB of video memory in a 4K video stream, with high edge deployment costs.

[0008] Although the above improvements have partially alleviated the problems, the contradiction that a single model is difficult to balance detection accuracy and efficiency remains unresolved: lightweight models sacrifice the detection ability of small targets, while cascaded architectures introduce additional computational overhead. Developing a detection model that can resist background interference, accurately identify small targets and occluded postures, and adapt to real-time computing on edge devices has become an urgent need in the field of escalator safety. Summary of the Invention

[0009] In view of the technical challenges of passenger posture detection in the escalator scenario, the present invention provides an escalator safety detection method based on a double-cascaded YOLOv8 architecture.

[0010] An escalator safety detection method based on a double-cascaded YOLOv8 architecture of the present invention includes the following steps:

[0011] Step 1: Construct the Yolov8-Escalator-DET network.

[0012] In the feature extraction stage, first construct the backbone network. By pruning, discard the P5-level features, retain the P2-P4 feature layers to construct a multi-scale perception system, and use the ordinary convolution module Conv for feature dimensionality reduction.

[0013] In the feature fusion stage, design a bidirectional weighted feature pyramid network BiFPN, dynamically adjust the fusion contribution degree of different-scale features through learnable feature layer weight coefficients, and use cross-level skip connections to achieve the interaction and enhancement of high-level semantic features and low-level texture features.

[0014] In the module reconstruction stage, the cross-stage partial connection module CSPPC is designed. By decomposing the standard convolution into a combined structure of grouped depth convolution and pointwise convolution, and combining the partial connection strategy in the channel dimension, 1 / 4 of the channels are retained in the PConv layer for full connection calculation, and skip connections are used for the remaining channels.

[0015] In the detection head optimization stage, a decoupled dual-branch prediction architecture is adopted. The spatial coordinate regression and object classification tasks are processed separately by independent position-sensitive convolutional kernels and class-aware convolutional kernels. A dynamic positive sample assignment strategy is introduced to adaptively adjust the anchor box density according to the feature map resolution. Finally, through a multi-task joint training strategy, combining the DIoU loss function and the Distribution Focal Loss classification loss, object detection in the elevator area is achieved.

[0016] Step 2: Construct the Yolov8-POSE-DET network.

[0017] In the feature extraction stage, first, a feature enhancement joint optimization architecture is constructed for the P2 level in the backbone network: multi-scale feature recombination is achieved by embedding the spatial pyramid depth convolution SPDConv after the P1 layer, the stride of the 3×3 convolution in the second layer is synchronously adjusted to 1 to suppress the loss of small target features, and the fine-grained channel attention mechanism FCA is introduced to generate global-local dual-path feature weight maps.

[0018] At the module reconstruction level, the C2fWTConv and C2fWTConv-REFM modules based on wavelet domain feature enhancement are constructed: the standard convolution is replaced by a 5×5 wavelet convolution WTConv, and the input / output channel constraints are balanced through the channel grouping strategy. The output channels of the first-layer convolution are evenly divided and then cross-group feature fusion is performed through the Bottleneck branch and skip connections respectively. Further, the Bottleneck is upgraded to the receptive field enhancement module REFM, integrating multi-branch dilation convolutional groups with dilation rates = 1, 2, 3; multi-scale context awareness is achieved to obtain the C2fWTConv-REFM module.

[0019] In the feature fusion stage, the adaptive scale feature pyramid ASFP is used to replace the original PAFPN. The high-resolution features of the P2 layer are extracted by SPDConv and bidirectionally weighted fused with the features of the P3-P5 layers, and the 3×3 wavelet convolution WTConv is combined to suppress background interference.

[0020] The detection head optimization inherits the decoupled dual-branch structure, and an SEAM attention mechanism layer is added before each detection head to enhance the learning of human pose features. At the same time, the dynamic WIoUv3 loss function is used to optimize the bounding box regression.

[0021] Step 3: Construct three datasets for the training of two networks: the escalator object detection dataset, the human pose dataset in conventional scenarios, and the human pose dataset in escalator scenarios.

[0022] Among them, the escalator object detection dataset contains 927 self - collected images. All samples are manually annotated, and the annotation category is the "lift" category; the human pose dataset in conventional scenarios contains 4748 images, which are collected from open - source network resources and completed annotation; the human pose dataset in escalator scenarios integrates 4235 images, which are jointly composed of open - source datasets and self - taken images. All data have been processed by standardized annotation; the annotation categories of the two pose datasets are all "fall", "bend", and "fall" categories.

[0023] Step 4: Use the escalator target dataset to train the Yolov8 - Escalator - DET network.

[0024] Randomly divide the training / validation / test sets in a ratio of 8:1:1 for escalator object detection training; implement a two - stage optimization strategy during the training process: freeze the backbone network in the first 50 epochs to retain pre - trained features and accelerate convergence, and then unfreeze the network and perform full - parameter fine - tuning with a cosine - annealing learning rate of 0.012→0.006, combined with 3 epochs of learning rate warm - up and disabling Mosaic data augmentation in the last 5 epochs; by setting an early - stopping monitoring period of 50 epochs and box / classification / depth supervision loss weights of 5.25 / 0.5 / 1.5, achieve a balance between the model convergence speed and generalization performance during 150 epochs of total training.

[0025] Step 5: First, use the human pose dataset in conventional scenarios to train the Yolov8 - POSE - DET network, and then fine - tune the best weights of the trained network using the human pose dataset in escalator scenarios.

[0026] The datasets for both trainings are randomly divided in the ratio of training dataset:validation dataset:test dataset = 8:1:1; the training parameters are the same as those of the Yolov8 - Escalator - DET network.

[0027] Step 6: Performance evaluation.

[0028] Evaluate the two trained models using the val mode. The evaluation includes the following parameter values:

[0029] P: Precision.

[0030] R: Recall.

[0031] F1: Considering both precision and recall comprehensively, the calculation formula is

[0032] mAP50: The mean average precision when the intersection over union (IoU) is 0.5, which is used to measure the comprehensive performance of the model in object detection at this IoU.

[0033] mAP75: The mean average precision when the intersection over union (IoU) is 0.75.

[0034] mAP50-95: The average of the mean average precisions calculated at IoUs from 0.5 to 0.95 with a step of 0.05.

[0035] Step 7: Develop a cascaded inference algorithm using a phased task decoupling strategy.

[0036] It mainly includes two main stages:

[0037] In the first stage, the original image is input into the model trained by the YOLOv8Escalator DET network to identify the region of interest (ROI) of the escalator operation; the model can output key information related to the ROI, including the coordinate positions of the upper left and lower right corners and the corresponding confidence values. Based on these coordinate information, the system will perform localization cropping on the original image to ensure that subsequent processing can focus on the elevator operation area.

[0038] In the second stage, the cropped image is input into the model trained by the high-precision YOLOv8POSE DET network, aiming to obtain the pose information of the passengers; the model will output the coordinates of the passenger pose category boxes and the corresponding confidence values.

[0039] During this process, the "stand" category is classified as "NORMAL", while the "bend" and "fall" categories are classified as "ABNORMAL"; after the output is obtained, the system will convert the coordinates back to the original image coordinate system and draw boxes of different poses in different colors in the original coordinate system for marking.

[0040] Furthermore, the backbone network integrating the fine-grained channel attention mechanism FCA is specifically as follows: First, the high-resolution feature map is multi-scale recombined through the Space-to-Depth transformation layer, and the feature preservation without downsampling is achieved by combining strided convolution; subsequently, the adaptive fine-grained channel attention FCA module is introduced, and its core lies in establishing a dual-path feature interaction mechanism - modeling the global channel correlation through the covariance matrix, and at the same time using the local receptive field convolution to capture the spatial context correlation to form a multi-grained feature weight mapping.

[0041] Furthermore, the construction of the C2fWTConv and C2fWTConv-REFM modules based on wavelet domain feature enhancement is specifically as follows:

[0042] The C2fWTConv and C2fWTConv-RFEM modules are used for architecture-level optimization. Among them, C2fWTConv-RFEM replaces the standard Bottleneck structure with the receptive field enhancement module RFEM to construct a multi-branch dilated convolution group. At the feature transformation level, the frequency domain decomposition method based on Haar wavelet is introduced into the feature preprocessing stage.

[0043] For a given input image X with size N w ×N w , a multi-level frequency domain sub-band is generated through the recursive wavelet transform WT: the low-frequency approximation component X LL , the horizontal high-frequency component X LH , the vertical high-frequency component X HL and the diagonal high-frequency component X HH ; its mathematical representation is:

[0044]

[0045] where i represents the current decomposition level.

[0046] In the implementation of WTConv, after the input feature map X is projected into the frequency domain space by the wavelet basis function, lightweight convolution operations are performed on each sub-band channel respectively, and this process is expressed as:

[0047] Y = IWT(Conv(W, WT(X)))(2)

[0048] where IWT represents the inverse wavelet transform, which recombines the transformed frequency components back into the spatial domain.

[0049] Furthermore, the optimization of the adaptive scale feature pyramid ASFP is specifically as follows:

[0050] The spatial pyramid depth convolution SPDConv is used to perform lightweight feature extraction on the P2 layer, and multi-scale fusion with the features of the P3 layer is achieved through the cascade fusion mechanism BiFPN_Concat3 of the bidirectional feature pyramid network to construct a composite feature space with complete spatial-semantic information.

[0051] The cross-stage partial pyramid convolution CSPPC module constructed by partial convolution completely replaces the original C2f module, while reducing the number of network parameters, maintaining the feature expression ability.

[0052] Furthermore, the formula for optimizing the bounding box regression with the dynamic WIoUv3 loss function is as follows:

[0053] L WIoUv3 = r × L WIoUv1 (3)

[0054]

[0055] L IoU = 1 - IoU ∈ [0, 1] (5)

[0056] Among them, the mapping of the outlier degree β and the gradient gain r is controlled by the hyperparameters α and δ, IoU is the overlap degree between the anchor box and the target box, and L WIoUv1 is the WIoU v1 loss function, which constructs an attention-based bounding box loss, while WIoU v3 additionally attaches a focusing mechanism by constructing a calculation method for the gradient gain on this basis; the formula of the WIoU v1 loss function is as follows:

[0057] L WIoUv1 = R WIoU ×L IoU (6)

[0058]

[0059] Among them, W g , H g is the size of the smallest enclosing box.

[0060] The beneficial technical effects of the present invention are as follows:

[0061] Aiming at the passenger pose detection task in the escalator scenario, the present invention solves the bottlenecks of traditional monitoring technologies in aspects such as background interference, missed detection of small targets, and performance degradation in occlusion scenarios. Through the design of a double-cascade architecture, the detection task is divided into two stages: in the first stage, the lightweight YOLOv8-Escalator-DET network is used to accurately locate the operating area of the escalator, effectively suppressing interference in adjacent areas; in the second stage, the high-precision YOLOv8-POSE-DET network, combined with multi-scale feature fusion and occlusion perception mechanism, completes the fine-grained parsing of the passenger pose. Experimental results show that the present invention significantly improves the detection ability of small targets and occluded targets while ensuring real-time performance. Through the optimization and innovation of the network structure, not only efficient object detection is achieved, but also the real-time computing requirements in resource-constrained environments are met, providing a novel and effective technical solution for the safety monitoring of escalators, and being able to provide strong support for the safety management of urban transportation hubs and commercial complexes. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic flow chart of the escalator safety detection method based on the double-cascade YOLOv8 architecture of the present invention.

[0063] Figure 2 is the architecture diagram of the lightweight improved Yolov8-Escalator-DET network of the present invention.

[0064] Figure 3 This is the high-precision Yolov8-POSE-DET network architecture diagram of the present invention.

[0065] Figure 4 These are the model parameters trained by the two networks of the present invention.

[0066] Figure 5 These are the results of the model performance evaluation of the present invention. Detailed implementation manners

[0067] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0068] An escalator safety detection method based on a dual-cascade YOLOv8 architecture of the present invention adopts a phased task decoupling strategy: in the first stage, the YOLOv8-Escalator-DET network improved by lightweight is used to locate the operation area (ROI) of the escalator, effectively suppressing the interference of adjacent non-elevator areas (such as fixed stairs, floor transition zones); in the second stage, the ROI area is input into the high-precision YOLOv8-POSE-DET network, and the fine-grained parsing of the passenger posture is completed through multi-scale feature fusion and occlusion perception mechanism. This dual-cascade architecture realizes the dual optimization of detection accuracy and real-time performance through a dynamic spatial filtering mechanism and a hierarchical feature focusing strategy: the first-stage network adopts a lightweight design to perform spatial domain suppression on non-elevator areas, significantly reducing the invalid calculation redundancy of the second-stage network; the second-stage network performs fine-grained posture parsing based on the ROI area, while improving the utilization rate of hardware resources, ensuring the detection accuracy of key targets. In order to verify the performance of the two networks in ROI area positioning and human posture detection, the present invention focuses on the complex background of the escalator, the diverse postures of passengers, and occlusion features, and constructs an escalator target dataset, a human posture dataset in ordinary scenes, and a human posture dataset in escalator scenes respectively. The model is trained and evaluated using these three datasets, and a cascade inference algorithm is developed to observe the overall performance of the model.

[0069] The flow of an escalator safety detection method based on a dual-cascade YOLOv8 architecture of the present invention is as Figure 1 shown, and specifically includes the following steps:

[0070] Step 1: Construct the Yolov8-Escalator-DET network.

[0071] The lightweight Yolov8-Escalator-DET network architecture of the present invention is as Figure 2As shown in the figure. This detection network is lightweight improved based on the YOLOv8 framework. Compared with the original YOLOv8 network, the number of parameters is reduced by 77.68%. It focuses on solving the problem of sensitivity to background interference in the human pose detection task in the escalator scenario. Its input is an RGB or BGR three-channel image, and the output is the target detection result (including bounding box coordinates and class confidence) in the escalator monitoring area. In the feature extraction stage, the improved lightweight backbone network adopts a hierarchical feature selection strategy, discards the P5-level features in the original network, and retains the P2-P4 three feature layers for multi-scale feature representation, effectively reducing computational redundancy while ensuring the coverage range of the receptive field. The feature fusion module uses a bidirectional feature pyramid network (BiFPN) to replace the original path aggregation network (PAFPN), and realizes the efficient fusion of high-level semantic information and low-level detail features through a bidirectional cross-scale connection mechanism. To further optimize the model complexity, the present invention reconstructs the C2f module in the original network into a CSPPC module based on cross-stage partial connection. It adopts the CSPNet (Cross Stage Partial Network) architecture design. By replacing the standard convolution with grouped partial convolution (Partial Convolution, PConv), the number of parameters is reduced by 37.6% while maintaining the feature expression ability. Finally, the multi-level optimized P2-P4 feature maps are input into the detection head (Detection Head), and the object localization and classification tasks are completed synchronously through a decoupled prediction mechanism.

[0072] Specific improvements include:

[0073] In the feature extraction stage, first construct the backbone network. Discard the P5-level features through pruning, retain the P2-P4 feature layers to construct a multi-scale perception system, and use the ordinary convolution module Conv for feature dimensionality reduction.

[0074] In the feature fusion stage, design a bidirectional weighted feature pyramid network BiFPN. Dynamically adjust the fusion contribution degree of different scale features through learnable feature layer weight coefficients, and use cross-level skip connections to realize the interactive enhancement of high-level semantic features and low-level texture features.

[0075] In the module reconstruction stage, design a cross-stage partial connection module CSPPC. By disassembling the standard convolution into a combined structure of grouped depth convolution and pointwise convolution, combined with the partial connection strategy in the channel dimension, retain 1 / 4 channels in the PConv layer for full connection calculation, and use skip connections for the remaining channels.

[0076] In the detection head optimization stage, a decoupled dual-branch prediction architecture is adopted. The spatial coordinate regression and object classification tasks are processed separately through independent position-sensitive convolutional kernels and class-aware convolutional kernels. A dynamic positive sample allocation strategy is introduced to adaptively adjust the anchor box density according to the feature map resolution. Finally, through a multi-task joint training strategy, combining the DIoU loss function and the Distribution Focal Loss classification loss, elevator area object detection is achieved.

[0077] Step 2: Construct the Yolov8-POSE-DET network.

[0078] The high-precision Yolov8-POSE-DET network architecture of the present invention is as Figure 3 shown. This network takes the cropped image of the escalator operation area (ROI) positioning coordinates output by the previous-level network as input, and the output is the human pose detection box and class confidence. Aiming at the defects of the original network in the insufficient recognition accuracy of small targets and the performance degradation in occlusion scenarios, the following core improvements are implemented in this architecture:

[0079] (1) In the feature extraction stage, first construct a feature enhancement joint optimization architecture for the P2 layer in the backbone network: realize multi-scale feature recombination by embedding a spatial pyramid depth convolution (SPDConv) after the P1 layer, synchronously adjust the second layer 3×3 convolution stride (stride = 1) to suppress the loss of small target features, and introduce a fine-grained channel attention mechanism FCA to generate a global-local dual-path feature weight map.

[0080] Aiming at the problem that small target features are easily lost in the escalator scenario, the present invention constructs a feature enhancement joint optimization architecture for the P2 layer of the backbone network: first, perform multi-scale recombination on the high-resolution feature map through a Space-to-Depth (SPD) transformation layer, and combine strided convolution to achieve non-downsampled feature preservation; subsequently, an adaptive fine-grained channel attention (FCA) module is innovatively introduced. The core of it is to establish a dual-path feature interaction mechanism - model the global channel correlation through the covariance matrix, and at the same time use the local receptive field convolution to capture the spatial context correlation to form a multi-grained feature weight map. Compared with the traditional fully-connected-based channel attention (such as SENet), FCA has achieved the following breakthroughs through dual-path feature interaction: ① explicit decoupling and fusion of global statistical features and local structural features; ② establish a channel-space joint attention mechanism, which increases the activation intensity of key feature channels by 2.1 times; ③ experiments on the COCO dataset show that this module can only increase 0.8M parameters to improve the small target detection mAP@0.5 by 4.7%.

[0081] (2) At the module reconstruction level, construct the C2fWTConv and C2fWTConv-REFM modules based on wavelet domain feature enhancement: Replace the standard convolution with a 5×5 wavelet convolution (WTConv), and balance the input / output channel constraints through a channel grouping strategy. After the output channels of the first-layer convolution are evenly divided, they are respectively subjected to cross-group feature fusion through the Bottleneck branch and the skip connection. Further upgrade the Bottleneck to the receptive field enhancement module REFM, and integrate multiple branches of dilation convolution groups (dilation rates = 1, 2, 3); achieve multi-scale context awareness to obtain the C2fWTConv-REFM module.

[0082] Aiming at the problem of spatial locality constraints in the complex image feature representation of the YOLOv8 core component C2f module, the present invention proposes a wavelet domain convolution reconstruction scheme: adopt the C2fWTConv and C2fWTConv-RFEM modules for architecture-level optimization. Among them, C2fWTConv-RFEM replaces the standard Bottleneck structure with a receptive field enhancement module (RFEM), constructs multiple branches of dilation convolution groups, and realizes the dynamic adaptation of the multi-scale receptive field of human targets on the premise of only increasing 8.7% of the number of parameters. At the feature transformation level, introduce the frequency domain decomposition method based on Haar wavelet into the feature preprocessing stage.

[0083] For an image X with a given input size of N w ×N w , generate multi-level frequency domain sub-bands through recursive wavelet transform (WT): low-frequency approximation component X LL , horizontal high-frequency component X LH , vertical high-frequency component X HL and diagonal high-frequency component X HH ; its mathematical representation is:

[0084]

[0085] where i represents the current decomposition level. The multi-resolution analysis framework constructed by cascaded wavelet transform can effectively capture the global contour (low frequency) and local detail (high frequency) features of the target object.

[0086] In the implementation of WTConv, after the input feature map X is projected into the frequency domain space by the wavelet basis function, lightweight convolution operations are performed on each sub-band channel respectively, and this process is expressed as:

[0087] Y = IWT(Conv(W, WT(X)))(2)

[0088] Among them, IWT represents the inverse wavelet transform, which recombines the transformed frequency components back into the spatial domain. This design enables the network to collaboratively learn in both the spatial-frequency dual domain, enhancing the model's sensitivity to fine-grained features such as clothing textures and limb edges.

[0089] (3) In the feature fusion stage, the Adaptive Scale Feature Pyramid (ASFP) is used to replace the original PAFPN. The high-resolution features (640×640) of the P2 layer are extracted through SPDConv and bidirectionally weighted fused with the features of the P3 - P5 layers. The 3×3 wavelet convolution WTConv is combined to suppress background interference.

[0090] The ASFP architecture optimizes performance by deeply integrating the P2 and P3 level features in the backbone network. Among them, the P2 layer, with its high spatial resolution characteristics, effectively retains the key detail features required for small target detection; while the P3 layer, while maintaining a moderate resolution, efficiently represents medium-scale targets through rich context semantic information. To balance computational resources and detection accuracy, the present invention uses Spatial Pyramid Depth Convolution (SPDConv) for lightweight feature extraction of the P2 layer, and realizes multi-scale fusion with the features of the P3 layer through the cascaded fusion mechanism of the Bidirectional Feature Pyramid Network (BiFPN_Concat3), constructing a composite feature space with complete spatial-semantic information. This design maintains the sensitivity of small target detection while reducing computational complexity. To further optimize the model efficiency, the present invention uses the Cross-Stage Partial Pyramid Convolution (CSPPC) module constructed by partial convolution to completely replace the original C2f module, maintaining the feature expression ability while reducing the number of network parameters.

[0091] (4) The detection head optimization inherits the decoupled double-branch architecture, and a SEAM attention mechanism layer is added before each detection head to enhance the learning of human pose features. At the same time, the dynamic WIoUv3 loss function is used to optimize the bounding box regression.

[0092] The SEAM attention mechanism (Scale-Aware Equalized Attention Mechanism) is an attention mechanism proposed by the YOLO-Face model. It can effectively learn the dependence relationship between features and enhance the feature expression ability. The present invention adds a SEAM attention mechanism layer before each detection head to enhance the learning of human pose features, especially for handling occlusion situations.

[0093] The specific optimization of the Adaptive Scale Feature Pyramid ASFP is as follows:

[0094] The spatial pyramid depth convolution SPDConv is used to perform lightweight feature extraction on the P2 layer, and the multi-scale fusion of features with the P3 layer is achieved through the cascaded fusion mechanism BiFPN_Concat3 of the bidirectional feature pyramid network, constructing a composite feature space with complete spatial-semantic information.

[0095] The cross-stage partial pyramid convolution CSPPC module constructed by partial convolution completely replaces the original C2f module, maintaining the feature expression ability while reducing the number of network parameters.

[0096] Furthermore, the formula for optimizing the bounding box regression with the dynamic WIoUv3 loss function is as follows:

[0097] L WIoUv3 = r × L WIoUv1 (3)

[0098]

[0099] L IoU = 1 - IoU ∈ [0,1] (5)

[0100] Among them, the mapping of the outlier degree β and the gradient gain r is controlled by the hyperparameters α and δ, IoU is the overlap degree between the anchor box and the target box, and L WIoUv1 is the WIoU v1 loss function, which constructs an attention-based bounding box loss, while WIoU v3 adds a focusing mechanism on this basis by constructing a calculation method for the gradient gain; the formula for the WIoU v1 loss function is as follows:

[0101] L WIoUv1 = R WIoU × L IoU (6)

[0102]

[0103] Among them, W g , H g is the size of the smallest enclosing box.

[0104] Step 3: Construct three datasets for the training of the two networks: the escalator target detection dataset, the regular scene human pose dataset, and the escalator scene human pose dataset.

[0105] Among them, the escalator target detection dataset contains 927 self - collected images (taken in the escalator scenes of subway stations and shopping malls). All samples are manually annotated, and the annotation category is the "lift" category; the conventional scene human pose dataset contains 4748 images, which are collected from open - source network resources and annotated; the escalator scene human pose dataset integrates 4235 images, which are composed of open - source datasets and self - taken images. All data have been processed with standardized annotation. The annotation categories of the two pose datasets are all "fall", "bend", and "fall" (it seems there is a repetition in the original text for the third category, but translated as is).

[0106] Step 4: Use the escalator target dataset to train the Yolov8 - Escalator - DET network.

[0107] Randomly divide the training / validation / test sets in a ratio of 8:1:1 for escalator target detection training; implement a two - stage optimization strategy during the training process: freeze the backbone network in the first 50 rounds to retain pre - trained features and accelerate convergence, and then unfreeze the network and perform full - parameter fine - tuning with a cosine - annealing learning rate that decays from 0.012 to 0.006, combined with 3 rounds of learning rate warm - up and disabling Mosaic data augmentation in the last 5 rounds; by setting an early - stopping monitoring period of 50 rounds and box / classification / depth supervision loss weights of 5.25 / 0.5 / 1.5, balance the model convergence speed and generalization performance during 150 rounds of total training.

[0108] Step 5: First, use the conventional scene human pose dataset to train the Yolov8 - POSE - DET network, and then fine - tune the best weights of the trained network using the escalator scene human pose dataset.

[0109] The datasets for both trainings are randomly divided in the ratio of training dataset:validation dataset:test dataset = 8:1:1; the training parameters are the same as those of the Yolov8 - Escalator - DET network.

[0110] Step 6: Performance evaluation.

[0111] The parameters of the two trained models are as Figure 4 shown. Evaluate the two trained models using the val mode. The evaluation includes the following parameter values:

[0112] P: Precision.

[0113] R: Recall.

[0114] F1: Considering both precision and recall comprehensively, the calculation formula is

[0115] mAP50: The mean average precision when the intersection - over - union is 0.5, which is used to measure the comprehensive performance of the model for target detection at this intersection - over - union ratio.

[0116] mAP75: The mean average precision when the intersection over union (IoU) is 0.75, which has a higher requirement for the accuracy of target localization compared to mAP50.

[0117] mAP50-95: The average of the mean average precisions calculated with an IoU step of 0.05 from 0.5 to 0.95, which can more comprehensively reflect the overall performance of the model under different degrees of overlap.

[0118] Through a comprehensive analysis of these parameters, it is possible to gain an in-depth understanding of the performance of the two models in aspects such as accuracy, recall ability, and the comprehensive object detection performance under different intersection over union ratios, providing a basis for the further optimization or selection of the models.

[0119] The evaluation results of the above indicators in this embodiment are as Figure 5 shown, where the lift category is recognized by the model trained by the Yolov8-Escalator-DET network (hereinafter referred to as the Yolov8-Escalator-DET model), and stand, bend, and fall are recognized by the model trained by the Yolov8-POSE-DET network (hereinafter referred to as the Yolov8-POSE-DET model). Comprehensive Figure 4 and Figure 5 , it can be analyzed that:

[0120] Yolov8-Escalator-DET model (detecting the lift category): In terms of the model structure, the number of parameters of this model is only 0.7M, and the size is 1.6M, belonging to a lightweight model, suitable for resource-constrained environments, with a moderate computational load (8.6 GFLOPs), suitable for real-time systems. In terms of evaluation indicators, the model performs extremely well in the detection of the lift category. In particular, the recall rate R reaches 1, indicating that almost all positive samples are correctly detected. The precision P is close to 0.98, and the false detection rate is low. The F1 score is 0.989, indicating a good balance between P and R. The mAP50 is as high as 0.995, indicating very accurate detection when IoU = 50%. Although mAP75 and mAP50-95 have decreased, for large targets such as "escalators", the positioning accuracy is sufficient. Overall, the overall performance of the Yolov8-Escalator-DET model is good.

[0121] Yolov8-POSE-DET model (detecting three categories: stand, bend, and fall): In terms of model parameters, this model has 315 layers and belongs to a deep structure, which can achieve the extraction of complex pose features. The model size is moderate (3.98M). Although the number of parameters (1.97M) is large, the computational cost (GFLOPs 5.5) is lower instead. It has good structural optimization and efficient computing characteristics, making it suitable for scenarios that require high accuracy but have available computing resources. In terms of evaluation metrics, all metrics for the stand and fall categories are high, showing excellent performance. In particular, mAP50 reaches 0.966, indicating accurate detection under a loose IoU, and mAP50-95 is 0.705, also showing relatively high localization accuracy. The metrics for the bend category are lower among all categories. This is because when annotating the dataset in the present invention, since the bend pose is between standing and falling and is not easy to define, unclear poses are classified into the bend category, resulting in a lower accuracy of the trained model in recognizing the bend pose. Generally speaking, the overall detection results of the Yolov8-POSE-DET model are relatively reliable.

[0122] Step 7: Develop a cascaded inference algorithm using a phased task decoupling strategy.

[0123] It mainly includes two main stages:

[0124] In the first stage, the original image is input into the model trained by the YOLOv8Escalator DET network to identify the Region of Interest (ROI) of the escalator operation; the model can output key information related to the ROI, including the coordinate positions of the upper left and lower right corners and the corresponding confidence values. Based on these coordinate information, the system will perform a localization cropping on the original image to ensure that subsequent processing can focus on the elevator operation area.

[0125] In the second stage, the cropped image is input into the model trained by the high-precision YOLOv8POSE DET network, aiming to obtain the pose information of the passengers; the model will output the coordinates (based on the coordinate system of the cropped image) of the passenger pose category boxes and the corresponding confidence values.

[0126] During this process, the "stand" category is classified as "NORMAL", while the "bend" and "fall" categories are classified as "ABNORMAL"; after the output is obtained, the system will convert the coordinates back to the original image coordinate system and draw boxes of different poses in different colors for marking in the original coordinate system.

[0127] Since the elevator operation area is in a fixed state in the monitoring system, the traditional approach of requiring elevator area recognition for each frame of the image will bring a large computational burden. The present invention adopts a dynamic update strategy based on local optimality, that is, it determines whether the current ROI area frame is better based on the recognition frame detected historically. If it is better, it will be updated, otherwise it will not be updated. Usually, the optimal area can be obtained in the first 30 frames, and there is no need to use the YOLOv8Escalator DET network in the future. The original image can be cut directly using the saved optimal ROI area coordinates, thereby significantly reducing the computational complexity and improving the frame rate.

[0128] Experimental verification shows that on a hardware platform equipped with NVIDIA GeForce RTX 4060 GPU, the present invention can achieve an average real-time processing frame rate of 30FPS. Its performance index has exceeded the 24FPS threshold required by the visual persistence characteristics of the human eye, and fully meets the requirements of the real-time video processing system for picture smoothness. The two improved models of the present invention can effectively achieve a good balance between real-time performance and detection accuracy, and are suitable for real-time monitoring systems with limited resources.

Claims

1. An escalator safety detection method based on a dual-cascade YOLOv8 architecture, characterized in that: The following steps are involved: Step 1: Build the Yolov8-Escalator-DET network; In the feature extraction stage, we first build a backbone network, discard the P5-level features through pruning, retain the P2-P4 feature layers to build a multi-scale perception system, and use the ordinary convolution module Conv to reduce the feature dimension; In the feature fusion stage, a bidirectional weighted feature pyramid network BiFPN is designed to dynamically adjust the fusion contribution of features of different scales through learnable feature layer weight coefficients, and use cross-level skip connections to achieve interactive enhancement of high-level semantic features and low-level texture features. In the module reconstruction stage, a cross-stage partial connection module CSPPC is designed. By decomposing the standard convolution into a combination of grouped depth convolution and point-by-point convolution, combined with the partial connection strategy in the channel dimension, 1 / 4 channels are reserved in the PConv layer for full connection calculation, and the remaining channels use skip connections; In the detection head optimization stage, a decoupled dual-branch prediction architecture is adopted. Independent position-sensitive convolution kernels and category-aware convolution kernels are used to process spatial coordinate regression and target classification tasks respectively. A dynamic positive sample allocation strategy is introduced to adaptively adjust the anchor box density according to the feature map resolution. Finally, a multi-task joint training strategy is used to combine the DIoU loss function and the Distribution Focal Loss classification loss to achieve elevator area target detection. Step 2: Build the Yolov8-POSE-DET network; In the feature extraction stage, we first construct a feature enhancement joint optimization architecture for the P2 layer in the backbone network: by embedding the spatial pyramid deep convolution SPDConv after the P1 layer to achieve multi-scale feature reorganization, and simultaneously adjust the stride of the second layer 3×3 convolution stride=1 to suppress the loss of small target features, and introduce the fine-grained channel attention mechanism FCA to generate a global-local two-way feature weight mapping; At the module reconstruction level, we construct C2fWTConv and C2fWTConv-REFM modules based on wavelet domain feature enhancement: replace the standard convolution with 5×5 wavelet convolution WTConv, balance the input / output channel constraints through the channel grouping strategy, where the first-layer convolution output channels are evenly divided and then cross-group feature fusion is performed through the Bottleneck branch and jump connection; further upgrade the Bottleneck to the receptive field enhancement module REFM, integrate multi-branch dilated convolution groups with dilation rates = 1, 2, 3; realize multi-scale context perception, and obtain the C2fWTConv-REFM module; In the feature fusion stage, the adaptive scale feature pyramid ASFP is used to replace the original PAFPN. The high-resolution features of the P2 layer are extracted through SPDConv, and bidirectional weighted fusion is performed with the features of the P3-P5 layers. The 3×3 wavelet convolution WTConv is combined to suppress background interference. The detection head optimization inherits the decoupled dual-branch architecture and adds a SEAM attention mechanism layer before each detection head to enhance the learning of human posture features. At the same time, the dynamic WIoUv3 loss function is used to optimize the bounding box regression. Step 3: Construct three datasets for training the two networks: escalator object detection dataset, regular scene human posture dataset, and escalator scene human posture dataset; Among them, the escalator object detection dataset contains 927 self-collected images, all samples are manually annotated, and the annotated category is lift; the conventional scene human posture dataset contains 4748 images, which are collected from open source network resources and annotated; the escalator scene human posture dataset integrates 4235 images, which are composed of open source datasets and self-photographed images, and all data are processed by standardized annotation; the annotated categories of the two posture datasets are fall, bend, and fall; Step 4: Train the Yolov8-Escalator-DET network using the escalator target dataset; The training / validation / test sets were randomly divided into 8:1:1 ratios for escalator object detection training. A two-stage optimization strategy was implemented during the training process: the backbone network was frozen for the first 50 rounds to retain the pre-training features and accelerate convergence, and then the network was unfrozen and all parameters were fine-tuned using a cosine decay learning rate of 0.012→0.006, with 3 rounds of learning rate warm-up and Mosaic data enhancement disabled for the last 5 rounds. By setting 50 rounds of early stopping monitoring cycles and box / classification / depth supervision loss weights of 5.25 / 0.5 / 1.5, a balance between model convergence speed and generalization performance was achieved in a total of 150 rounds of training. Step 5: First, train the Yolov8-POSE-DET network using the human posture dataset of the conventional scene, and then fine-tune the trained optimal weights using the human posture dataset of the escalator scene; The datasets for both trainings were randomly divided into training dataset: validation dataset: test dataset = 8:1:1 ratio; the training parameters were the same as those of the Yolov8-Escalator-DET network; Step 6: Performance evaluation; Use the val mode to evaluate the two trained models. The evaluation includes the following parameter values: P: accuracy; R: recall rate; F1: Comprehensively considers the accuracy and recall rate, and the calculation formula is: mAP50: The average precision when the intersection-over-union ratio is 0.5, which is used to measure the comprehensive performance of the model for target detection under this intersection-over-union ratio; mAP75: Mean average precision when the intersection-over-union ratio is 0.75; mAP50-95: the average of the mean average precisions calculated from the intersection-over-union ratio from 0.5 to 0.95 with a step size of 0.05; Step 7: Develop a cascade reasoning algorithm using a staged task decoupling strategy; It mainly consists of two main stages: In the first stage, the original image is input into the model trained by the YOLOv8 Escalator DET network to identify the escalator operation area ROI; the model can output key information related to the ROI, including the coordinate positions of the upper left corner and the lower right corner and the corresponding confidence values. Based on these coordinate information, the system will positionally crop the original image to ensure that subsequent processing can focus on the elevator operation area; In the second stage, the cropped images are fed into a model trained by a high-precision YOLOv8 POSE DET network to obtain the passenger’s posture information; the model outputs the coordinates of the passenger’s posture category box and the corresponding confidence value; During this process, the "stand" category is classified as "NORMAL", while the "bend" and "fall" categories are classified as "ABNORMAL"; after the output is obtained, the system converts the coordinates back to the original image coordinate system and draws boxes of different postures in different colors in the original coordinate system for marking.

2. According to claim 1, an escalator safety detection method based on a dual-cascade YOLOv8 architecture is characterized in that: The backbone network integrates the fine-grained channel attention mechanism FCA as follows: first, the high-resolution feature map is reorganized at multiple scales through the Space-to-Depth transformation layer, and the strided convolution is combined to achieve feature preservation without downsampling; then the adaptive fine-grained channel attention FCA module is introduced, the core of which is to establish a two-way feature interaction mechanism - modeling the global channel correlation through the covariance matrix, and using the local receptive field convolution to capture the spatial context association, forming a multi-granularity feature weight mapping.

3. According to claim 1, an escalator safety detection method based on a dual-cascade YOLOv8 architecture is characterized in that: The C2fWTConv and C2fWTConv-REFM modules constructed based on wavelet domain feature enhancement are specifically as follows: The C2fWTConv and C2fWTConv-RFEM modules are used for architecture-level optimization. C2fWTConv-RFEM replaces the standard Bottleneck structure with the receptive field enhancement module RFEM to construct a multi-branch dilated convolution group. At the feature transformation level, the frequency domain decomposition method based on Haar wavelet is introduced into the feature preprocessing stage. For a given input size N w ×N w The image X is transformed by recursive wavelet transform WT to generate multi-level frequency domain subbands: low-frequency approximate component X LL , horizontal high frequency component X LH , vertical high frequency component X HL And the diagonal high frequency component X HH ; Its mathematical representation is: Where i represents the current decomposition level; In the WTConv implementation, after the input feature map X is projected into the frequency domain space by the wavelet basis function, each sub-band channel performs a lightweight convolution operation. The process is expressed as: Y=IWT(Conv(W,WT(X)))(2) where IWT stands for inverse wavelet transform, which recombines the transformed frequency components back into the spatial domain.

4. The escalator safety detection method based on a dual-cascade YOLOv8 architecture according to claim 1 is characterized in that: The adaptive scale feature pyramid ASFP optimization is specifically as follows: The spatial pyramid deep convolution SPDConv is used to extract lightweight features from the P2 layer, and the multi-scale fusion with the P3 layer features is achieved through the cascade fusion mechanism of the bidirectional feature pyramid network BiFPN Concat3 to construct a composite feature space with complete spatial-semantic information; The cross-stage partial pyramid convolution CSPPC module constructed using partial convolution completely replaces the original C2f module, reducing the number of network parameters while maintaining the feature expression capability.

5. The escalator safety detection method based on the dual cascade YOLOv8 architecture according to claim 1 is characterized in that: The formula for optimizing bounding box regression using the dynamic WIoUv3 loss function is as follows: L WIoUv3 =r×L WIoUv1 (3) L IoU =1-IoU∈[0,1] (5) Among them, the mapping between outlier degree β and gradient gain r is controlled by hyperparameters α and δ, IoU is the overlap between the anchor box and the target box, and L WIoUv1 is the WIoU v1 loss function, which constructs an attention-based bounding box loss, while WIoU v3 adds a focusing mechanism by constructing a gradient gain calculation method on this basis; the WIoU v1 loss function formula is as follows: L WIoUv1 =R WIoU ×L IoU (6) Among them, W g , H g is the size of the smallest enclosing box.

Citation Information

Cited By

  • Pavement defect analysis and detection method and system

    CN120563521A

  • Forestry remote sensing data ontology feature recognition method and system for blockchain right confirmation

    CN120874126A

  • Fabric quality grade evaluation method and system based on surface defects

    CN121169895A