Method and system for detecting foreign matters and defects of comb plate of escalator
By improving the YOLOV5 object detection algorithm, combining the channel and space attention modules and Shape-IoU and NWD loss functions, the problems of high false alarm rate and high miss detection of small object detection of escalator comb plates are solved, and a higher accuracy and robust detection effect is achieved.
Patent Information
- Application Number
- CN202510602338.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-15
AI Technical Summary
The existing escalator comb plate detection technology has problems such as high false alarm rate, high miss detection rate and insufficient model generalization ability in small target detection, especially the poor recognition effect of comb plate defects and foreign objects.
The improved YOLOV5 object detection algorithm is adopted to form the SPPFA module by adding channel attention and spatial attention to the SPPF module of the backbone network, and weighting it with Shape-IoU, NWD and IoU loss functions to improve the model's capture ability and anti-interference ability of key features and improve the small object detection effect.
The detection accuracy and robustness of foreign objects and defects of the escalator comb tooth plate are significantly improved, the false alarm rate is reduced, the positioning accuracy of irregular defects and small targets is enhanced, and the generalization ability of the model is improved.
Smart Images

Figure CN120495596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of escalator abnormality monitoring, in particular to a method and system for detecting foreign matter and defects on an escalator comb plate. Background Art
[0002] The comb plate of an escalator is the core safety component that connects the step treads to the floor. Its structural integrity and operational reliability are key to ensuring passenger safety and efficient equipment operation. If the comb plate has defects such as tooth groove fracture, deformation and wear, or foreign objects such as gravel and coins are stuck, it may cause multiple levels of risks: at the micro level, it may cause misalignment, increased equipment vibration and abnormal noise; at the macro level, it may cause major safety accidents such as embedded shoes and limbs, and even cause systemic mechanical damage. Therefore, comb plate defect and foreign object detection technology is not only an important barrier to public safety, but also a strategic need to extend the service life of equipment and promote the construction of intelligent operation and maintenance systems.
[0003] The current mainstream visual inspection technology mainly presents a two-track development path: the first is based on traditional image processing methods, which realize defect recognition through median filtering noise reduction combined with edge detection, line matching and other algorithms. Although it has the advantage of real-time performance, it has problems such as high missed detection rate of small features and fluctuating false alarm rate. The second is an intelligent detection solution based on deep learning, which relies on the target detection network to achieve high-precision positioning. However, it is limited by the micro-scale characteristics of comb plate defects (usually less than 2mm) and the scarcity of industry-specific data sets. In actual deployment, it faces the challenge of insufficient model generalization ability. For example, the existing target detection algorithm does not pay enough attention to key position information and cannot effectively use limited computing resources to focus on the most important features. It is generally ineffective for detecting small targets and targets with complex and changeable shapes, which can easily lead to false alarms or missed alarms. Summary of the Invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a method and system for detecting foreign matter and defects on escalator comb plates that can improve the detection effect of small targets.
[0005] Technical solution: The method for detecting foreign matter and defects in escalator comb plates of the present invention comprises the following steps: constructing a sample data set, wherein the sample data set includes images and annotation information of broken teeth of the comb plate, and images and annotation information of foreign matter;
[0006] Training a detection model using the sample data set;
[0007] Acquire an image of the comb plate to be inspected, identify broken teeth and foreign objects using the trained detection model, and obtain the categories and bounding boxes of the broken teeth and / or foreign objects, which are used to calibrate the positions of the broken teeth and / or foreign objects in the image of the escalator comb plate;
[0008] The detection model uses YOLOV5 as the framework, including a backbone network, a neck network and a detection head. In the backbone network, channel attention and spatial attention are added to the SPPF module to form an SPPFA module;
[0009] The SPPFA module adds a channel attention module after each pooling layer of the SPPF module. The output of each channel attention module is spliced and then input into the spatial attention module. After one convolution layer, the output of the SPPFA module is obtained.
[0010] Furthermore, the loss function of the detection model is a weighted sum of the Shape-IoU loss function, the NWD loss function, and the IoU loss function.
[0011] Furthermore, the channel attention module pools the input feature map F of size H×W×C to obtain two feature maps of size 1×1×C, which are added and activated after passing through the fully connected layer to obtain the channel feature weight M c , and finally M c Multiplying with F gets the output of the channel attention module.
[0012] Furthermore, the spatial attention module pools the input feature map F of size H×W×C to obtain two feature maps of size H×W×1, which are concatenated in the channel dimension and then subjected to convolution and pooling operations to obtain the spatial feature weight M. s , and finally M s Multiplying with F gets the output of the spatial attention module.
[0013] The escalator comb plate foreign matter and defect detection system of the present invention comprises:
[0014] A model construction unit is used to establish a detection model. The detection model uses YOLOV5 as a framework and includes a backbone network, a neck network, and a detection head. In the backbone network, channel attention and spatial attention are added to the SPPF module to form an SPPFA module. The SPPFA module adds a channel attention module after each pooling layer of the SPPF module. The output of each channel attention module is spliced and then input into the spatial attention module. After passing through a convolutional layer, the output of the SPPFA module is obtained.
[0015] A model training unit is used to construct a sample data set, wherein the sample data set includes images and annotation information of broken teeth of a comb plate, and images and annotation information of foreign objects; and train a detection model using the sample data set;
[0016] The escalator comb plate foreign body and defect detection unit is used to collect the comb plate image to be detected, use the trained detection model to identify broken teeth and foreign objects, obtain the category and bounding box of the broken teeth and / or foreign objects, and calibrate the position of the broken teeth and / or foreign objects in the escalator comb plate image.
[0017] Furthermore, the loss function of the detection model is a weighted sum of the Shape-IoU loss function, the NWD loss function, and the IoU loss function.
[0018] Furthermore, the channel attention module pools the input feature map F of size H×W×C to obtain two feature maps of size 1×1×C, which are added and activated after passing through the fully connected layer to obtain the channel feature weight M c , and finally M c Multiplying with F gets the output of the channel attention module.
[0019] Furthermore, the spatial attention module pools the input feature map F of size H×W×C to obtain two feature maps of size H×W×1, which are concatenated in the channel dimension and then subjected to convolution and pooling operations to obtain the spatial feature weight M. s , and finally M s Multiplying with F gets the output of the spatial attention module.
[0020] The electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the computer program is loaded into the processor, the method for detecting foreign matter and defects in the escalator comb plate is implemented.
[0021] The computer-readable storage medium of the present invention stores a computer program, and when the computer program is executed by a processor, the method for detecting foreign matter and defects in the escalator comb plate is implemented.
[0022] Beneficial effects: Compared with the prior art, the advantages of the present invention are: (1) The present invention improves the SPPF module in the YOLOV5 target detection algorithm through spatial attention and channel attention, thereby improving the model's ability to capture key features and anti-interference ability, and improving the small target detection effect; among them, the channel attention mechanism can suppress redundant channels (such as background noise) and strengthen key channels related to the task (such as the target's contour and texture) by analyzing the importance of each channel feature; the spatial attention mechanism locates important areas (such as defect locations and foreign body contours) in the spatial dimension, weakens the interference of irrelevant areas, and ultimately improves the model's ability to capture key features and anti-interference ability. (2) The present invention introduces Shape-IoU, NWD and YOLOV5's own IoU loss function for weighting. By combining shape refinement modeling with distribution distance measurement, the limitations of traditional IoU in complex shape and small target scenes are solved, and the robustness and accuracy of the detector are significantly improved. In order to break through the rectangular box constraint, enhance the sensitivity to deformed targets, and alleviate the optimization deviation of large and small targets, the Shape-IoU loss function is used to improve the model's positioning accuracy for irregular defects (such as wear and deformation) and reduce the false alarm rate. In order to solve the IoU failure problem of small targets, improve the anti-interference ability of boundary fuzzy, and enhance the generalization ability across scales, the NWD loss function is used to provide a smooth gradient, which is conducive to ensuring the convergence stability of small targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a diagram of the detection model architecture of the present invention.
[0024] Figure 2 This is a diagram of the SPPFA module architecture of the present invention.
[0025] Figure 3 This is a diagram of the channel attention module architecture of an embodiment of the present invention.
[0026] Figure 4 This is a diagram of the spatial attention module architecture of an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0028] like Figure 1 As shown in the figure, the present invention improves the YOLOV5 target detection network structure, and integrates the channel attention module (Cam) and spatial attention module (Sam) mechanisms between the Maxpool layer and the Concat layer and between the Concat layer and the convolution layer in the SPPF structure of the backbone network to form an SPPFA structure. Figure 2As shown in the figure, the SPPFA module adds a channel attention mechanism after each maximum pooling layer in the SPPF module structure, and then fuses the features that have not passed through the pooling layer, passed through the pooling layer once + Cam, passed through the pooling layer twice + Cam, and passed through the pooling layer three times + Cam through the Concat layer, and then passes through the spatial attention mechanism and one more convolution to obtain the output of the SPPFA module.
[0029] The channel attention and spatial attention used in this embodiment are from the CBAM attention mechanism. In practice, other channel attention and spatial attention mechanisms can also be used, for example, the SENet channel attention mechanism can be used.
[0030] like Figure 3 As shown in the figure, the channel attention mechanism first performs global maximum pooling and global average pooling on the input feature map F. The size of the input feature map F changes from the original H×W×C to two 1×1×C feature maps. Then the two feature maps are input into two fully connected layers MLP, and finally two 1×1×C feature maps are output. Then they are added and the Sigmoid activation function is used to limit their values between 0 and 1 to obtain a channel feature weight M with a size of 1×1×C. c Finally, the input feature map needs to be combined with M c The output of the Cam module is obtained by multiplication. The whole process can be expressed by equations (1) and (2), where σ represents the Sigmoid activation function:
[0031] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))) (1)
[0032] Cam(F)=F×M c (F) (2)
[0033] like Figure 4 As shown in the figure, the spatial attention mechanism is similar to the channel attention mechanism in that it also performs global maximum pooling and global average pooling on the input feature map, but the difference is that it is performed in the channel dimension. The feature map F' with an input size of H×W×C is subjected to global maximum pooling and global average pooling to obtain two feature maps with a size of H×W×1. The two feature maps are then spliced in the channel dimension to obtain a feature map of H×W×2. Then, a 7×7 convolution operation is performed to change the size of the feature map to H×W×1. Finally, the Sigmoid activation function is also used to limit its value between 0 and 1 to obtain a spatial feature weight M with a size of H×W×1. s Finally, we need to combine the input feature map with M s The output result of the Sam module is obtained by multiplication. The whole process can be expressed by equations (3) and (4), where f7×7 Indicates that a convolution kernel of size 7×7 is used:
[0034] M s (F')=σ(f 7×7 ([AvgPool(F');MaxPool(F')])) (3)
[0035] Sam(F')=F'×M c (F') (4)
[0036] In addition, considering that the contour shape of the escalator comb plate defect has great uncertainty, and the IoU loss function has great limitations in the scenarios of complex shape matching and small target detection, the present invention performs a weighted combination of Shape-IoU, NWD and YOLOV5's own IoU loss function to enhance the model's detection capabilities for irregularly shaped comb plate defects and comb plate inclusions of small foreign objects. Among them, Shape-IoU introduces shape information (such as contours, key points or aspect ratios) on the basis of the traditional IoU loss function, pays more attention to the alignment of the target's geometric structure, reduces missed detection or false detection of irregular or fuzzy shaped targets, and reduces sensitivity to slight boundary offsets, avoiding evaluation biases caused by shape asymmetry. NWD models the detection box as a Gaussian distribution, measures similarity by distribution distance, is more robust to slight position deviations of small targets, avoids the gradient vanishing problem of IoU, and in small target scenarios, even if the predicted box does not overlap with the true box, NWD can still provide effective gradient updates.
[0037] In the YOLOV5 target detection algorithm, the IoU loss function is defined as shown in formula (5), where B and B gt Represent the model prediction box and the true box respectively:
[0038]
[0039] The definition of Shape-IoU loss function is as follows, where w in equations (6) and (7) is gt and h gt represents the width and height of the real box, w and h represent the width and height of the anchor box, and scale is the scale factor. c 、y c 、 Represent the center point coordinates of the anchor box and the ground truth box, c is the diagonal length of the minimum bounding box, ww and hh are the shape weight coefficients from equations (6) and (7). w With ω his the shape difference weight of width and height, θ=4 is an exponential parameter used to control the sensitivity of the penalty term (the larger the value, the stronger the penalty for the difference). Formula (11) is the final loss function, where 1-IoU is the basic loss term, distance shape is the shape-aware center point distance penalty term, 0.5×Ω shape is the weighted value of the shape similarity penalty term (the coefficient 0.5 is a balance parameter).
[0040]
[0041]
[0042] The definition of NWD is as follows, where weight = 2 and C is a coefficient associated with the data set, in this embodiment C = 3.
[0043]
[0044] The Shape-IoU, NWD and IoU loss functions are weighted, and the final weighted loss function Loss is shown in formula (14), where α and β are the weights of the loss function. In this embodiment, α = 0.1 and β = 0.4.
[0045] Loss = α(1-IoU) + β(1-L Shape-IoU )+(1-α-β)(1-NWD) (14)
[0046] The method of the present invention is verified by specific experiments below.
[0047] In this experiment, images of defects and foreign matter on the comb plates at the entrance and exit of escalators were obtained as initial data input. The images were then rotated and other operations were performed to further enhance the dataset, which consisted of 1,214 images in the training set and 339 images in the validation set. After adjusting the network structure and loss function, a comparative experiment was conducted to verify the effectiveness of the improvement. Four models were trained from scratch using the same dataset and training parameters. The experimental results are shown in Table 1. The four models are: the original YOLOV5 model; a model obtained by replacing the SPPF in YOLOV5 with SPPFA; a model obtained by replacing the IoU loss function in YOLOV5 with the Loss loss function of the present invention; and a model obtained by replacing the SPPF in YOLOV5 with SPPFA and the IoU loss function in YOLOV5 with the Loss loss function of the present invention.
[0048] Table 1 Comparative experimental results
[0049]
[0050]
[0051] The experimental results in Table 1 show that the detection model of the present invention improves the mAP_0.5 index from 87.54% to 90.10%, showing higher overall detection accuracy.
[0052] The escalator comb plate foreign matter and defect detection system of the present invention comprises:
[0053] A model construction unit is used to establish a detection model. The detection model uses YOLOV5 as a framework and includes a backbone network, a neck network, and a detection head. In the backbone network, channel attention and spatial attention are added to the SPPF module to form an SPPFA module. The SPPFA module adds a channel attention module after each pooling layer of the SPPF module. The output of each channel attention module is spliced and then input into the spatial attention module. After passing through a convolutional layer, the output of the SPPFA module is obtained.
[0054] A model training unit is used to construct a sample data set, wherein the sample data set includes images and annotation information of broken teeth of a comb plate, and images and annotation information of foreign objects; and train a detection model using the sample data set;
[0055] The escalator comb plate foreign body and defect detection unit is used to collect the comb plate image to be detected, use the trained detection model to identify broken teeth and foreign objects, obtain the category and bounding box of the broken teeth and / or foreign objects, and calibrate the position of the broken teeth and / or foreign objects in the escalator comb plate image.
[0056] The electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the computer program is loaded into the processor, the method for detecting foreign matter and defects in the escalator comb plate is implemented.
[0057] The computer-readable storage medium of the present invention stores a computer program, and when the computer program is executed by a processor, the method for detecting foreign matter and defects in the escalator comb plate is implemented.
[0058] The computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store program code in the form of instructions or data structures and that can be accessed by a computer.
[0059] The processor is configured to execute the computer program stored in the memory to implement the various steps in the method involved in the above embodiment.
Claims
1. A method for detecting foreign matter and defects in escalator comb plates, characterized in that: The method comprises the following steps: constructing a sample data set, wherein the sample data set includes images and annotation information of broken teeth of a comb plate, and images and annotation information of foreign matter; Training a detection model using the sample data set; Acquire an image of the comb plate to be inspected, identify broken teeth and foreign objects using the trained detection model, and obtain the categories and bounding boxes of the broken teeth and / or foreign objects, which are used to calibrate the positions of the broken teeth and / or foreign objects in the image of the escalator comb plate; The detection model uses YOLOV5 as the framework, including a backbone network, a neck network and a detection head. In the backbone network, channel attention and spatial attention are added to the SPPF module to form an SPPFA module; The SPPFA module adds a channel attention module after each pooling layer of the SPPF module. The output of each channel attention module is spliced and then input into the spatial attention module. After one convolution layer, the output of the SPPFA module is obtained.
2. The method for detecting foreign matter and defects in escalator comb plates according to claim 1, characterized in that: The loss function of the detection model is the weighted sum of the Shape-IoU loss function, the NWD loss function, and the IoU loss function.
3. The method for detecting foreign matter and defects in escalator comb plates according to claim 1, characterized in that: The channel attention module pools the input feature map F of size H×W×C to obtain two feature maps of size 1×1×C, and then adds and activates them after passing through the fully connected layer to obtain the channel feature weight M. c , and finally M c Multiplying with F gets the output of the channel attention module.
4. The method for detecting foreign matter and defects in escalator comb plates according to claim 1, characterized in that: The spatial attention module pools the input feature map F of size H×W×C to obtain two feature maps of size H×W×1, and then concatenates the two in the channel dimension and performs convolution and pooling operations to obtain the spatial feature weight M s , and finally M s Multiplying with F gets the output of the spatial attention module.
5. A system for detecting foreign matter and defects in escalator comb plates, characterized in that: include: A model construction unit is used to establish a detection model. The detection model uses YOLOV5 as a framework and includes a backbone network, a neck network, and a detection head. In the backbone network, channel attention and spatial attention are added to the SPPF module to form an SPPFA module. The SPPFA module adds a channel attention module after each pooling layer of the SPPF module. The output of each channel attention module is spliced and then input into the spatial attention module. After passing through a convolutional layer, the output of the SPPFA module is obtained. A model training unit is used to construct a sample data set, wherein the sample data set includes images and annotation information of broken teeth of a comb plate, and images and annotation information of foreign objects; and train a detection model using the sample data set; The escalator comb plate foreign body and defect detection unit is used to collect the comb plate image to be detected, use the trained detection model to identify broken teeth and foreign objects, obtain the category and bounding box of the broken teeth and / or foreign objects, and calibrate the position of the broken teeth and / or foreign objects in the escalator comb plate image.
6. The escalator comb plate foreign matter and defect detection system according to claim 5 is characterized in that: The loss function of the detection model is the weighted sum of the Shape-IoU loss function, the NWD loss function, and the IoU loss function.
7. The escalator comb plate foreign matter and defect detection system according to claim 5, characterized in that: The channel attention module pools the input feature map F of size H×W×C to obtain two feature maps of size 1×1×C, and then adds and activates them after passing through the fully connected layer to obtain the channel feature weight M. c , and finally M c Multiplying with F gets the output of the channel attention module.
8. The escalator comb plate foreign matter and defect detection system according to claim 5, characterized in that: The spatial attention module pools the input feature map F of size H×W×C to obtain two feature maps of size H×W×1, and then concatenates the two in the channel dimension and performs convolution and pooling operations to obtain the spatial feature weight M s , and finally M s Multiplying with F gets the output of the spatial attention module.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is loaded into a processor, the method for detecting foreign matter and defects in an escalator comb plate according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting foreign matter and defects in an escalator comb plate according to any one of claims 1 to 4 is implemented.
Citation Information
Cited By
Supervision-based and non-supervision-based defect detection method, model establishment method and device
CN120931630A
Supervised and unsupervised defect detection methods, model building methods and devices
CN120931630B