Method and system for high-altitude operation safety belt wearing detection

By improving the YOLOv11 model and combining data enhancement technology, a two-stage target detection network is built, which solves the problems of insufficient feature extraction capabilities and low model adaptability in the detection of high-altitude work safety belt wear, and achieves higher-precision safety belt wear detection.

CN120296344APending Publication Date: 2025-07-11ANHUI UNIV

Patent Information

Application Number
CN202510349111.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art methods for high-altitude work safety belt wear detection have insufficient feature extraction capabilities, low model adaptability, and low accuracy of output detection results.

Method used

Using the improved YOLOv11 model, a two-stage object detection network is built by inserting the MixAM attention module in front of the SPPF layer of the backbone network and adding the MixAM attention module behind each C3K2 layer of the neck network, and a two-stage object detection network is built, combining data augmentation technology to train and detect the seat belt wear of high-altitude workers.

Benefits of technology

It improves the feature extraction ability and adaptability of the model, improves the accuracy of seat belt wear detection, and can better identify and classify various incorrect seat belt wear behaviors, reducing safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296344A_ABST
    Figure CN120296344A_ABST
Patent Text Reader

Abstract

The invention discloses a high-altitude operation safety belt wearing detection method and system. The method comprises the following steps: obtaining a data set DataSet1; a data set DataSet2 is obtained; inserting a MixAM attention module in front of an SPPF layer of a backbone network of the YOLOv11 model, adding the MixAM attention module in a neck network of the YOLOv11 model, and constructing a first target detection network; the MixAM attention module is formed by connecting space attention and channel attention in parallel; inputting the data set DataSet1 into a first target detection network to train the first target detection network; selecting another YOLOv11 model as a second target detection network, and inputting the data set DataSet2 into the second target detection network to train the second target detection network; the picture of the actual working scene is input into a first target detection network, an image of a person leaving the ground is cut out, the image is input into a second target detection network, whether a safety belt is worn or not is detected, and if it is detected that the safety belt is not worn, the system gives an alarm; the method has the advantages of strong feature extraction capability, strong model adaptability and high target detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-altitude operation safety, and particularly to a method and system for detecting the wearing of safety belts for high-altitude operations. Background Art

[0002] High-altitude power grid operation refers to operations such as high-altitude maintenance, inspection, installation, and demolition on power facilities. Such operations usually involve working on high-voltage transmission lines, and require operators to climb to relatively high positions, such as transmission towers, utility poles, etc., for equipment maintenance and inspection work. Therefore, safety protection must be well done, which involves the wearing of safety belts. High-altitude power grid operation is an important link in the operation and maintenance of the power system, and its safety and efficiency directly affect the stability and reliability of power supply. By strict safety specifications and technical requirements, the operation risks can be effectively reduced and the safety of operators can be guaranteed.

[0003] Chinese Patent Publication No. CN119068515A discloses a method and system for detecting illegal wearing of safety belts for high-altitude operations. The method includes: obtaining an image of a high-altitude operation to be detected, and preprocessing the image of the high-altitude operation to be detected to obtain a target image to be detected; inputting the target image to be detected into a target safety belt illegal wearing detection model for feature extraction and semantic guidance to obtain the prediction probability of illegal items. By performing multi-level feature extraction and semantic guidance on the images in the high-altitude operation scenario, various incorrect safety belt wearing behaviors can be effectively identified and classified, thereby improving the safety monitoring level of high-altitude operations and reducing potential safety hazards. Although it mentions feature extraction, no corresponding design is carried out for the feature extraction module, so the feature extraction ability needs to be improved, the adaptability of the model is not high, and the accuracy of the output detection results is not high. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that the feature extraction ability of the existing method for detecting the wearing of safety belts for high-altitude operations needs to be improved, the adaptability of the model is not high, and the accuracy of the output detection results is not high.

[0005] The present invention solves the above technical problems by the following technical means: A method for detecting the wearing of safety belts for high-altitude operations, including:

[0006] Step A: Collect the original images of high-altitude operators working in the situation of being off the ground and not off the ground, and screen and perform data augmentation on them to obtain a data set DataSet1; collect the original images of wearing and not wearing safety belts, and screen and perform data augmentation on them to obtain a data set DataSet2;

[0007] Step B: Insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module to the neck network of the YOLOv11 model to construct the first object detection network; the MixAM attention module consists of a spatial attention and a channel attention in parallel;

[0008] Step C: Input the dataset DataSet1 into the first object detection network for training;

[0009] Step D: Select another YOLOv11 model as the second object detection network, and input the dataset DataSet2 into the second object detection network for training;

[0010] Step E: Input the picture of the actual working scene into the first object detection network, crop out the image of the person above the ground, input it into the second object detection network, detect whether the seat belt is worn, and if it is detected that the seat belt is not worn, the system will issue an alarm for warning.

[0011] Further, step A includes:

[0012] Collect pictures of aerial work personnel including ground work and aerial work on the network as the first dataset; collect pictures including wearing seat belts and not wearing seat belts as the second dataset; remove the repeated, damaged, and unavailable ones in the first dataset and the second dataset; perform data augmentation on the first dataset to establish the dataset DataSet1, and perform data augmentation on the second dataset to establish the dataset DataSet2. The data augmentation includes cropping, rotating, and mosaic to increase the data volume and using methods such as adjusting hue, saturation, and brightness to enhance the complexity.

[0013] Further, step B includes:

[0014] Insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module behind each C3K2 layer of the neck network of the YOLOv11 model as a new C3K2 layer to form the first object detection network.

[0015] Further, the spatial attention of the MixAM attention module divides the input feature X along the horizontal and vertical directions. The features divided in the horizontal and vertical directions are respectively subjected to average pooling. The features in the horizontal and vertical directions after average pooling are both subjected to fully connected and 2D convolution, and then normalized using a batch normalization layer and a non-linear neural network respectively. The results of the features in the horizontal and vertical directions after normalization are both subjected to 2D convolution and the Sigmoid activation function to obtain the weights in the horizontal and vertical directions; the channel attention of the MixAM attention module performs global average pooling on the input feature X, and then obtains the weights in the channel dimension through 1D convolution and the Sigmoid activation function. The weights in the horizontal and vertical directions and the weights in the channel dimension are multiplied to obtain the final weight, and finally a residual connection is made to obtain the final output. The residual connection refers to multiplying the final weight by the input feature X and then adding it to X to obtain the final output.

[0016] Further, step E includes:

[0017] Input the picture of the actual working scenario into the detection system. First, all the people in the picture are detected by the first object detection network, and those on the ground and not on the ground are separated; then all the people on the ground are cropped out to form a new picture and input into the second object detection network; finally, the second object detection network detects whether the people in the input picture wear seat belts. If it is detected that the seat belts are not worn, the system will issue an alarm for warning.

[0018] The present invention also provides a system for detecting the wearing of seat belts for high-altitude operations, including:

[0019] A data processing unit, configured to collect the original images of high-altitude operation personnel working in the situation of being on the ground and not on the ground, screen and perform data augmentation on them to obtain a data set DataSet1; collect the original images of wearing seat belts and not wearing seat belts and perform data screening and augmentation on them to obtain a data set DataSet2;

[0020] A model building unit, configured to insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module to the neck network of the YOLOv11 model to build a first object detection network; the MixAM attention module is composed of spatial attention and channel attention in parallel;

[0021] A first training unit, configured to input the data set DataSet1 into the first object detection network for training;

[0022] The second training unit is used to select another YOLOv11 model as the second target detection network, and input the data set DataSet2 into the second target detection network for training;

[0023] The target detection unit is used to input the picture of the actual working scene into the first target detection network, crop out the image of the person off the ground, input it into the second target detection network, and detect whether the seat belt is worn. If it is detected that the seat belt is not worn, the system will issue an alarm for warning.

[0024] Furthermore, the data processing unit is also used for:

[0025] Collect pictures of aerial work personnel including ground work and aerial work on the network as the first data set; collect pictures including wearing seat belts and not wearing seat belts as the second data set; remove the repeated, damaged, and unavailable ones in the first data set and the second data set; perform data augmentation on the first data set to establish the data set DataSet1, and perform data augmentation on the second data set to establish the data set DataSet2. The data augmentation includes cropping, rotation, and mosaic to increase the data volume, and using methods such as adjusting hue, saturation, and brightness to enhance the complexity.

[0026] Furthermore, the model building unit is also used for:

[0027] Insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module behind each C3K2 layer of the neck network of the YOLOv11 model as a new C3K2 layer to form the first target detection network.

[0028] Furthermore, the spatial attention of the MixAM attention module divides the input feature X along the horizontal and vertical directions. The features divided in the horizontal and vertical directions are respectively subjected to average pooling. The features in the horizontal and vertical directions after average pooling are both subjected to fully connected and 2D convolution, and then normalized using a batch normalization layer and a non-linear neural network respectively. The results of the features in the horizontal and vertical directions after normalization are both subjected to 2D convolution and Sigmoid activation function to obtain the weights in the horizontal and vertical directions; the channel attention of the MixAM attention module performs global average pooling on the input feature X, and then obtains the weights in the channel dimension through 1D convolution and Sigmoid activation function. Multiply the weights in the horizontal and vertical directions and the weights in the channel dimension to obtain the final weight, and finally perform a residual connection to obtain the final output. The residual connection refers to multiplying the final weight by the input feature X and then adding it to X to obtain the final output.

[0029] Furthermore, the target detection unit is also used for:

[0030] Input the pictures of the actual working scenarios into the detection system. First, all the people in the pictures are detected by the first object detection network, and those on the ground and not on the ground are distinguished. Then, all the people on the ground are cropped out to form new pictures and input into the second object detection network. Finally, the second object detection network detects whether the people in the input pictures wear safety belts. If it detects that a safety belt is not worn, the system will issue an alarm for warning.

[0031] The advantages of the present invention are as follows:

[0032] The MixAM attention module of the present invention is composed of spatial attention and channel attention in parallel. Thus, the MixAM attention module can enable the network to simultaneously learn the features in the horizontal direction, vertical direction, and channel direction. Compared with YOLOv11, the improved YOLOv11-MixAM model can better capture the spatial relationships in the horizontal and vertical directions due to the introduction of the MixAM module. By learning the spatial dependence relationships, the ability to extract features is enhanced, the adaptability of the model is improved, and the accuracy of the detected results output is higher.

[0033] The MixAM attention module of the present invention uses an adaptive convolution kernel size, which can dynamically adjust the size of the convolution kernel according to the input features. Thus, it can more flexibly capture the dependence relationships between channels, can process features of different sizes, and improves the adaptability of the model. Description of the Drawings

[0034] Figure 1 It is a flowchart of a method for detecting the wearing of safety belts in high-altitude operations disclosed in an embodiment of the present invention;

[0035] Figure 2 It is a schematic diagram of the MixAM attention module in a method for detecting the wearing of safety belts in high-altitude operations disclosed in an embodiment of the present invention;

[0036] Figure 3 It is a schematic diagram of the YOLOv11-MixAM model in a method for detecting the wearing of safety belts in high-altitude operations disclosed in an embodiment of the present invention;

[0037] Figure 4 It is a schematic diagram of the system operation interface of the detection system composed of the first object detection network and the second object detection network in a method for detecting the wearing of safety belts in high-altitude operations disclosed in an embodiment of the present invention;

[0038] Figure 5 It is a schematic diagram of some data in the data set DataSet1 in a method for detecting the wearing of safety belts in high-altitude operations disclosed in an embodiment of the present invention

[0039] Figure 6Schematic diagram of part of the data in DataSet2 in a method for detecting the wearing of a safety belt for high-altitude work disclosed in an embodiment of the present invention. Detailed implementation mode

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] Embodiment 1

[0042] Embodiment 1 of the present invention provides a method for detecting the wearing of a safety belt for high-altitude work, using YOLov11 for the safety belt detection task. The system is a two-stage network. Since the actual working scenario is relatively complex and there are many interfering objects, if a first-order network is used, it will simultaneously detect people on the ground and people in the air, and the difficulty of detecting the safety belt is relatively high. Some interfering objects will be detected as safety belts, such as trees and wires. Therefore, a two-stage network is designed as Figure 1 , which includes a human detection model, intermediate processing, and a safety belt detection model. The human detection model is used to detect and distinguish people in the picture. The input is the picture to be detected, and the output is the prediction box; the intermediate processing crops the part of the person off the ground from the original picture according to the label of the prediction box to form a new picture; the safety belt detection model is used to detect whether the safety belt is worn. The input is the cropped picture, and the output is the prediction box of the safety belt. The system can well avoid these interfering objects. The method includes the following steps:

[0043] S1. Collect the original images of high-altitude workers working in the air and on the ground from the Internet as the first data set; collect the original images of wearing and not wearing safety belts as the second data set. The specific process of S1 is as follows:

[0044] S101. Collect 2546 pictures of high-altitude workers including ground work and high-altitude work on the Internet as the first data set; collect 2184 pictures including wearing and not wearing safety belts as the second data set.

[0045] S102. Remove the repeated, damaged, and unavailable ones in the first data set; remove the repeated, damaged, and unavailable ones in the second data set.

[0046] S2. Perform data augmentation on the first dataset to establish dataset DataSet1 for training the first object detection network, and perform data augmentation on the second dataset to establish dataset DataSet2 for training the second object detection network. The specific content of S2 is as follows:

[0047] Since the amount of collected data is relatively small and there is a gap with the actual scenario, data augmentation is performed. Specifically, methods such as cropping, rotation, and mosaic are used to increase the data volume; hue, saturation, brightness, etc. are used to enhance the complexity.

[0048] S3. Improve the YOLOv11 model to obtain the YOLOv11-MixAM model as the first object detection network;

[0049] S301. The YOLOv11-MixAM model of the present invention is improved based on the existing YOLOv11 model. The YOLOv11 model belongs to the existing well-known technology. For relevant records, reference can be made to the literature "Khanam R, Hussain M. Yolov11: An overview of the key architectural enhancements[J]. arXiv preprint arXiv:2410.17725, 2024.", or the YOLOv11 model recorded in the link https: / / zhuanlan.zhihu.com / p / 23251465868. Based on the YOLOv11 model, the MixAM attention module designed by the present invention is inserted in front of the SPPF layer of the backbone network (BackBone), and the MixAM attention module designed by the present invention is added behind each C3K2 layer of the neck network (neck) as a new C3K2 layer.

[0050] S302. The MixAM attention module is composed of spatial attention and channel attention in parallel. Spatial attention divides the input feature X along the horizontal and vertical directions, then performs average pooling (Avg Pool), then Concat and 2D convolution, and then uses the batch normalization layer (BatchNorm) and non-linear neural network (Non-linear) for normalization processing respectively. The results of the two-way normalization processing are both subjected to 2D convolution and the Sigmoid activation function to obtain the weights in the horizontal and vertical directions.

[0051] Channel attention performs global average pooling on the feature X, then obtains the weights in the channel dimension through 1D convolution and the Sigmoid activation function, multiplies the weights in the horizontal and vertical directions and the weights in the channel dimension to obtain the final weights, and finally makes a residual connection to get the final output. The residual connection means multiplying the final weights by the input feature X and then adding X to get the final output. The MixAM attention module of the present invention enables the network to simultaneously learn the features in the horizontal, vertical, and channel directions. Compared with the existing YOLOv11 model, the improved YOLOv11-MixAM model of the present invention can better capture the spatial relationships in the horizontal and vertical directions due to the introduction of the MixAM attention module, enhance the ability to extract features by learning spatial dependency relationships, and use an adaptive convolution kernel size, which can dynamically adjust the size of the convolution kernel according to the input features, thereby more flexibly capturing the dependencies between channels, being able to process features of different sizes, and improving the adaptability of the model. The MixAM attention module designed by the present invention is as Figure 2 shown, Figure 2 where C×H×W, etc. represent the feature dimensions, and the YOLOv11-MixAM model is as Figure 3 shown.

[0052] S4. Input the data set DataSet1 into the first target detection network for training;

[0053] S401. Divide the data set DataSet1 into a training set and a test set according to 8:2.

[0054] S402. Use the pre-trained network weights as the initial weights and adopt the method of transfer learning to train the YOLOv11-MixAM model using the training set. The pre-trained weights are iterated for 200 epochs, the learning rate is set to 0.01, the decay rate is 0.005, the probabilities of cropping and rotation are 0.5, and the mosaic augmentation is 1.0.

[0055] S5. Test the YOLOv11-MixAM model and update the learning parameters of the YOLOv11-MixAM model;

[0056] S501. Use the test set in DataSet1 to test the YOLOv11-MixAM model. The results are shown in Table 1, where Para represents the number of model parameters in MB, GFLOPS represents the model's computational volume in GB, Offgraoud_P represents the detection accuracy of personnel off the ground in %, and Groud_P represents the detection accuracy of personnel on the ground in %. YOLOv11+Augment means adding data augmentation operations to the YOLOv11 model, and this data augmentation operation includes random vertical / horizontal flipping, random scaling, etc.; YOLOv11+Augment+MixAM means adding data augmentation operations and the MixAM attention module to the YOLOv11 model. From the test results in Table 1, it can be seen that the YOLOv11-MixAM model proposed by the present invention has the highest detection accuracy of personnel off the ground and the detection accuracy of personnel on the ground.

[0057] Table 1 Comparison results after adding the MixAM attention module

[0058]

[0059]

[0060] S6. Use the unimproved YOLOv11 model as the second object detection network. The unimproved YOLOv11 model is the YOLOv11 model of the prior art described above. The YOLOv11 model directly used in the second stage of the network has not undergone data augmentation and network architecture improvement.

[0061] S7. Input DataSet2 into the second object detection network for training.

[0062] S8. Use the two-stage network designed by the present invention to detect the pictures of the high-altitude working scenarios of the power grid to be detected.

[0063] S801. Combine the trained first object detection network and the second object detection network to form a complete detection system.

[0064] S802. Input the pictures of the actual working scenarios into the detection system. First, all the people in the pictures are detected by the first object detection network, and those on the ground and off the ground are separated. Then, all the people off the ground are cropped out to form new pictures and input into the second-stage network, that is, the second object detection network. Finally, the second object detection network detects whether the people in the input pictures wear safety belts. If it is detected that the safety belts are not worn, the system will issue an alarm for warning.

[0065] Figure 4 Schematic diagram of the system operation interface of the detection system composed of the first object detection network and the second object detection networkFigure 5 Schematic diagram of partial data in DataSet1 Figure 6 Schematic diagram of partial data in DataSet2; It can be seen from the figure that the present invention constructs a detection system by collecting pictures of high-altitude workers working on the ground and at high altitudes, as well as pictures of wearing and not wearing safety belts to train the first target detection network and the second target detection network. Through this detection system, all people in the picture can be detected, and those on the ground and not on the ground can be distinguished, and whether they wear safety belts can be detected. If it is detected that a safety belt is not worn, the system will issue an alarm for warning.

[0066] Embodiment 2

[0067] Based on Embodiment 1, Embodiment 2 of the present invention further provides a system for detecting the wearing of safety belts for high-altitude operations, including:

[0068] A data processing unit, configured to collect original images of high-altitude workers working in the situation of being on the ground and not on the ground, screen and perform data augmentation on them to obtain DataSet1; collect original images of wearing and not wearing safety belts and perform screening and data augmentation on them to obtain DataSet2;

[0069] A model building unit, configured to insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module to the neck network of the YOLOv11 model to construct the first target detection network; the MixAM attention module is composed of a spatial attention and a channel attention in parallel;

[0070] A first training unit, configured to input DataSet1 into the first target detection network for training;

[0071] A second training unit, configured to select another YOLOv11 model as the second target detection network, and input DataSet2 into the second target detection network for training;

[0072] A target detection unit, configured to input a picture of an actual working scenario into the first target detection network, crop out the image of a person on the ground, input it into the second target detection network, and detect whether a safety belt is worn. If it is detected that a safety belt is not worn, the system will issue an alarm for warning.

[0073] Specifically, the data processing unit is further configured to:

[0074] Collect pictures of aerial work personnel including ground work and aerial work on the network as the first dataset; collect pictures of wearing safety belts and not wearing safety belts as the second dataset; remove the repeated, damaged, and unavailable ones from the first and second datasets; perform data augmentation on the first dataset to establish dataset DataSet1, and perform data augmentation on the second dataset to establish dataset DataSet2. Data augmentation includes cropping, rotation, and mosaic to increase the data volume, and adjusting the tone, saturation, and brightness to enhance the complexity.

[0075] Specifically, the model building unit is further used for:

[0076] Insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module behind each C3K2 layer of the neck network of the YOLOv11 model as a new C3K2 layer to form the first object detection network.

[0077] Specifically, the spatial attention of the MixAM attention module divides the input feature X horizontally and vertically. The features divided horizontally and vertically are respectively subjected to average pooling. The features in the horizontal and vertical directions after average pooling are both subjected to fully connected and 2D convolution, and then normalized using a batch normalization layer and a non-linear neural network respectively. The results of the features in the horizontal and vertical directions after normalization are both subjected to 2D convolution and a Sigmoid activation function to obtain the weights in the horizontal and vertical directions; the channel attention of the MixAM attention module performs global average pooling on the input feature X, and then obtains the weights in the channel dimension through 1D convolution and a Sigmoid activation function. Multiply the weights in the horizontal and vertical directions and the weights in the channel dimension to obtain the final weight, and finally perform a residual connection to obtain the final output. The residual connection refers to multiplying the final weight by the input feature X and then adding it to X to obtain the final output.

[0078] Specifically, the object detection unit is further used for:

[0079] Input the pictures of the actual work scene into the detection system. First, all the people in the pictures are detected by the first object detection network, and the people on the ground and those not on the ground are distinguished; then all the people off the ground are cropped out to form new pictures and input into the second object detection network; finally, the second object detection network detects whether the people in the input pictures wear safety belts. If it detects that the safety belt is not worn, the system will issue an alarm for warning.

[0080] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for detecting the wearing of a safety belt for working at heights, characterized in that, Including: Step A: Collect the original images of high-altitude workers working in the situations of being off the ground and on the ground, and screen and perform data augmentation on them to obtain the data set DataSet1; Collect the original images of workers wearing safety belts and not wearing safety belts, and screen and perform data augmentation on them to obtain the data set DataSet2; Step B: Insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module to the neck network of the YOLOv11 model to construct the first object detection network; The MixAM attention module is composed of spatial attention and channel attention in parallel; Step C: Input the data set DataSet1 into the first object detection network for training; Step D: Select another YOLOv11 model as the second object detection network, and input the data set DataSet2 into the second object detection network for training; Step E: Input the picture of the actual working scene into the first object detection network, crop out the image of the person off the ground, input it into the second object detection network, detect whether the safety belt is worn, and if it is detected that the safety belt is not worn, the system will issue an alarm for warning.

2. The method for detecting the wearing of a safety belt for working at height according to claim 1, wherein Step A includes: Collect pictures of high-altitude workers including ground work and high-altitude work on the network as the first data set; collect pictures including wearing safety belts and not wearing safety belts as the second data set; remove the repeated, damaged, and unavailable ones in the first data set and the second data set; perform data augmentation on the first data set to establish the data set DataSet1, and perform data augmentation on the second data set to establish the data set DataSet2. The data augmentation includes cropping, rotating, and using mosaic to increase the data volume, and using methods such as adjusting hue, saturation, and brightness to enhance the complexity.

3. A method for detecting the wearing of a safety belt for working at height according to claim 1, characterized in that, Step B includes: Insert a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add a MixAM attention module behind each C3K2 layer of the neck network of the YOLOv11 model as the new C3K2 layer to form the first object detection network.

4. A method for detecting the wearing of a safety belt for working at height according to claim 1, characterized in that, The spatial attention of the MixAM attention module divides the input feature X horizontally and vertically. The features divided horizontally and vertically are respectively subjected to average pooling. The features in the horizontal and vertical directions after average pooling are both subjected to fully connected and 2D convolution, and then normalized using a batch normalization layer and a non-linear neural network respectively. The results of the features in the horizontal and vertical directions after normalization are both subjected to 2D convolution and Sigmoid activation function to obtain the weights in the horizontal and vertical directions; the channel attention of the MixAM attention module performs global average pooling on the input feature X, and then obtains the weights in the channel dimension through 1D convolution and Sigmoid activation function. The weights in the horizontal and vertical directions and the weights in the channel dimension are multiplied to obtain the final weight, and finally a residual connection is made to obtain the final output. The residual connection refers to multiplying the final weight by the input feature X and then adding it to X to obtain the final output.

5. A method for detecting the wearing of a safety belt for working at heights according to claim 1, characterized in that, Step E includes: Input the picture of the actual working scene into the detection system. First, all the people in the picture are detected by the first object detection network, and those on the ground and not on the ground are distinguished; then all the people on the ground are cropped out to form a new picture and input into the second object detection network; finally, the second object detection network detects whether the people in the input picture wear seat belts. If it is detected that the seat belts are not worn, the system will issue an alarm for warning.

6. A system for detecting the wearing of a safety belt for working at height, characterized in that, It includes: A data processing unit, which is used to collect the original images of high-altitude workers working in the situations of on the ground and not on the ground, and screen and data augment them to obtain the data set DataSet1; Collect the original images of wearing seat belts and not wearing seat belts and screen and data augment them to obtain the data set DataSet2; A model building unit, which is used to insert the MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and add the MixAM attention module to the neck network of the YOLOv11 model to build the first object detection network; The MixAM attention module is composed of spatial attention and channel attention in parallel; A first training unit, which is used to input the data set DataSet1 into the first object detection network for training; A second training unit, which is used to select another YOLOv11 model as the second object detection network, and input the data set DataSet2 into the second object detection network for training; An object detection unit, which is used to input the picture of the actual working scene into the first object detection network, crop out the images of the people on the ground, input them into the second object detection network, detect whether the seat belts are worn, and if it is detected that the seat belts are not worn, the system will issue an alarm for warning.

7. The system for detecting the wearing of a safety belt for working at height according to claim 6, characterized in that, The data processing unit is also used for: Collect pictures of high-altitude workers including ground work and high-altitude work on the network as the first data set; collect pictures including wearing seat belts and not wearing seat belts as the second data set; remove the repeated, damaged and unavailable ones in the first data set and the second data set; Data augmentation is performed on the first data set to establish the data set DataSet1, and data augmentation is performed on the second data set to establish the data set DataSet2. The data augmentation includes cropping, rotation, and mosaic to increase the data volume, and adjusting the hue, saturation, and brightness to enhance the complexity.

8. The system for detecting the wearing of a safety belt for working at height according to claim 6, characterized in that, The model building unit is also used for: Inserting a MixAM attention module in front of the SPPF layer of the backbone network of the YOLOv11 model and adding a MixAM attention module behind each C3K2 layer of the neck network of the YOLOv11 model as a new C3K2 layer to form a first object detection network.

9. The system for detecting the wearing of a safety belt for working at height according to claim 6, characterized in that, The spatial attention of the MixAM attention module divides the input feature X along the horizontal and vertical directions. The features divided in the horizontal and vertical directions are respectively subjected to average pooling. The features in the horizontal and vertical directions after average pooling are both subjected to fully connected and 2D convolution, and then normalized using a batch normalization layer and a non-linear neural network respectively. The results of the features in the horizontal and vertical directions after normalization are both subjected to 2D convolution and a Sigmoid activation function to obtain the weights in the horizontal and vertical directions; the channel attention of the MixAM attention module performs global average pooling on the input feature X, and then obtains the weights in the channel dimension through 1D convolution and a Sigmoid activation function. The weights in the horizontal and vertical directions and the weights in the channel dimension are multiplied to obtain the final weight, and finally a residual connection is made to obtain the final output. The residual connection refers to multiplying the final weight by the input feature X and then adding it to X to obtain the final output.

10. A system for detecting the wearing of a safety belt for working at heights according to claim 6, characterized in that, The object detection unit is also used for: Inputting the picture of the actual working scene into the detection system. First, all the people in the picture are detected by the first object detection network, and the people on the ground and not on the ground are distinguished; then all the people on the ground are cropped out to form a new picture and input into the second object detection network; finally, the second object detection network detects whether the people in the input picture wear seat belts. If it is detected that the seat belts are not worn, the system will issue an alarm for warning.

Citation Information

Patent Citations

  • Safety belt illegal wearing detection method and system for high-altitude operation

    CN119068515A

Cited By

  • Safety rope operation state monitoring method and system based on dual-scale differential entropy characteristics

    CN121071564A

  • Safety rope operation state monitoring method and system based on double-scale differential entropy features

    CN121071564B