Worker helmet target detection method based on YOLOv7 algorithm

By improving the YOLOv7 algorithm, using the EfficientVit model and CSE module to enhance feature extraction, combined with the ICIoULoss function training model, the problems of helmet detection accuracy and speed in the existing technology are solved, and efficient and real-time detection is achieved in smart construction sites.

CN120259706APending Publication Date: 2025-07-04CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311604029.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient detection accuracy, poor robustness and slow detection speed when detecting workers' helmets. Especially when small target detection is poor, it is difficult to meet the real-time detection needs in actual construction site scenarios.

Method used

The improved YOLOv7 algorithm is adopted to replace the CBS, MP and ELAN modules in the Backbone stage using the EfficientVit model, and add CSE modules in the Head stage, and combine the ICIoULoss function for model training to achieve end-to-end helmet detection.

Benefits of technology

It improves the feature extraction ability of small targets, realizes high-precision helmet detection, and can complete detection in real time in actual scenarios, which is suitable for real-time alarms at smart construction sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259706A_ABST
    Figure CN120259706A_ABST
Patent Text Reader

Abstract

The invention discloses a worker helmet target detection method based on a YOLOv7 algorithm, and the method comprises the following steps: S1, collecting data and constructing a helmet detection data set: collecting pictures of a person wearing a helmet and a person not wearing the helmet under various scenes and angles, marking the collected pictures, and guaranteeing that each image is marked with the position and category information of the helmet; s2, constructing a helmet detection model of a YOLOv7 algorithm, wherein the YOLOv7 algorithm comprises an image input stage, a Backbone stage, a Head stage and a Prediction stage, and a CBS module, an MP module and an ELAN module in the Backbone stage are replaced by using an OfficientVit model; a CSE module is additionally arranged behind the ELAN-W module in the Head stage; s3, using a PyTorch framework to train the model, and using an ICIoU Loss function to calculate the distance deviation between the prediction target frame and the real target frame. According to the invention, the requirement of real-time detection is met while high precision of helmet detection in an actual scene is ensured, and technical support is provided for realizing intelligent construction site construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of security technology, and particularly relates to a method for detecting worker helmets based on the YOLOv7 algorithm. Background Art

[0002] With the rapid development and progress of science and technology, convolutional neural networks have been proposed and successfully applied to various computer vision tasks, such as image classification, object detection, semantic segmentation, etc. Object detection is a basic computer vision task that requires detecting the objects to be detected in images or videos and locating their coordinates. Due to the rapid development of technologies such as deep learning and convolutional neural networks, object detection technologies based on convolutional neural networks have now been successfully applied in multiple industries, such as smart cities, smart construction sites, etc.

[0003] The safety of electric power construction has always been one of the hot topics of current social concern, and it is becoming increasingly important to ensure the safety of electric power workers who transmit electric energy. Existing research shows that the main cause of most electric power safety accidents is mainly due to the non-standard behavior of workers during construction, and safety helmets are the basic equipment to ensure the safety of construction personnel. When an emergency occurs, safety helmets can effectively protect the safety of construction personnel's heads. Therefore, detecting the safety helmets of construction personnel in construction scenes has great significance for maintaining personnel safety.

[0004] Nowadays, object detection technology has been successfully applied to the detection of worker safety helmets. Zhang Mingyuan et al. proposed a method for detecting the wearing of safety helmets using deep learning, constructed the operating environment of the deep learning network model, used the deep learning network model to obtain the characteristics of construction personnel wearing safety helmets, and used an evaluation method to achieve the detection results of the wearing of safety helmets by construction personnel. However, the evaluation indicators selected by this method are all the results of expert voting, and its subjectivity leads to poor final detection results. Liu Xin et al. [7] proposed a method for detecting the wearing of safety helmets, which inputs construction scene images into a convolutional neural network, continuously iterates the network to extract the characteristics of construction personnel's safety helmets, and outputs the classification results of the wearing of safety helmets. However, this method is affected by blurred construction scene images and occlusion situations, and its monitoring results are not accurate enough. Wu Minyu et al. added deconvolution to enhance the expression ability of the model for small safety helmet targets. Compared with simple upsampling factors, deconvolution can learn target feature information during model training to enhance the model's ability to extract target features. However, these methods are limited by detection accuracy and detection time, resulting in limited application ranges and scenarios.

[0005] At the same time, the helmet is often presented as a small target in pictures or videos, which increases the difficulty of detection. Existing models often have poor detection accuracy and robustness for small targets. At the same time, if too many small targets are introduced in the training stage, it will also reduce the model's ability to learn the features of the target. In addition, most of the previous methods have the problem of slow detection speed. When deployed to the actual construction site scenario, it is often necessary to perform real-time detection on the video and issue an alarm, which requires the detection model to have fewer model parameters while ensuring good detection accuracy. To address these two major problems, the present invention designs a lightweight but feature-rich target detection algorithm framework, which has fewer model parameters while ensuring efficient target extraction ability. At the same time, for the convenience of deployment in the actual scenario, the target detection algorithm framework designed by the present invention has a simple and efficient structure, and adopts an end-to-end method to directly output the target result by inputting a picture, and can realize the real-time detection and alarm of workers' helmets in the real construction site scenario. Summary of the Invention

[0006] The object of the present invention is to provide a method for detecting workers' helmet targets based on the improved YOLOv7 algorithm in view of the above deficiencies in the technology, which meets the requirements of real-time detection while ensuring high-precision helmet detection in the actual scenario, and provides technical support for the realization of the construction of a smart construction site.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for detecting workers' helmet targets based on the YOLOv7 algorithm, comprising the following steps: S1: Collect data and construct a helmet detection data set: Collect pictures of helmeted and un-helmeted people in various scenarios and angles, and then mark the collected pictures, ensuring that the position and category information of the helmet are marked in each image; S2: Construct a helmet detection model of the YOLOv7 algorithm: The YOLOv7 algorithm includes an image input, a Backbone stage, a Head stage, and a Prediction stage. The CBS, MP, and ELAN modules in the Backbone stage are replaced with the EfficientVit model; a CSE module is added after the ELAN-W module in the Head stage; S3: Use the PyTorch framework to train the model, and at the same time use the I CIoU Loss function to calculate the distance deviation between the predicted target box and the real target box.

[0008] The helmet detection method in step S2 includes the following steps: Step 1: Extract pictures collected by the video or camera frame by frame; Step 2: Send the frame-by-frame extracted images into the YOLOv7 algorithm for detection to obtain the network output results. The network output results are the target category and bounding box coordinate information predicted for each pixel point on the output feature map. Step 3: In the second step, the detection results obtained by passing the input images through the YOLOv7 algorithm are obtained. Then, these detection results are mapped back to the original images. The mapping factor is obtained by dividing the size of the output feature map by the size of the original input image. Finally, first filter out the invalid detection boxes according to the detection confidence score, and then use the non-maximum suppression post-processing algorithm to filter out the overlapping target boxes. Finally, the final detection results obtained by passing the input images through an end-to-end helmet detection network are obtained.

[0009] The technical effects of the present invention are as follows: 1. The algorithm has strong feature extraction ability, especially for small targets.

[0010] 2. A more accurate method is proposed to calculate the position deviation between the predicted target box and the true target box.

[0011] 3. The algorithm is relatively concise and can complete the detection of safety helmets in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present invention will be further described below in conjunction with the drawings and embodiments: Figure 1 It is the YOLOv7 algorithm model diagram of the present invention; Figure 2 It is the YOLOv7 algorithm model diagram of the present invention; Figure 3 It is the YOLOv7 algorithm model diagram of the present invention; Figure 4 It is the input image and detection result diagram of the present invention in a real scene; Figure 5 It is the input image and detection result diagram of the present invention in a real scene; Figure 6 It is the algorithm flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0013] A method for detecting worker helmets based on the YOLOv7 algorithm includes the following steps: S1: Collect data and construct a helmet detection data set To adapt to helmet detection in real scenarios, the present invention first collects multiple scene pictures, which should cover various scenes and angles, including pictures of people wearing helmets and not wearing helmets. Specifically, it is obtained through the following ways: (1) Network search: Use a search engine or image database to search for images of helmets and related scenes. Ensure that the images contain different lighting conditions, backgrounds, and people; (2) Camera capture: Install cameras at construction sites, stadiums, or specific places to capture real-time images of people wearing helmets and not wearing helmets; (3) Data augmentation: Perform data augmentation on the collected images, such as rotation, flipping, cropping, and changing brightness / contrast, to increase data diversity.

[0014] After obtaining the pictures, the present invention annotates the collected images to ensure that the position and category information of the helmets are marked on each image. To accumulate a sufficient number and variety of helmet image samples and provide a data basis for model training.

[0015] S2: Build a helmet detection model based on the YOLOv7 algorithm The overall architecture of the YOLOv7 algorithm is as Figure 1 shown. The overall architecture includes image input, Backbone stage, Head stage, and Prediction stage. To better extract image features, the present invention uses the EfficientVit model to replace the CBS, MP, and ELAN modules in the Backbone stage of the existing YOLOv7 algorithm, so that while reducing the computational amount, it can more effectively extract target features. In addition, in the Head stage, the present invention adds a CSE module after the ELAN-W module. The overall structure of the CSE module is as Figure 2 shown. The CSE module inputs the intermediate layer feature map and outputs the enhanced feature map, noting that the size of the input feature map of this module is the same as the size of the output feature map. The CSE module first divides the input feature map into four sub-parts (i.e., upper left, upper right, lower left, and lower right), and then the upper left and upper right sub-feature maps are respectively sent to the SENet to extract features, and the lower left and lower right are respectively sent to a convolutional layer to extract features. Finally, these four enhanced sub-feature maps are spliced to obtain the final enhanced feature map. SENet is a channel attention mechanism model, and its overall structure is as Figure 3 shown, which weights the channel information of the feature map to increase the spatial information corresponding to the important channels of the feature map, thereby highlighting and extracting the target features.

[0016] S3: Use the PyTorch framework to train the model, and at the same time use the I CIoU Loss function to calculate the distance deviation between the predicted target box and the real target box During model training, it is necessary to construct a loss function to provide a correct guiding direction for the optimization of model parameters. In the target bounding box loss function used in the training stage of the present invention, ICIoU Loss (Improved Complete Intersection over Union Loss) is proposed to more accurately calculate the distance deviation between the predicted target box and the ground truth target box. The overall expression of ICIoU Loss is as follows in Equation (1).

[0017]

[0018] Where IoU is the intersection over union between the predicted target box and the ground truth target box, ρ represents the Euclidean distance between the two rectangular boxes, c is the diagonal distance of the closure region of the two target boxes, b and b gt represent the centers of the two rectangular boxes respectively, w gt and h gt represent the width and height of the ground truth target box and the width and height of the predicted box respectively. β is a metric parameter used to measure the width-height consistency of the anchor box.

[0019] In the model training stage of the present invention, the PyTorch framework is used to train the model, and a server equipped with an Nvidia GeForce RTX 2080TI graphics card is used to train the helmet detection network. The central processing unit of the server is an Inter i9-12900HK, and the memory is 32G.

[0020] The helmet detection method in step S2 includes the following steps: Step 1: Image, video or camera input In the inference stage of the present invention, image, video or camera input is accepted. When video and camera are used as input, the present invention extracts the video or camera frame by frame according to the frame rate, and then sends each frame into the next step.

[0021] Step 2: Detection by an end-to-end helmet detection network In the first step above, the input images frame by frame are obtained, and then the pictures are sent into the worker helmet target detection method of the YOLOv7 algorithm for detection to obtain the network output results, that is, the target category and bounding box coordinate information predicted for each pixel point on the output feature map.

[0022] Step 3: Detection post-processing and result output In the second step, the detection results obtained by the worker helmet target detection method using the YOLOv7 algorithm for the input image are obtained. Next, in this step, these detection results will first be mapped back to the original image from the output image, that is, a coordinate mapping transformation is performed. The mapping factor is obtained by dividing the size of the output feature map by the size of the original input image. Then, the present invention first filters out invalid detection boxes according to the detection confidence score (default set to 0.30), and then uses the non-maximum suppression post-processing algorithm to filter out overlapping target boxes. Finally, the final prediction results obtained by the input image passing through an end-to-end helmet detection network are obtained. Figure 4 is the input image in the real scene, Figure 5 is the detection result image.

[0023] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A method for detecting worker helmets based on the YOLOv7 algorithm, characterized in that, The steps are as follows: S1: Collect data and construct a helmet detection dataset: Collect pictures of helmeted and un-helmeted people in various scenarios and angles, and then label the collected pictures to ensure that the position and category information of the workers wearing helmets are marked in each image; S2: Construct a helmet detection model for the YOLOv7 algorithm: The overall architecture of the YOLOv7 algorithm includes an image input, a Backbone stage, a Head stage, and a Prediction stage. Replace the CBS, MP, and ELAN modules in the Backbone stage with the EfficientVit model; add a CSE module after the ELAN-W module in the Head stage; S3: Use the PyTorch framework to train the model, and at the same time use the ICIoU Loss function to calculate the distance deviation between the predicted target box and the true target box.

2. The method for detecting worker helmets based on the YOLOv7 algorithm according to claim 1, characterized in that, The helmet detection method in step S2 includes the following steps: Step 1: Extract pictures captured by the video or camera frame by frame; Step 2: Send the pictures extracted frame by frame into the YOLOv7 algorithm for detection to obtain the network output result, which is the target category and bounding box coordinate information predicted for each pixel point on the output feature map; Step 3: For the detection results obtained by the input pictures passing through the YOLOv7 algorithm in the second step, then map these detection results back to the original picture. The mapping factor is obtained by dividing the size of the output feature map by the size of the original input image. Finally, first filter out invalid detection boxes according to the detection confidence score, and then use the non-maximum suppression post-processing algorithm to filter out overlapping target boxes, and finally obtain the final detection results of the input pictures passing through an end-to-end helmet detection network.

3. A method for detecting worker helmets based on the YOLOv7 algorithm according to claim 1 or 2, characterized in that, The overall expression of the ICIoU Loss function is: Among them, IoU is the intersection over union between the predicted target box and the ground truth target box, ρ represents the Euclidean distance between the two rectangular boxes, c is the diagonal distance of the closure region of the two target boxes, and b and b gt respectively represent the center points of the two rectangular boxes, w gt , h gt , w and h respectively represent the width and height of the ground truth target box and the width and height of the predicted box, and β is a metric parameter used to measure the width-height consistency of the anchor box.