Intelligent student protection equipment detection method for vocational education training workshop
Through the improved YOLOV8 model and data enhancement technology, combined with the attention mechanism, the automation and real-time detection of student protective equipment in the vocational education training workshop was achieved, and the problems of low detection efficiency and low accuracy in the existing technology were solved, which significantly improved the safety guarantee ability.
Patent Information
- Application Number
- CN202510072686.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology has problems such as low efficiency, low accuracy and great influence of human factors in the detection of student protective equipment in vocational education training workshops, making it difficult to achieve automated and intelligent testing.
The improved YOLOV8 model is used to combine data enhancement technology and attention mechanism, and the surveillance camera takes images of students wearing protective equipment in real time, perform automated detection and evaluation, and generate evaluation results.
It realizes automation and real-time monitoring, significantly improves detection efficiency and accuracy, reduces human errors, and enhances the safety guarantee capabilities of the workplace.
Smart Images

Figure CN120107996A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an intelligent detection method for student protective equipment in a vocational education training workshop, belonging to the technical field of personal safety protection equipment detection. Background Art
[0002] In the early days, factories relied on manual inspections and video surveillance to ensure that employees wore personal protective equipment correctly. These traditional methods have obvious limitations, such as environmental interference, misjudgment, and potential safety hazards in high-risk operations. Video review is time-consuming and requires high concentration, and employees' self-inspection is easily affected by their personal safety awareness, increasing risks. As the scale of factories and the complexity of the environment increase, manual monitoring becomes impractical and difficult to respond in real time. Traditional methods are increasingly limited when faced with the detection needs of multiple protective equipment. Therefore, there is an urgent need to promote the automation and intelligence of personal protective equipment detection to improve efficiency and accuracy and reduce vulnerabilities caused by human factors. Sensor-based personal protective equipment detection technology has made progress in the field of industrial safety. By deploying sensors and analyzing signals, it can determine whether employees are wearing protective equipment correctly. These technologies improve safety monitoring efficiency and enhance safety management. However, widespread application still faces challenges, such as high economic costs, the impact of environmental factors on detection accuracy, and reliance on technology that may weaken employees' safety awareness. Therefore, before promoting application, it is necessary to comprehensively consider costs, technical limitations, and the impact on employee behavior to find more efficient and economical solutions.
[0003] With the continuous development of computer vision detection technology, personal protective equipment detection technology based on deep learning has shown significant advantages in improving detection accuracy and speed. Through continuous technological innovation, this technology is actively responding to various challenges in industrial environments and providing strong technical support for ensuring workplace safety. Compared with traditional automation equipment, the detection method based on deep learning greatly reduces costs, while effectively ensuring safety and having relatively low requirements for the production environment. Summary of the invention
[0004] The present invention provides an intelligent detection method for student protective equipment in vocational education training workshops. Through intelligent and automated detection and evaluation means, the detection efficiency, accuracy and safety are significantly improved, human errors are effectively reduced, and the safety assurance capability of the workplace is enhanced to solve the problems existing in the prior art.
[0005] An intelligent detection method for student protective equipment in a vocational education training workshop comprises the following steps: Step 1: Take images of the protective equipment worn by students in the training workshop, collect image data of protective equipment including helmets, masks, goggles, and reflective vests, and build a data set for model training; Step 2: Use LabelImage to classify and annotate the collected images, indicating whether the students are wearing protective equipment (helmets, masks, goggles, and safety clothing); Step 3: Use data enhancement technology (adjust the brightness, contrast and saturation of the image, add Gaussian noise, and perform horizontal flipping, vertical flipping and angle rotation) to process the annotated images, expand the data set, and re-annotate and divide the data set; expand the diversity of the data set. At the same time, divide the data set into training set, test set and validation set according to 8:1:1; Step 4: Input the divided data set into the improved YOLOV8 network for training. During the training process, set the epoch to 400 and use the pre-trained model to accelerate the convergence of the network. After the training is completed, save the model; Step 5: The model detects students' wearing conditions in real time: the trained model is loaded into the safety training system, and the model is called to detect students' wearing conditions of protective equipment in the training workshop in real time; Step 6: Intelligent evaluation: Based on the recognition results output by the model and the preset safety specifications, the compliance of each protective equipment is automatically determined and an evaluation result is generated. The evaluation is performed based on the recognition results of the model. When the personal protective equipment worn by the student meets the safety standards, the evaluation results show that the student's protective equipment meets the requirements. At this time, the corresponding equipment will be activated and the indicator light will turn green, indicating that the safety inspection has been passed.
[0006] The step 1 specifically includes: The data set was obtained through surveillance cameras installed in various vocational education training workshops. The cameras used are of the HIKVISION brand, with a 4-megapixel high-definition resolution and an image sampling frequency of one frame every 5 seconds to ensure that the image is clear and contains sufficient details. These images come from multiple different vocational education training environments to ensure that the data is widely representative and practical. After completing the data collection, through strict screening, 3,000 high-quality images that meet the research needs were finally collected.
[0007] The step 2 specifically includes: The LabelImage tool is used to annotate the collected images, and the annotation content includes eight different protective equipment states: wearing a helmet (helmet), wearing safety goggles (goggles), wearing a safety vest (vest), wearing a mask (mask), not wearing a helmet (no_helmet), not wearing safety goggles (no_goggles), not wearing a safety vest (Safety Vest), and not wearing a mask (no_mask).
[0008] The step three specifically includes: Data augmentation operations include adjusting the brightness, contrast, and saturation of the image, adding Gaussian noise, and performing horizontal flipping, vertical flipping, and angular rotation. These data augmentation techniques simulate the various viewing conditions that may be encountered in the actual environment by increasing the visual changes in the image, thereby improving the model's recognition ability in complex scenes. In addition, data augmentation can help the model learn more generalized features by introducing visual perturbations and deformations, thereby effectively alleviating the overfitting problem that may occur when the number of training samples is limited.
[0009] The step 4 specifically includes: By introducing the SCDown module to optimize downsampling, the model has been effectively improved in terms of lightweight, meeting the performance and computing resource requirements of the workshop environment. At the same time, the DWR module enhances the information fusion capability of the backbone network at different scales, improves the processing capability of multi-scale features, and solves the difficult problem that the complex background and obstructions in the training workshop put forward high requirements for target detection. In order to improve the robustness of the model, the GAM attention mechanism is adopted, so that the model pays more attention to key features in complex backgrounds, thereby achieving more stable performance. In addition, the use of the NWD loss function enhances the network's attention to small targets, helping the predicted box to fit the real box more accurately, thereby reducing position offset. The step six specifically includes: Before the training begins, the cameras in the training room will automatically turn on to conduct real-time safety checks on the personal protective equipment worn by students. Only when the personal protective equipment worn by students meets safety standards will the corresponding equipment be activated and the indicator light on the equipment will turn green, indicating that the safety check has passed.
[0010] The present invention has the following beneficial effects and advantages: Through the improved YOLOv8 model, automation and real-time monitoring are achieved, which can efficiently detect the protective equipment worn by students, avoid omissions and delays in manual inspections, and ensure the standardized wearing of protective equipment. By introducing the SCDown module to optimize downsampling, the lightweight performance of the model is improved, and the demand for computing resources in the workshop environment is met. The DWR module enhances the multi-scale feature fusion capability and effectively copes with the challenges brought by complex backgrounds and occlusions. The GAM attention mechanism is used to improve the robustness of the model, making it more accurate in focusing on key features in complex environments, thereby enhancing stability. Finally, the NWD loss function enhances the network's attention to small targets, helping the predicted box to fit the real box more accurately, thereby reducing position offset, which is particularly suitable for the high-precision requirements in personal protective equipment detection. In addition, through data enhancement technologies such as image color inversion and brightness adjustment, the robustness and generalization ability of the system are improved, effectively coping with different light and occlusion problems. The system implements intelligent evaluation, automatically judges the compliance of students' protective equipment, reduces interference from human factors, and ensures that the corresponding equipment will be activated only after the protective equipment meets safety standards. In summary, the present invention significantly improves detection efficiency, accuracy and safety through intelligent and automated detection and evaluation methods, effectively reduces human errors, and enhances the safety assurance capability of the workplace. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0012] Figure 1 It is an algorithm flow chart of the intelligent detection method of student protective equipment for vocational education training workshops of the present invention; Figure 2 It is the helmet category diagram of the present invention; Figure 3 It is the goggles category diagram of the present invention; Figure 4 It is the vest category diagram of the present invention; Figure 5 It is the mask category diagram of the present invention; Figure 6 It is the no_goggles category diagram of the present invention; Figure 7 It is the no_vest category diagram of the present invention; Figure 8 It is the no_mask category diagram of the present invention; Fig. 9It is the no_helmet category diagram of the present invention; Fig.10 It is the improved YOLOv8 network structure diagram of the present invention; Fig.11 is a picture of a student wearing personal protective equipment to be tested, taken in Example 1 of the present invention; Fig.12 This is a recognition effect diagram of Example 1 of the present invention. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0014] Example Reference Figure 1-12 , an intelligent detection method for student protective equipment in vocational education training workshops, comprising the following steps: Step 1: Take images of the protective equipment worn by students in the training workshop, collect image data of protective equipment including helmets, masks, goggles, and reflective vests, and build a data set for model training; In a vocational education training factory environment, accurate and efficient monitoring of students wearing personal protective equipment is crucial to ensure safety. This measure helps to take necessary safety precautions in a timely manner, thereby reducing potential safety risks. However, there is currently no public, dedicated dataset covering multiple types of protective equipment. The primary task of the present invention is to construct a dataset that can accurately reflect the situation of students wearing personal protective equipment in a training factory environment.
[0015] The present invention focuses on monitoring the wearing of protective equipment by students in practical training factories, and constructs a data set including whether safety helmets are worn in accordance with regulations, whether safety goggles are worn in accordance with regulations, whether safety vests are worn in accordance with regulations, and whether masks are worn in accordance with regulations. Data collection is carried out by acquiring long-range video images through surveillance cameras installed in factory workshops. All cameras are HIKVISION brand 4-megapixel high-definition devices, and the image sampling frequency is set to extract one frame every 5 seconds to ensure image quality and details. These images were collected in multiple vocational education and training factory environments to ensure the diversity and wide applicability of the data. After strict screening, 3,000 high-quality images that met the research needs were finally collected.
[0016] Step 2: Use LabelImage to classify and annotate the collected images, indicating whether the students are wearing protective equipment (helmets, masks, goggles, and safety clothing); In this step, LabelImage software is used to divide the dataset of students' personal safety protection equipment into eight categories: wearing a helmet (helmet), wearing safety goggles (goggles), wearing a safety vest (vest), wearing a mask (mask), not wearing a helmet (no_helmet), not wearing safety goggles (no_goggles), not wearing a safety vest (no_vest), and not wearing a mask (no_mask). During the training process, the wearing of personal protective equipment is crucial to the safety of students. Constructing and improving the dataset can help to identify and correct improper wearing behavior in a timely manner, and fundamentally ensure the safety of students.
[0017] Step 3: Use data enhancement technology (adjust the brightness, contrast and saturation of the image, add Gaussian noise, and perform horizontal flipping, vertical flipping and angle rotation) to process the annotated images, expand the data set, and re-annotate and divide the data set; expand the diversity of the data set. At the same time, divide the data set into training set, test set and validation set according to 8:1:1; For the personal protective equipment detection task, a series of data enhancement methods were used, including adjusting the brightness, contrast and saturation of the image, adding Gaussian noise, and performing horizontal flipping, vertical flipping and angular rotation. These data enhancement methods increase the visual variation in the image and simulate the various viewing conditions that may be encountered in the actual environment, thereby improving the model's recognition ability in complex scenes. In addition, data enhancement helps the model learn more generalizable features by introducing visual perturbations and deformations, thereby alleviating the overfitting phenomenon that occurs when the number of training samples is limited.
[0018] In step S3, in order to improve the robustness of the personal protective equipment detection task, the present invention adopts a series of data enhancement methods. First, brightness adjustment simulates different lighting conditions by randomly selecting a brightness factor (ranging from 0.7 to 1.3), thereby enhancing the model's adaptability to lighting changes. Contrast adjustment changes the difference between the brightest and darkest areas in the image, and randomly adjusts by a factor between 0.8 and 1.2, thereby improving the clarity and detail recognition of the image. Saturation adjustment enhances the vividness of the color or weakens the color by randomly selecting a factor between 0.5 and 1.5, so that the model can adapt to different visual environments. In order to simulate shooting conditions in low light or high ISO environments, the present invention also adds Gaussian noise, and enhances the robustness of the model to noise interference by adding random noise (standard deviation ranges from 10 to 30) to the image. In addition, horizontal flip and vertical flip (50% probability each) help the model adapt to target detection in different directions, especially the wearing state of personal protective equipment, by changing the direction of the target. Finally, angle rotation randomly rotates the image between -30° and 30°, allowing the model to better recognize objects at different angles. These data enhancement methods combined significantly improve the performance of the model under various complex conditions and enhance its generalization ability.
[0019] Through these enhancement operations, the number of images in the dataset was expanded to 4,000, and they were divided into training set, validation set and test set in a ratio of 8:1:1, containing 3,200, 400 and 400 images respectively. This division strategy aims to strictly evaluate the generalization ability and practical application effect of the model through independent validation set and test set. This not only ensures the extensiveness and depth of model training, but also promotes its reliability and stability evaluation in unseen environments.
[0020] Step 4: Input the divided data set into the improved YOLOV8 network for training. During the training process, set the epoch to 400 and use the pre-trained model to accelerate the convergence of the network. After the training is completed, save the model; In the target detection task, although the YOLOv8 model performs well, its complex network structure and a large number of convolutional layers lead to increased model complexity and high demand for computing resources.
[0021] Improvement 1: Lightweight model design In order to cope with the high requirements for accuracy and real-time performance of personal protective equipment detection, as well as the resource consumption and deployment issues of the model, the present invention optimizes the YOLOv8 model. Specifically, the present invention uses the SCDown module to replace the traditional convolution operation. SCDown is an efficient convolutional network structure that generates partial feature maps through a small amount of convolution calculations and further streamlines the calculation process using simple linear operations. Finally, these feature maps are merged through a series operation, which significantly reduces the amount of calculation and the number of parameters of the model, while improving efficiency without sacrificing performance.
[0022] Improvement 2: Introducing the attention mechanism In the detection of personal protective equipment, complex backgrounds and small targets are often the main reasons for missed detection. To solve this problem, the present invention introduces an attention mechanism, which can effectively improve the performance of the model in changing and complex environments. The attention mechanism is inspired by human visual attention. It emphasizes important feature areas in the input and ignores irrelevant parts, thereby improving the efficiency and accuracy of feature extraction. The present invention adopts a parameter-free GAM (Global Attention Mechanism) and embeds it into the end of the backbone network of the model. The purpose is to improve the detection accuracy of key features of personal protective equipment without increasing the computational burden.
[0023] Improvement 3: DWR-C2f architecture optimization When processing multi-dimensional features, the DWR (Dynamic Weighting Refinement) mechanism significantly improves the model's feature extraction capabilities. The present invention innovatively combines DWR with the C2f architecture to form a DWR-C2f model. This architecture optimizes the feature extraction process and improves the effect of feature fusion by dynamically adjusting the focus on features of different scales. C2f itself has excellent multi-scale feature fusion capabilities, allowing each output layer to extract rich semantic information, while the DWR mechanism further enhances the accuracy of the model, especially in complex scenarios, where the detection performance of small targets has been significantly improved.
[0024] Improvement 4: Introducing NWD loss function In the detection task of student personal protective equipment, especially the positioning problem of small targets, the traditional loss function based on IoU (Intersection over Union) measurement is often sensitive to position deviation, resulting in insufficient small target detection performance. In order to improve this problem, the present invention proposes a loss function based on Wasserstein distance (W-Distance) - NWD (Normalized Wasserstein Distance). The NWD measurement method avoids the limitation of traditional methods relying on bounding box overlap by modeling the target box as a two-dimensional Gaussian distribution. Specifically, the method models the bounding box as a two-dimensional Gaussian distribution, in which the pixel weight at the center of the bounding box is the highest, and the weight value decreases with the distance from the center to the edge. This modeling method allows NWD to more finely distinguish the distribution differences between the target object information and background information when processing the detection process. Bounding box ,in is the center point coordinate, and is the width and height. Modeled as a Gaussian distribution , defined as: Then the Wasserstein distance is used to calculate the distribution distance between the two Gaussian distributions. and , and The second-order Wasserstein distance between is defined as: Normalizing it to a value between 0 and 1 gives NWD: in is a constant, and the NWD indicator is designed as a loss function as shown below.
[0025] in, is the Gaussian distribution model for the prediction box P, is the Gaussian distribution model of the prediction box G. By taking advantage of the fact that NWD is insensitive to target size, the network can focus more on small targets, making the predicted frame closer to the real frame and reducing the position offset during prediction. Therefore, using the NWD loss function instead of the original loss function is more suitable for the needs of personal protective equipment detection.
[0026] Result graph:
[0027] The experimental results show that the SDGF-YOLO model performs well in key performance indicators such as Precision, Recall, F1 and mAP, all of which are better than the above models. In addition, in terms of indicators such as Parameters, FLOPs and FPS, the SDGF-YOLO model also performs well, showing significant advantages and proving its lightweight and high efficiency. These results fully demonstrate the excellent performance of the SDGN-YOLO model, ensuring its efficient execution of personal protective equipment detection tasks in the vocational education training workshop environment.
[0028] Step 5: The model detects students' wearing conditions in real time: Load the trained model into the safety training system, and call the model to detect students' wearing conditions of protective equipment (such as helmets, masks, goggles, and safety clothing) in the training workshop in real time; Step 6: Intelligent evaluation: Based on the recognition results output by the model and the preset safety specifications, the compliance of each protective equipment is automatically determined and an evaluation result is generated. The evaluation is performed based on the recognition results of the model. When the personal protective equipment worn by the student meets the safety standards, the evaluation results show that the student's protective equipment meets the requirements. At this time, the corresponding equipment will be activated and the indicator light will turn green, indicating that the safety inspection has been passed.
[0029] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.
Claims
1. An intelligent detection method for student protective equipment in vocational education training workshops, characterized by: The following steps are involved: Step 1: Take images of the protective equipment worn by students in the training workshop, collect image data of protective equipment including helmets, masks, goggles, and reflective vests, and build a data set for model training; Step 2: Use LabelImage to classify and label the collected images; Step 3: Use data augmentation technology to process the annotated images, expand the data set, and re-annotate and divide the data set; Step 4: Input the divided data set into the improved YOLOV8 network for training; Step 5: The model detects students' wearing conditions in real time: the trained model is loaded into the safety training system, and the model is called to detect students' wearing conditions of protective equipment in the training workshop in real time; Step 6: Intelligent evaluation: Based on the recognition results output by the model and combined with the preset safety regulations, the compliance of each item of protective equipment is automatically determined and an evaluation result is generated.
2. According to claim 1, a method for intelligent detection of student protective equipment for vocational education training workshops is characterized by: The step 1 specifically includes: The data set was obtained through surveillance cameras installed in each vocational education training workshop. The image sampling frequency of the surveillance cameras is one frame every 5 seconds to ensure that the image is clear and contains sufficient details. The images come from multiple different vocational education training environments. After completing the data collection, they are strictly screened and finally high-quality images that meet the research needs are collected.
3. According to claim 1, a method for intelligent detection of student protective equipment for vocational education training workshops is characterized by: The step 2 specifically includes: The LabelImage tool is used to annotate the collected images, and the annotation content includes eight different protective equipment states: wearing a hard hat, wearing safety goggles, wearing a safety vest, wearing a mask, not wearing a hard hat, not wearing safety goggles, not wearing a safety vest, and not wearing a mask.
4. According to claim 1, a method for intelligent detection of student protective equipment for vocational education training workshops is characterized by: The step three specifically includes: The data enhancement techniques include adjusting the brightness, contrast and saturation of the image, adding Gaussian noise, and performing horizontal flipping, vertical flipping and angular rotation processing.
5. According to claim 1, a method for intelligent detection of student protective equipment for vocational education training workshops is characterized by: The improved YOLOV8 network in step 4 specifically includes: S4.1: The SCDown module is used to replace the traditional convolution operation. SCDown generates some feature maps through a small amount of convolution calculations, and uses simple linear operations to further simplify the calculation process. The feature maps are merged through serial operations, which significantly reduces the amount of calculation and the number of parameters of the model, while improving efficiency without sacrificing performance. S4.2: The attention mechanism is introduced to effectively improve the performance of the model in changing and complex environments. The parameter-free GAM is adopted and embedded into the end of the backbone network of the model to improve the detection accuracy of key features of personal protective equipment without increasing the computational burden. S4.3: DWR is combined with C2f architecture to form the DWR-C2f model, which optimizes the feature extraction process and improves the effect of feature fusion by dynamically adjusting the focus on features of different scales; S4.4: A loss function based on Wasserstein distance, NWD, is proposed. The NWD measurement method models the target box as a two-dimensional Gaussian distribution. By modeling the bounding box as a two-dimensional Gaussian distribution, where the pixel weight at the center of the bounding box is the highest and the weight value decreases as the distance from the center to the edge increases, NWD can more finely distinguish the distribution differences between the target object information and the background information when processing the detection process. ,in is the center point coordinate, and For width and height, Modeled as a Gaussian distribution , defined as: Then the Wasserstein distance is used to calculate the distribution distance between two Gaussian distributions. and , and The second-order Wasserstein distance between is defined as: Normalizing it to a value between 0 and 1 gives NWD: in is a constant, and the NWD indicator is designed as a loss function as follows: in, is the Gaussian distribution model for the prediction box P, is the Gaussian distribution model of the prediction box G, Taking advantage of the fact that NWD is insensitive to target size, the network's attention to small targets is strengthened, the predicted box is closer to the real box, and the position offset caused by prediction is reduced. Using the NWD loss function instead of the original loss function is more suitable for the needs of personal protective equipment detection.
6. According to claim 1, a method for intelligent detection of student protective equipment for vocational education training workshops is characterized by: The step six specifically includes: Before the training begins, the camera in the training room will automatically turn on to conduct real-time safety checks on the students' personal protective equipment. When the personal protective equipment worn by the students meets the safety standards, the device will be activated and the indicator light on the device will turn green, indicating that the safety check has passed.