Method and system for detecting wearing of individual protective equipment of superfine dust operating personnel
Through the intelligent detection methods of deep learning and machine vision, combined with the fusion of ConvNeXt and BiFPN features, the YOLOv5s model is optimized, and the real-time and management efficiency of individual protective equipment detection in the fire explosive production process is solved, high-precision identification and real-time detection of individual protective equipment are achieved, and safety management efficiency is improved.
Patent Information
- Application Number
- CN202510267403.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art In the process of fire explosives production, the wear detection of individual protective equipment has problems such as poor real-time performance, high labor consumption, low management efficiency and difficulty in recording and traceability, resulting in the inability to effectively protect the health of staff.
The intelligent detection methods of deep learning and machine vision are adopted, and the real-time detection of individual protective equipment is carried out through the YOLOv5s detection model, combined with the fusion of ConvNeXt network and BiFPN feature, and the target detection is optimized using the DIOU_NMS method to build a wearable detection system for individual protective equipment for ultra-fine dust operators.
It realizes high-precision identification and real-time detection of individual protective equipment, improves safety protection awareness, reduces manual inspection workload, provides complete data analysis and traceability, reduces labor costs and accident risks, and improves management efficiency.
Smart Images

Figure CN120375414A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of ultrafine dust operations, and particularly relates to a method and system for detecting the wearing of personal protective equipment for workers in ultrafine dust operations. Background Art
[0002] Currently, during the production process of explosives, the working site environment is complex and changeable. When workers come into contact with harmful factors such as flammable and explosive dust and harmful gases in each process, there is a situation where the personal protective equipment is not targeted and effective enough. Accidents often occur where personnel are not effectively protected due to incomplete or non-standard wearing of personal protective equipment. The main management methods currently adopted are posting reminder signs, safety education, manual inspections, etc., and there are the following problems: 1. Poor real-time inspection: Manual inspection cannot cover all time periods.
[0003] 2. High labor consumption: Regular inspections by special personnel are required.
[0004] 3. Low management efficiency: It is difficult to cope with the complex and changeable working environment.
[0005] 4. Difficult record tracing: Lack of systematic management means. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for detecting the wearing of personal protective equipment for workers in ultrafine dust operations, realizing real-time intelligent detection of the wearing status of personal protective equipment and protecting the physical health of workers.
[0007] The purpose of the present invention is achieved by the following technical means. A method for detecting the wearing of personal protective equipment for workers in ultrafine dust operations includes the following steps. Collect images of protective equipment, construct a recognition data set of protective equipment, and label the protective equipment information on the images of protective equipment. Establish a detection model, establish a YOLOv5s detection model, including an input end, a backbone network, a bottleneck layer, and an output layer. Input end: Input the images of part of the recognition data set of protective equipment into the detection model as the training set, perform Mosaic data augmentation, automatically calculate anchor boxes, and image scaling. Backbone network, that is, a feature extraction network, is used to extract the feature information of the image. First, through slicing operations, the input image is copied four times, and each pixel takes values at every other pixel. Finally, the images are fused to obtain the underlying feature map; the underlying feature map is further extracted by ConvNexT to obtain the processed feature map. Through spatial pyramid pooling, the output of ConvNexT is pooled using three different pooling kernels for downsampling, and then spliced and fused to obtain features of various scales. Using a feature fusion network, the features are fused. The features of various scales extracted by the backbone network are subjected to multi-scale feature fusion through the cross-scale feature fusion structure BiFPN to obtain the finally fused multi-scale feature map.
[0008] The output layer generates the final detection result. The output layer performs object detection on the finally fused multi-scale feature map according to the DIOU_NMS method and the loss function to generate the final detection result.
[0009] The labeled protective equipment information includes the target category, bounding box coordinates, and target ID.
[0010] After collecting the protective equipment images, data augmentation is also performed on the images. Through the Copy-Pasting data augmentation strategy, oversampling and data pasting are performed on the small targets in the dataset to provide sufficient small targets to match the anchors, thereby improving the performance of small target detection. Specifically, oversampling is performed on the small target sample data, and then two images are mixed together through Copy-Pasting to form a new image, which is added to the protective equipment recognition dataset.
[0011] The backbone network uses the ConvNexT network for feature extraction. Specifically, in the first stage, the preprocessed input image is downsampled using a convolutional layer with the same convolutional kernel size and stride, and 3 ConvNeXt Blocks are stacked to output a feature map. In the second stage: The input is the feature map output in the first stage. A convolutional layer with the same convolutional kernel size and stride is used for downsampling to further reduce the size of the feature map. 3 ConvNeXt Blocks are stacked to continue feature extraction and fusion, and a feature map is output. In the third stage: The input is the feature map output in the second stage. A convolutional layer with the same convolutional kernel size and stride is used for downsampling again. 9 ConvNeXt Blocks are stacked in this stage to output a feature map. In the fourth stage: The input is the feature map output in the third stage. A convolutional layer with the same convolutional kernel size and stride is used for downsampling, and 3 ConvNeXt Blocks are stacked to finally obtain a feature map with a size of 7×7 and 768 channels. After global average pooling and the LN layer, it is finally output through a linear classifier.
[0012] The ConvNeXt Block adopts an inverted bottleneck layer design. First, it uses a 3×3 depthwise separable convolution for feature extraction, then uses a 1×1 convolution to increase the number of channels from 96 to 384 for dimensionality increase, and then uses a 1×1 convolution again to reduce the number of channels from 384 back to 96 for dimensionality reduction to reduce the loss of high-dimensional information.
[0013] The specific DIOU_NMS method is as follows:
[0014] Among them: a is the predicted bounding box, and a gt is the ground truth bounding box; ρ 2 (a, a gt ) is the distance between the centers of the predicted bounding box and the ground truth bounding box; is the diagonal length of the smallest enclosing rectangle of the two bounding boxes; The candidate set for model detection is d i , the maximum score bounding box M, then the update formula of the DIOU-NMS method is:
[0015] Among them: Q i is the score of different classifications; R DIOU (M, d i ) is the value of R DIOU with respect to M and d i ; ε is the threshold manually set for the NMS operation; IOU is the intersection over union; i is the number of anchor boxes corresponding to the grid.
[0016] The loss function is GIoU.
[0017] An individual protective equipment wearing detection system for ultrafine dust workers includes an image collection module for monitoring on-site images; an environmental monitoring module for monitoring on-site environmental indicators through sensors; a protective equipment detection module for identifying the types of protective equipment worn by workers based on the on-site images collected by the image collection module; a controller module for judging the types of protective equipment required on-site based on the data collected by the environmental monitoring module and comparing it with the recognition results of the protective equipment detection module. When the target recognition result does not have the required protective equipment, it issues an alarm and notification.
[0018] The beneficial effects of the present invention are as follows: 1. Introduce ConvNeXt to optimize the backbone and neck networks of YOLOv5, obtaining stronger feature expression capabilities. Compared with the original backbone network of YOLOv5, it can extract features such as the texture and shape of images more effectively, especially for target features in complex scenarios. With more efficient computational performance, ConvNeXt optimizes computational resources in its design, reducing the amount of computation and memory occupancy while maintaining high accuracy. This enables a faster inference speed under the same hardware conditions when using YOLOv5 for object detection, improving the overall detection efficiency.
[0019] 2. Improve the Neck neck module through BiFPN to achieve simpler and faster multi-scale feature fusion. In addition, introduce learnable weights in BiFPN so that it can learn the importance of different input features and repeatedly apply top-down and bottom-up multi-scale feature fusion.
[0020] 3. Use the improved DIOU_NMS method instead of the traditional NMS method, comprehensively considering the distance, overlapping area, and aspect ratio between the predicted box and the ground truth box. It is considered that two predicted boxes with a greater center point distance may be located on different targets. Combining the intersection over union (IoU) of the two boxes and the center point distance can optimize the IoU loss and guide the learning of the center point, enabling more accurate regression of the predicted box. Brief Description of the Drawings
[0021] Figure 1 is the original YOLOv5 network architecture diagram; Figure 2 is the ConvNeXt network architecture diagram; Figure 3 is the ConvNeXt Block network architecture diagram; Figure 4 is the downsampling network architecture diagram; Figure 5 is the flowchart of the detection system for the individual protective equipment worn by ultrafine dust workers; The present invention will be further described in detail below with reference to the drawings and embodiments. Detailed Embodiment
[0022] The method for detecting the wearing of individual protective equipment by ultrafine dust workers includes the following steps: Collect images of protective equipment, construct a dataset for identifying protective equipment, and label the protective equipment information on the images of protective equipment; The labeled protective equipment information includes the target category, bounding box coordinates, and target ID.
[0023] Deep learning is a data-driven model that requires a large amount of sample data, and the coverage of the sample data for scenarios needs to be complete enough. The performance of the deep learning model depends on the performance of the sample data. The sample data situation will affect the generalization of the model. In order to complete the task of intelligent identification of individual protective equipment for ammunition production practitioners, therefore, the protective equipment identification dataset needs to collect multi-scenario samples, and the collection scenarios include indoor, outdoor, strong light, weak light, etc. By collecting 15,000 images with different backgrounds, it is used to make an individual protective equipment identification dataset, and the individual protective equipment includes: safety helmets, full-face masks, dust-proof work clothes, and protective gloves.
[0024] Data annotation, the dataset includes images of size 640*640 pixels (the above 4 types of individual protective equipment), and the dataset is divided into a training set and a validation set in a ratio of 4:1. The collected image data needs to be manually annotated and then used to train the model. The dataset uses the LabelImg image annotation software to annotate the information of the individual protective equipment contained in all the above images. The annotation information includes the target category, bounding box coordinates, and target ID.
[0025] Detection model establishment, establish a YOLOv5s detection model, including an input end, a backbone network, a bottleneck layer, and an output layer; Input end, use part of the images of the protective equipment identification dataset as the training set to input into the detection model, perform Mosaic data augmentation, automatically calculate anchor boxes, and image scaling; After collecting the protective equipment images, data augmentation is also performed on the images. Through the Copy-Pasting data augmentation strategy, oversampling and data pasting are performed on the small targets in the dataset to provide enough small targets to match the anchors, thereby improving the performance of small target detection; Specifically, oversampling is performed on the small target sample data, and then two images are mixed together through Copy-Pasting to form a new image, which is added to the protective equipment identification dataset.
[0026] As Figure 1 shown, Input (input end) includes three parts, namely Mosaic data augmentation, automatic calculation of anchor boxes, and image scaling: The first part, Mosaic data augmentation, scales, crops, and arranges the images containing individual protective equipment, increases data, enriches small targets, and improves the detection ability of small targets in individual protective equipment images; The second part is the adaptive anchor box calculation. According to the initial anchor box, a prediction box is output, and then compared with the true box, and updated backward according to the difference to update the parameters. The third part is the image scaling function, that is, self and image filling, which scales and fills the input image so that the unified size of the input image is 608×608×3.
[0027] In computer vision tasks such as object detection, small objects often have a small proportion in the image and little feature information, making it difficult for the model to accurately detect them.
[0028] Therefore, the Copy-Pasting data augmentation strategy is adopted to oversample, copy, and paste small objects in the dataset, so as to provide enough small objects to match the anchors, thereby improving the performance of small object detection. This augmentation method improves the performance of small object detection by oversampling the small object sample data and then performing Copy-Pasting operations on the small objects in the data samples. Oversampling the images of small objects to improve the problem of fewer small object images, and repeatedly training the pictures containing small objects through multiple times. This method is simple and effective. Among them, the number of copies is the oversampling rate.
[0029] Copy the sample data containing small objects to increase the frequency of their appearance in the training dataset. If the number of images containing small objects in the original dataset is small, by copying these images, the proportion of small object samples in the dataset is increased, so that the model has more opportunities to learn the features of small objects.
[0030] After performing the oversampling operation, then through hybrid pasting (Copy-Pasting), the following formula is used to mix two images together.
[0031] I 1 ×α+I 2 ×(1-α) In the above formula: I 1 represents the pasted object image, I 2 is the object image to be pasted, α is the mask. Simply put, it is to cut out the pixels in the mask part of I 1 and randomly paste them into I 2.
[0032] After oversampling, select two images, cut out the small object area in one image, and then paste it into a suitable position in the other image to form a new image. This process can simulate the appearance of small objects in different backgrounds and further enrich the training samples of small objects. The new image obtained after hybrid pasting is usually added to the training dataset as a new training sample for subsequent model training. This can increase the diversity of training data and improve the detection performance of the model for small objects.
[0033] Differences between Mix and Paste (Copy-Pasting) and Mosaic data augmentation at the input end: 1. Data fusion method Mix and Paste crops out small target regions from one image and directly pastes them into another image, mainly focusing on the transfer and fusion of small targets and only involving two images.
[0034] Mosaic data augmentation stitches together four different images in a certain way to form a new image. During the stitching process, operations such as scaling and cropping are performed on these four images, and then they are combined into a new training sample.
[0035] 2. Focus of data augmentation purpose The main purpose of Mix and Paste is to solve the problem of insufficient small target samples. By copying and pasting small targets into different background images, the training samples of small targets in different scenarios are increased, and the detection ability of the model for small targets is improved.
[0036] Mosaic data augmentation: focuses on increasing the diversity and richness of data. By stitching together four different images, the model can see more different target and background information in one training, enhancing the generalization ability of the model.
[0037] 3. Operation complexity Mix and Paste is relatively simple, mainly involving cropping and pasting operations. It is necessary to focus on adjusting the position and size of small targets to ensure that the pasted image looks natural.
[0038] Mosaic data augmentation operation is relatively complex. It is necessary to scale, crop, and stitch four images, and also consider adjusting the size of the stitched image and the position of the target.
[0039] The backbone network, that is, the feature extraction network, is used to extract the feature information of the image. First, through slicing operations, the input image is copied four times, and every other pixel is taken from each copy. Finally, the images are fused to obtain the underlying feature map; the underlying feature map is further extracted by ConvNexT to obtain the processed feature map. Through spatial pyramid pooling, the output of ConvNexT is pooled using three different pooling kernels for downsampling, and then stitched and fused to obtain features of various scales; The existing Bacbone (backbone network), i.e., the feature extraction network, includes Focus, Conv, CSP, and SPP. Among them, Focus first copies the input image four times through slicing operations, takes pixel values at every other pixel for each copy, and finally fuses the images, reducing the model's computational load, reducing the number of layers, and improving the inference speed; the CSP module splits the underlying feature map into two parts by channel in CSPNet. One part passes through a dense block (composed of multiple fully connected layers) and a transition layer (usually a convolutional layer with a kernel size of 1×1), and the other part combines with the transmitted feature map, not only reducing the computational load but also improving the inference speed and accuracy. SPP is spatial pyramid pooling, which performs pooling operations using three different pooling kernels for downsampling, and then splicing and fusion, which can increase the receptive field of the feature map.
[0040] Due to the small size of the personal protective equipment, features are easily lost during feature extraction. To address this problem, the ConvNexT network is used to replace the CSP module for feature extraction. Because the CSP module in the original network is mainly used for feature fusion and reducing computational load, and the ConvNeXt module performs better in this regard. ConvNexT is a pure convolutional network. Under the same computing power conditions, its inference speed and accuracy are higher.
[0041] The ConvNexT network has a stronger feature expression ability. ConvNexT adopts modern convolutional design concepts, such as depthwise separable convolutions and large convolutional kernels. These designs can capture the feature information of images at different scales. Compared with the original backbone network of YOLOv5, it can more effectively extract features such as the texture and shape of images, especially for target features in complex scenarios.
[0042] With more efficient computing performance, ConvNexT optimizes the computing resources in its design. While maintaining high accuracy, it reduces the computational load and memory occupancy. This enables a faster inference speed and improved overall detection efficiency when using YOLOv5 for object detection under the same hardware conditions.
[0043] As Figure 2 shown Figure 4 in the figure, the backbone network uses the ConvNexT network for feature extraction. Specifically, in the first stage, for the input preprocessed image, assuming the number of input channels is 3, a convolutional layer with the same convolutional kernel size and stride is used for downsampling to adjust the number of channels to 96, and 3 ConvNeXt Blocks are stacked to output the feature map. Second stage: The input is the feature map output from the first stage. A convolutional layer with the same convolutional kernel size and stride is used for downsampling to further reduce the size of the feature map. Three ConvNeXt Blocks are stacked to continue feature extraction and fusion, maintaining the number of channels at 96, and the output feature map is obtained. Third stage: The input is the feature map output from the second stage. A convolutional layer with the same convolutional kernel size and stride is used for downsampling again. Nine ConvNeXt Blocks are stacked in this stage, which is the most stacked among the four stages, for in-depth feature extraction and fusion. The number of channels remains 96, and the output feature map is obtained. Fourth stage: The input is the feature map output from the third stage. A convolutional layer with the same convolutional kernel size and stride is used for downsampling. Three ConvNeXt Blocks are stacked, and finally a feature map with a size of 7×7 and 768 channels is obtained. After global average pooling and the LN layer, it is finally output through a linear classifier.
[0044] The ConvNexT network adjusts the stacking times in ResNet50 from (3, 4, 6, 3) to (3, 3, 9, 3). Downsampling uses a convolutional layer with the same convolutional kernel size and stride. The balance precision is achieved through depthwise separable convolutions, and the number of channels is adjusted to 96. The ConvNeXt Block is designed as an inverted bottleneck layer. Feature extraction is completed using 3×3 depthwise separable convolutions. The number of channels is increased from 96 to 384 through 1×1 convolutions for dimensionality increase, and then the number of channels is reduced from 384 to 96 using 1×1 convolutions for dimensionality reduction to reduce the loss of high-dimensional information. The ConvNeXt network obtains a 7×7×768 feature map through feature extraction in four stages, and then through global average pooling and the LN layer, and finally outputs through a linear classifier.
[0045] The said ConvNeXt Block adopts an inverted bottleneck layer design. First, 3×3 depthwise separable convolutions are used for feature extraction, then the number of channels is increased from 96 to 384 through 1×1 convolutions for dimensionality increase, and then the number of channels is reduced from 384 back to 96 using 1×1 convolutions for dimensionality reduction to reduce the loss of high-dimensional information.
[0046] ConvNeXt adopts an inverted bottleneck layer structure, reducing the loss of small target feature information for individual protection in downsampling. ConvNeXt adds the LN layer before and after downsampling and after global average pooling. Through such a normalization strategy, the stability of the model is improved. Since ConvNeXt reduces the use of activation functions and BN layers, the ConvNeXt model not only improves the accuracy of individual protection target detection but also reduces the number of network operations.
[0047] Using a feature fusion network, the features are fused. The features of each scale extracted by the backbone network are subjected to multi-scale feature fusion through the cross-scale feature fusion structure BiFPN to obtain the final fused multi-scale feature map.
[0048] Neck is a feature fusion network that fuses features and then transmits them to the output end. It includes the FPN (Feature Pyramid Networks) + PAN (Path Aggregation Network) structure, where: the FPN is a top-down feature transmission structure that uses upsampling (Upsample) to transmit and fuse feature information to obtain a feature map; the PAN adopts a bottom-up transmission and fusion method to enhance the feature extraction ability.
[0049] YOLOV5s has good detection effect on significant targets. In the production site of explosives, the background is relatively complex, and the detection target is relatively small personal protective equipment. Using the original YOLOv5s will cause the information of the shallow and deep layers to interfere with each other, which will affect the detection accuracy, and its recognition performance is average. Therefore, YOLOV5s is improved.
[0050] BiFPN (Bidirectional Feature Pyramid Network) is a new cross-scale feature fusion structure that extends the design concept of the traditional FPN. By introducing learnable weighted feature aggregation and bidirectional path enhancement (including top-down and bottom-up feature flows), it realizes the efficient fusion of multi-scale features. Its core innovation lies in the adoption of a weighted feature fusion mechanism based on fast normalization (Fast Normalized Fusion). The input features of each node are adaptively weighted by learnable scalar weights, and residual connections are combined to maintain the stable propagation of gradients. This design not only improves the feature reuse rate, but also, by removing redundant connections and optimizing the network topology structure, maintains a powerful feature expression ability while reducing the computational complexity, showing significant performance advantages in dense prediction tasks such as object detection. Using the Bidirectional Feature Pyramid Network (BiFPN) can achieve simple and fast multi-scale feature fusion. It uses cross-scale connections to remove the nodes in PANet that contribute less to feature fusion and adds additional connections between the input and output nodes at the same level. We use BiFPN to improve the Neck module and use a one-layer BiFPN structure to improve the training efficiency of the model. Using BiFPN to improve the neck of YOLOv5s can achieve simpler and faster multi-scale feature fusion. In addition, learnable weights are introduced into BiFPN so that it can learn the importance of different input features and repeatedly apply top-down and bottom-up multi-scale feature fusion.
[0051] Specifically, BiFPN (Bidirectional Feature Pyramid Network) is a network structure for efficiently fusing multi-scale features. The process of fusing multi-scale features mainly includes the following key steps: 1. Input feature preparation: Usually, feature maps of different scales are extracted by the backbone network. These feature maps have different spatial resolutions and semantic information. First, adjust the number of channels of the input feature maps of different scales. Use a 1×1 convolutional kernel to unify the number of channels of all feature maps to the same value, which is convenient for subsequent feature fusion operations.
[0052] 2. Construct bidirectional feature fusion paths: (1) Top-down path: Starting from the highest-level feature map, gradually increase its resolution and fuse it with lower-level feature maps.
[0053] (2) Starting from the lowest-level feature map, gradually reduce its resolution and fuse it with higher-level feature maps.
[0054] 3. Repeat fusion multiple times (optional) To further enhance the effect of feature fusion, the above bidirectional feature fusion process can be repeated multiple times. Each repetition allows features of different scales to better communicate and complement each other, so that the finally fused feature maps contain richer and more comprehensive semantic and detailed information.
[0055] 4. Output the fused features After bidirectional feature fusion (possibly repeated multiple times), the finally fused multi-scale feature maps are obtained. These feature maps can be used for subsequent object detection, semantic segmentation and other tasks.
[0056] BiFPN effectively integrates the information of feature maps of different scales through bidirectional feature fusion paths and learnable weight mechanisms, improving the model's detection and recognition capabilities for objects of different sizes.
[0057] The output layer generates the final detection results. The output layer performs object detection on the finally fused multi-scale feature maps according to the DIOU_NMS method and the loss function, generating the final detection results.
[0058] Output (the output layer) is the last stage of the object detection network, responsible for generating the final detection results. Different-scale detection results are output through Conv (convolution module), including NMS (Non—Maximum Suppression) and the loss function. GIoU is used as the loss function in the original YOLOv5s model.
[0059] In YOLOV5s, non-maximum-suppression (NMS) is used to filter and delete multiple target boxes. Since only overlapping areas are analyzed, it is easy to miss detections in cases of occlusion or close target distances.
[0060] To solve this problem, the NMS method is improved, and the improved DIOU_NMS method replaces the traditional method.
[0061] The DIOU_NMS method is specifically as follows:
[0062] Among them: a is the prediction box, a gt is the ground-truth box; 2 (a,a gt ) is the distance between the center of the predicted box and the center of the true box; is the diagonal length of the minimum circumscribed rectangle of the two boxes; The candidate set for model checking is d i , the maximum score frame M, then the DIOU-NMS method update formula is:
[0063] Where: Q i is the score of different categories; R DIOU (M,d i ) is R DIOU About M and d i ; ε is the manually set threshold for the NMS operation; IOU is the intersection-over-union ratio; i is the number of anchor boxes corresponding to the grid.
[0064] The loss function is GIoU.
[0065] The DIOU-NMS method comprehensively considers the distance, overlapping area and aspect ratio between the predicted box and the true box, and believes that two predicted boxes with farther center point distances may be located on different targets. Combining the intersection of the two boxes and the center point distance can optimize the IOU loss, guide the learning of the center point, and regress the predicted box more accurately.
[0066] like Figure 5 As shown, the personal protective equipment wearing detection system for ultra-fine dust workers includes: An image collection module, used to monitor on-site images; The environmental monitoring module monitors on-site environmental indicators through sensors; for example, the workplace environment is monitored through temperature and humidity sensors, dust sensors, toxic gas sensors, etc.
[0067] The protective equipment detection module identifies the types of protective equipment worn by the staff according to the on-site images collected by the image collection module; The controller module judges the types of protective equipment that need to be worn on-site according to the data collected by the environmental monitoring module, and compares it with the recognition result of the protective equipment detection module. When the target recognition result does not have the required protective equipment, an alarm and notification are made.
[0068] This system uploads the sensor data collected by the workshop environment detection system to the upper computer through the environmental monitoring module. Through judgment and decision-making, it determines the types of personal protective equipment that need to be worn, determines what kind of protective equipment needs to be worn in this environment, and feeds back the determined equipment type to the recognition algorithm.
[0069] The protective equipment detection module and the controller module detect the personal protective equipment of personnel in this environment of the workshop. If it is detected that the personal protective equipment is not worn or worn abnormally, etc., the abnormal signal will be transmitted to the upper computer and the alarm arranged in the workshop for alarm, so as to remind the operators to wear or correct the wearing of personal protective equipment in time, and ensure the occupational safety and health of the operators.
[0070] This solution adopts intelligent detection methods of deep learning and machine vision, which can achieve high-precision recognition of personal protective equipment, and adaptively detect the wearing of personal protective equipment of personnel in real time in the on-site monitoring video, so as to improve the awareness of wearing safety protection articles for employees, reduce personal injuries. The research and development of this technology has important theoretical value and practical significance. This invention performs excellently in terms of detection performance, and the recognition accuracy rate, recall rate and F1 score reach 92.55%, 95.15% and 93.32% respectively, and meet the real-time detection requirements (≥25fps). The system has strong environmental adaptability, can be applicable to various lighting conditions, and has excellent anti-interference ability and stable operation ability. By realizing 24-hour uninterrupted monitoring, the workload of manual inspection is greatly reduced, and complete data analysis and traceability capabilities are provided, effectively improving the management efficiency. At the same time, this invention can significantly reduce labor costs, reduce accident risks, improve management efficiency, and has obvious economic benefits. Experiments show that this method can provide reliable technical support for the wearing detection of personal protective equipment for ultra-fine dust practitioners.
Claims
1. A detection method for the wearing of personal protective equipment for workers exposed to ultrafine dust, characterized in that, It includes the following steps: Collect images of protective equipment, construct a dataset for identifying protective equipment, and label the information of protective equipment on the images of protective equipment; Establish a detection model, establish a YOLOv5s detection model, including an input end, a backbone network, a bottleneck layer, and an output layer; Input end: Use the images of part of the dataset for identifying protective equipment as the training set to input into the detection model, perform Mosaic data augmentation, automatically calculate anchor boxes, and resize the images; Backbone network, that is, a feature extraction network, used to extract the feature information of the image. First, through slicing operations, the input image is copied four times, and each pixel takes values at every other pixel. Finally, the images are fused to obtain the underlying feature map; the underlying feature map is further extracted by ConvNexT to obtain the processed feature map. Through spatial pyramid pooling, the output of ConvNexT is pooled using three different pooling kernels for downsampling, and then spliced and fused to obtain features of various scales; Use a feature fusion network to fuse the features. The features of various scales extracted by the backbone network are subjected to multi-scale feature fusion through the cross-scale feature fusion structure BiFPN to obtain the finally fused multi-scale feature map. The output layer generates the final detection result. The output layer performs object detection on the finally fused multi-scale feature map according to the DIOU_NMS method and the loss function to generate the final detection result.
2. The individual protection equipment wearing detection method for ultrafine dust workers according to claim 1, wherein: The information of the labeled protective equipment includes the target category, bounding box coordinates, and target ID.
3. The individual protective equipment wearing detection method for ultra-fine dust workers according to claim 1, characterized in that: After collecting the images of protective equipment, data augmentation is also performed on the images. Through the Copy-Pasting data augmentation strategy, oversampling and data pasting are performed on the small targets in the dataset to provide sufficient small targets to match the anchors, thereby improving the performance of small target detection; Specifically, oversampling is performed on the small target sample data, and then two images are mixed together through Copy-Pasting to form a new image, which is added to the dataset for identifying protective equipment.
4. A method for detecting the wearing of personal protective equipment for ultra-fine dust workers according to claim 1, characterized in that: The backbone network uses the ConvNexT network for feature extraction, Specifically, in the first stage, the preprocessed input image is downsampled using a convolutional layer with the same convolutional kernel size and stride, and 3 ConvNeXt Blocks are stacked to output a feature map; Second stage: The input is the feature map output in the first stage. It is downsampled using a convolutional layer with the same convolutional kernel size and stride to further reduce the size of the feature map. 3 ConvNeXt Blocks are stacked to continue feature extraction and fusion, and a feature map is output; Third stage: The input is the feature map output in the second stage. It is downsampled again using a convolutional layer with the same convolutional kernel size and stride. 9 ConvNeXt Blocks are stacked in this stage to output a feature map; Stage 4: The input is the feature map output from Stage 3. Downsampling is performed using a convolutional layer with the same convolutional kernel size and stride. Three ConvNeXt Blocks are stacked, and finally a feature map with a size of 7×7 and 768 channels is obtained. After global average pooling and the LN layer, it is finally output through a linear classifier.
5. A method for detecting the wearing of personal protective equipment for ultra-fine dust workers according to claim 4, characterized in that: The ConvNeXt Block adopts an inverted bottleneck layer design. First, a 3×3 depthwise separable convolution is used for feature extraction, then the number of channels is increased from 96 to 384 through a 1×1 convolution for dimensionality increase, and then the number of channels is reduced back from 384 to 96 through a 1×1 convolution for dimensionality reduction to reduce high-dimensional information loss.
6. The individual protection equipment wearing detection method for ultrafine dust workers according to claim 1, characterized in that: The specific DIOU_NMS method is as follows: Among them: a is the predicted box, a gt is the ground truth box; ρ 2 (a, a gt ) is the distance between the centers of the predicted box and the ground truth box; is the diagonal length of the minimum bounding rectangle of the two boxes; The candidate set for model detection is d i , the maximum score box M, and the update formula of the DIOU-NMS method is: Where: Q i is the score of different classifications; R DIOU (M, d i ) is the value of R DIOU with respect to M and d i ; ε is the threshold manually set for the NMS operation; IOU is the intersection over union; i is the number of anchor boxes corresponding to the grid.
7. A method for detecting the wearing of personal protective equipment for ultra-fine dust workers according to claim 1, characterized in that: The loss function is GIoU.
8. An individual protection equipment wearing detection system for ultra-fine dust workers, characterized in that, It includes: An image collection module for monitoring on-site images; An environmental monitoring module for monitoring on-site environmental indicators through sensors; A protective equipment detection module for identifying the types of protective equipment worn by staff based on the on-site images collected by the image collection module; A controller module for judging the types of protective equipment required on-site based on the data collected by the environmental monitoring module and comparing it with the recognition results of the protective equipment detection module. When the target recognition result does not have the required protective equipment, an alarm and notification are issued.