Safety equipment wearing target detection method and system based on YOLOv5s

By introducing SIoU loss function and dynamic hierarchical adjustment strategy in the YOLOv5s network model, combined with the adaptive feature fusion mechanism, efficient detection of dense and small targets in port safety operations is achieved, solving the problem of insufficient detection accuracy in the shipping industry by existing algorithms, and improving detection accuracy and recall rate.

CN120340029APending Publication Date: 2025-07-18COSCO SHIPPING TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510401619.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing object detection algorithms cannot accurately identify dense targets and smaller targets in the shipping industry, resulting in the inability to achieve practical results in the detection accuracy.

Method used

The wearable object detection method of security equipment based on YOLOv5s is adopted. By introducing the SIoU loss function to replace the GIoU loss function, combined with dynamic hierarchical adjustment strategy and adaptive feature fusion mechanism, the number of neural network layers of the FPN structure is dynamically adjusted, and two-stage object detection is carried out, including small-target screening and large-target screening.

Benefits of technology

It significantly improves the detection accuracy and recall rate of small and large targets, reduces model training time, and improves the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340029A_ABST
    Figure CN120340029A_ABST
Patent Text Reader

Abstract

The invention provides a YOLOv5s-based safety equipment wearing target detection method and system, and the method comprises the steps: firstly obtaining a port safety operation image, carrying out the safety equipment wearing compliance marking according to a preset marking criterion, introducing an SIoU loss function on the basis of a YOLOv5s network model to replace a GIoU loss function in the YOLOv5s network model, and obtaining a new YOLOv5s network model; a new YOLOv5s network model neck network is constructed, a dynamic hierarchy adjustment strategy is combined with an adaptive feature fusion mechanism, the number of neural network layers of an FPN structure in the new YOLOv5s network model neck network is automatically and dynamically adjusted, an improved FPNX structure is obtained, an improved YOLOv5s network model is further obtained, then first-stage small target screening detection and second-stage large target screening detection are sequentially carried out, and a new YOLOv5s network model is obtained. And obtaining the second target detection result containing all the safety equipment in the target image to complete the wearing detection of the safety equipment, so that the detection capability of small targets and large targets can be remarkably improved to adapt to different detection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent shipping, and particularly relates to a safety equipment wearing target detection method and system based on YOLOv5s. Background Art

[0002] In the wave of global economic integration, ports, as important nodes of international trade, the level of their safe operation is directly related to the stable development of the national economy and social harmony. With the continuous growth of port cargo throughput and the increasingly complex operation environment, port safe operation has become a major issue that cannot be ignored. In this context, the importance of personal protective equipment, especially safety helmets and work clothes, has become increasingly prominent, becoming a key factor in ensuring the life safety of port workers. At the same time, effective detection and management of the behaviors of not wearing safety helmets and not wearing work clothes are also an important part of improving the port safety management level. Therefore, establishing an effective detection mechanism to timely discover and correct the behaviors of not wearing safety helmets and not wearing work clothes is also an important measure to improve the port safety management level. However, the cost of manual operation and maintenance is relatively high, and it requires people to detect the corresponding scenarios for a long time. And after a long time of work, errors are inevitable in manual operation and maintenance. With the increasing development of shipping trade, the cost of manual operation and maintenance will gradually increase, and the possibility of accidents will also increase accordingly. Therefore, the high cost and low efficiency of daily safety management operation and maintenance have become a difficult problem that the current industry urgently needs to explore.

[0003] In recent years, great breakthroughs have been achieved in the target detection algorithms of image recognition using deep learning. Especially in the field of target detection, for example: YOLO, CNN, and RNN, etc. have been extremely widely applied in various industrial fields. The safety management mechanism in the shipping industry can also be combined with artificial intelligence technology to be further optimized and improved. The detection mechanism can include various means such as daily inspections, video monitoring, and intelligent recognition. By strengthening the intensity of daily inspections, ensuring that each worker can correctly wear a safety helmet and work clothes; using video monitoring technology to comprehensively cover the operation area and timely discover and correct violations; adopting an intelligent recognition system to automatically detect the personnel entering and leaving the port to ensure that the wearing rate of personal protective equipment reaches 100%. The implementation of these measures will greatly improve the efficiency and level of port safety management and provide strong guarantee for the safe production of ports.

[0004] With the continuous breakthrough of deep learning technology, the artificial intelligence industry has changed with each passing day. Among them, the one-stage object detection algorithms represented by the YOLO (You Only Look Once) series and the two-stage object detection algorithms represented by the R-CNN (Region-Based Convolutional Neural Network) series are the most widely used. It has extremely important value in important industries, including motor vehicle autonomous driving, camera face recognition, and the identification of special situations such as fires and traffic accidents. Therefore, combining the safety management mechanism with artificial intelligence object detection technology in the shipping field can make up for the deficiencies of high cost and low efficiency brought by manual management. This technology has extremely important theoretical and practical significance both in maintaining long-term monitoring and ensuring the accuracy and integrity of monitoring results.

[0005] These current mainstream neural network detection algorithms are all based on deep learning object detection network algorithms. However, there are still the following problems in the detection of these mainstream object detection algorithms in the current safety management mechanism of the shipping industry: 1) The targets in the detection area often exist in an extremely dense phenomenon, resulting in the mainstream object detection network algorithms being unable to accurately identify each independent target, and often identifying the entire dense group as a single target. 2) The targets in the detection area often have the phenomenon that the pixel points of the targets to be detected in the data are small and the pixel color difference is small, resulting in the mainstream object detection network algorithms being unable to identify the targets, causing frequent missed detection phenomena.

[0006] To sum up, the current mainstream object detection network algorithms cannot achieve a practical effect in the recognition accuracy of dense targets and small targets, resulting in the current mainstream object detection network algorithms being unable to adapt to the requirements of relevant scenarios in the industry. Summary of the Invention

[0007] To solve the problems that the existing object detection algorithms cannot achieve a practical effect in the recognition accuracy of dense targets and small targets and cannot accurately identify each independent target, the present invention provides a method for detecting the wearing of safety equipment based on YOLOv5s, which can ensure the quality of the labeled data, greatly accelerate the convergence speed of the YOLOv5s algorithm, and significantly improve the detection ability for small and large targets to adapt to different detection scenarios. The present invention also relates to a system for detecting the wearing of safety equipment based on YOLOv5s.

[0008] The technical solution of the present invention is as follows:

[0009] A method for detecting the wearing of safety equipment based on YOLOv5s, characterized by comprising the following steps:

[0010] Steps for image acquisition and annotation: Obtain multiple images of port safety operations to be measured and historical port safety operation images for training, and label the compliance of safety equipment wearing for personnel in each historical port safety operation image according to the preset annotation criteria to obtain multiple annotated historical port safety operation images;

[0011] Steps for model improvement: Introduce the SIoU loss function on the basis of the YOLOv5s network model to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model; and adopt a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model to obtain an improved FPNX structure, and then obtain an improved YOLOv5s network model;

[0012] Steps for small target screening in the first stage: Use multiple annotated historical port safety operation images as training set samples to train the improved YOLOv5s network model to obtain a trained personnel detection model, and input the images of port safety operations to be measured into the personnel detection model for small target screening detection in the first stage to obtain the first target detection results of all personnel in the port safety operation images, and automatically crop the cropping areas containing personnel according to the first target detection results of each personnel;

[0013] Steps for large target screening in the second stage: Use a super-resolution network based on a convolutional neural network to magnify the cropping areas containing personnel to obtain magnified target images, use a part of the target images among all the target images as training set samples to train the improved YOLOv5s network model to obtain a trained safety equipment detection model, and input another part of the target images among all the target images into the safety equipment detection model for large target screening detection of safety equipment in the second stage to obtain the second target detection results of all safety equipment in the target images to complete the wearing detection of safety equipment.

[0014] Preferably, in the steps for model improvement, the SIoU loss function optimizes the prediction bounding box regression process and accelerates the convergence of the new YOLOv5s network model by introducing penalty terms for angular cost, distance cost, and shape cost, and comprehensively considering the geometric relationship between the predicted bounding box and the ground truth bounding box;

[0015] Adopting a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model specifically includes:

[0016] The dynamic hierarchical adjustment strategy is adopted to add a high-resolution neural network layer and a low-resolution neural network layer respectively on the basis of the FPN structure with multiple neural networks in the new YOLOv5s network model, so that the resolution of the added high-resolution neural network layer is greater than that of the neural network layer with the highest resolution in the multiple neural networks of the FPN structure, and the resolution of the added low-resolution neural network layer is less than that of the neural network layer with the lowest resolution in the multiple neural networks of the FPN structure;

[0017] In the first-stage small target screening step, when performing the first-stage small target detection, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the high-resolution neural network layer, and then the high-resolution neural network layer after the highest weight is assigned is used to perform the first-stage small target screening detection;

[0018] In the second-stage large target screening step, when performing the second-stage detection of the enlarged target, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the low-resolution neural network layer, and then the low-resolution neural network layer after the highest weight is assigned is used to perform the second-stage large target screening detection of safety equipment.

[0019] Preferably, in the image acquisition and annotation step, the safety equipment includes safety helmets and work clothes. The compliance annotation of the wearing of safety equipment by people in the port safety operation image according to the preset annotation criteria specifically includes:

[0020] Judge whether the people in the port safety operation image wear safety helmets according to the preset annotation criteria. If so, mark that they wear safety helmets. If not, then judge whether the head has a blurred contour. If so, do not annotate. If not, then judge whether the head is blocked. If so, do not annotate. If not, then judge whether they wear other types of hats other than safety helmets. If so, mark that they wear other hats. If not, mark that they do not wear safety helmets;

[0021] And judge whether the people in the port safety operation image wear work vests / straps / one-piece suits according to the preset annotation criteria. If so, mark that they wear work clothes; if not, then judge whether the color of the upper body is solid. If not, mark that they do not wear work clothes; if so, then judge whether the pants are shorts. If so, mark that they do not wear work clothes; if not, mark that they wear work clothes.

[0022] Preferably, after the second-stage large target screening step, it further includes a two-stage detection result comparison and analysis step: comparing the first-stage detection result output by the personnel detection model with the second-stage detection result output by the safety equipment detection model. If the target is not detected in the first-stage detection result but is detected in the second-stage detection result, then the second-stage detection result is used as the final detection result; if the target is detected in both the first-stage detection result and the second-stage detection result, then the second-stage detection result is used as the final detection result; if the target is detected in the first-stage detection result but not detected in the second-stage detection result, then the first-stage detection result is used as the final detection result.

[0023] Preferably, both the first target detection result and the second target detection result include a bounding box, a class label, and a confidence score. The bounding box score is calculated using the SIoU loss function, the class probability score is calculated using the class prediction loss function, and the confidence score is calculated using the confidence prediction loss function.

[0024] A safety equipment wearing target detection system based on YOLOv5s, characterized in that it includes an image acquisition and annotation module, a model improvement module, a first-stage small target screening module, and a second-stage large target screening module connected in sequence.

[0025] The image acquisition and annotation module acquires multiple to-be-detected port safety operation images and historical port safety operation images for training, and performs compliance annotation on the safety equipment worn by personnel in each historical port safety operation image according to a preset annotation criterion to obtain multiple annotated historical port safety operation images.

[0026] The model improvement module introduces the SIoU loss function on the basis of the YOLOv5s network model to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model; and adopts a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model to obtain an improved FPNX structure, and further obtains an improved YOLOv5s network model.

[0027] The first-stage small target screening module uses multiple annotated historical port safety operation images as training set samples to train the improved YOLOv5s network model to obtain a trained personnel detection model, and inputs the to-be-detected port safety operation images into the personnel detection model for the first-stage small target screening detection to obtain a first target detection result containing all personnel in the port safety operation image, and automatically crops the cropping area containing the personnel according to the first target detection result of each personnel.

[0028] The second-stage large target screening module uses a super-resolution network based on a convolutional neural network to magnify the cropped area containing personnel to obtain a magnified target image. A part of the target images among all the target images are used as training set samples to train the improved YOLOv5s network model, obtaining a trained safety equipment detection model. Another part of the target images among all the target images are input into the safety equipment detection model for second-stage large target screening detection of safety equipment, obtaining the second target detection results of all safety equipment in the target image to complete the wearing detection of safety equipment.

[0029] Preferably, in the model improvement module, a dynamic hierarchical adjustment strategy is combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers in the FPN structure of the neck network of the new YOLOv5s network model. Specifically, it includes:

[0030] Using the dynamic hierarchical adjustment strategy, a high-resolution neural network layer and a low-resolution neural network layer are respectively added on the basis of the FPN structure with multiple neural networks in the new YOLOv5s network model, so that the resolution of the added high-resolution neural network layer is greater than the neural network layer with the highest resolution in the multiple neural networks of the FPN structure, and the resolution of the added low-resolution neural network layer is less than the neural network layer with the lowest resolution in the multiple neural networks of the FPN structure;

[0031] In the first-stage small target screening module, when performing the first-stage small target detection, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the high-resolution neural network layer, and then the first-stage small target screening detection is performed using the high-resolution neural network layer after the highest weight is assigned;

[0032] In the second-stage large target screening module, when performing the second-stage detection of the magnified target, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the low-resolution neural network layer, and then the second-stage large target screening detection of safety equipment is performed using the low-resolution neural network layer after the highest weight is assigned.

[0033] Preferably, in the image acquisition and annotation module, the safety equipment includes safety helmets and work clothes. The compliance annotation of the wearing of safety equipment by people in the port safety operation image is carried out according to the preset annotation criteria. Specifically, it includes:

[0034] Judge whether the person in the port safety operation image wears a safety helmet according to the preset annotation criteria. If so, annotate that the safety helmet is worn. If not, then judge whether the head has a blurred contour. If so, do not annotate. If not, then judge whether the head is blocked. If so, do not annotate. If not, then judge whether the person wears other types of hats besides the safety helmet. If so, annotate that other hats are worn. If not, annotate that the safety helmet is not worn;

[0035] And judge whether the person in the port safety operation image wears a work vest / strap / one-piece suit according to the preset annotation criteria. If so, annotate that the work clothes are worn. If not, then judge whether the color of the upper garment is a solid color. If not, annotate that the work clothes are not worn. If so, then judge whether the trousers are shorts. If so, annotate that the work clothes are not worn. If not, annotate that the work clothes are worn.

[0036] Preferably, it further includes a two-stage detection result comparison and analysis module. The two-stage detection result comparison and analysis module is connected to the second-stage large target screening module, and is used to compare the first-stage detection result output by the personnel detection model with the second-stage detection result output by the safety equipment detection model. If the first-stage detection result does not detect the target but the second-stage detection result detects the target, then use the second-stage detection result as the final detection result; if both the first-stage detection result and the second-stage detection result detect the target, then use the second-stage detection result as the final detection result; if the first-stage detection result detects the target but the second-stage detection result does not detect the target, then use the first-stage detection result as the final detection result.

[0037] Preferably, both the first target detection result and the second target detection result include a bounding box, a class label, and a confidence score. The SIoU loss function is used to calculate the bounding box score, the class prediction loss function is used to calculate the class probability score, and the confidence prediction loss function is used to calculate the confidence score.

[0038] The beneficial effects of the present invention are:

[0039] A method for detecting the wearing of safety equipment based on YOLOv5s provided by the present invention first obtains multiple images of port safety operations to be measured and historical port safety operation images for training, and labels the compliance of personnel wearing safety equipment in each historical port safety operation image according to a preset annotation criterion to obtain multiple annotated historical port safety operation images. By formulating a complete and clear data annotation criterion, the data required for model training can be standardized and normalized, and the situation where the model training data is contaminated due to blurred annotation boundaries, resulting in the model being unable to converge, can be avoided. The phenomenon of false alarms of the model can be effectively avoided, ensuring the high quality of the annotated image data, and effectively improving the overall detection accuracy and recall rate of the YOLOv5s algorithm. And on the basis of the YOLOv5s network model, the SIoU loss function is introduced to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model. By using the SIoU loss function instead of the GIoU loss function, the mismatch between the expected true box (true bounding box) and the predicted box (predicted bounding box) is taken into consideration, and the penalty metric is redefined, considering the vector angle between the expected regressions. On the premise of ensuring the advantages of the GIoU algorithm, the convergence speed of the YOLOv5s algorithm is greatly accelerated. The time required for model training convergence can be significantly reduced while improving the accuracy, and then the number of rounds and time required for model training are greatly reduced. Then, a dynamic hierarchical adjustment strategy is combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model to obtain an improved FPN structure, and then an improved YOLOv5s network model is obtained. By dynamically adjusting the FPN structure according to the different input image resolutions and the different target recognition accuracies in the image, different detection scenarios can be adapted, thereby optimizing the multi-scale target detection performance. Through hierarchical expansion and adaptive feature fusion, the detection accuracy of small and large targets is significantly improved, making the detection of small and large targets by the model more accurate and complete. Finally, the images of port safety operations to be measured are input into the trained human detection model for the first-stage small target screening detection to obtain small target images with people, and a super-resolution network based on a convolutional neural network is used to magnify the target images with people. It can greatly improve the image resolution and retain or enhance the image details, effectively reducing the blur problem caused by magnification. Then, the magnified target images are input into the trained safety equipment detection model for the second-stage large target screening detection to obtain operation images containing safety equipment to complete the wearing detection of safety equipment, effectively improving the detection accuracy.

[0040] The present invention preliminarily screens the labeled port safety operation images by using a trained personnel detection model, and detects almost all the personnel in the port safety operation images. At this time, since the personnel in the port safety operation images account for a relatively small proportion in the whole image, the small target detection of the personnel targets in the whole image is carried out first, and safety equipment such as safety helmets and work clothes is not detected. Then, a super-resolution network based on a convolutional neural network is used to magnify the target image with people, and then the trained safety equipment detection model is used to further carry out fine detection of large targets such as safety helmets and work clothes. Since it is the detection after magnifying the personnel cropping area, it is large target detection, accurately detecting the operation images containing safety equipment, greatly improving the detection rate and accuracy. By distinguishing the large and small targets in two stages and detecting them in sequence, it not only further ensures the detection accuracy and recall rate of the model, but also makes the overall detection result of the model more reliable.

[0041] The present invention also relates to a safety equipment wearing target detection system based on YOLOv5s. This system corresponds to the above-mentioned safety equipment wearing target detection method based on YOLOv5s, and can be understood as a system for implementing the above-mentioned safety equipment wearing target detection method based on YOLOv5s. It includes an image acquisition and annotation module, a model improvement module, a first-stage small target screening module, and a second-stage large target screening module connected in sequence. Each module works in coordination with each other. By formulating complete, clear, and easy-to-understand data annotation criteria, the data required for model training can be standardized and normalized, and it will not cause the model training data to be contaminated due to fuzzy annotation boundaries, resulting in the model being unable to converge, effectively avoiding the phenomenon of model false alarms; by replacing the original GIoU loss function with the SIoU loss function, this loss function can significantly reduce the time required for model training to converge on the premise of improving accuracy, and further significantly reduce the number of rounds and time required for model training. The dynamic hierarchical adjustment strategy is combined with the adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers in the FPN structure of the neck network of the new YOLOv5s network model, making the detection of small and large targets by the model more accurate and complete, and greatly improving the detection rate and accuracy through two-stage segmented detection. Brief Description of the Drawings

[0042] Figure 1 It is a flowchart of the safety equipment wearing target detection method based on YOLOv5s of the present invention.

[0043] Figure 2 It is a schematic diagram of the safety helmet annotation process of the present invention.

[0044] Figure 3 It is a schematic diagram of the work clothes annotation process of the present invention.

[0045] Figure 4 It is a schematic diagram comparing FPN and FPNX of the present invention.

[0046] Figure 5 It is a schematic diagram of the present invention after enlarging the cropped area including the person. DETAILED DESCRIPTION

[0047] The present invention will be described below in conjunction with the accompanying drawings.

[0048] The present invention relates to a safety equipment wearing target detection method based on YOLOv5s. In order to solve the technical problem that the existing mainstream target detection algorithm cannot accurately identify targets without helmets and work clothes under the safety management mechanism of the shipping industry due to the fact that the targets are often dense and the pixels are small, the present invention uses data annotated according to the annotation criteria as input, based on YOLOv5_6.0 (hereinafter referred to as YOLOv5), replaces the original FPN structure with the FPNX structure, uses the SIoU loss function to replace the GIoU loss function, and after obtaining the detection results, the second stage small model is used to re-check the first stage detection results. The flowchart of the method is as follows Figure 1 As shown, the following steps are included in sequence:

[0049] 1. Image acquisition and labeling steps: obtain multiple port safety operation images to be tested and historical port safety operation images used for training, and label the compliance of safety equipment wearing of personnel in each historical port safety operation image according to preset labeling criteria to obtain multiple labeled historical port safety operation images.

[0050] At present, when manually labeling images of port safety operations, the image scenes are often inconsistent, and the standard rules are not unified and clear. This makes data labelers confused about whether the targets in the image data should be labeled as helmets and work clothes, which ultimately leads to image data labeling errors and contamination of the image data sent to training, affecting the overall detection accuracy and recall rate of the YOLOv5 algorithm. Therefore, in order to standardize the image data labeling rules and ensure the quality of the labeled data, labeling criteria and processes for labeling helmets and work clothes targets have been formulated, and the compliance of the safety equipment (i.e., helmets and work clothes) worn by people in each historical port safety operation image is labeled according to the preset labeling criteria, and multiple labeled historical port safety operation images are obtained. Specifically, Figure 2 and 3As shown in the figure, the labels for safety helmets and work clothes in the present invention are specifically divided into 5 categories, and all safety helmets and work clothes targets are labeled according to these 5 categories: 1) Wearing a safety helmet; 2) Wearing other hats; 3) Not wearing a safety helmet; 4) Wearing work clothes; 5) Not wearing work clothes. It can be seen that there is a category of "wearing other hats" among the categories of safety helmets. Since the targets of the safety helmet-related categories are small and the features are relatively blurred, the YOLOv5 algorithm is prone to misjudgment. Therefore, the present invention decides to add the category of "wearing other hats" as an adversarial label for safety helmet-related labels to improve the detection accuracy of the YOLOv5s algorithm.

[0051] As Figure 2 shown, this process is the labeling process for defining safety helmet targets. When a person appears in the port safety operation image, it starts to judge whether to label the wearing of a safety helmet. First, according to the preset labeling criteria, it judges whether the person target in the port safety operation image wears a safety helmet (that is, a dome-shaped reflective plastic safety helmet, similar to the safety helmet worn at the construction site). If so (wearing a safety helmet), then label wearing a safety helmet; if not (not wearing), then enter the next process. Secondly, it judges whether the head of the person target has a blurred contour and the human eyes cannot recognize clearly. If so (the head contour of the person target is not clear), then no labeling is carried out; if not (the contour is clear), then enter the next process. Thirdly, it judges whether the top of the head of the person target is blocked. If so (the top of the head of the person target is blocked), then no labeling is carried out; if not (not blocked), then enter the next process. Finally, it judges whether the person target wears other types of hats except the safety helmet (that is, non-safety helmet hats). If so (the person target wears other types of hats), then label wearing other hats; if not (not wearing), then label not wearing a safety helmet.

[0052] Similarly, as Figure 3 shown, this process is the labeling process for defining work clothes-related targets. When a person appears in the port safety operation image, it starts to judge whether to label the wearing of work clothes. First, according to the preset labeling criteria, it judges whether the person target in the port safety operation image wears work clothes (standard crew work clothes, including coveralls, reflective vests, and suspenders). If so (wearing work clothes), then label wearing work clothes; if not (not wearing), then enter the next process. Secondly, it judges whether the color of the upper body of the person target's body part is a solid color. If not (not a solid color), then label not wearing work clothes; if so (is a solid color), then enter the next process. Finally, it judges whether the pants of the person target's body part are shorts. If so (shorts), then label not wearing work clothes; if not (not shorts), then label wearing work clothes.

[0053] II. Model improvement steps: On the basis of the YOLOv5s network model, introduce the SIoU loss function to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model; and adopt a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model to obtain an improved FPNX structure, and then obtain an improved YOLOv5s network model.

[0054] Specifically, the loss function in the YOLOv5 algorithm consists of three parts: bounding box regression loss, class prediction loss, and confidence prediction loss. Among them, the Generalized Intersection over Union Loss (GIoU Loss) function is used to calculate the bounding box score, the class prediction loss is used to calculate the class probability score, and the confidence prediction loss function is used to calculate the objectness score.

[0055] Although the GIoU loss function can solve the problem of ineffective recognition of overlapping targets compared with the loss functions of previous versions, it is very dependent on the closed box of IoU, resulting in the need for more iterations and the inability of the YOLOv5s algorithm to converge well. Therefore, the present invention introduces the SIoU loss function on the basis of the YOLOv5s network model to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model; by using the SIoU loss function instead of the GIoU loss function, the situation of mismatch between the expected true box and the predicted box is taken into consideration, and the penalty metric is redefined, considering the vector angle between the expected regressions. On the premise of ensuring the advantages of the GIoU algorithm, the convergence speed of the YOLOv5s algorithm is greatly accelerated.

[0056] Furthermore, the introduced SIoU loss function consists of four parts: Angle Cost, Distance Cost, Shape Cost, and IoU Cost. Among them, Angle Cost calculates the angular difference between the center points of the ground truth box and the predicted box, Distance Cost calculates the distance difference between the center points of the ground truth box and the predicted box, Shape Cost calculates the shape difference between the ground truth box and the predicted box, and IoU Cost calculates the intersection over union of the ground truth box and the predicted box. That is to say, by introducing penalty terms for Angle Cost, Distance Cost, and Shape Cost, the SIoU loss function comprehensively considers the geometric relationship between the predicted bounding box and the ground truth bounding box (such as angular difference, center point difference, aspect ratio shape difference), thereby optimizing the regression process of the predicted bounding box. Compared with GIoU, SIoU can accelerate the model convergence speed and improve the detection accuracy.

[0057] The FPN structure in the neck part of the new YOLOv5s network model is called the Feature Pyramid, and its structure consists of a series of bottom-up processes, top-down processes, and lateral connections. Through these operations, multiple images with different scales are obtained to distinguish and detect objects of different sizes with different features in the images. In the present invention, by adopting a dynamic hierarchical adjustment strategy, a high-resolution neural network layer and a low-resolution neural network layer are respectively added on the basis of the FPN structure with multiple neural networks in the new YOLOv5s network model to obtain an improved FPNX structure, so that the resolution of the added high-resolution neural network layer is greater than that of the neural network layer with the highest resolution in the multiple neural networks of the FPN structure, and the resolution of the added low-resolution neural network layer is less than that of the neural network layer with the lowest resolution in the multiple neural networks of the FPN structure; and when performing the first-stage small object detection subsequently, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the high-resolution neural network layer, and then the high-resolution neural network layer after the highest weight is assigned is used for the first-stage small object screening detection; while when performing the second-stage detection of the enlarged objects, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the low-resolution neural network layer, and then the low-resolution neural network layer after the highest weight is assigned is used for the second-stage screening detection of large objects of safety equipment.

[0058] Among them, such as Figure 4As shown in the figure, when the improved FPNX structure has 5 layers, the input layer has a size of 640*640. After downsampling by 1 / 2, layer 1 is obtained, that is, the size of layer 1 is 320*320. And so on until layer 5, the size of layer 5 is 20*20. Layer 6 is directly copied from layer 5. By upsampling (i.e., layer 5) by a factor of 2, layer 6 is obtained, that is, the size of layer 6 is 40*40. Then, a 1*1 convolution is used to reduce the dimension of layer 4, reducing the number of convolution kernels in layer 4 to reduce the data calculation amount. Then, the corresponding elements of layer 4 and layer 6 are fused to summarize the high and low level features, obtaining layer 7, that is, the size of layer 7 is still 40*40. Then, a 1*1 convolution is used to reduce the dimension of layer 3, and layer 3 is fused with layer 7 enlarged by a factor of 2, obtaining layer 8, that is, the size of layer 8 is 80*80. And so on to obtain layer 9 and layer 10. The size of layer 9 is 160*160, and the size of layer 10 is 320*320. Among them, the specific layer numbers of layer 1 to layer 5 correspond to the formula:

[0059]

[0060] In the above formula, n0 is 5, p s is the size of the input layer, designed as 640, wh is the network size of the current layer. For example, for layer 1, it is 320*320, and the corresponding n value for layer 1 is 4. n is the remaining block layer number of the current network.

[0061] III. First-stage small target screening steps: Use multiple labeled historical port safety operation images as training set samples to train the improved YOLOv5s network model to obtain a trained personnel detection model. Then, since the personnel in the to-be-detected port safety operation image occupy a relatively small proportion in the entire image, it is necessary to screen and detect the personnel targets in the to-be-detected port safety operation image as small targets. Therefore, the to-be-detected port safety operation image is input into the personnel detection model for the first-stage small target screening detection, obtaining the first target detection result containing all personnel in the port safety operation image, and automatically cropping the cropping area containing the personnel according to the first target detection result of each personnel.

[0062] For example, in the original FPN structure, the resolutions of the multi-layer neural networks from high to low are the three neural networks of C3, C4, and C5. First, a dynamic hierarchical adjustment strategy is adopted to add a C2 high-resolution neural network layer before C3 to detect small targets, and the resolution of C2 is higher than that of C3. A C6 low-resolution neural network layer is added after C5 to detect large targets, and the resolution of C6 is lower than that of C5. Then, when detecting small targets, the attention mechanism in the adaptive feature fusion mechanism is used to assign different weights to C2, C3, C4, C5, and C6 respectively, and the highest weight is assigned to the C2 high-resolution neural network layer. Furthermore, the C2 high-resolution neural network layer with the highest weight assigned is used for small target screening and detection.

[0063] IV. Steps for screening large targets in the second stage: Although the improved YOLOv5s network model has greatly improved in detection accuracy and recall rate, it still cannot achieve a good detection effect in some complex and difficult scenarios. In these scenarios, in the small target screening and detection in the first stage, human targets can generally be recognized, but safety helmet and work clothes targets related to people cannot be recognized. Even in general scenarios, due to algorithm errors, there will inevitably be a small number of false alarms and missed detections. Therefore, first, a super-resolution network based on a convolutional neural network is used to magnify the cropped area containing people to obtain the magnified target image. Then, a part of the target images among all the target images is used as the training set sample to train the improved YOLOv5s network model to obtain a trained safety equipment detection model. Then, since the target image is a magnified image of a person and the person in the target image wears safety equipment, another part of the target images among all the target images is input into the safety equipment detection model for the second-stage safety equipment large target screening and detection to obtain the second target detection result containing all the safety equipment in the target image, so as to complete the wearing detection of the safety equipment.

[0064] For example, in the original FPN structure, the resolutions of the multi-layer neural networks from high to low are the three neural networks of C3, C4, and C5. First, a dynamic hierarchical adjustment strategy is adopted to add a high-resolution neural network layer C2 before C3 to detect small targets, and the resolution of C2 is higher than that of C3. And a low-resolution neural network layer C6 is added after C5 to detect large targets, and the resolution of C6 is lower than that of C5. When detecting the enlarged target image, the attention mechanism in the adaptive feature fusion mechanism is used to assign different weights to C2, C3, C4, C5, and C6 respectively, and the highest weight is assigned to the low-resolution neural network layer C6. Then, the low-resolution neural network layer C6 with the highest weight assigned is used for large target screening and detection. By combining the dynamic hierarchical adjustment strategy with the adaptive feature fusion mechanism, the improved FPNX structure can be dynamically adjusted to adapt to different detection scenarios, significantly improving the detection accuracy of small and large targets. At the same time, the target of not wearing safety helmets and work clothes can be effectively distinguished and recognized from non-such targets, improving the detection accuracy and recall rate of the YOLOv5 algorithm.

[0065] Preferably, both the first target detection result and the second target detection result include a bounding box, a class label, and a confidence score. The bounding box score is calculated using the SIoU loss function, the class probability score is calculated using the class prediction loss function, and the confidence score is calculated using the confidence prediction loss function.

[0066] Since the training and test data used in the second-stage YOLOv5s are all enlarged personnel targets, the features of each target image are extracted after being enlarged compared with the original image, thus obtaining more and more complete feature information. Therefore, the second-stage safety equipment detection model with more feature information has higher detection accuracy in terms of detection effect compared with the first-stage personnel detection model. As Figure 5 shown, the person target marked in the original image data will be cropped out and enlarged to obtain the second-stage detection data on the right. It can be clearly seen from this that the person target after cropping and enlargement has more definite, clearer, and richer feature information compared with the person target in the original image data. Furthermore, the second-stage target detection algorithm (i.e., the safety equipment detection model) can obtain more accurate results through more information.

[0067] V. Steps for comparative analysis of two-stage detection results. This step is an optional step. After obtaining the second target detection result, the first-stage detection result output by the personnel detection model is also compared with the second-stage detection result output by the safety equipment detection model. If the target is not detected in the first-stage detection result but is detected in the second-stage detection result, the second-stage detection result is used as the final detection result; if the target is detected in both the first-stage detection result and the second-stage detection result, the second-stage detection result is used as the final detection result; if the target is detected in the first-stage detection result but not detected in the second-stage detection result, the first-stage detection result is used as the final detection result.

[0068] Example:

[0069] This experiment was conducted on the Linux operating system, with 4 NVIDA A100 GPUs each with 40GB of video memory, and the deep learning framework was PyTorch 2.1.0. In order to meet the deployment requirements in places such as ship terminals and ports where high-quality hardware is not available, and to maintain relatively stable and high-quality detection accuracy, the pre-trained weights selected in the experiment of the present invention were YOLOv5s. The Mosaic data augmentation strategy built into YOLOv5 was used in the experiment, and the key hyperparameters included in the experiment are shown in Table 1 below:

[0070] Table 1

[0071] Experiment hyperparameters Experiment hyperparameter values Experiment hyperparameters Experiment hyperparameter values Initial learning rate 0.01 Training image size 960 Learning rate momentum 0.937 Number of training epochs 200 IoU training threshold 0.2 Batch-size 300 IoU loss coefficient 0.05 Number of networks passed in each time 64

[0072] In the experiment, the classification of the real situation and the prediction situation is shown in Table 2. Among them, True Positive (TP) represents that both the real situation and the prediction situation are positive examples, False Positive (FP) represents that the real situation is a negative example while the prediction is a positive example, False Negative (FN) represents that the real situation is a positive example while the prediction is a negative example, and True Negative (TN) represents that both the real situation and the prediction situation are negative examples.

[0073] Table 2

[0074]

[0075] The measurement metrics used in the experiments of this invention include Precision (P), Recall (R), and mean Average Precision (mAP). Precision is the proportion of correctly predicted positive examples among all predicted positive examples; Recall is the proportion of correctly predicted positive examples among all true positive examples; mean Average Precision is obtained from the mean of the Average Precision (AP) greater than the IoU threshold, and Average Precision is the average of the Precision values over the area enclosed by the Precision-Recall curve (PR curve) and the coordinate axes. Generally, it is calculated using the following integral method:

[0076] Precision = TP / (TP + FP)

[0077] Recall = TP / (TP + FN)

[0078]

[0079] Training Results and Comparative Analysis:

[0080] The training results of the improved YOLOv5s model algorithm in the second stage are shown in Table 3 as follows.

[0081] Table 3

[0082]

[0083]

[0084] In Table 3, "YOLOv5s-HC" is the YOLOv5s model improved in the second stage of this invention, which is specifically used to detect targets without wearing a helmet and without wearing work clothes. After the original YOLOv5s object detection algorithm (YOLOv5s) is optimized by this invention and combined with the second-stage re-inspection, the mAP for targets without wearing a helmet and without wearing work clothes can reach 97.6% and 87.2% respectively. To further verify the effectiveness of the optimization of the object detection algorithm by this invention, the algorithm of this invention and the original YOLOv5s algorithm are now tested on the same test set, and the comparison results of various performance indicators are shown in Table 4.

[0085] Table 4

[0086]

[0087] As can be seen from Table 4, the optimized object detection algorithm (YOLOv5s-HC) of the present invention has better detection effects on those not wearing safety helmets and work clothes. The mAP has been improved by 4-5 percentage points, and the test results are relatively ideal, proving that the optimization strategy of the present invention, including the optimization of the YOLOv5 algorithm and the retest of the object detection algorithm in the second stage, can effectively improve the detection effects of the YOLOv5 object detection algorithm on safety helmets and work clothes. Its model attributes with size s and significantly improved detection accuracy can be integrated into a complete safety inspection framework, can be widely deployed in various production sites, and have strong application and promotion value.

[0088] The present invention also relates to a safety equipment wearing object detection system based on YOLOv5s. This system corresponds to the above-mentioned safety equipment wearing object detection method based on YOLOv5s and can be understood as a system for implementing the above method. The system includes an image acquisition and annotation module, a model improvement module, a first-stage small object screening module, and a second-stage large object screening module that are connected in sequence. Specifically,

[0089] The image acquisition and annotation module acquires multiple to-be-tested port safety operation images and historical port safety operation images for training, and performs compliance annotation on the safety equipment worn by personnel in each historical port safety operation image according to a preset annotation criterion to obtain multiple annotated historical port safety operation images;

[0090] The model improvement module introduces the SIoU loss function on the basis of the YOLOv5s network model to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model; and adopts a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model to obtain an improved FPNX structure, and further obtains an improved YOLOv5s network model;

[0091] The first-stage small object screening module uses multiple annotated historical port safety operation images as training set samples to train the improved YOLOv5s network model to obtain a trained personnel detection model, and inputs the to-be-tested port safety operation images into the personnel detection model for first-stage small object screening detection to obtain a first object detection result containing all personnel in the port safety operation image, and automatically crops the cropping area containing the personnel according to the first object detection result of each personnel;

[0092] The second-stage large target screening module uses a super-resolution network based on a convolutional neural network to magnify the cropped area containing personnel to obtain a magnified target image. A part of the target images among all the target images are used as training set samples to train an improved YOLOv5s network model, obtaining a trained safety equipment detection model. Another part of the target images among all the target images are input into the safety equipment detection model for second-stage large target screening detection of safety equipment, obtaining the second target detection results of all safety equipment in the target image to complete the wearing detection of safety equipment.

[0093] Preferably, in the model improvement module, a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism is used to automatically dynamically adjust the number of neural network layers in the FPN structure of the neck network of the new YOLOv5 network model. Specifically, it includes:

[0094] Using the dynamic hierarchical adjustment strategy, on the basis of the FPN structure with multiple neural networks in the new YOLOv5s network model, a high-resolution neural network layer and a low-resolution neural network layer are respectively added, so that the resolution of the added high-resolution neural network layer is greater than the neural network layer with the highest resolution in the multiple neural networks of the FPN structure, and the resolution of the added low-resolution neural network layer is less than the neural network layer with the lowest resolution in the multiple neural networks of the FPN structure;

[0095] In the first-stage small target screening module, when performing the first-stage small target detection, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the high-resolution neural network layer, and then the first-stage small target screening detection is performed using the high-resolution neural network layer after the highest weight is assigned;

[0096] In the second-stage large target screening module, when performing the second-stage detection of the magnified target, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the low-resolution neural network layer, and then the second-stage large target screening detection of safety equipment is performed using the low-resolution neural network layer after the highest weight is assigned.

[0097] Preferably, in the image acquisition and annotation module, the safety equipment includes safety helmets and work clothes. The compliance annotation of the wearing of safety equipment by people in the port safety operation image is carried out according to the preset annotation criteria. Specifically, it includes:

[0098] Judge whether the person in the port safety operation image wears a safety helmet according to the preset annotation criteria. If so, annotate that the person wears a safety helmet. If not, then judge whether the head has a blurred outline. If so, do not annotate. If not, then judge whether the head is blocked. If so, do not annotate. If not, then judge whether the person wears other types of hats besides the safety helmet. If so, annotate that the person wears other hats. If not, annotate that the person does not wear a safety helmet;

[0099] And judge whether the person in the port safety operation image wears a work vest / strap / one-piece suit according to the preset annotation criteria. If so, annotate that the person wears work clothes; If not, then judge whether the color of the upper garment is a solid color. If not, annotate that the person does not wear work clothes; If so, then judge whether the trousers are shorts. If so, annotate that the person does not wear work clothes; If not, annotate that the person wears work clothes.

[0100] Preferably, it further includes a two-stage detection result comparison and analysis module. The two-stage detection result comparison and analysis module is connected to the second-stage large target screening module and is used to compare the first-stage detection result output by the personnel detection model with the second-stage detection result output by the safety equipment detection model. If the first-stage detection result does not detect the target but the second-stage detection result detects the target, then use the second-stage detection result as the final detection result; If both the first-stage detection result and the second-stage detection result detect the target, then use the second-stage detection result as the final detection result; If the first-stage detection result detects the target but the second-stage detection result does not detect the target, then use the first-stage detection result as the final detection result.

[0101] Preferably, both the first target detection result and the second target detection result include a bounding box, a class label, and a confidence score. The bounding box score is calculated using the SIoU loss function, the class probability score is calculated using the class prediction loss function, and the confidence score is calculated using the confidence prediction loss function.

[0102] The present invention provides an objective and scientific method and system for detecting the wearing of safety equipment based on YOLOv5s. By formulating complete, clear, and easy-to-understand data annotation criteria, the data required for model training can be standardized and normalized, preventing the model training data from being contaminated due to ambiguous annotation boundaries, which could cause the model to fail to converge and effectively avoiding the occurrence of model false alarms. By replacing the original GIoU loss function with the SIoU loss function, this loss function can significantly reduce the time required for model training to converge while improving the accuracy, thereby greatly reducing the number of epochs and time required for model training. The dynamic hierarchical adjustment strategy combined with the adaptive feature fusion mechanism is used to automatically dynamically adjust the number of neural network layers in the FPN structure of the neck network of the new YOLOv5s network model, making the model more accurate and complete in detecting small and large targets, and significantly improving the detection rate and accuracy through two-stage segmented detection.

[0103] It should be noted that the above specific implementation manners can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the present invention can still be modified or equivalently replaced. In short, all technical solutions and their improvements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the patent of the present invention.

Claims

1. A safety equipment wearing target detection method based on YOLOv5s, characterized in that, It includes the following steps: Image acquisition and annotation step: Acquire multiple images of port safety operations to be measured and historical port safety operation images for training, and annotate the compliance of personnel wearing safety equipment in each historical port safety operation image according to the preset annotation criteria to obtain multiple annotated historical port safety operation images; Model improvement step: On the basis of the YOLOv5s network model, introduce the SIoU loss function to replace the GIoU loss function in the YOLOv5s network model to obtain a new YOLOv5s network model; and adopt a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers in the FPN structure of the neck network of the new YOLOv5s network model to obtain an improved FPNX structure, and then obtain an improved YOLOv5s network model; First-stage small target screening step: Use multiple annotated historical port safety operation images as training set samples to train the improved YOLOv5s network model to obtain a trained personnel detection model, and input the images of port safety operations to be measured into the personnel detection model for the first-stage small target screening detection to obtain the first target detection result containing all personnel in the port safety operation image, and automatically crop the cropping area containing the personnel according to the first target detection result of each personnel; Second-stage large target screening step: Use a super-resolution network based on a convolutional neural network to magnify the cropping area containing the personnel to obtain an enlarged target image, use a part of the target images in all the target images as training set samples to train the improved YOLOv5s network model to obtain a trained safety equipment detection model, and input another part of the target images in all the target images into the safety equipment detection model for the second-stage safety equipment large target screening detection to obtain the second target detection result containing all the safety equipment in the target image to complete the wear detection of the safety equipment.

2. The method for detecting the wearing target of safety equipment based on YOLOv5s according to claim 1, characterized in that, In the model improvement step, the SIoU loss function optimizes the prediction bounding box regression process and accelerates the convergence of the new YOLOv5s network model by introducing penalty terms for angle cost, distance cost, and shape cost, and comprehensively considering the geometric relationship between the predicted bounding box and the true bounding box; Adopting a dynamic hierarchical adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers in the FPN structure of the neck network of the new YOLOv5s network model specifically includes: Adopting a dynamic hierarchical adjustment strategy to add a high-resolution neural network layer and a low-resolution neural network layer respectively on the basis of the FPN structure with multiple neural networks in the new YOLOv5s network model, so that the resolution of the added high-resolution neural network layer is greater than the neural network layer with the highest resolution in the multiple neural networks of the FPN structure, and the resolution of the added low-resolution neural network layer is less than the neural network layer with the lowest resolution in the multiple neural networks of the FPN structure; In the first-stage small target screening step, when performing the first-stage small target detection, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the high-resolution neural network layer, and then the high-resolution neural network layer with the highest weight assigned is used to perform the first-stage small target screening detection; In the second-stage large target screening step, when performing the second-stage enlarged target detection, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the low-resolution neural network layer, and then the low-resolution neural network layer with the highest weight assigned is used to perform the second-stage large target screening detection of safety equipment.

3. The method for detecting the wearing target of safety equipment based on YOLOv5s according to claim 1, wherein, In the image acquisition and annotation step, the safety equipment includes safety helmets and work clothes. The compliance annotation of the safety equipment worn by people in the port safety operation image according to the preset annotation criteria specifically includes: Judge whether the person in the port safety operation image wears a safety helmet according to the preset annotation criteria. If so, mark wearing a safety helmet. If not, then judge whether the head is a blurred contour. If so, do not annotate. If not, then judge whether the head is blocked. If so, do not annotate. If not, then judge whether other types of hats other than the safety helmet are worn. If so, mark wearing other hats. If not, mark not wearing a safety helmet; And judge whether the person in the port safety operation image wears a work vest / strap / one-piece suit according to the preset annotation criteria. If so, mark wearing work clothes; if not, then judge whether the color of the upper body is solid. If not, mark not wearing work clothes; if so, then judge whether the pants are shorts. If so, mark not wearing work clothes; if not, mark wearing work clothes.

4. The method for detecting the wearing target of safety equipment based on YOLOv5s according to any one of claims 1 to 3, characterized in that, After the second-stage large target screening step, it further includes a two-stage detection result comparison and analysis step: comparing the first-stage detection result output by the personnel detection model with the second-stage detection result output by the safety equipment detection model. If the first-stage detection result does not detect a target but the second-stage detection result detects a target, then the second-stage detection result is used as the final detection result; If both the first-stage detection result and the second-stage detection result detect a target, then the second-stage detection result is used as the final detection result; If the first-stage detection result detects a target but the second-stage detection result does not detect a target, then the first-stage detection result is used as the final detection result.

5. The method for detecting the wearing target of safety equipment based on YOLOv5s according to any one of claims 1 to 3, characterized in that, Both the first target detection result and the second target detection result include a bounding box, a class label, and a confidence score. The SIoU loss function is used to calculate the bounding box score, the class prediction loss function is used to calculate the class probability score, and the confidence prediction loss function is used to calculate the confidence score.

6. A safety equipment wearing target detection system based on YOLOv5s, characterized in that, It includes an image acquisition and annotation module, a model improvement module, a first-stage small target screening module, and a second-stage large target screening module connected in sequence, The image acquisition and annotation module acquires multiple images of port safety operations to be measured and historical port safety operation images for training, and annotates the compliance of safety equipment wearing for personnel in each historical port safety operation image according to a preset annotation criterion, obtaining multiple annotated historical port safety operation images; The model improvement module introduces the SIoU loss function on the basis of the YOLOv5s network model to replace the GIoU loss function in the YOLOv5s network model, obtaining a new YOLOv5s network model; and adopts a dynamic layer adjustment strategy combined with an adaptive feature fusion mechanism to automatically dynamically adjust the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model, obtaining an improved FPNX structure, and further obtaining an improved YOLOv5s network model; The first-stage small target screening module uses multiple annotated historical port safety operation images as training set samples to train the improved YOLOv5s network model, obtaining a trained personnel detection model, and inputs the images of port safety operations to be measured into the personnel detection model for the first-stage small target screening detection, obtaining a first target detection result containing all personnel in the port safety operation image, and automatically cropping out the cropping area containing the personnel according to the first target detection result of each personnel; The second-stage large target screening module uses a super-resolution network based on a convolutional neural network to magnify the cropping area containing the personnel, obtaining magnified target images, uses a part of the target images among all the target images as training set samples to train the improved YOLOv5s network model, obtaining a trained safety equipment detection model, and inputs another part of the target images among all the target images into the safety equipment detection model for the second-stage safety equipment large target screening detection, obtaining a second target detection result containing all safety equipment in the target image to complete the wear detection of the safety equipment.

7. The safety equipment wearing target detection system based on YOLOv5s according to claim 6, characterized in that, In the model improvement module, the automatic dynamic adjustment of the number of neural network layers of the FPN structure in the neck network of the new YOLOv5s network model by using a dynamic layer adjustment strategy combined with an adaptive feature fusion mechanism specifically includes: Using a dynamic layer adjustment strategy to respectively add a high-resolution neural network layer and a low-resolution neural network layer on the basis of the FPN structure with multiple neural network layers in the new YOLOv5s network model, so that the resolution of the added high-resolution neural network layer is greater than the neural network layer with the highest resolution in the multiple neural network layers of the FPN structure, and the resolution of the added low-resolution neural network layer is less than the neural network layer with the lowest resolution in the multiple neural network layers of the FPN structure; In the first-stage small target screening module, when performing the first-stage small target detection, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the high-resolution neural network layer, and then the first-stage small target screening detection is performed by using the high-resolution neural network layer with the highest weight assigned; When performing the second-stage enlarged object detection in the second-stage large target screening module, the attention mechanism in the adaptive feature fusion mechanism is used to assign the highest weight to the low-resolution neural network layer, and then the low-resolution neural network layer with the highest weight assigned is used to perform the second-stage large target screening detection of safety equipment.

8. The safety equipment wearing target detection system based on YOLOv5s according to claim 6, characterized in that, In the image acquisition and annotation module, the safety equipment includes safety helmets and work clothes. The compliance annotation of the wearing of safety equipment by people in the port safety operation image according to the preset annotation criteria specifically includes: Judging whether the people in the port safety operation image wear safety helmets according to the preset annotation criteria. If so, annotate wearing a safety helmet. If not, then judge whether the head has a blurred contour. If so, do not annotate. If not, then judge whether the head is blocked. If so, do not annotate. If not, then judge whether other types of hats other than safety helmets are worn. If so, annotate wearing other hats. If not, annotate not wearing a safety helmet; And judging whether the people in the port safety operation image wear a work vest / strap / one-piece suit according to the preset annotation criteria. If so, annotate wearing work clothes; if not, then judge whether the color of the upper body is a solid color. If not, annotate not wearing work clothes; if so, then judge whether the pants are shorts. If so, annotate not wearing work clothes; if not, annotate wearing work clothes.

9. The YOLOv5s-based safety equipment wearing target detection system according to any one of claims 6 to 8, characterized in that, It also includes a two-stage detection result comparison and analysis module. The two-stage detection result comparison and analysis module is connected to the second-stage large target screening module, and is used to compare the first-stage detection result output by the personnel detection model with the second-stage detection result output by the safety equipment detection model. If the first-stage detection result does not detect a target but the second-stage detection result detects a target, then the second-stage detection result is used as the final detection result; If both the first-stage detection result and the second-stage detection result detect a target, then the second-stage detection result is used as the final detection result; If the first-stage detection result detects a target but the second-stage detection result does not detect a target, then the first-stage detection result is used as the final detection result.

10. The safety equipment wearing target detection system based on YOLOv5s according to any one of claims 6 to 8, characterized in that, Both the first target detection result and the second target detection result include a bounding box, a class label, and a confidence score. The SIoU loss function is used to calculate the bounding box score, the class prediction loss function is used to calculate the class probability score, and the confidence prediction loss function is used to calculate the confidence score.

Citation Information

Cited By

  • Cabin hatch intelligent identification and screening method based on point cloud projection and deep learning

    CN121213855A