Safety helmet standard wearing detection method based on improved YOLOv5

By introducing serpentine convolution, transformer layer and improved feature fusion network in YOLOv5, problems such as receptive field fixation and information loss in safety helmet standard wear detection in the prior art are solved, and higher detection accuracy and generalization capabilities are achieved.

CN120126068APending Publication Date: 2025-06-10SHANGHAI WONDERTEK SOFTWARE CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510094128.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the standardized wear detection of safety helmets, there are problems such as fixed convolution operation receptive field, poor ability to handle non-uniform changes, weak long-distance dependency modeling capabilities, and loss of feature enhancement network information.

Method used

By modifying the feature extraction neural network CSPDarknet in YOLOv5, introducing serpentine convolution and transformer layers, improving the feature fusion pyramid network, using WIoU loss function, and proposing a soft-mosaic data enhancement method.

Benefits of technology

The model's adaptability to nonlinear strip objects and long-distance dependency modeling capabilities are enhanced, information loss is reduced, and detection accuracy and generalization capabilities are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126068A_ABST
    Figure CN120126068A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial detection, in particular to an improved YOLOv5-based safety helmet standard wearing detection method, which comprises the following steps of S1, establishing a safety helmet wearing data set with labels; s2, modifying an original feature extraction neural network CSPDarknet to obtain a pre-trained neural network, and performing secondary training on the pre-trained neural network according to the training set to obtain a trained neural network; s3, verifying the neural network training effect according to the test set to obtain a final neural network; and S4, obtaining a safety helmet wearing detection result according to the working video of the safety production area and the final neural network. According to the method, the adaptability of the non-rigid linear object is enhanced through the snakelike convolutional network; the features of the semantic layer are made to pay attention to global information through a transform layer; the optimization feature enhancement module is an AFPN to realize the characterization capability of the overall enhanced network; for a standard wearing scene of the safety helmet, soft-mosaic is proposed in data enhancement to enrich sample diversity, so that the model is more robust and has generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial inspection, and particularly relates to a safety helmet standard wearing detection method, device and storage medium based on improved YOLOv5. Background Art

[0002] In many industries such as construction, manufacturing, mining, and power, due to the complexity and danger of their working environments, there are various potential hazards such as falling objects from height, mechanical equipment collisions, and electrical accidents. This makes work safety a core concern in the process of enterprise management and operation. To ensure the life safety of employees in such high-risk environments, the state and various industries have formulated strict and detailed safety standards and regulations, aiming to prevent and reduce the occurrence of work-related injuries through various measures such as standardizing work processes and strengthening safety protection.

[0003] As the most basic and crucial personal protective equipment, safety helmets are widely equipped and used in various construction sites and production areas. It has a variety of important protective functions. For example, it can effectively absorb impact energy. When being hit by an external object, through its own structural and material characteristics, it buffers and disperses the impact force, avoiding direct strong impact on the head and getting injured; it can prevent sharp objects from penetrating, providing a reliable physical barrier for the head; at the same time, some safety helmets also have electrical insulation performance, which can block electric current in an electrical working environment and prevent electric shock accidents; in addition, safety helmets can also block direct sunlight and rain to a certain extent, creating a relatively comfortable and safe working condition for employees. The core role of safety helmets is to comprehensively protect the heads of employees from external potential risk factors, thus largely avoiding serious consequences caused by accidental injuries and being one of the key defenses for ensuring the life safety of employees.

[0004] Currently, the detection algorithm for the standard wearing of safety helmets in the field of industrial inspection optimizes the network structure, such as adding more branch structures in the backbone network and the feature enhancement network, which optimizes the performance of the model to a certain extent. However, the following two points are often not considered: (1) The receptive field of the standard convolution operation is relatively fixed, and its ability to process non-uniform changes and spatial adaptability are relatively poor, especially for objects with irregular outer contours such as safety helmets and their buckles (especially the buckles); (2) When using the fully convolutional structure with small convolutional kernels for feature extraction in the backbone features of YOLOv5, it usually cannot effectively capture the relationship between any two positions in the image, which makes its modeling ability weak in the case of long-range dependencies; (3) The PAFPN structure used in its feature enhancement network cannot directly interact with the features of non-adjacent layers during the feature enhancement strategy, which may lead to information loss or degradation; (4) Existing classification methods for standards usually roughly use binary classification to judge, and do not further refine the standard and non-standard behaviors for actual problems, which makes it difficult for the model to converge when there are oppositions between the refined category domains in a certain large category. Summary of the Invention

[0005] The purpose of the present invention is to solve the defects existing in the prior art, and provide a method for detecting the standard wearing of safety helmets based on improved YOLOv5, including the following steps: S1: Collect and organize several initial data sets containing human bodies and safety helmets, process the initial data sets and obtain a safety helmet wearing data set with annotations; S2: Modify the original feature extraction neural network CSPDarknet in YOLOv5 to obtain a pre-trained neural network, pre-train the pre-trained neural network through an open-source data set and retain the generated pre-trained model weights, and perform secondary training on the pre-trained neural network according to the training set of the safety helmet wearing data set and the pre-trained model weights until the network converges to obtain a trained neural network; S3: Verify the effect of the trained neural network according to the test set of the safety helmet wearing data set, and further fine-tune the training parameters of the trained neural network in combination with the verification effect to obtain a final neural network; S4: Obtain the safety helmet wearing detection result, confidence probability, and number of detections per second according to the work video in the safe production area and the final neural network.

[0006] Preferably, in step S1, processing the initial data set and obtaining a safety helmet wearing data set with annotations is further as follows: Collect and organize several initial data sets containing human bodies and safety helmets, and uniformly extract pictures in different scenarios, scales, and states from the initial data sets by means of video frame extraction to obtain extracted pictures; Data label the extracted images through the open-source image annotation tool labelimg to obtain labeled images; Convert the labeled images into a data format for YOLOV5 training to obtain converted images. The converted images regenerate the size and quantity of the prior boxes according to clustering to obtain a safety helmet wearing dataset.

[0007] Preferably, data labeling the extracted images through the open-source image annotation tool labelimg to obtain labeled images further includes: Draw a rectangular box through the image annotation tool labelimg, label the coordinate information of the rectangular box, and select four annotation types: not wearing a safety helmet, wearing a safety helmet without a frontal face, wearing a safety helmet with a frontal face but not wearing it properly, and wearing a safety helmet with a frontal face and wearing it properly to obtain the labeled images.

[0008] Preferably, in step S2, modifying the original feature extraction neural network CSPDarknet to obtain a pre-trained neural network is further as follows: Change the standard convolutional layer Conv in the fifth and seventh layers of the backbone network to the serpentine convolution DSConcv; Change the third CSP module in the backbone network to a transforme self-attention module; Change the feature fusion pyramid network PAEPN to an asymptotic feature pyramid network AFPN; Change the CIOU loss function in the original rectangular box regression to a WIoU loss function; Change the mosaic data augmentation to soft-mosaic data augmentation.

[0009] Preferably, pre-training the pre-trained neural network through an open-source dataset and retaining the generated pre-trained model weights further includes: Input the open-source dataset into the pre-trained neural network to obtain an input feature map; The serpentine convolution DSConcv samples the output feature map through a regular grid to obtain sampling values; The serpentine convolution DSConcv obtains the value at each position on the feature map according to the sampling values and the learned offsets.

[0010] Preferably, the serpentine convolution DSConcv obtaining the value at each position on the feature map according to the sampling values and the learned offsets is further as follows: The formula for the serpentine convolution DSConcv to obtain the value at each position on the feature map according to the sampling values and the learned offsets is as follows: Where, is the true pixel coordinate, defines the field of view size of the convolutional kernel, is the learned offset, is the convolutional kernel, is the input feature map, is the output feature map. In , the offset of adjacent kernel elements is constrained between [-1, 1].

[0011] Preferably, the training effect of the neural network is verified according to the test set, and the training parameters of the training neural network are further fine-tuned in combination with the verification effect to obtain the final neural network, which further includes: The test set inputs the training neural network to calculate the performance evaluation index. The training neural network is analyzed according to the performance evaluation index, and the final neural network is obtained by adjusting the learning rate, data augmentation strategy, and loss function weight; Among them, if the overall performance of the training neural network does not meet the expectation and the detection results fluctuate greatly, it is adjusted by reducing the learning rate. If the convergence speed of the training neural network is too slow and the training loss has been at a high level, it is adjusted by increasing the learning rate; If the training neural network has poor detection of safety helmets in certain specific scenarios, angles, or scales, it is adjusted by specifically increasing the corresponding data augmentation methods; If the detection effects between different categories in the annotation types are unbalanced, the weights of the losses of different categories in the WIoU loss function are adjusted.

[0012] Preferably, the WIoU loss function includes: The calculation formula of the WIoU loss function is as follows: Among them, ( is the center point coordinate of the prediction box, is the center point coordinate of the true box, the width of the bounding box of the prediction box and the true box, the height of the bounding box of the true box is the intersection over union of the regions of the prediction box and the true box.

[0013] A computer device, characterized in that it includes a memory and a processor. When the computer-readable instructions stored in the memory are executed by the processor, the processor executes the steps of the method for detecting the standard wearing of safety helmets based on the improved YOLOv5 in an embodiment of the present invention.

[0014] A storage medium storing computer-readable instructions, characterized in that when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for detecting the standard wearing of safety helmets based on the improved YOLOv5 in an embodiment of the present invention.

[0015] Compared with the prior art, the beneficial effects of the present invention are: The present invention processes the initial data set to obtain a data set of safety helmet wearings with annotations, making the distance between safety helmet wearing categories as large as possible.

[0016] The present invention modifies the original feature extraction neural network CSPDarknet to obtain a pre-trained neural network, realizes the feature extraction of non-linear strip objects, enhances the network's ability to extract semantic features and the ability to model long-distance dependencies, and reduces the information loss that may be caused by the indirect fusion of non-adjacent features.

[0017] The present invention enhances the adaptability of the model to non-rigid linear objects by introducing serpentine convolution and transformer layer in the backbone network, introduces the Asymptotic Feature Pyramid Network AFPN in the feature enhancement part to avoid the information loss problem caused by the inability of non-adjacent layer features to directly interact during feature fusion, overall enhances the representation ability of the network, adds a transformer layer to enable the semantic layer features to pay attention to global information, and refines the safety helmet standard wearing categories for the safety helmet standard wearing scenario, and proposes soft-mosaic in the YOLOV5 data augmentation to enrich the sample diversity, making the model more robust and generalizable, and effectively improving the accuracy of the model for detecting the standard wearing of safety helmets on the premise of ensuring real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.

[0019] Figure 1 It is a flowchart of the method for detecting the standard wearing of safety helmets based on the improved YOLOv5 of the present invention; Figure 2 It is another flowchart of the method for detecting the standard wearing of safety helmets based on the improved YOLOv5 of the present invention; Figure 3 It is a schematic diagram of data annotation of the method for detecting the standard wearing of safety helmets based on the improved YOLOv5 of the present invention; Figure 4 It is a schematic diagram of network structure optimization of the method for detecting the standard wearing of safety helmets based on the improved YOLOv5 of the present invention; Figure 5 This is the schematic diagram of AFPN for the helmet standard wearing detection method based on the improved YOLOv5 of the present invention; Figure 6 This is the schematic diagram of the transformer operation for the helmet standard wearing detection method based on the improved YOLOv5 of the present invention; Figure 7 This is the comparison schematic diagram of the standard convolution kernel and the serpentine convolution for the helmet standard wearing detection method based on the improved YOLOv5 of the present invention; Figure 8 This is the schematic diagram of data augmentation for the helmet standard wearing detection method based on the improved YOLOv5 of the present invention, where Figure 8 (a) is the diagram without retaining the label box, Figure 8 (b) is the diagram with the label box retained; Figure 9 This is the schematic diagram of WIoU for the helmet standard wearing detection method based on the improved YOLOv5 of the present invention. Detailed implementation manners

[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Obviously, the described embodiments are a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art in the present application without creative efforts shall fall within the scope of protection of the present application.

[0021] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0022] The first embodiment In high-risk industries such as construction, manufacturing, mining, and power, work safety has always been the core of enterprise management and operation. The operating environments in these industries usually have many potential hazards, such as falling objects from height, collisions between mechanical equipment, electrical accidents, etc. To effectively prevent and reduce work-related injuries and ensure the safety of employees' lives, the state and the industry have formulated strict safety standards and regulations. Among them, safety helmets, as one of the most basic personal protective equipment, are widely used in various construction sites and production areas. Safety helmets can absorb impact energy, prevent penetration injuries, provide electrical insulation, block sunlight and rain, etc. Their main function is to protect the head from being hit and impacted by external objects and prevent serious consequences caused by accidental injuries.

[0023] Although safety helmets can theoretically provide effective head protection, if worn improperly, their protective effect will be greatly reduced and may even pose greater safety hazards. Such as: Protection failure: Not worn or worn loosely: If an employee does not wear a safety helmet or wears it too loosely, once an object falls or there is a collision, the safety helmet cannot play its due protective role, which may lead to head injuries or even traumatic brain injuries.

[0024] Improper wearing position: The safety helmet should fit tightly on the top of the head, but some employees wear it on the back of the head or the forehead. This incorrect wearing method will weaken its buffering and impact absorption ability and increase the risk of injury.

[0025] Risk of electrical accidents Impaired insulation performance: In an electrical operation environment, the insulation performance of a safety helmet is crucial. If there are scratches, damages or moisture on the surface of the safety helmet, its insulation effect will be greatly reduced, making it easy to cause electric shock accidents.

[0026] Exposure of metal accessories: Some safety helmets are equipped with metal decorations or accessories. If worn improperly, these metal parts may come into contact with live equipment, increasing the risk of electric shock.

[0027] Obstructed vision: Blocking vision: If the safety helmet is worn too low or too tightly, it may block the employee's vision, affecting their observation of the surrounding environment and increasing the risk of operation errors and collision accidents.

[0028] Interference of reflective lenses: Some safety helmets are equipped with reflective lenses. If worn incorrectly, they may reflect light and interfere with the employee's visual judgment, especially in a strong light environment.

[0029] 4. Psychological paralysis: False sense of security: Improperly wearing a safety helmet may give employees an illusion of "I have worn a safety helmet, so I don't have to worry about safety issues", thus relaxing their vigilance and ignoring other safety measures, ultimately leading to accidents.

[0030] 5. Legal and economic responsibilities: Legal responsibility: If an industrial accident occurs due to an employee not wearing a safety helmet properly, the enterprise may face legal responsibilities and compensation risks. According to relevant laws and regulations, the enterprise has the obligation to provide employees with personal protective equipment that meets the standards and ensure its correct use.

[0031] Economic losses: Industrial accidents will not only bring physical and mental pain to employees, but also cause huge economic losses to the enterprise, including medical expenses, production suspension losses, insurance claims, etc.

[0032] In summary, wearing a safety helmet properly is one of the important measures to ensure production safety. Improper wearing of a safety helmet will not only weaken its protective effect, but also may bring serious safety hazards. Therefore, the development and application of a safety helmet proper wearing detection system based on the improved YOLOv5 algorithm has important practical significance. Through means such as real-time monitoring, data analysis, and safety training, this system can not only effectively prevent accidents and protect employees' health, but also help enterprises meet the requirements of laws and regulations, reduce legal and economic risks, and comprehensively improve the level of safety production management.

[0033] Please refer to Figure 1 and Figure 2 As shown, the method for detecting the proper wearing of a safety helmet based on the improved YOLOv5 provided in this embodiment includes the following steps: S1: Collect and organize several initial datasets containing human bodies and safety helmets, process the initial datasets, and obtain a safety helmet wearing dataset with annotations.

[0034] Preferably, in step S1, processing the initial datasets and obtaining a safety helmet wearing dataset with annotations is further as follows: Collect and organize several initial datasets containing human bodies and safety helmets, and uniformly extract pictures in different scenarios, scales, and states from the initial datasets by means of video frame extraction to obtain extracted pictures. Specifically, in this embodiment, about 50,000 real work scene pictures and 100 simulated experiment videos of wearing safety helmets with a duration of 5 minutes are collected; Perform data annotation on the extracted pictures through the open-source image annotation tool labelimg to obtain annotated pictures; The labeled images are converted into a data format for YOLOV5 training to obtain converted images. The converted images are used to regenerate the size and number of prior boxes according to clustering, resulting in a safety helmet wearing dataset. Specifically, in this embodiment, 100 simulated videos are divided according to the ratio of training set:validation set = 9:1, and the divided videos are respectively frame-extracted (1 frame per second) to obtain 27,000 simulated experimental safety helmet standard wearing images for training and 3,000 simulated experimental safety helmet standard wearing images for validation. Approximately 50,000 open-source real safety helmet standard wearing images are also divided into a training set and a validation set according to the ratio of 9:1.

[0035] Please refer to Figure 3 As shown in the figure, the extracted images are data-labeled by the open-source image annotation tool labelimg to obtain labeled images, which further includes: Using the image annotation tool labelimg to draw a rectangular box, annotating the coordinate information of the rectangular box, and selecting four annotation types: not wearing a safety helmet, wearing a safety helmet without a frontal face, wearing a safety helmet with a frontal face but not in a standard manner, and wearing a safety helmet with a frontal face in a standard manner, so as to obtain labeled images.

[0036] S2: Modify the original feature extraction neural network CSPDarknet in YOLOv5 to obtain a pre-trained neural network. The pre-trained neural network is pre-trained through an open-source dataset and the generated pre-trained model weights are retained. The pre-trained neural network is secondarily trained according to the training set of the safety helmet wearing dataset and the pre-trained model weights until the network converges to obtain a trained neural network.

[0037] Please refer to Figure 4 As shown in the figure, in step S2, modifying the original feature extraction neural network CSPDarknet to obtain a pre-trained neural network is further as follows: Change the standard convolutional layers Conv in the fifth and seventh layers of the backbone network to serpentine convolutional layers DSConcv. Different from deformable convolutional layers that are completely freely learned to deform and offset, their receptive fields tend to deviate from the target. Especially when dealing with linear structures, the serpentine convolutional layer uses an iterative strategy to sequentially select the next position to observe for each target to be processed, thereby ensuring the continuity of attention and not spreading the receptive field too far due to large deformation offsets, making the convolutional kernel continuous in the two-dimensional space. Specifically, in this embodiment, please refer to Figure 4 As shown in the figure, replace the last two standard convolutions in the original YOLOv5 backbone network with serpentine convolutions, that is, introduce an accumulated offset on the basis of convolution. Taking a 3×3 convolutional kernel with a stride of 1 as an example.

[0038] The third CSP module in the backbone network is changed to a Transformer self-attention module, so that when extracting semantic information from high-level features, it can better focus on global information. Specifically, in this embodiment, it can calculate the degree of association between each position and all other positions, so that it can easily capture long-range dependencies in the image.

[0039] The feature fusion pyramid network PAEPN is changed to an asymptotic feature pyramid network AFPN to reduce information loss during the feature fusion process. Specifically, in this embodiment, please refer to Figure 5 and Figure 4 As shown by the blue dashed box, when fusing the low-level, middle-level, and high-level features extracted by the backbone network, the fusion of the low-level and high-level features requires the transition of the middle-level feature fusion and cannot be directly fused, which causes a certain degree of information loss during feature fusion. On the contrary, the asymptotic feature pyramid network (AFPN) module can support direct interaction between non-adjacent layers. It starts by fusing two adjacent low-level features and gradually incorporates high-level features into the fusion process. In this way, a large semantic gap between non-adjacent levels can be avoided.

[0040] Please refer to Figure 6 As shown, for convolution, it performs local feature extraction on a large feature map by sliding a small convolution kernel, and cannot achieve long-distance information interaction. While Transformer flattens the feature map into a sequence and then obtains the attention features through the attention calculation formula. By changing the third CSP module in the backbone network to a Transformer self-attention module, after the AFPN outputs the predicted feature map, operations such as flattening the spatial dimension of the image features and adjusting the channel dimension are performed to convert them into a format suitable for self-attention calculation for dimension adjustment to meet the input requirements of the Transformer self-attention module, so that when extracting semantic information from high-level features, it can better focus on global information at the appropriate position in the model.

[0041] Compared with the local window information extracted by convolution, QKV all contain all the information of the input. In subsequent matrix calculations, the value of each position interacts with the global features, so that it can capture features with long-range dependencies, which performs better than standard convolution in feature representation for flame detection. The Transformer self-attention module uses its ability to capture long-range dependencies to re-weight the input features and optimize the feature representation, focusing on more critical regions and feature information in the image, further improving the quality of the features, and outputting the feature map after self-attention processing for subsequent operations.

[0042] Change the CIOU loss function in the original rectangular box regression to the WIoU loss function to increase the influence of the width and height of the prediction box in loss calculation. Specifically, in this embodiment, the CIOU loss function is calculated as follows: where c is the diagonal length of the minimum circumscribed rectangle of the prediction box and the ground truth box, is the Euclidean distance between the centers of the prediction box and the ground truth box, w, h, wgt, hgt are the widths and heights of the prediction box and the ground truth box respectively, and IOU represents the intersection over union of the areas of the prediction box and the ground truth box. Although the overlapping area, center point distance, and aspect ratio of the bounding box regression are considered, the difference in aspect ratio reflected by v in its formula, rather than the true difference between the width and height and their respective confidences, sometimes hinders the effective optimization of similarity by the model.

[0043] Change the mosaic data augmentation to soft-mosaic data augmentation to avoid the problem of sample class shift caused by the existing mosaic method during data augmentation. Specifically, in this embodiment, there is no means to augment similar pictures in the original data augmentation of YOLOv5. In actual training samples, there are many similar pictures extracted from consecutive video frames. To further enrich the diversity of training data, mosaic data augmentation is added in the data augmentation, and mosaic is optimized for the task of detecting the standard wearing of safety helmets. Please refer to Figure 8 As shown, the original mosaic data augmentation first aligns four pictures with a certain center point and then performs random cropping to obtain new samples. Through this method, the diversity of samples can be effectively enriched. However, for the task of detecting the standard wearing of safety helmets, random cropping may cause the labels of the samples to change, making it difficult for the classification loss to converge. Among them Figure 8 (a) The person's head in the lower right corner has its label information changed from standard to without face after cropping. If it is retained, it will lead to inaccurate labels, so it needs to be deleted. Similarly Figure 8 (b) The two person's heads on the right side can be retained because their label information remains unchanged after cropping. In soft-mosaic, if Hcrop / Horigin>=0.8 and Wcrop / Worigin>=0.8, the labeled boxes with crop after augmentation are retained, otherwise they are removed and the corresponding areas are randomly filled.

[0044] Preferably, pre-training the pre-trained neural network with an open-source dataset and retaining the generated pre-trained model weights further includes: The open-source dataset is input into the pre-trained neural network to obtain an input feature map; The serpentine convolution DSConcv samples the output feature map through a regular grid to obtain sampled values; The serpentine convolution DSConcv obtains the value at each position on the feature map based on the sampled values and the learned offsets.

[0045] Preferably, the serpentine convolution DSConcv obtaining the value at each position on the feature map based on the sampled values and the learned offsets is further as follows: The formula for the serpentine convolution DSConcv to obtain the value at each position on the feature map based on the sampled values and the learned offsets is as follows: Where, is the true pixel coordinate, defines the field of view size of the convolution kernel, is the learned offset, is the convolution kernel, is the input feature map, is the output feature map. In , the offsets of adjacent kernel elements are constrained between [-1, 1], so that the serpentine convolution not only ensures the continuity of attention but also does not cause the perceptual range to be too dispersed due to large deformation offsets. Specifically, in this embodiment, the offset at the (-1, -1) position is offset based on the offset at the (-1, 0) position, starting from the (0, 0) point, and advancing two learning offsets at a time. When the deformable convolution learns the offset, no other constraints are added, which makes its perceptual field tend to deviate from the target, especially when dealing with linear structures such as the straps of safety helmets. The serpentine dynamic convolution is different. The convolution kernel is flattened and extended from the center point of the convolution kernel to both ends step by step in the horizontal and vertical directions. The learned offset at a certain position is obtained through an iterative strategy. Each is restricted between [-1, 1]. The learned offset at a certain position is obtained by adding the offset at the previous position to the offset at this position . Since is usually a decimal, when calculating the value of the feature map x through , bilinear interpolation is used for calculation. In this way, its convolution kernel is linearly continuous in space, can achieve a larger receptive field compared to the standard convolution, and ensures the continuity of the extracted features in space compared to the deformable convolution, which is also consistent with the need to extract the linear features of the straps when wearing a safety helmet properly.

[0046] In this embodiment, since the network structure of the feature extraction neural network (CSPDarknet) is modified, it is impossible to use the original YOLOv5 to fine-tune on the new dataset based on the pre-trained model weights, and the trained neural network is verified according to the test set.

[0047] S3: Verify the effect of the trained neural network according to the test set of the safety helmet wearing dataset, and further fine-tune the training parameters of the trained neural network in combination with the verification effect to obtain the final neural network. Specifically, in this embodiment, the trained YOLO v5 detection model is used to verify the model effect on the test set, and the training parameters of the model are further fine-tuned in combination with the detection effect and continue to train to achieve the optimal performance of the model, and the final model weights are retained. Finally, input the video of wearing a safety helmet simulated by the experiment, load the optimal model weights to detect the video, locate and classify the human head area on each frame of the image and output and display.

[0048] Preferably, verifying the effect of the trained neural network according to the test set and further fine-tuning the training parameters of the trained neural network in combination with the verification effect to obtain the final neural network further includes: The test set is input into the trained neural network to calculate the performance evaluation index, the trained neural network is analyzed according to the performance evaluation index, and the final neural network is obtained by adjusting the learning rate, data augmentation strategy and loss function weight; Among them, if the overall performance of the trained neural network does not meet the expectation and the detection results fluctuate greatly, it is adjusted by reducing the learning rate. If the convergence speed of the trained neural network is too slow and the training loss has been at a high level, it is adjusted by increasing the learning rate; If the trained neural network has poor detection of safety helmets in certain specific scenarios, angles or scales, it is adjusted by specifically increasing the corresponding data augmentation method; If the detection effects between different categories in the annotation types are unbalanced, adjust the weights of the losses of different categories in the WIoU loss function.

[0049] More preferably, calculating the performance evaluation index according to the accuracy rate, recall rate and mean average precision value, and further: The statistical accuracy is the number of actual positive examples among the statistically predicted positive examples (i.e., the helmet-wearing situation is correctly detected to conform to the corresponding category), denoted as True Positive (TP). For example, if it is predicted that 50 cases of correct helmet wearing are correct, these 50 are TP; the number of statistically predicted positive examples that are actually negative examples is denoted as False Positive (FP). For example, if 10 cases of actual non-compliant wearing are wrongly judged as compliant, these 10 are FP; the accuracy is calculated according to the formula Precision = TP / (TP + FP). In the above example, the accuracy is 50 / (50 + 10) = 5 / 6.

[0050] The statistical recall rate Recall is the number of actually positive examples that are correctly predicted as positive examples, that is, the TP mentioned above; the number of actually positive examples that are not detected is denoted as False Negative (FN). For example, if there are actually 60 cases of correct helmet wearing, but the model only detects 50 cases, then FN is 10; the recall rate is calculated according to the formula Recall = TP / (TP + FN). In this example, the recall rate is 50 / 60 = 5 / 6.

[0051] The statistical mean average precision mAP (Mean Average Precision) is to calculate the precision at different recall rates for each category and draw a precision-recall curve (PR curve). For example, for the category of "correct frontal face wearing with a helmet", at different confidence thresholds, different combinations of precision and recall rate will be obtained, and the PR curve of this category is drawn through these combinations. Calculate the area under the PR curve to obtain the average precision (Average Precision, AP) of this category. Average the APs of all categories (not wearing a helmet, wearing a helmet without a frontal face, wearing a helmet with an incorrect frontal face wearing, wearing a helmet with a correct frontal face wearing) to obtain mAP.

[0052] If the accuracy is low, it indicates that there are many misdetection situations in the model. The possible reasons include: The data of some categories in the dataset is unbalanced, resulting in poor discrimination ability of the model for minority categories. For example, the number of samples in the category of "wearing a helmet with an incorrect frontal face wearing" is too small, and the model does not learn enough.

[0053] The model is too complex and overfitting occurs. It overfits the noise or special situations in the training data and cannot generalize well to the test data. For example, the number of layers or parameters of the neural network is too large, resulting in good performance on the training set but a decline in performance on the test set.

[0054] Insufficient or inappropriate feature extraction does not fully capture the key information of helmet wearing features. For example, the convolution kernel size and step size in the network are not set reasonably, and the helmet features of different scales cannot be effectively extracted.

[0055] If recall is low, it could be because: The model does not learn the features of some helmet wearing situations deeply enough, resulting in some positive examples not being detected. For example, the model has difficulty recognizing helmets at special angles or under occlusion.

[0056] There are some difficult-to-detect samples in the dataset, and the number is large, which affects the overall performance of the model. For example, some samples with dim light and blurred images, the model did not fully learn how to deal with such situations during training.

[0057] If the overall performance of the model does not meet expectations and the test results fluctuate greatly, the learning rate can be appropriately reduced. For example, multiply the current learning rate by a smaller factor such as 0.1 or 0.5 to allow the model to adjust parameters more finely and avoid skipping the optimal solution during training.

[0058] If the model converges too slowly and the training loss remains at a high level, you can increase the learning rate appropriately, but be careful to avoid unstable training. Generally, you can try multiplying the learning rate by a smaller multiple such as 1.5 or 2 to observe the changes in model performance.

[0059] Data enhancement strategy adjustment: Based on the analysis of the detection results, if it is found that the model has poor detection of helmets in certain scenarios, angles or scales, the corresponding data enhancement methods can be added in a targeted manner.

[0060] For example, if the detection accuracy of small-scale helmets is low, data enhancement operations such as image scaling and cropping can be added to enrich the data of small-scale targets. The specific operation can be to randomly scale the image in the data preprocessing stage with a scaling ratio between 0.8 and 1.2, and then randomly crop out fixed-size image blocks to increase the number and diversity of small-scale helmet samples.

[0061] If there are problems with helmet detection under different lighting conditions, you can add data enhancement for lighting changes to simulate different lighting scenarios. For example, use an image processing library (such as OpenCV) to randomly adjust the brightness, contrast, saturation, etc. of the image so that the model can learn the characteristics of the helmet under different lighting conditions.

[0062] Loss function weight adjustment: If the detection effects between different categories are unbalanced, consider adjusting the weights of different category losses in the loss function.

[0063] For example, for the category of "wearing a safety helmet with an improper frontal wearing, which has a poor detection effect", its weight in the classification loss can be appropriately increased. Assuming that the weights of all categories in the original classification loss are 1, the weight of the category of "wearing a safety helmet with an improper frontal wearing" can be adjusted to 2 or 3, so that the model pays more attention to the correct classification of this category during training, thereby improving the detection performance of this category.

[0064] Please refer to Figure 4 and Figure 7 As shown, by adding a learnable offset in the standard convolution operation, a convolutional kernel with a regular shape can extract information from non-rigid regions.

[0065] The standard 2D convolution calculation process includes two steps: Sampling on the input feature map x using a regular grid R (used to define the receptive field and dilation factor); Summing the sampled values weighted by w. Taking a 3×3 convolutional kernel with a dilation factor of 1 as an example, its R = [(-1, -1), (-1, 0), (-1, 1), (0, -1), (0, 0), (0, 1), (1, -1), (1, 0), (1, 1)]. For each position p0 on the output feature map y, its calculation formula is as follows: where the positions in R are enumerated.

[0066] Since the training data inevitably contains low-quality examples, geometric metrics such as distance and aspect ratio will exacerbate the penalty for low-quality examples, thus degrading the generalization performance of the model. A good loss function should weaken the penalty of geometric metrics when the anchor box and the target box coincide well, and not overly interfere with the training to enable the model to have better generalization ability. Therefore, the bounding box regression loss in YOLOv5 object detection is optimized to WIoU.

[0067] Please refer to Figure 9 As shown, the WIoU loss function includes: The WIoU function can avoid the problem that geometric metrics such as distance and aspect ratio in the training data, which inevitably contain low-quality examples, exacerbate the penalty for low-quality examples and thus degrade the generalization performance of the model. The calculation formula of the WIoU loss function is as follows: where, ( is the center point coordinate of the predicted box, is the center point coordinate of the ground truth box, is the width of the bounding box of the predicted box and the ground truth box, The height of the bounding box of the predicted box and the ground truth box is the intersection over union (IoU) of the regions of the predicted box and the ground truth box, is the predicted box, is the ground truth box, [1, e), which will significantly amplify the (when it comes to the normal-quality anchor box, is the intermediate value, and the normal-quality anchor box fits well with the target box. Compared with the high-quality anchor box, will be larger, and the gain given to the normal anchor box will be more, will be larger, and thus more attention will be paid to the normal-quality anchor box), , which will significantly reduce the of the high-quality anchor box, and significantly reduce its attention to the distance from the center point when the anchor box coincides well with the target box (the high-quality anchor box fits well with the target box, and thus will be closer to 1 according to the calculation formula. Compared with the normal-quality anchor box, will be smaller. When two anchor boxes have the same IoU with a target box, it means that is the same, but the of the anchor box with a farther center point distance will be larger, will be larger, and thus more attention will be paid to the anchor box with a farther center point distance).

[0068] S4: Obtain the helmet-wearing detection results, confidence probability, and the number of detections per second based on the work video of the safety production area and the final neural network.

[0069] In this embodiment, several improvements are made to the feature extraction neural network of YOLO v5, including the introduction of serpentine convolution, transformer, and AFPN modules. The last two standard convolutions of the backbone network are replaced with serpentine convolutions with the same kernel size, enabling the model to learn the representation ability of non-linear objects. The third CSP module is replaced with a transformer layer to enable attention to global information when improving high-level semantic information. The original feature enhancement network PA-FPN is replaced with a progressive pyramid feature fusion network to increase the computational cost brought by a small amount of feature fusion to reduce the information loss problem that may be caused during adjacent layer feature fusion. In addition, the WIoU loss, which is more suitable for non-rigid objects such as helmet specification wearing, is introduced into the loss function to improve the anti-interference ability of the loss function during training, thereby improving the detection effect of the model. In addition to the improvements to the network architecture and loss function, this embodiment also proposes a soft-mosaic method for data augmentation, which can greatly enrich the diversity of samples on a limited dataset without contaminating the original labels, and to a certain extent improves the generalization and robustness of the trained model.

[0070] Second Embodiment In some embodiments of the present application, a computer device is further provided, including a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the method for detecting the standard wearing of a safety helmet based on improved YOLOv5 in an embodiment of the present invention.

[0071] The present invention further provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for detecting the standard wearing of a safety helmet based on improved YOLOv5 in an embodiment of the present invention.

[0072] It can be understood that for the aforementioned method for detecting the standard wearing of a safety helmet based on improved YOLOv5, if it is implemented in the form of software functional modules and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., all kinds of media that can store program codes.

[0073] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0074] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A helmet standard wearing detection method based on improved YOLOv5, characterized in that: The following steps are involved: S1: Collect and organize several initial data sets containing human bodies and helmets, process the initial data sets and obtain annotated helmet wearing data sets; S2: modifying the original feature extraction neural network CSPDarknet in YOLOv5 to obtain a pre-trained neural network, pre-training the pre-trained neural network through an open source data set and retaining the generated pre-trained model weights, and performing secondary training on the pre-trained neural network according to the training set of the helmet wearing data set and the pre-trained model weights until the network converges to obtain a trained neural network; S3: verifying the effect of training the neural network according to the test set of the helmet wearing data set, and further fine-tuning the training parameters of the neural network based on the verification effect to obtain the final neural network; S4: Obtain helmet wearing detection results, confidence probability, and detection frames per second based on the safe production area work video and the final neural network.

2. The helmet standard wearing detection method based on improved YOLOv5 according to claim 1 is characterized in that: In step S1, the initial data set is processed to obtain a helmet wearing data set with annotations, further comprising: Collect and organize several initial data sets containing human bodies and helmets, and evenly extract pictures of different scenes, scales, and states from the initial data sets by video frame extraction to obtain extracted pictures; Perform data labeling on the extracted image using the open source image labelimg tool to obtain a labeled image; The labeled image is converted into a data format for YOLOV5 training to obtain a converted image, and the converted image regenerates the size and number of the priori box according to clustering to obtain a helmet wearing data set.

3. The helmet standard wearing detection method based on improved YOLOv5 according to claim 2 is characterized in that: The extracted image is labeled with data using the open source image labeling tool labelimg, and obtaining the labeled image further includes: Draw a rectangular box through the image annotation tool labelimg, annotate the coordinate information of the rectangular box, and select four annotation types: not wearing a safety helmet, wearing a safety helmet but not facing the front, wearing a safety helmet with the front face worn improperly, and wearing a safety helmet with the front face worn properly, to obtain the annotated image.

4. The helmet standard wearing detection method based on improved YOLOv5 according to claim 3 is characterized in that: In step S2, the original feature extraction neural network CSPDarknet is modified to obtain a pre-trained neural network further as: The standard convolutional layers Conv in the fifth and seventh layers of the backbone network are changed to snake convolution DSConcv; The third CSP module in the backbone network is changed to a transforme self-attention module; The feature fusion pyramid network PAEPN is changed to the asymptotic feature pyramid network AFPN; Change the CIOU loss function in the original rectangular box regression to the WIoU loss function; Change mosaic data enhancement to soft-mosaic data enhancement.

5. The helmet standard wearing detection method based on improved YOLOv5 according to claim 4 is characterized in that: Pretraining the pretrained neural network through an open source dataset and retaining the generated pretrained model weights further comprises: The open source data set is input into the pre-trained neural network to obtain an input feature map; The snake convolution DSConcv samples the output feature map through a regular grid to obtain sampling values; The snake convolution DSConcv obtains each position value on the feature map according to the sampling value and the learned offset.

6. The helmet standard wearing detection method based on improved YOLOv5 according to claim 5 is characterized in that: The snake convolution DSConcv obtains each position value on the feature map according to the sampling value and the learned offset, which is further: The calculation formula of the snake convolution DSConcv to obtain each position value on the feature map according to the sampling value and the learned offset is as follows: in, is the real pixel coordinate, Defines the field of view size of the convolution kernel, is the learned offset, is the convolution kernel, is the input feature map, is the output feature map, In , the offsets of adjacent kernel elements are constrained to be between [-1, 1].

7. According to the method for detecting the standard wearing of helmets based on the improved YOLOv5 of claim 6, verifying the effect of training the neural network according to the test set, and further fine-tuning the training parameters of the training neural network in combination with the verification effect to obtain the final neural network further comprises: The test set is input into the training neural network to calculate the performance evaluation index, the training neural network is analyzed according to the performance evaluation index, and the final neural network is obtained by adjusting the learning rate, data enhancement strategy and loss function weight; If the overall performance of the training neural network does not meet expectations and the test results fluctuate greatly, the training is adjusted by reducing the learning rate; if the training neural network converges too slowly and the training loss is always at a high level, the training is adjusted by increasing the learning rate; If the trained neural network performs poorly on helmet detection in certain scenes, angles or scales, adjustments are made by adding corresponding data enhancement methods in a targeted manner; If the detection effects between different categories in the annotation type are unbalanced, the weights of the losses of different categories in the WIoU loss function are adjusted.

8. The helmet standard wearing detection method based on improved YOLOv5 according to claim 7 is characterized in that: The WIoU loss function includes: The calculation formula of WIoU loss function is as follows: in,( is the coordinate of the center point of the prediction box, is the center point coordinate of the real frame, is the width of the bounding box of the predicted box and the real box, is the height of the bounding box of the predicted box and the real box, It is the intersection-over-union ratio of the predicted box and the true box area.

9. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the helmet standard wearing detection method based on the improved YOLOv5 as claimed in any one of claims 1 to 8.

10. A storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the helmet standard wearing detection method based on improved YOLOv5 as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Safety helmet construction environment detection method and system based on image recognition

    CN121259737A