Helmet Wearing Detection Method Based on Improved YOLOv5

By improving the feature extraction and fusion module of the YOLOv5 model, using the multi-layer Swin transformer Block network and self-attention mechanism, the accuracy of safety helmet wear detection in occlusion and low-light environments is solved, and higher detection accuracy is achieved.

CN114973122BActive Publication Date: 2025-07-22SHAOGUAN COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210467457.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-07-22
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In the prior art, the hard helmet wear detection method based on YOLOv5 is difficult to identify the detection targets that are blocked and in dark places, and the missed detection rate is high, so it is impossible to effectively supervise the wearing of the hard helmet at the construction site.

Method used

Using the improved YOLOv5 model, features are extracted through a multi-layer Swin transformer Block network, and combined with the self-attention mechanism and feature fusion module, the detection targets with occlusion and low brightness are identified, enhancing the detection accuracy of the model.

Benefits of technology

It improves the ability to identify detection targets in occlusion and low-light environments, reduces the false detection rate, and improves the accuracy of safety helmet wear detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973122B_ABST
    Figure CN114973122B_ABST
Patent Text Reader

Abstract

The present invention relates to a safety helmet wearing detection method based on improved YOLOv5, which includes the steps of: obtaining a to-be-detected image containing detection targets; inputting the to-be-detected image into an improved YOLOv5 model for target detection to obtain the position and size information and the category to which the detection targets belong. Compared with the prior art, the present invention provides a safety helmet wearing detection method based on improved YOLOv5. By using a multi-layer Swin transformer Block network to extract features, the feature extraction ability of the model for the to-be-detected image is enhanced. By extracting features through the self-attention mechanism, image features that contribute greatly to the recognition of detection targets can be obtained. And through multi-level feature extraction, richer image features can be obtained, so that detection targets that are occluded and those with lower brightness can be recognized. In addition, objects similar in shape to the detection targets can also be distinguished, with a low false detection and missed detection rate and high detection accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of safety helmet wearing detection, and in particular to a safety helmet wearing detection method based on improved YOLOv5. Background Art

[0002] As the world's largest manufacturing country, China attaches great importance to the safety issues at construction sites. Wearing a safety helmet correctly can effectively prevent or reduce the head injuries of workers in the process of work. However, in actual production activities, although it is clearly required that construction workers must wear safety helmets to enter the construction site, it is still difficult to prevent some individuals from wearing safety helmets irregularly during work due to luck or other reasons. At present, the supervision of safety helmet wearing mainly relies on manual labor, which is inefficient and labor-consuming. Moreover, it is difficult for manual labor to focus on monitoring for a long time, and there will be certain oversights. An existing safety helmet wearing detection method based on YOLOv5 can automatically detect whether a worker wears a safety helmet by inputting the image of the construction site into the YOLOv5 model. However, due to the dense population of workers at the construction site, the detection targets are easily blocked, and the lighting environment in some construction sites is relatively dark. This detection method is difficult to identify the blocked and dark detection targets, resulting in a high false detection and missed detection rate. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a safety helmet wearing detection method based on improved YOLOv5, which can identify blocked and small-sized detection targets, with a low false detection and missed detection rate and a high recognition accuracy.

[0004] The present invention is realized through the following technical solutions: A safety helmet wearing detection method based on improved YOLOv5, comprising the steps of: obtaining a to-be-detected image containing a detection target; inputting the to-be-detected image into an improved YOLOv5 model for object detection to obtain the position and size information and the category to which the detection target belongs. Among them, the improved YOLOv5 model includes a feature extraction module, a feature fusion module, and a result prediction module. The feature extraction module includes a tile segmentation sub-module, a linear embedding sub-module, a first Swin-T sub-module, a first tile splicing sub-module, a second Swin-T sub-module, a second tile splicing sub-module, a third Swin-T sub-module, a third tile splicing sub-module, a fourth Swin-T sub-module, a fourth tile splicing sub-module, and a fifth Swin-T sub-module. When the feature extraction module extracts features from the to-be-detected image, it includes the steps of: inputting the to-be-detected image into the image segmentation sub-module for image segmentation; inputting the segmented to-be-detected image into the linear embedding sub-module for linear transformation; inputting the linearly transformed to-be-detected image into the first Swin-T sub-module for feature extraction to obtain a first Swin-T feature map; inputting the first Swin-T feature map into the first tile splicing sub-module for downsampling to obtain a first-level feature map; inputting the first-level feature map into the second Swin-T sub-module for feature extraction to obtain a second Swin-T feature map; inputting the second Swin-T feature map into the second tile splicing sub-module for downsampling to obtain a second-level feature map; inputting the second-level feature map into the third Swin-T sub-module for feature extraction to obtain a third Swin-T feature map; inputting the third Swin-T feature map into the third tile splicing sub-module for downsampling to obtain a third-level feature map; inputting the third-level feature map into the fourth Swin-T sub-module for feature extraction to obtain a fourth Swin-T feature map; inputting the fourth Swin-T feature map into the fourth tile splicing sub-module for downsampling to obtain a fourth-level feature map; inputting the fourth-level feature map into the fifth Swin-T sub-module for feature extraction to obtain a fifth Swin-T feature map. Among them, the first Swin-T sub-module, the second Swin-T sub-module, the fourth Swin-T sub-module, and the fifth Swin-T sub-module each include two Swin transformer Block networks, and the third Swin-T sub-module includes six of the Swin transformer Block networks. The Swin transformer Block network is used for extracting image features from the input feature map;The feature fusion module is used to fuse the fifth Swin-T feature map, the fourth Swin-T feature map, the third Swin-T feature map, and the second Swin-T feature map to obtain output feature maps with multiple different grid sizes; the result prediction module is used to predict the position and size information and the category to which the detection target belongs according to the output feature maps with multiple different grid sizes.

[0005] Compared with the prior art, the present invention provides a safety helmet wearing detection method based on improved YOLOv5. The features are extracted through a multi-layer Swin transformer Block network, which enhances the feature extraction ability of the model for the image to be detected. By extracting features through the self-attention mechanism, image features that contribute greatly to the identification of the detection target can be obtained. Moreover, through multi-level feature extraction, richer image features can be obtained, so that occluded detection targets and detection targets with low brightness can be identified. In addition, objects similar in shape to the detection target can be distinguished, with a low false detection and missed detection rate and high detection accuracy of the model.

[0006] Furthermore, the Swin transformer Block network includes four LayerNorm layers, one multi-head self-attention layer, two MLP layers, one shifted window multi-head self-attention layer, four DropPath layers, and four residual connection layers. When the Swin transformer Block network extracts image features from the input feature map, the steps are as follows: input the input feature map into the LayerNorm layer for normalization processing; input the normalized input feature map into the multi-head self-attention layer for multi-head self-attention feature extraction to obtain a multi-head self-attention feature map; input the multi-head self-attention feature map into the DropPath layer for random inactivation; input the multi-head self-attention feature map output by the DropPath layer and the input feature map into the residual connection layer for residual connection to obtain a first intermediate feature map; input the first intermediate feature map into the LayerNorm layer for normalization processing; input the normalized first intermediate feature map into the MLP layer for linear transformation to obtain a first transformed feature map; input the first transformed feature map into the DropPath layer for random inactivation; input the first transformed feature map output by the DropPath layer and the first intermediate feature map into the residual connection layer for residual connection to obtain a second intermediate feature map; input the second intermediate feature map into the LayerNorm layer for normalization processing; input the normalized second intermediate feature map into the shifted window multi-head self-attention layer for pixel-shifted multi-head self-attention feature extraction to obtain a shifted multi-head self-attention feature map; input the shifted multi-head self-attention feature map into the DropPath layer for random inactivation; input the shifted multi-head self-attention feature map output by the DropPath layer and the second intermediate feature map into the residual connection layer for residual connection to obtain a third intermediate feature map; input the third intermediate feature map into the LayerNorm layer for normalization processing; input the normalized third intermediate feature map into the MLP layer for linear transformation to obtain a second transformed feature map; input the second transformed feature map into the DropPath layer for random inactivation; input the second transformed feature map output by the DropPath layer and the third intermediate feature map into the residual connection layer for residual connection to obtain the feature map as the output of the Swin transformer Block network.

[0007] Further, the first tile splicing sub-module, the second tile splicing sub-module, the third tile splicing sub-module, and the fourth tile splicing sub-module all include a tile segmentation layer, a concat layer, a LayerNorm layer, and a fully connected layer. Among them, the tile segmentation layer is used to divide adjacent pixels with an interval of 2 in the input feature map with dimensions [H, W, C] into multiple tiles; the concat layer is used to perform concat splicing on the divided tiles to obtain a feature map with dimensions changed to [H / 2, W / 2, 4C]; the LayerNorm layer is used to normalize the feature map output by the concat layer; the fully connected layer is used to linearly transform the number of channels of the feature map output by the LayerNorm layer to obtain a feature map with dimensions [H / 2, W / 2, 2C].

[0008] Furthermore, the feature fusion module includes a first CONV layer, a first UP layer, a first Concat layer, a first C3-Ghost layer, a second CONV layer, a second UP layer, a second Concat layer, a second C3-Ghost layer, a third CONV layer, a third UP layer, a third Concat layer, a third C3-Ghost layer, a fourth CONV layer, a fourth Concat layer, a fourth C3-Ghost layer, a fifth CONV layer, a fifth Concat layer, a fifth C3-Ghost layer, a sixth CONV layer, a sixth Concat layer, and a sixth C3-Ghost layer. When the feature fusion module fuses the fifth Swin-T feature map, the fourth Swin-T feature map, the third Swin-T feature map, and the second Swin-T feature map to obtain output feature maps of multiple different grid sizes, it includes the steps of: obtaining the fifth Swin-T feature map and inputting it into the first CONV layer for convolution processing to obtain a first convolutional feature map; inputting the first convolutional feature map into the first UP layer for upsampling operation; obtaining the fourth Swin-T feature map and jointly inputting it with the feature map output by the first UP layer into the first Concat layer for Concat splicing; inputting the feature map output by the first Concat layer into the first C3-Ghost layer for convolution processing to obtain a first output feature map; inputting the first output feature map into the second CONV layer for convolution processing to obtain a second convolutional feature map; inputting the second convolutional feature map into the second UP layer for upsampling operation; obtaining the third Swin-T feature map and jointly inputting it with the feature map output by the second UP layer into the second Concat layer for Concat splicing; inputting the feature map output by the second Concat layer into the second C3-Ghost layer for convolution processing to obtain a second output feature map; inputting the second output feature map into the third CONV layer for convolution processing to obtain a third convolutional feature map; inputting the third convolutional feature map into the third UP layer for upsampling operation; obtaining the second Swin-T feature map and jointly inputting it with the feature map output by the third UP layer into the third Concat layer for Concat splicing; inputting the feature map output by the third Concat layer into the third C3-Ghost layer for convolution processing to obtain a third output feature map; inputting the third output feature map into the fourth CONV layer for convolution processing to obtain a fourth convolutional feature map; jointly inputting the fourth convolutional feature map and the third convolutional feature map into the fourth Concat layer for Concat splicing; inputting the feature map output by the fourth Concat layer into the fourth C3-Ghost layer for convolution processing to obtain the fourth output feature map; inputting the fourth output feature map into the fifth CONV layer for convolution processing to obtain a fifth convolutional feature map;The fifth convolutional feature map and the second convolutional feature map are jointly input into the fifth Concat layer for Concat splicing; the feature map output by the fifth Concat layer is input into the fifth C3-Ghost layer for convolutional processing to obtain a fifth output feature map; the fifth output feature map is input into the sixth CONV layer for convolutional processing to obtain a sixth convolutional feature map; the sixth convolutional feature map and the first convolutional feature map are jointly input into the sixth Concat layer for Concat splicing; the feature map output by the sixth Concat layer is input into the sixth C3-Ghost layer for convolutional processing to obtain a sixth output feature map.;

[0009] Furthermore, the feature fusion module includes a first CONV layer, a first UP layer, a first Concat layer, a first C3-Ghost layer, a second CONV layer, a second UP layer, a second Concat layer, a second C3-Ghost layer, a third CONV layer, a third UP layer, a third Concat layer, a third C3-Ghost layer, a fourth CONV layer, a fourth Concat layer, a fourth C3-Ghost layer, a fifth CONV layer, a fifth Concat layer, a fifth C3-Ghost layer, a sixth CONV layer, a sixth Concat layer, and a sixth C3-Ghost layer. When the feature fusion module fuses the fifth Swin-T feature map, the fourth Swin-T feature map, the third Swin-T feature map, and the second Swin-T feature map to obtain output feature maps of multiple different grid sizes, it includes the steps of: obtaining the fifth Swin-T feature map and inputting it into the first CONV layer for convolution processing to obtain a first convolutional feature map; inputting the first convolutional feature map into the first UP layer for upsampling operation; obtaining the fourth Swin-T feature map and jointly inputting it with the feature map output by the first UP layer into the first Concat layer for Concat splicing; inputting the feature map output by the first Concat layer into the first C3-Ghost layer for convolution processing to obtain a first output feature map; inputting the first output feature map into the second CONV layer for convolution processing to obtain a second convolutional feature map; inputting the second convolutional feature map into the second UP layer for upsampling operation; obtaining the third Swin-T feature map and jointly inputting it with the feature map output by the second UP layer into the second Concat layer for Concat splicing; inputting the feature map output by the second Concat layer into the second C3-Ghost layer for convolution processing to obtain a second output feature map; inputting the second output feature map into the third CONV layer for convolution processing to obtain a third convolutional feature map; inputting the third convolutional feature map into the third UP layer for upsampling operation; obtaining the second Swin-T feature map and jointly inputting it with the feature map output by the third UP layer into the third Concat layer for Concat splicing; inputting the feature map output by the third Concat layer into the third C3-Ghost layer for convolution processing to obtain a third output feature map; inputting the third output feature map into the fourth CONV layer for convolution processing to obtain a fourth convolutional feature map; jointly inputting the third Swin-T feature map, the fourth convolutional feature map, and the third convolutional feature map into the fourth Concat layer for Concat splicing; inputting the feature map output by the fourth Concat layer into the fourth C3-Ghost layer for convolution processing to obtain the fourth output feature map;Input the fourth output feature map into the fifth CONV layer for convolution processing to obtain a fifth convolutional feature map; input the fourth Swin-T feature map, the fifth convolutional feature map, and the second convolutional feature map into the fifth Concat layer for Concat splicing; input the feature map output by the fifth Concat layer into the fifth C3-Ghost layer for convolution processing to obtain a fifth output feature map; input the fifth output feature map into the sixth CONV layer for convolution processing to obtain a sixth convolutional feature map; input the sixth convolutional feature map and the first convolutional feature map into the sixth Concat layer for Concat splicing; input the feature map output by the sixth Concat layer into the sixth C3-Ghost layer for convolution processing to obtain a sixth output feature map.

[0010] Further, when the first C3-Ghost layer, the second C3-Ghost layer, the third C3-Ghost layer, the fourth C3-Ghost layer, the fifth C3-Ghost layer, and the sixth C3-Ghost layer perform convolution processing on the input feature map, it includes the steps of: performing a standard convolution operation on the input feature map to compress the number of channels, and performing feature extraction through N cascaded Ghost Bottleneck modules to obtain a first C3-Ghost feature map; performing another standard convolution operation on the input feature map to obtain a second C3-Ghost feature map; concatenating and superimposing the first C3-Ghost feature map and the second C3-Ghost feature map along the channel dimension, and performing feature fusion through convolution to obtain the output feature map; wherein, the steps when the Ghost Bottleneck module performs feature extraction include: inputting the input feature map into the first layer of Ghost module for convolution operation, and processing it through a BN layer and a Relu activation function with sparsity; inputting the feature map processed by the BN layer and the Relu activation function into the second layer of Ghost module for convolution operation, and processing it through another BN layer; the steps when the Ghost module performs convolution operation include: performing pointwise convolution on the input feature map through a 1×1 convolution kernel, compressing the number of channels of the input feature map through a scaling factor, normalizing it through a BatchNorm2d layer at the same time, and processing it through a SiLU activation function to obtain a concentrated feature map; performing layer-by-layer convolution operation on the concentrated feature map, normalizing it through a BatchNorm2d layer again, and processing it through a SiLU activation function to obtain a redundant feature map; concatenating and superimposing the concentrated feature map and the redundant feature map along the channel dimension, and outputting the superimposed result.

[0011] Further, the scaling factor is 2.

[0012] Furthermore, the steps of the result prediction module predicting the position and size information of the detection target and its belonging category based on the output feature maps of multiple different grid sizes include: at each spatial point of the third output feature map, the fourth output feature map, the fifth output feature map, and the sixth output feature map, predicting through four prior anchor boxes of corresponding sizes respectively to obtain the coordinate offsets (t x , t y ) of the predicted detection target bounding box, width t w , height t h , probability value, and confidence of the predicted category; according to the coordinate offsets (t x , t y ), width t w , and height t h to obtain the position coordinates and width and height of the detection target. The expression for the position coordinates (b x , b y ) of the detection target is:

[0013] b x = σ(t x ) + C x , b y = σ(t y ) + C y

[0014] In the formula, C x , C y are respectively the coordinates of the upper left corner of the grid where the detection target is located;

[0015] The expressions for the width b w and height b h of the detection target are:

[0016]

[0017] In the formula, p w , p h are respectively the width and height of the prior anchor box;

[0018] After non-maximum suppression processing, the final position coordinates, width, height, and confidence of the predicted category of the detection target are obtained, and the predicted category with high confidence is determined as the belonging category of the corresponding detection target.

[0019] Furthermore, the sizes of the prior anchor boxes are updated according to the dataset used to train the improved YOLOv5 model, including the steps: one, randomly selecting the actual bounding boxes in the dataset as the initial clustering centers according to the probability by the roulette wheel algorithm; two, calculating the distance Loss between each actual bounding box and the current clustering center. The expression for the Loss distance is:

[0020]

[0021] In the formula, Box i is the area of the i-th actual border among n actual borders, and Center j is the area of the j-th clustering center among k clustering centers;

[0022] Third, divide the actual border into the clustering category to which the clustering center with the shortest distance belongs; Fourth, calculate the median of the actual border coordinates in each clustering category, and update the clustering center of the corresponding category with this median. Repeat steps two to four until k clustering centers with stable positions are obtained, and determine the obtained clustering centers with stable positions as the prior anchor boxes; Fifth, calculate the size error degrees between each actual border and each prior anchor box, obtain the average value of the minimum size error degree values corresponding to each actual border, determine this average value as the fitness of the prior anchor box, and determine the prior anchor box with the highest fitness as the updated prior anchor box.

[0023] Further, the size of the prior anchor box is updated according to the dataset used to train the improved YOLOv5 model, and it further includes step six: perform a linear transformation on the updated prior anchor box, transform the minimum width of the updated prior anchor box to 0.8 times, and the maximum width to 1.5 times, and keep the width-to-height ratio unchanged.

[0024] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0025] Figure 1 is a schematic flowchart of the helmet wearing detection method based on the improved YOLOv5 in the embodiment;

[0026] Figure 2 is Figure 1 a schematic structural diagram of the improved YOLOv5 model in step S2 shown;

[0027] Figure 3 is a schematic structural diagram of the Swin transformer Block network in the embodiment;

[0028] Figure 4 is a schematic structural diagram of the multi-head self-attention layer in the embodiment;

[0029] Figure 5 is a schematic flowchart of the first multi-head self-attention module in the embodiment;

[0030] Figure 6 is a schematic structural diagram of the MLP layer in the embodiment;

[0031] Figure 7 It is a schematic structural diagram of the shifted window multi-head self-attention layer in the embodiment;

[0032] Figure 8 It is a schematic process diagram of the second multi-head self-attention module in the embodiment;

[0033] Figure 9 It is a schematic algorithm flow diagram of the first C3-Ghost layer, the second C3-Ghost layer, the third C3-Ghost layer, the fourth C3-Ghost layer, the fifth C3-Ghost layer, and the sixth C3-Ghost layer in the embodiment;

[0034] Figure 10 It is the detection result image output by the traditional YOLOv5 model in experiment (1);

[0035] Figure 11 It is the detection result image output by the improved YOLOv5 model in experiment (1);

[0036] Figure 12 It is the detection result image output by the traditional YOLOv5 model in experiment (2);

[0037] Figure 13 It is the detection result image output by the improved YOLOv5 model in experiment (2);

[0038] Figure 14 It is the detection result image output by the traditional YOLOv5 model in experiment (3);

[0039] Figure 15 It is the detection result image output by the improved YOLOv5 model in experiment (3). Detailed implementation manners

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0041] It should be clear that the described embodiments are only part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0043] In the description of this application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not necessarily describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise specified, "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0044] Please refer to Figure 1 , which is a schematic flowchart of the safety helmet wearing detection method based on the improved YOLOv5 in this embodiment. The method includes the steps:

[0045] S1: Obtain the image to be detected containing the detection target;

[0046] S2: Input the image to be detected into the improved YOLOv5 model for object detection to obtain the position and size information and the category to which the detection target belongs.

[0047] In step S1, the image to be detected can be collected at the construction site by a camera. After the camera collects the image, it is transmitted to a processor capable of executing a computer program through a data cable or network. The image collected by the camera can be picture data or video stream data. When the collected image is video stream data, the frame image is extracted as the image to be detected.

[0048] The image to be detected can also be a dataset for training and testing the improved YOLOv5 model. The images in the dataset are images of people wearing safety helmets at the construction site. The dataset can be an existing publicly available safety helmet open-source dataset. In addition, to improve the recognition effect of the model, the existing publicly available safety helmet open-source dataset can also be expanded by web crawling or field collection of images of people wearing safety helmets in different lighting environments at the construction site. Among them, the LableIMG annotation tool is used to make the labels of the dataset of images of people wearing safety helmets collected by web crawling or field collection, including the actual borders of the areas where the safety helmets are located. In this embodiment, the dataset is randomly divided into a training set and a test set according to a ratio of 9:1.

[0049] The detection target in the image to be detected can be image information such as an image of a person wearing a safety helmet or an image of a person's head without a safety helmet that needs to be recognized.

[0050] Please refer to Figure 2, which is a schematic structural diagram of the improved YOLOv5 model described in step S2. The model includes a feature extraction module, a feature fusion module, and a result prediction module. The feature extraction module is used to extract image features from the image to be detected; the feature fusion module is used to fuse the image features extracted by the feature extraction module and output output feature maps of multiple different grid sizes; the result prediction module is used to predict the position and size information and the category to which the detection target belongs according to the output feature maps of multiple different grid sizes.

[0051] Specifically, the feature extraction module includes a Patch Partition sub-module, a Linear Embedding sub-module, a first Swin-T (Swin transformer) sub-module, a first Patch Merging sub-module, a second Swin-T sub-module, a second Patch Merging sub-module, a third Swin-T sub-module, a third Patch Merging sub-module, a fourth Swin-T sub-module, a fourth Patch Merging sub-module, and a fifth Swin-T sub-module. When the feature extraction module extracts features from the image to be detected, it includes the following steps:

[0052] Input the image to be detected into the image segmentation sub-module for image segmentation;

[0053] Input the segmented image to be detected into the linear embedding sub-module for linear transformation;

[0054] Input the linearly transformed image to be detected into the first Swin-T sub-module for feature extraction to obtain a first Swin-T feature map;

[0055] Input the first Swin-T feature map into the first Patch Merging sub-module for downsampling to obtain a first-level feature map;

[0056] Input the first-level feature map into the second Swin-T sub-module for feature extraction to obtain a second Swin-T feature map;

[0057] Input the second Swin-T feature map into the second Patch Merging sub-module for downsampling to obtain a second-level feature map;

[0058] Input the second-level feature map into the third Swin-T sub-module for feature extraction to obtain a third Swin-T feature map;

[0059] Input the third Swin-T feature map into the third Patch Merging sub-module for downsampling to obtain a third-level feature map;

[0060] Input the third-level feature map into the fourth Swin-T sub-module for feature extraction to obtain a fourth Swin-T feature map;

[0061] The fourth Swin-T feature map is input into the fourth tile splicing sub-module for downsampling to obtain the fourth-level feature map;

[0062] The fourth-level feature map is input into the fifth Swin-T sub-module for feature extraction to obtain the fifth Swin-T feature map.

[0063] Among them, when the image segmentation sub-module performs image segmentation on the image to be detected, every a×a adjacent pixels in the input image to be detected with the dimension of [H, W, CH] are segmented into a tile, and are unfolded along the channel direction to obtain the image to be detected with the dimension of [H / a, W / a, a 2 CH], where H is the height of the image to be detected, W is the width of the image to be detected, and CH is the number of channels of the image to be detected. In a specific implementation, a is set to 4.

[0064] When the linear embedding sub-module performs a linear transformation on the segmented image to be detected, a linear transformation is performed on the channel data of each pixel of the segmented image to be detected to obtain the image to be detected with the dimension of [H / a, W / a, C], where C is a hyperparameter for adjusting the number of image channels to adapt to the feature fusion module. In a specific implementation, the hyperparameter C is set to 64.

[0065] Both the first Swin-T sub-module, the second Swin-T sub-module, the fourth Swin-T sub-module, and the fifth Swin-T sub-module include two Swin transformer Block networks for extracting image features from the input feature map, and the third Swin-T sub-module includes six of the said Swin transformer Block networks. Please refer to Figure 3 , which is the structural schematic diagram of the Swin transformer Block network. The Swin transformer Block network includes four LayerNorm layers, one multi-head self-attention (W-MSA) layer, two MLP layers, one shifted window multi-head self-attention (SW-MSA) layer, four DropPath layers, and four residual connection layers. When the Swin transformer Block network extracts image features, it includes the steps of:

[0066] Input the input feature map Z l-1 Input it into the LayerNorm layer for normalization processing; input the normalized input feature map Z l-1Input into the multi-head self-attention layer for multi-head self-attention feature extraction to obtain the multi-head self-attention feature map; input the multi-head self-attention feature map into the DropPath layer to randomly deactivate the multi-branch paths in the Swin Transformer Block, which is a regularization strategy to improve the generalization ability of the model and prevent overfitting; combine the multi-head self-attention feature map output by the DropPath layer with the feature map Z l-1 Input into the residual connection layer for residual connection to obtain the first intermediate feature map

[0067] The first intermediate feature map obtained by residual connection Input into the LayerNorm layer for normalization; the normalized first intermediate feature map Input into the MLP layer for linear transformation to obtain the first transformed feature map; input the first transformed feature map into the DropPath layer for random inactivation; combine the first transformed feature map output by the DropPath layer with the first intermediate feature map Input into the residual connection layer for residual connection to obtain the second intermediate feature map Z l ;

[0068] The second intermediate feature map Z l Input into the LayerNorm layer for normalization; the normalized second intermediate feature map Z l Input into the shifted window multi-head self-attention layer for pixel-shifted multi-head self-attention feature extraction to obtain the shifted multi-head self-attention feature map; input the shifted multi-head self-attention feature map into the DropPath layer for random inactivation; combine the shifted multi-head self-attention feature map output by the DropPath layer with the second intermediate feature map Z l Input into the residual connection layer for residual connection to obtain the third intermediate feature map

[0069] The third intermediate feature map Input into the LayerNorm layer for normalization; the normalized third intermediate feature map Input into the MLP layer for linear transformation to obtain the second transformed feature map; input the second transformed feature map into the DropPath layer for random inactivation; combine the second transformed feature map output by the DropPath layer with the third intermediate feature map Input into the residual connection layer for residual connection to obtain the output feature map Z l+1 。

[0070] For more details, please refer to Figure 4, which is a schematic structural diagram of a multi-head self-attention layer. The multi-head self-attention layer includes a first window partition module, a first multi-head self-attention module, and a first window reverse module. Among them, the first window partition module is used to partition the input feature map into non-overlapping independent windows of M*M adjacent pixels, that is, to partition the feature map into multiple patch vectors, so as to limit the calculation of the first multi-head self-attention module within each independent window, thereby reducing the computational complexity.

[0071] The first multi-head self-attention module is used to perform multi-head scaled dot-product attention calculations on each independent window respectively to obtain the multi-head self-attention features corresponding to each independent window. Specifically, please refer to Figure 5 , which is a schematic flow diagram of the first multi-head self-attention module. The steps in the first multi-head self-attention module include: performing a linear transformation on the patch vectors of each independent window in the channel dimension to double the number of channels, and at the same time dividing them into h subspaces in the feature dimension, where h is the number of attention heads; passing through h different parameter matrices W Q 、W K 、W V respectively perform linear transformations on the query Q (Quary), key K (Key), and value V (Value) of each pixel in the h subspaces, and perform scaled dot-product attention calculations; input the h calculation results into the Concat module and the Linear module, and perform splicing and fusion through the learnable weight matrix W O to jointly combine the feature information learned from different subspaces to obtain multi-head self-attention features. Among them, the expression of the scaled dot-product attention calculation result head i of the i-th attention head is:

[0072] head i =Attention(QW i Q ,KW i K ,VW i V )

[0073] In the formula, W i Q is the i-th parameter matrix W Q ; W i K is the i-th parameter matrix W K ; W i V is the i-th parameter matrix W V; Attention() is a normalized scaled dot product model, and its expression is:

[0074]

[0075] In the formula, QK T is the process of information interaction between different pixel points, and the similarity between different pixel points is calculated through the dot product; d is the dimension of a query and key vector, and by dividing by performing a scaling operation can ensure the stability of the gradient; B is the learnable relative position encoding (Relative Position Bias),

[0076] The splicing fusion expression of the multi-head self-attention features is:

[0077] MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O

[0078] The first window reorganization module is used to restore and splice the multi-head self-attention features of each independent window to obtain a complete multi-head self-attention feature map.

[0079] Please refer to Figure 6 , which is the structural schematic diagram of the MLP layer. The MLP layer includes two Linear layers, one GELU layer, and two Dropout layers. When the MLP layer performs a linear transformation on the input feature map, it includes the steps: linearly transforming the input feature map through the first Linear layer to obtain a feature map with four times the original number of channels; activating the feature map with four times the number of channels through the GELU layer using a non-linear activation function to increase the non-linearity of the network model; randomly inactivating the feature map output by the GELU layer through the first Dropout layer to avoid the network over-relying on certain local features and enhance the generalization of the model; linearly transforming the feature map output by the first Dropout layer through the second Linear layer to obtain a feature map with the same number of channels as the feature map input to the MLP layer; after randomly inactivating the feature map output by the second Linear layer through the second Dropout layer, it is determined as the output feature map.

[0080] Please refer to Figure 7, which is a schematic structural diagram of a shifted window multi-head self-attention layer. The shifted window multi-head self-attention layer includes a Cyclic Shift module, a second window splitting module, a second multi-head self-attention module, a second window recombination module, and a Reverse Cyclic Shift module. Among them, the Cyclic Shift module is used to move the top M / 2 rows of pixels of the input feature map to the bottom and the leftmost M / 2 columns of pixels to the rightmost.

[0081] The second window splitting module is used to split the shifted feature map into multiple non-overlapping independent windows of M*M adjacent pixels.

[0082] The second multi-head self-attention module is used to perform multi-head scaled dot-product attention calculations on each independent window respectively to obtain the shifted multi-head self-attention features corresponding to each independent window. Please refer to Figure 8 , which is a schematic process diagram of the second multi-head self-attention module. The steps in the second multi-head self-attention module include: linearly transforming the patch vectors of each independent window in the channel dimension to double the number of channels, and at the same time splitting them into h subspaces in the feature dimension, where h is the number of attention heads; passing through h different parameter matrices W Q 、W K 、W V respectively perform linear transformations on the query Q, key K, and value V of the pixel points in the h subspaces, and perform scaled dot-product attention calculations. A masking mechanism is added to the calculations to set the attention weight coefficients between the pixel points in non-adjacent regions before shifting in the independent window to 0, so as to isolate the ineffective information exchange between the pixel points in non-adjacent regions of the input feature map. In a specific implementation, subtract 100 from the similarity results of the pixel points in non-adjacent regions before shifting in the independent window, and then the result after softmax normalization is 0; input the h calculation results into the Concat module and the Linear module, and perform splicing and fusion through the learnable weight matrix W O to obtain the shifted multi-head self-attention features.

[0083] The second window recombination module is used to restore and splice the shifted multi-head self-attention features of each independent window.

[0084] The Reverse Cyclic Shift module is used to move the rightmost M / 2 columns of pixels of the restored and spliced feature map to the leftmost and the bottom M / 2 rows of pixels to the top, so as to restore the pixel positions of the feature map after cyclic shift and obtain the shifted multi-head self-attention feature map.

[0085] The first tile splicing sub-module, the second tile splicing sub-module, the third tile splicing sub-module, and the fourth tile splicing sub-module all include a tile segmentation layer, a concat layer, a LayerNorm layer, and a fully connected layer. When the first tile splicing sub-module, the second tile splicing sub-module, the third tile splicing sub-module, and the fourth tile splicing sub-module perform downsampling on the input feature map, the adjacent pixels with an interval of 2 in the input feature map with dimensions [H, W, C] are divided into multiple tiles by the tile segmentation layer; the divided tiles are concatenated through the concat layer, so that the dimension of the input feature map becomes [H / 2, W / 2, 4C]; the feature map is normalized through the LayerNorm layer; the number of channels of the feature map is linearly transformed through the fully connected layer, so that the dimension of the input feature map becomes [H / 2, W / 2, 2C].

[0086] The feature fusion module includes a first CONV layer, a first UP layer, a first Concat layer, a first C3-Ghost layer, a second CONV layer, a second UP layer, a second Concat layer, a second C3-Ghost layer, a third CONV layer, a third UP layer, a third Concat layer, a third C3-Ghost layer, a fourth CONV layer, a fourth Concat layer, a fourth C3-Ghost layer, a fifth CONV layer, a fifth Concat layer, a fifth C3-Ghost layer, a sixth CONV layer, a sixth Concat layer, and a sixth C3-Ghost layer. When the feature fusion module fuses the image features extracted by the feature extraction module, it includes the steps:

[0087] Obtain the fifth Swin-T feature map and input it into the first CONV layer for convolution processing to obtain the first convolution feature map; input the first convolution feature map into the first UP layer for upsampling operation; obtain the fourth Swin-T feature map and input it together with the feature map output by the first UP layer into the first Concat layer for Concat splicing; input the feature map output by the first Concat layer into the first C3-Ghost layer for convolution processing to obtain the first output feature map;

[0088] Input the first output feature map into the second CONV layer for convolution processing to obtain the second convolution feature map; input the second convolution feature map into the second UP layer for upsampling operation; obtain the third Swin-T feature map and input it together with the feature map output by the second UP layer into the second Concat layer for Concat splicing; input the feature map output by the second Concat layer into the second C3-Ghost layer for convolution processing to obtain the second output feature map;

[0089] The second output feature map is input into the third CONV layer for convolution processing to obtain the third convolution feature map; the third convolution feature map is input into the third UP layer for upsampling operation; the second Swin-T feature map is obtained and jointly input into the third Concat layer with the feature map output by the third UP layer for Concat splicing; the feature map output by the third Concat layer is input into the third C3-Ghost layer for convolution processing to obtain the third output feature map;

[0090] The third output feature map is input into the fourth CONV layer for convolution processing to obtain the fourth convolution feature map; the fourth convolution feature map and the third convolution feature map are jointly input into the fourth Concat layer for Concat splicing; the feature map output by the fourth Concat layer is input into the fourth C3-Ghost layer for convolution processing to obtain the fourth output feature map;

[0091] The fourth output feature map is input into the fifth CONV layer for convolution processing to obtain the fifth convolution feature map; the fifth convolution feature map and the second convolution feature map are jointly input into the fifth Concat layer for Concat splicing; the feature map output by the fifth Concat layer is input into the fifth C3-Ghost layer for convolution processing to obtain the fifth output feature map;

[0092] The fifth output feature map is input into the sixth CONV layer for convolution processing to obtain the sixth convolution feature map; the sixth convolution feature map and the first convolution feature map are jointly input into the sixth Concat layer for Concat splicing; the feature map output by the sixth Concat layer is input into the sixth C3-Ghost layer for convolution processing to obtain the sixth output feature map.

[0093] In a specific implementation, the first UP layer, the second UP layer, and the third UP layer perform upsampling operations through the nearest neighbor interpolation algorithm.

[0094] In a preferred embodiment of the feature fusion module for fusing the image features extracted by the feature extraction module, the step of jointly inputting the fourth convolution feature map and the third convolution feature map into the fourth Concat layer for Concat splicing can be replaced with: jointly inputting the third Swin-T feature map, the fourth convolution feature map, and the third convolution feature map into the fourth Concat layer for Concat splicing; the step of jointly inputting the fifth convolution feature map and the second convolution feature map into the fifth Concat layer for Concat splicing can be replaced with: jointly inputting the fourth Swin-T feature map, the fifth convolution feature map, and the second convolution feature map into the fifth Concat layer for Concat splicing. In this preferred embodiment, by adding horizontal skip connections between the original input nodes and output nodes at the same level, the feature maps at the same level can share each other's semantic information, which can strengthen feature fusion to improve the model accuracy.

[0095] Please refer to Figure 9 which is a schematic diagram of the algorithm flow of the first C3-Ghost layer, the second C3-Ghost layer, the third C3-Ghost layer, the fourth C3-Ghost layer, the fifth C3-Ghost layer, and the sixth C3-Ghost layer. When the first C3-Ghost layer, the second C3-Ghost layer, the third C3-Ghost layer, the fourth C3-Ghost layer, the fifth C3-Ghost layer, and the sixth C3-Ghost layer perform convolution processing on the input feature map, it includes the steps:

[0096] Perform a standard convolution operation on the input feature map to compress the number of channels, and perform feature extraction through N cascaded GhostBottleneck modules to obtain the first C3-Ghost feature map;

[0097] At the same time, perform another standard convolution operation on the input feature map to obtain the second C3-Ghost feature map;

[0098] Concat and stack the first C3-Ghost feature map and the second C3-Ghost feature map in the channel dimension, and perform feature fusion through convolution to obtain the output feature map.

[0099] More specifically, when the Ghost Bottleneck module performs feature extraction on the input feature map, it includes the steps:

[0100] Input the input feature map into the first layer of Ghost module for convolution operation, and process it through the BN (Batch Normalization) layer and the Relu activation function with sparsity, where the BN layer is used to ensure that the input of each layer of the network has the same distribution, and the Relu activation function is used to avoid the phenomenon of gradient disappearance in backpropagation;

[0101] Input the feature map processed by the BN layer and the Relu activation function into the second layer of Ghost module for convolution operation, and process it through another BN layer. At this time, the Relu activation function is not used because the hard saturation to 0 in the negative half-axis of the ReLU activation function will make the output data distribution not zero-mean, resulting in neuron inactivation and thus reducing the performance of the network.

[0102] Among them, when the Ghost module performs convolution operation on the input feature map, it includes the steps:

[0103] The input feature map is subjected to pointwise convolution through a 1×1 convolution kernel, and then the number of channels of the input feature map is compressed by a scaling factor ratio. At the same time, normalization is performed through the BatchNorm2d layer, and it is processed through the SiLU activation function to obtain a concentrated feature map. In this embodiment, the scaling factor ratio is 2, and the number of channels of the input feature map is compressed to half of the original;

[0104] Perform layer-by-layer convolution on the concentrated feature map, then perform normalization through the BatchNorm2d layer, and process it through the SiLU activation function to obtain a redundant feature map;

[0105] Concat and stack the concentrated feature map and the redundant feature map along the channel dimension, and output the stacked result.

[0106] The result prediction module obtains the position and size information and the category to which the detection target belongs according to the third output feature map, the fourth output feature map, the fifth output feature map, and the sixth output feature map with different grid sizes. Among them, the output feature map with the largest grid size is used to detect small targets, and the feature map with the smallest grid size is used to detect large targets. It specifically includes the steps:

[0107] At each spatial point of the third output feature map, the fourth output feature map, the fifth output feature map, and the sixth output feature map with different grid sizes, predictions are respectively made through four prior anchor boxes with corresponding sizes to obtain the coordinate offsets (t x , t y ) of the predicted target bounding box, width t w , height t h , probability value, and confidence of the predicted category; according to the coordinate offsets (t x , t y ) of the predicted target bounding box, the width t w and height t h of the target bounding box, obtain the position coordinates and width and height of the detection target. The expression for the position coordinates (b x , b y ) of the detection target is:

[0108] b x = σ(t x ) + C x , b y = σ(t y ) + C y

[0109] In the formula, C x , C y are respectively the coordinates of the upper left corner of the grid where the detection target is located in the feature map,

[0110] The width b of the detection targetw 、 The height b h The expression for is:

[0111]

[0112] In the formula, p w 、 p h are the width and height of the prior anchor box respectively;

[0113] Finally, through non-maximum suppression processing, the position coordinates, width and height of the final detected object and the confidence of the predicted class are obtained, and the predicted class with high confidence is determined as the class to which the corresponding detected object belongs. In a specific implementation, the position coordinates, width and height of the final detected object and the class to which the detected object belongs will be marked on the image to be detected and output as a detection result image. The classes to which the detected objects belong are wearing a safety helmet correctly and not wearing a safety helmet.

[0114] When training the improved YOLOv5 model of this embodiment, the position size information and the class to which the detected object in the image to be detected output by the model are input into the loss function to calculate the gap between the predicted data and the actual data, and the parameters of the model are adjusted through the loss function.

[0115] In a preferred embodiment, the prior anchor box size in the improved YOLOv5 model is updated according to the dataset used to train the improved YOLOv5 model, which specifically includes the steps of:

[0116] First, in the dataset, the actual bounding boxes are selected as the initial clustering centers according to the probability by the roulette wheel algorithm;

[0117] Second, calculate the distance Loss between each actual bounding box in the dataset and the current clustering center. The expression for the Loss distance is:

[0118]

[0119] In the formula, Box i is the area of the i-th actual bounding box among n actual bounding boxes, and Center j is the area of the j-th clustering center among k clustering centers;

[0120] Third, divide the actual bounding boxes into the clustering categories to which the clustering centers with the shortest distance to them belong;

[0121] Fourth, calculate the median of the actual bounding box coordinates in each clustering category, update the clustering center of the corresponding category with this median, and repeat steps two to four until k clustering centers with stable positions are selected, and the obtained clustering centers are determined as the prior anchor boxes;

[0122] 5. Calculate the dimensional error degrees between each actual bounding box and each prior anchor box respectively, obtain the average value of the minimum dimensional error degree values corresponding to each of the actual bounding boxes, determine this average value as the fitness of the prior anchor box, and determine the prior anchor box with the highest fitness as the updated prior anchor box;

[0123] To better utilize the multi-scale object detection ability of the improved YOLOv5 model, it further includes Step 6: perform a linear transformation on the prior anchor box, transform the minimum width of the prior anchor box to 0.8 times, and the maximum width to 1.5 times, while keeping the width-to-height ratio unchanged.

[0124] The following is a detection comparison experiment between the improved YOLOv5 model and the traditional YOLOv5 model:

[0125] (1) Input the to-be-detected images with the detection targets occluded into the traditional YOLOv5 model and the improved YOLOv5 model respectively to obtain the detection result images. Please refer to Figure 10 and Figure 11 , where Figure 10 is the detection result image output by the traditional YOLOv5 model, Figure 11 is the detection result image output by the improved YOLOv5 model. It can be seen that the improved YOLOv5 model can still detect the occluded detection target, while the traditional YOLOv5 model cannot detect this occluded detection target.

[0126] (2) Input the to-be-detected images containing circular controller images into the traditional YOLOv5 model and the improved YOLOv5 model respectively to obtain the detection result images. Please refer to Figure 12 and Figure 13 , where Figure 12 is the detection result image output by the traditional YOLOv5 model, Figure 13 is the detection result image output by the improved YOLOv5 model. It can be seen that the traditional YOLOv5 model misidentifies the circular controller as a person's head without a safety helmet, while the improved YOLOv5 model can distinguish that the circular controller in the image is not a person's head and has higher detection accuracy.

[0127] (3) Input the to-be-detected images obtained under low light conditions into the traditional YOLOv5 model and the improved YOLOv5 model respectively to obtain the detection result images. Please refer to Figure 14 and 15 , where Figure 14 is the detection result image output by the traditional YOLOv5 model, Figure 15It is the detection result image output by the improved YOLOv5 model. It can be seen that there are missed detections in the traditional YOLOv5 model, while the improved YOLOv5 model can identify all detection targets in the image to be detected and still maintain a high detection accuracy under the influence of the lighting environment.

[0128] In addition, ablation experiments were also carried out on the improved YOLOv5 model. The experimental results are shown in Table 1. Among them, "traditional YOLOv5" is the traditional YOLOv5 model with four detection scales; "traditional YOLOv5 + C3 - Ghost" is the model in which the C3 module of the traditional YOLOv5 model is replaced by the C3 - Ghost module of the embodiment; "traditional YOLOv5 + improved feature fusion" is the model with a lateral skip connection added between the original input node and the output node of the same level in the feature fusion module of the traditional YOLOv5 model with four detection scales; "traditional YOLOv5 + C3 - Ghost + improved feature fusion" is the model in which the C3 module of the traditional YOLOv5 model is replaced by the C3 - Ghost module of the embodiment and a lateral skip connection is added between the original input node and the output node of the same level in the feature fusion module; "traditional YOLOv5 + SwinTransformer" is the model using Swin Transformer as the backbone feature extraction network of the traditional YOLOv5; "traditional YOLOv5 + Swin Transformer + C3Ghost" is the model using Swin Transformer as the backbone feature extraction network of the traditional YOLOv5 and replacing the C3 module of the traditional YOLOv5 model with the C3 - Ghost module of the embodiment; "improved YOLOv5" is the improved YOLOv5 model in the embodiment; P is the accuracy of the model, R is the recall rate of the model, mAP@.5 is the average of the AP (Average Precision) values of each category when the IoU threshold is 0.5; mAP@.5:.95 represents the average mAP corresponding to the IoU threshold starting from 0.5 and increasing in steps of 0.05 to 0.95.

[0129] Table 1

[0130]

[0131] Among them, the number of parameters of the traditional YOLOv5 model is 7.17×10 6 , and the number of parameters of the traditional YOLOv5 + C3 - Ghost is 6.14×10 6, It can be seen that, while keeping the mAP@.5 value almost unchanged, the number of parameters of the traditional YOLOv5+C3-Ghost is reduced by 14.4% compared with the traditional YOLOv5, which proves that the C3-Ghost module in this embodiment can effectively reduce the model parameters and computational complexity. The mAP@.5 value of the traditional YOLOv5+improved feature fusion is increased by 0.5% compared with the traditional YOLOv5. The traditional YOLOv5+Swin Transformer model is improved by 2.1% in the mAP@.5:.95 metric compared with the traditional YOLOv5 model. The traditional YOLOv5+Swin Transformer+C3Ghost model is improved by 1.9% in the mAP@.5:.95 metric. The improved YOLOv5 network model has higher feature extraction ability due to the feature extraction based on Swin Transformer. At the same time, it has both the computational simplicity brought by the C3-Ghost module and the high accuracy brought by the improved feature fusion. It can be seen from Table 1 that the improved YOLOv5 model in this embodiment is improved by 2.3% in the mAP@.5:.95 metric compared with the traditional YOLOv5 model, that is, it obviously has higher detection accuracy.

[0132] The present application may be implemented in the form of a computer program product on one or more storage media including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc., which contain program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device.

[0133] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one Figure One process or multiple processes and / or boxes Figure One or multiple boxes.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure One One process or a plurality of processes and / or blocks Figure One Steps for realizing the functions specified in one block or a plurality of blocks.

[0135] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising an ……" does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0136] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and the present invention is also intended to cover these modifications and improvements.

Claims

1. An improved YOLOv5-based safety helmet wearing detection method, characterized in that, Including the steps: Obtain a to-be-detected image containing a detection target; Input the to-be-detected image into an improved YOLOv5 model for object detection to obtain the position and size information and the category to which the detection target belongs. Among them, the improved YOLOv5 model includes a feature extraction module, a feature fusion module, and a result prediction module. The feature extraction module includes a tile segmentation sub-module, a linear embedding sub-module, a first Swin-T sub-module, a first tile stitching sub-module, a second Swin-T sub-module, a second tile stitching sub-module, a third Swin-T sub-module, a third tile stitching sub-module, a fourth Swin-T sub-module, a fourth tile stitching sub-module, and a fifth Swin-T sub-module. When the feature extraction module extracts features from the to-be-detected image, it includes the steps: Input the to-be-detected image into the image segmentation sub-module for image segmentation; Input the segmented to-be-detected image into the linear embedding sub-module for linear transformation; Input the linearly transformed to-be-detected image into the first Swin-T sub-module for feature extraction to obtain a first Swin-T feature map; Input the first Swin-T feature map into the first tile stitching sub-module for downsampling to obtain a first-level feature map; Input the first-level feature map into the second Swin-T sub-module for feature extraction to obtain a second Swin-T feature map; Input the second Swin-T feature map into the second tile stitching sub-module for downsampling to obtain a second-level feature map; Input the second-level feature map into the third Swin-T sub-module for feature extraction to obtain a third Swin-T feature map; Input the third Swin-T feature map into the third tile stitching sub-module for downsampling to obtain a third-level feature map; Input the third-level feature map into the fourth Swin-T sub-module for feature extraction to obtain a fourth Swin-T feature map; Input the fourth Swin-T feature map into the fourth tile stitching sub-module for downsampling to obtain a fourth-level feature map; Input the fourth-level feature map into the fifth Swin-T sub-module for feature extraction to obtain a fifth Swin-T feature map; Among them, the first Swin-T sub-module, the second Swin-T sub-module, the fourth Swin-T sub-module, and the fifth Swin-T sub-module each include two Swin transformer Block networks, and the third Swin-T sub-module includes six of the Swin transformer Block networks. The Swin transformer Block network is used for image feature extraction of the input feature map; The feature fusion module is used to fuse the fifth Swin-T feature map, the fourth Swin-T feature map, the third Swin-T feature map, and the second Swin-T feature map by using an improved C3-Ghost module to obtain output feature maps with multiple different grid sizes; wherein, the improved C3-Ghost module includes six layers of C3-Ghost, and each layer of C3-Ghost includes N serially connected Ghost Bottleneck modules; The result prediction module is used to predict the position and size information and the category to which the detection target belongs based on the output feature maps with multiple different grid sizes; The feature fusion module includes a first CONV layer, a first UP layer, a first Concat layer, a first C3-Ghost layer, a second CONV layer, a second UP layer, a second Concat layer, a second C3-Ghost layer, a third CONV layer, a third UP layer, a third Concat layer, a third C3-Ghost layer, a fourth CONV layer, a fourth Concat layer, a fourth C3-Ghost layer, a fifth CONV layer, a fifth Concat layer, a fifth C3-Ghost layer, a sixth CONV layer, a sixth Concat layer, and a sixth C3-Ghost layer. When the feature fusion module fuses the fifth Swin-T feature map, the fourth Swin-T feature map, the third Swin-T feature map, and the second Swin-T feature map to obtain output feature maps with multiple different grid sizes, it includes the steps of: Obtain the fifth Swin-T feature map and input it into the first CONV layer for convolution processing to obtain a first convolution feature map; input the first convolution feature map into the first UP layer for upsampling operation; obtain the fourth Swin-T feature map and input it together with the feature map output by the first UP layer into the first Concat layer for Concat splicing; input the feature map output by the first Concat layer into the first C3-Ghost layer for convolution processing to obtain a first output feature map; Input the first output feature map into the second CONV layer for convolution processing to obtain a second convolution feature map; input the second convolution feature map into the second UP layer for upsampling operation; obtain the third Swin-T feature map and input it together with the feature map output by the second UP layer into the second Concat layer for Concat splicing; input the feature map output by the second Concat layer into the second C3-Ghost layer for convolution processing to obtain a second output feature map; Input the second output feature map into the third CONV layer for convolution processing to obtain a third convolution feature map; input the third convolution feature map into the third UP layer for upsampling operation; obtain the second Swin-T feature map and input it together with the feature map output by the third UP layer into the third Concat layer for Concat splicing; input the feature map output by the third Concat layer into the third C3-Ghost layer for convolution processing to obtain a third output feature map; Input the third output feature map into the fourth CONV layer for convolution processing to obtain a fourth convolution feature map; input the fourth convolution feature map and the third convolution feature map together into the fourth Concat layer for Concat splicing; input the feature map output by the fourth Concat layer into the fourth C3-Ghost layer for convolution processing to obtain a fourth output feature map; Input the fourth output feature map into the fifth CONV layer for convolution processing to obtain a fifth convolution feature map; input the fifth convolution feature map and the second convolution feature map together into the fifth Concat layer for Concat splicing; input the feature map output by the fifth Concat layer into the fifth C3-Ghost layer for convolution processing to obtain a fifth output feature map; Input the fifth output feature map into the sixth CONV layer for convolution processing to obtain a sixth convolution feature map; input the sixth convolution feature map and the first convolution feature map together into the sixth Concat layer for Concat splicing; input the feature map output by the sixth Concat layer into the sixth C3-Ghost layer for convolution processing to obtain a sixth output feature map; Alternatively, the feature fusion module includes a first CONV layer, a first UP layer, a first Concat layer, a first C3-Ghost layer, a second CONV layer, a second UP layer, a second Concat layer, a second C3-Ghost layer, a third CONV layer, a third UP layer, a third Concat layer, a third C3-Ghost layer, a fourth CONV layer, a fourth Concat layer, a fourth C3-Ghost layer, a fifth CONV layer, a fifth Concat layer, a fifth C3-Ghost layer, a sixth CONV layer, a sixth Concat layer and a sixth C3-Ghost layer. When the feature fusion module fuses the fifth Swin-T feature map, the fourth Swin-T feature map, the third Swin-T feature map, and the second Swin-T feature map to obtain output feature maps of multiple different grid sizes, it includes the steps: Obtain the fifth Swin-T feature map and input it into the first CONV layer for convolution processing to obtain the first convolutional feature map; input the first convolutional feature map into the first UP layer for upsampling operation; obtain the fourth Swin-T feature map and input it together with the feature map output by the first UP layer into the first Concat layer for Concat splicing; input the feature map output by the first Concat layer into the first C3-Ghost layer for convolution processing to obtain the first output feature map; Input the first output feature map into the second CONV layer for convolution processing to obtain the second convolutional feature map; input the second convolutional feature map into the second UP layer for upsampling operation; obtain the third Swin-T feature map and input it together with the feature map output by the second UP layer into the second Concat layer for Concat splicing; input the feature map output by the second Concat layer into the second C3-Ghost layer for convolution processing to obtain the second output feature map; Input the second output feature map into the third CONV layer for convolution processing to obtain the third convolutional feature map; input the third convolutional feature map into the third UP layer for upsampling operation; obtain the second Swin-T feature map and input it together with the feature map output by the third UP layer into the third Concat layer for Concat splicing; input the feature map output by the third Concat layer into the third C3-Ghost layer for convolution processing to obtain the third output feature map; Input the third output feature map into the fourth CONV layer for convolution processing to obtain the fourth convolutional feature map; input the third Swin-T feature map, the fourth convolutional feature map and the third convolutional feature map together into the fourth Concat layer for Concat splicing; input the feature map output by the fourth Concat layer into the fourth C3-Ghost layer for convolution processing to obtain the fourth output feature map; Input the fourth output feature map into the fifth CONV layer for convolution processing to obtain the fifth convolutional feature map; input the fourth Swin-T feature map, the fifth convolutional feature map and the second convolutional feature map together into the fifth Concat layer for Concat splicing; input the feature map output by the fifth Concat layer into the fifth C3-Ghost layer for convolution processing to obtain the fifth output feature map; Input the fifth output feature map into the sixth CONV layer for convolution processing to obtain the sixth convolutional feature map; input the sixth convolutional feature map and the first convolutional feature map together into the sixth Concat layer for Concat splicing; input the feature map output by the sixth Concat layer into the sixth C3-Ghost layer for convolution processing to obtain the sixth output feature map.

2. The method according to claim 1, characterized in that: The Swin transformer Block network includes four LayerNorm layers, one multi-head self-attention layer, two MLP layers, one shifted window multi-head self-attention layer, four DropPath layers, and four residual connection layers. When the Swin transformer Block network performs image feature extraction on the input feature map, it includes the following steps: Input the input feature map into the LayerNorm layer for normalization processing; input the normalized input feature map into the multi-head self-attention layer for multi-head self-attention feature extraction to obtain a multi-head self-attention feature map; input the multi-head self-attention feature map into the DropPath layer for random inactivation; input the multi-head self-attention feature map output by the DropPath layer and the input feature map into the residual connection layer for residual connection to obtain a first intermediate feature map; Input the first intermediate feature map into the LayerNorm layer for normalization processing; input the normalized first intermediate feature map into the MLP layer for linear transformation to obtain a first transformed feature map; input the first transformed feature map into the DropPath layer for random inactivation; input the first transformed feature map output by the DropPath layer and the first intermediate feature map into the residual connection layer for residual connection to obtain a second intermediate feature map; Input the second intermediate feature map into the LayerNorm layer for normalization processing; input the normalized second intermediate feature map into the shifted window multi-head self-attention layer for pixel-shifted multi-head self-attention feature extraction to obtain a shifted multi-head self-attention feature map; input the shifted multi-head self-attention feature map into the DropPath layer for random inactivation; input the shifted multi-head self-attention feature map output by the DropPath layer and the second intermediate feature map into the residual connection layer for residual connection to obtain a third intermediate feature map; Input the third intermediate feature map into the LayerNorm layer for normalization processing; input the normalized third intermediate feature map into the MLP layer for linear transformation to obtain a second transformed feature map; input the second transformed feature map into the DropPath layer for random inactivation; input the second transformed feature map output by the DropPath layer and the third intermediate feature map into the residual connection layer for residual connection to obtain the feature map as the output of the Swin transformer Block network.

3. The method according to claim 1, wherein: The first tile splicing submodule, the second tile splicing submodule, the third tile splicing submodule and the fourth tile splicing submodule all include a tile segmentation layer, a concat layer, a LayerNorm layer and a fully connected layer, wherein the tile segmentation layer is used to divide adjacent pixels with an interval of 2 in the input feature map of dimension [H, W, C] into multiple tiles; the concat layer is used to concat the segmented tiles to obtain a feature map with a dimension of [H / 2, W / 2, 4C]; the LayerNorm layer is used to normalize the feature map output by the concat layer; the fully connected layer is used to linearly transform the number of channels of the feature map output by the LayerNorm layer to obtain a feature map with a dimension of [H / 2, W / 2, 2C].

4. The method according to claim 1, characterized in that: When the first C3-Ghost layer, the second C3-Ghost layer, the third C3-Ghost layer, the fourth C3-Ghost layer, the fifth C3-Ghost layer, and the sixth C3-Ghost layer perform convolution processing on the input feature map, the steps include: Perform a standard convolution operation on the input feature map to compress the number of channels, and extract features through N series-connected GhostBottleneck modules to obtain the first C3-Ghost feature map; Performing another standard convolution operation on the input feature map to obtain a second C3-Ghost feature map; Concat the first C3-Ghost feature map and the second C3-Ghost feature map according to the channel dimension, and perform feature fusion through convolution to obtain an output feature map; The steps of feature extraction in the Ghost Bottleneck module include: The input feature map is input into the first layer Ghost module module for convolution operation, and processed by the BN layer and the Relu activation function with sparsity; The feature map processed by the BN layer and the Relu activation function is input into the second layer Ghost module module for convolution operation, and processed by another BN layer; The steps of the Ghost module when performing the convolution operation include: The input feature map is convolved point by point through a 1×1 convolution kernel, and the number of channels of the input feature map is compressed by a scaling factor. At the same time, it is normalized through a BatchNorm2d layer and processed by a SiLU activation function to obtain a concentrated feature map. The concentrated feature map is convolved layer by layer, normalized by a BatchNorm2d layer, and processed by a SiLU activation function to obtain a redundant feature map; The concentrated feature map and the redundant feature map are concat-superimposed according to the channel dimension, and the superposition result is output.

5. The method according to claim 4, characterized in that: The scaling factor is 2.

6. The method according to claim 1, wherein The step of predicting the position and size information of the detection target and the category to which it belongs based on the output feature maps of multiple different grid sizes by the result prediction module comprises: At each spatial point of the third output feature map, the fourth output feature map, the fifth output feature map, and the sixth output feature map, predictions are respectively made through four prior anchor boxes of corresponding sizes to obtain the coordinate offsets (t x , t y ) of the predicted detection target bounding box, the width t w , the height t h , the probability value, and the confidence of the predicted class; according to the coordinate offsets (t x , t y ), the width t w , and the height t h , the position coordinates and width and height of the detection target are obtained. The expression for the position coordinates (b x , b y ) of the detection target is: b x = σ(t x ) + C x , b y = σ(t y ) + C y where C x and C y are the coordinates of the upper left corner of the grid where the detection target is located, respectively; The width b of the detection target w and the height b h are expressed as: where p w and p h are the width and height of the prior anchor box, respectively; After non-maximum suppression processing, the position coordinates, width, height, and confidence of the predicted category of the final detected target are obtained, and the predicted category with a high confidence is determined as the category to which the corresponding detected target belongs.

7. The method according to claim 6, wherein: The size of the prior anchor box is updated according to the dataset used to train the improved YOLOv5 model, including the steps of: First, the actual bounding boxes in the dataset are selected as the initial clustering centers according to the probability by the roulette wheel algorithm. Second, calculate the distance Loss between each actual bounding box and the current clustering center. The expression of the Loss distance is: Wherein, Box i is the area of the i-th actual border among n actual borders, and Center j is the area of the j-th cluster center among k cluster centers; Third, divide the actual bounding boxes into the clustering categories to which the clustering centers with the shortest distance belong. Fourth, calculate the median of the actual bounding box coordinates in each clustering category, and update the clustering center of the corresponding category with this median. Repeat steps two to four until k position-stable clustering centers are obtained, and determine the obtained position-stable clustering centers as the prior anchor boxes. Fifth, calculate the size error degrees between each actual bounding box and each prior anchor box respectively, obtain the average value of the minimum size error degree values corresponding to each actual bounding box, determine this average value as the fitness of the prior anchor box, and determine the prior anchor box with the highest fitness as the updated prior anchor box.

8. The method according to claim 7, wherein The size of the prior anchor box is updated according to the dataset used to train the improved YOLOv5 model, and also includes step six: perform a linear transformation on the updated prior anchor box, transform the minimum width of the updated prior anchor box to 0.8 times, and the maximum width to 1.5 times, and keep the aspect ratio unchanged.