Helmet Detection Method Based on Noise Cancellation Training and Dimensionality Reduction Attention Mechanism
By introducing noise cancellation training and dimensionality reduction attention mechanisms into the DETR object detection algorithm and building a lightweight feature extraction network, the problem of slow training convergence speed and large memory usage in DETR algorithm is solved, and the model is lightweight and efficient detection is achieved.
Patent Information
- Application Number
- CN202310816812.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-07-05
AI Technical Summary
The existing DETR object detection algorithm has problems such as slow training convergence speed and large video memory usage, and the model parameters and calculation amount are huge, making it difficult to operate effectively in edge devices.
Using a method based on noise cancellation training and dimensionality reduction attention mechanism, we will reduce the model size and improve the convergence speed by building a noise cancellation training framework and improving the multi-head attention mechanism. At the same time, a lightweight feature extraction network is built using sub-feature fusion and cross-layer perception enhancement convolution modules, and joins the DETR network as a backbone network.
It effectively improves the convergence speed of the model, reduces the video memory usage during model training, and reduces the model size, while achieving accurate detection of the wearing of the safety helmet.
Smart Images

Figure CN116758485B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition and computer vision, and particularly relates to a safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism. Background Art
[0002] In a construction site, construction workers must wear safety helmets to ensure their work safety. During industrial production, the most common cause of traumatic brain injury accidents is that workers fall from heights or are hit on the head by falling objects. Wearing a safety helmet can minimize the harm to the head. Research shows that most safety accidents can be avoided if workers wear safety helmets during work.
[0003] Currently, the recognition of the above safety helmet wearing situation mainly relies on manual inspection. Manual inspection is easily interfered by various factors, reducing efficiency, and cannot be monitored all day long, wasting human resources. Therefore, this method has low efficiency and cannot fully meet the actual needs of each safety supervision department. By using devices such as surveillance cameras and applying detection algorithms based on machine vision, such as the DETR object detection algorithm, to automatically identify the safety helmet wearing situation of workers, real-time monitoring can be achieved, avoiding problems such as missed detection and false detection caused by various subjective factors in manual detection, and improving the supervision efficiency and automation level while saving human resources. Although DETR has proposed a new paradigm for object detection, compared with object detectors based on convolutional neural networks, its training convergence speed is very slow. Due to the discreteness of bipartite graph matching and the randomness of model training in DETR, the matching of ground-truth becomes a dynamic and unstable process, greatly slowing down the convergence speed of the training model. The multi-head attention mechanism used in DETR is usually affected by high computational and memory costs, occupying a large amount of GPU memory during training, which may lead to the inability to use higher-resolution input images or other functions due to insufficient memory. In addition, the feature extraction network occupies most of the parameters and computational volume of the DETR model, and good detection performance cannot be continued to be exerted on edge devices with less memory and computational resources. Summary of the Invention
[0004] In view of the defects and deficiencies existing in the prior art, the present invention proposes a lightweight safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism.
[0005] First, obtain the safety helmet data pictures and preprocess them to construct a safety helmet detection dataset. Then, construct a noise cancellation training framework to replace the bipartite graph matching algorithm, add noise to the ground-truth information of each target object, introduce an attention mask, group the noise objects, and finally input the noise objects into the Transformer decoder. In addition, the present invention improves the multi-head attention mechanism in the Transformer architecture, decouples the key features into a one-dimensional row feature and a one-dimensional column feature through one-dimensional global average pooling, then performs row attention and column attention in sequence, and finally performs weighted summation. The present invention uses a lightweight feature extraction network constructed by sub-feature fusion and cross-layer perception enhanced convolution modules, and adds it as a backbone network to the DETR network based on the Transformer architecture to obtain a lightweight DETR detection network. The present invention can effectively reduce the model size of the detection model and accurately detect whether the worker wears a safety helmet.
[0006] It can effectively improve the model convergence speed and reduce the video memory occupation during model training, while reducing the model size of the detection model and accurately detecting whether the worker wears a safety helmet.
[0007] The technical solution specifically adopted by the present invention to solve its technical problems is as follows:
[0008] A safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism, characterized by comprising the following steps;
[0009] Step S1: Obtain the safety helmet data pictures and preprocess them to construct a safety helmet detection dataset;
[0010] Step S2: Construct a noise cancellation training framework, add noise to the ground-truth information of each target object, introduce an attention mask, group the noise objects, and finally input the noise objects into the Transformer decoder;
[0011] Step S3: Improve the multi-head attention mechanism in the Transformer architecture, decouple the key features into a one-dimensional row feature and a one-dimensional column feature through one-dimensional global average pooling, then perform row attention and column attention in sequence, and finally perform weighted summation on the two parts of one-dimensional attention results;
[0012] Step S4: Use a lightweight feature extraction network constructed by sub-feature fusion and cross-layer perception enhanced convolution modules, and add it as a backbone network to the DETR network based on the Transformer architecture to obtain a lightweight DETR detection network;
[0013] Step S5: Tune the training hyperparameters of the DETR object detection algorithm, train the DETR object detection network using the safety helmet detection dataset to obtain a safety helmet detection model, and use the safety helmet detection model to detect the input image.
[0014] Further, step S1 is specifically as follows:
[0015] Step S11; Obtain data pictures related to safety helmets, screen the pictures and uniformly name them;
[0016] Step S12: Process the pictures using neighborhood denoising and median filtering;
[0017] Step S13: Determine the object categories in the safety helmet pictures, use the LabelImg annotation tool to annotate the preprocessed data pictures, and obtain and save the annotation information;
[0018] Step S14: Make a dataset according to the requirements of the DETR model, divide all the data into a training set, a validation set and a test set according to a ratio, and generate the training set, validation set and test set json annotation files required for training the model according to the xml files containing the annotation information of the picture data.
[0019] Further, step S2 is specifically as follows:
[0020] Step S21: Add noise to the category information and coordinate box information of each detected object in the DETR detection network, and add the loss of the noise cancellation module to the loss function;
[0021] Step S22: Add an attention mask to the self-attention module of the DETR detection network, divide the noise objects into multiple denoising groups, the attention mask is divided into a denoising part and a matching part, and ensure that there is no information interaction between the matching part and the denoising part, and between the denoising groups; The calculation method of the attention mask is specifically as follows:
[0022]
[0023] Among them, P and M respectively represent the number of denoising groups and ground-truth objects, N is the number of queries in the matching part; The first P×M rows and columns of the attention mask represent the denoising part, and the rest represent the matching part; a ij = 1 means that the i-th query cannot see the j-th query, otherwise set a ij to 0;
[0024] Step S23: Input the noise objects into the Transformer decoder, and predict the object category and coordinate box by learning the noise objects and ground-truth; where ground-truth represents the true information of the object category and coordinate box.
[0025] Furthermore, step S21 is specifically as follows:
[0026] Step S211: Add class noise, and randomly replace the true class of the target object with any other class at a given probability;
[0027] Step S212: Add bounding box noise, which is divided into two categories: center point displacement and scale scaling; for center point displacement, ensure that the center point remains inside the true box and perform random movement; for scale scaling, randomly scale the length and width of the bounding box;
[0028] Step S213: Add the loss of the noise elimination module to the loss function, and combine it with the bounding box loss, class loss, and other losses as the final loss value. The formula of the loss function is as follows:
[0029]
[0030]
[0031] Among them, x, y, w, h represent the predicted coordinate information of the target object detection box, is the true coordinate corresponding to the target object detection box; λ coord is the bounding box prediction loss weight, and λ noobj is the confidence prediction loss weight of the prediction box that does not contain the target object; indicates whether the detection object appears in the i-th grid. If it is, it is 1; otherwise, it is 0; represents whether the j-th candidate box in the i-th grid is responsible for predicting the current object. If it is, it is 1; otherwise, it is 0; is the opposite of in meaning; represents the classification probability, that is, the probability that the target in the current prediction box belongs to a certain class; represents the confidence of the j-th candidate box in the i-th grid in the prediction result, is the true value; T loss represents the network denoising training information; during detection, the image is divided into S×S grids, and B candidate boxes are generated in each grid. c represents the target category, and class represents the total number of categories.
[0032] Furthermore, step S3 is specifically as follows:
[0033] Step S31: Decouple the two-dimensional feature into a one-dimensional row feature and a one-dimensional column feature through one-dimensional global average pooling;
[0034] Step S32: Perform row-level attention and column-level attention processing on the decoupled one-dimensional row features and one-dimensional column features respectively. The attention weight maps are specifically as follows:
[0035]
[0036]
[0037]
[0038] where softmax represents the normalized exponential function, e represents the base of the natural logarithm function, z i is the output value of the i-th node, C is the number of output nodes, that is, the number of classification categories; Att row and Att col represent the attention weight maps at the row and column levels respectively, N q represents the pre-set number of queries, Q row and Q col represent the Query features at the row level and column level respectively, and represent the transposes of the Key features at the row level and column level respectively, D K represents the dimension of the Key feature. Query, Key, and Value are the feature matrices in the self-attention mechanism;
[0039] Step S33: Perform weighted summation on the results after row attention and column attention processing according to the row and column dimensions. Specifically:
[0040]
[0041] where, and represent weighted summation of the features according to the column dimension and row dimension respectively, and V represents the Value feature.
[0042] Furthermore, Step S4 is specifically as follows:
[0043] Step S41: Use a convolution module, a pointwise convolution module, and a depth convolution module to construct three types of sub-feature fusion and cross-layer perception enhancement modules: SFCE-Block-A, SFCE-Block-B, and SFCE-Block-C;
[0044] where SFCE-Block-A includes a 3×3 convolution module and a 1×1 pointwise convolution module;
[0045] SFCE-Block-B sequentially includes a 1×1 convolution module, a 3×3 convolution module, a 3×3 depth convolution module, and a 1×1 pointwise convolution module;
[0046] The SFCE-Block-C sequentially includes a 3×3 depth convolution module, a 1×1 convolution module, a 3×3 convolution module, a 1×1 convolution module, a 3×3 convolution module, and a 3×3 depth convolution module;
[0047] It is used to combine channel information and spatial information to enhance the network's ability to capture details;
[0048] Step S42: Construct a lightweight feature extraction network using the sub-feature fusion and cross-layer perception enhancement module and the STEM module. The first layer of the feature extraction network uses the STEM module, followed by 1 SFCE-Block-A, 11 SFCE-Block-Bs, and 7 SFCE-Block-Cs:
[0049] Step S43: Construct a lightweight DETR object detection network using the lightweight feature extraction network and the decoder and encoder of the Transformer.
[0050] Further, step S5 is specifically as follows:
[0051] Step S51: Obtain the optimal values of the hyperparameters according to preset experiments. Tuning the hyperparameters will make the training model of the lightweight DETR object detection network reach the optimal state;
[0052] Step S52: Set the number of training iterations to N, set the data picture reading batch size batch-sizes, and obtain a safety helmet detection model after training;
[0053] Step S53: Use the safety helmet detection model to perform detection tasks. The input data picture is processed by the feature extraction network and the encoder and decoder to obtain a prediction result;
[0054] Step S54: Plot the obtained prediction result in the original image and output the final result image.
[0055] Compared with the prior art, the present invention and its preferred solutions have the following beneficial effects:
[0056] 1. The safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism constructed by the present invention is more lightweight than other existing DETR detection methods and can achieve precise detection of safety helmets;
[0057] 2. Regarding the problems of slow convergence speed and large video memory occupation of the DETR object detection model, the present invention proposes a training method based on noise cancellation and a dimensionality reduction attention mechanism, which can effectively improve the model convergence speed and reduce the video memory occupation during model training;
[0058] 3. For the problem of the huge number of parameters and computational complexity of the DETR target detection model in the present invention, a lightweight feature extraction network is constructed through sub-feature fusion and cross-layer perception enhancement modules, which can achieve almost the same detection effect as the original network with a smaller model size. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0060] Figure 1 It is a schematic diagram of the design route and principle of the embodiment of the present invention.
[0061] Figure 2 It is a schematic diagram of the specific structure of the sub-feature fusion and cross-layer perception enhancement module of the embodiment of the present invention.
[0062] Figure 3 It is a schematic diagram of the specific structure of the feature extraction network of the embodiment of the present invention.
[0063] Figure 4 It is a schematic diagram of the lightweight DETR target detection network of the embodiment of the present invention. SPECIFIC EMBODIMENTS
[0064] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0065] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0066] Please refer to Figure 1 , the present invention provides a safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism, including the following steps:
[0067] Step S1: Obtain safety helmet data pictures, preprocess them, and construct a safety helmet detection data set;
[0068] Step S2: Construct a noise cancellation training framework to replace the bipartite graph matching algorithm, add noise to the ground-truth information of each target object, introduce an attention mask, group the noise objects, and finally input the noise objects into the Transformer decoder;
[0069] Step S3: Improve the multi-head attention mechanism in the Transformer architecture, decouple the key features into one-dimensional row features and one-dimensional column features through one-dimensional global average pooling, then perform row attention and column attention in sequence, and finally perform weighted summation on the one-dimensional attention results of the two parts;
[0070] Step S4: Construct a lightweight feature extraction network using the sub-feature fusion and cross-layer perception enhanced convolutional module, and add it as the backbone network to the DETR network based on the Transformer architecture to obtain a lightweight DETR detection network;
[0071] Step S5: Optimize the training hyperparameters of the DETR object detection algorithm, train the DETR object detection network using the safety helmet detection dataset to obtain a safety helmet detection model, and use the safety helmet detection model to detect the input image.
[0072] In this embodiment, step S1 is specifically:
[0073] Step S11: Obtain data pictures related to safety helmets, screen the pictures and uniformly name them;
[0074] Step S12: Process the pictures using neighborhood denoising and median filtering;
[0075] Step S13: Determine the object categories in the safety helmet pictures, use the LabelImg annotation tool to annotate the preprocessed data pictures, and obtain and save the annotation information;
[0076] Step S14: Make a dataset according to the requirements of the DETR model, divide all the data into a training set, a validation set and a test set according to a ratio, and generate the training set, validation set and test set json annotation files required for training the model according to the xml file containing the annotation information of the picture data.
[0077] In this embodiment, step S2 is specifically:
[0078] Step S21: Add noise to the category information and coordinate box information of each detected object in the DETR detection network. First, add category noise, and randomly replace the true category of the target object with any other category with a certain probability; then add coordinate box noise, which can be divided into two categories, center point displacement and scale scaling. For the center point displacement, ensure that the center point is still inside the true box and perform random movement. For the scale scaling, randomly scale the length and width of the coordinate box; finally, add the loss of the noise elimination module to the loss function, and combine it with the coordinate box loss, category loss and other losses as the final loss value;
[0079] Among them, the formula of the loss function is as follows:
[0080]
[0081] Among them, x, y, w, h represent the predicted coordinate information of the target object detection box, is the true coordinate of the object detection box. λ coord is the bounding box prediction loss weight, λnoobj It is the confidence prediction loss weight for the prediction box that does not contain the target object. Indicates whether the detected object appears in the i-th grid. If so, it is 1; otherwise, it is 0. Represents whether the j-th candidate box in the i-th grid is responsible for predicting the current object. If so, it is 1; otherwise, it is 0. On the contrary. Represents the classification probability, that is, the probability that the target in the current prediction box belongs to a certain category. Represents the confidence of the j-th candidate box in the i-th grid in the prediction result. Is the ground truth value. T loss Represents the network denoising training information. During detection, the image is divided into S×S grids, and B candidate boxes are generated in each grid. c represents the target category, and class represents the total number of categories.
[0082] Step S22: Add an attention mask in the self-attention module of the DETR detection network to divide the noise objects into multiple denoising groups. The attention mask is divided into a denoising part and a matching part to ensure that there is no information interaction between the matching part and the denoising part, and between the denoising groups. The specific calculation method of the attention mask is as follows:
[0083]
[0084] Among them, P and M represent the number of denoising groups and ground-truth objects respectively, and N is the number of queries in the matching part. The first P×M rows and columns of the attention mask represent the denoising part, and the rest represent the matching part. a ij a = 1 means that the i-th query cannot see the j-th query, otherwise set a ij to 0;
[0085] Step S23: Input the noise objects into the Transformer decoder to predict the object category and coordinate box by learning the noise objects and ground-truth. Among them, ground-truth represents the true information of the object category and coordinate box.
[0086] In this embodiment, step S3 is specifically as follows:
[0087] Step S31: Decouple the two-dimensional feature into a one-dimensional row feature and a one-dimensional column feature through one-dimensional global average pooling;
[0088] Step S32: Perform row-level attention and column-level attention on the decoupled one-dimensional row feature and one-dimensional column feature respectively. The attention weight map is specifically as follows:
[0089]
[0090]
[0091]
[0092] Among them, softmax represents the normalized exponential function, e represents the base of the natural logarithm function, z i is the output value of the i-th node, C is the number of output nodes, that is, the number of classification categories. Att row and Att col represent the attention weight maps at the row and column levels, N q represents the pre-set number of queries, Q row and Q col respectively represent the Query features at the row level and the column level, and respectively represent the transposes of the Key features at the row level and the column level, D K represents the dimension of the Key feature, and Query, Key, and Value are the feature matrices in the self-attention mechanism;
[0093] Step S33: Weightedly sum the results after row attention and column attention processing according to the row and column dimensions, specifically:
[0094]
[0095] Among them, and respectively represent weighted summation of the features according to the column dimension and the row dimension, and V represents the Value feature.
[0096] In this embodiment, step S4 specifically includes the following steps:
[0097] Step S41: Use a convolutional module, a pointwise convolutional module, and a depth convolutional module to construct three types of sub-feature fusion and cross-layer perception enhancement modules: SFCE-Block-A, SFCE-Block-B, and SFCE-Block-C;
[0098] Among them, SFCE-Block-A includes a 3×3 convolutional module and a 1×1 pointwise convolutional module;
[0099] SFCE-Block-B sequentially includes a 1×1 convolutional module, a 3×3 convolutional module, a 3×3 depth convolutional module, and a 1×1 pointwise convolutional module;
[0100] SFCE-Block-C sequentially includes a 3×3 depth convolutional module, a 1×1 convolutional module, a 3×3 convolutional module, a 1×1 convolutional module, a 3×3 convolutional module, and a 3×3 depth convolutional module;
[0101] It is used to combine channel information and spatial information to enhance the network's ability to capture details. The specific structure of the sub-feature fusion and cross-layer perception enhancement module is as Figure 2 shown. This module has three different types of structures, corresponding to (a), (b), and (c) in the figure respectively. Conv represents the convolutional module, Pwise represents the pointwise convolutional module, and Dwise represents the depth convolutional module.
[0102] Step S42: Use the sub-feature fusion and cross-layer perception enhancement module and the STEM module to construct a lightweight feature extraction network. The first layer of the feature extraction network uses the STEM module, followed by 1 SFCE-Block-A, 11 SFCE-Block-B (a total of three groups: 2 + 3 + 6), and 7 SFCE-Block-C. The specific structure of the feature extraction network is as Figure 3 shown.
[0103] Step S43: Further use the lightweight feature extraction network and the decoder and encoder of the Transformer to construct a lightweight DETR object detection network. The specific structure is as Figure 4 shown. Among them, FFN represents the feed-forward neural network, and Class and Box represent the object category and coordinate information respectively.
[0104] In this embodiment, step S5 is specifically as follows:
[0105] Step S51: Obtain the optimal values of the hyperparameters according to the preset experiment. Tuning the hyperparameters will make the training model reach the optimal state;
[0106] Step S52: Set the number of training iterations to N, set the data picture reading batch size batch-sizes, and obtain the safety helmet detection model after training;
[0107] Step S53: Use the safety helmet detection model to perform the detection task. The input data picture is processed by the feature extraction network and the encoder and decoder to obtain the prediction result;
[0108] Step S54: Plot the obtained prediction result in the original image and output the final result image.
[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0111] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0113] This patent is not limited to the above best implementation manner. Anyone can derive various other forms of hard hat detection methods based on noise cancellation training and dimensionality reduction attention mechanisms under the inspiration of this patent. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. A safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism, characterized in that, it includes the following steps; Step S1: Obtain safety helmet data pictures, preprocess them, and construct a safety helmet detection data set; Step S2: Construct a noise cancellation training framework, add noise to the ground-truth information of each target object, introduce an attention mask, group the noise objects, and finally input the noise objects into the Transformer decoder; Step S3: Improve the multi-head attention mechanism in the Transformer architecture, decouple the key features into one-dimensional row features and one-dimensional column features through one-dimensional global average pooling, then perform row attention and column attention in sequence, and finally perform weighted summation on the one-dimensional attention results of the two parts; Step S4: Use a sub-feature fusion and cross-layer perception enhanced convolutional module to construct a lightweight feature extraction network, and add it as a backbone network to the DETR network based on the Transformer architecture to obtain a lightweight DETR detection network; Step S5: Tune the training hyperparameters of the DETR object detection algorithm, use the safety helmet detection data set to train the DETR object detection network to obtain a safety helmet detection model, and use the safety helmet detection model to detect the input image; Step S2 is specifically: Step S21: Add noise to the category information and coordinate box information of each detected object in the DETR detection network, and add the loss of the noise cancellation module to the loss function; Step S22: Add an attention mask to the self-attention module of the DETR detection network, divide the noise objects into multiple denoising groups, the attention mask is divided into a denoising part and a matching part, to ensure that there is no information interaction between the matching part and the denoising part, and between the denoising groups; the calculation method of the attention mask is specifically: Among them, P and M respectively represent the number of denoised groups and ground-truth objects, and N is the number of queries in the matching part; the first P×M rows and columns of the attention mask represent the denoised part, and the rest represent the matching part; a ij = 1 indicates that the i-th query cannot see the j-th query, otherwise set a ij to 0; Step S23: Input the noise objects into the Transformer decoder, and predict the object category and coordinate box by learning the noise objects and the ground-truth; where the ground-truth represents the true information of the object category and coordinate box; Step S21 is specifically: Step S211: Add category noise, and randomly replace the true category of the target object with any other category with a given probability; Step S212: Add coordinate box noise, which is divided into two categories: center point displacement and scale scaling; for the center point displacement, ensure that the center point is still inside the true box and perform random movement; for the scale scaling, randomly scale the length and width of the coordinate box; Step S213: Add the loss of the noise cancellation module to the loss function, combine it with the coordinate box loss, category loss, and other losses as the final loss value, and the formula of the loss function is as follows: Among them, x, y, w, h represent the predicted coordinate information of the target object detection box, are the true coordinates corresponding to the target object detection box; λ coord is the boundary box prediction loss weight, and λ noobj is the confidence prediction loss weight of the prediction box that does not contain the target object; indicates whether the detected object appears in the i-th grid. If it is, it is 1; otherwise, it is 0; represents whether the j-th candidate box in the i-th grid is responsible for predicting the current object. If it is, it is 1; otherwise, it is 0; is the opposite of in meaning; represents the classification probability, that is, the probability that the target in the current prediction box belongs to a certain category; represents the confidence in the prediction result of the j-th candidate box in the i-th grid, is the true value; T loss represents the network denoising training information; during detection, the image is divided into S×S grids, and B candidate boxes are generated in each grid. c represents the target category, and class represents the total number of categories; Step S3 is specifically: Step S31: Decouple the two-dimensional features into one-dimensional row features and one-dimensional column features through one-dimensional global average pooling; Step S32: Perform row-level attention and column-level attention processing on the one-dimensional row features and one-dimensional column features obtained after decoupling respectively, and the attention weight map is specifically: Among them, softmax represents the normalized exponential function, e represents the base of the natural logarithm function, and z i is the output value of the i-th node, C is the number of output nodes, that is, the number of classification categories; Att row and Att col represent the attention weight maps at the row and column levels respectively, N q represents the pre-set number of queries, Q row and Q col represent the Query features at the row and column levels respectively, and represent the transposes of the Key features at the row and column levels respectively, D K represents the dimension of the Key feature, and Query, Key, and Value are the feature matrices in the self-attention mechanism; Step S33: Perform weighted summation on the results processed by row attention and column attention in the row and column dimensions, specifically: Among them, and respectively represent weighted summation of features in the column dimension and row dimension, and V represents the Value feature; Step S4 is specifically as follows: Step S41: Use a convolution module, a pointwise convolution module, and a depth convolution module to construct three types of sub-feature fusion and cross-layer perception enhancement modules: SFCE-Block-A, SFCE-Block-B, and SFCE-Block-C; Among them, SFCE-Block-A includes a 3×3 convolution module and a 1×1 pointwise convolution module; SFCE-Block-B sequentially includes a 1×1 convolution module, a 3×3 convolution module, a 3×3 depth convolution module, and a 1×1 pointwise convolution module; SFCE-Block-C sequentially includes a 3×3 depth convolution module, a 1×1 convolution module, a 3×3 convolution module, a 1×1 convolution module, a 3×3 convolution module, and a 3×3 depth convolution module; It is used to combine channel information and spatial information to enhance the network's ability to capture details; Step S42: Use the sub-feature fusion and cross-layer perception enhancement module and the STEM module to construct a lightweight feature extraction network. The first layer of the feature extraction network uses the STEM module, followed by 1 SFCE-Block-A, 11 SFCE-Block-B, and 7 SFCE-Block-C; Step S43: Use the lightweight feature extraction network and the decoder and encoder of the Transformer to construct a lightweight DETR object detection network.
2. The safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism according to claim 1, characterized in that: Step S1 is specifically as follows: Step S11; Obtain data pictures related to safety helmets, screen the pictures and uniformly name them; Step S12: Process the pictures using neighborhood denoising and median filtering; Step S13: Determine the object categories in the safety helmet pictures, and use the LabelImg annotation tool to annotate the preprocessed data pictures to obtain and save the annotation information; Step S14: Make a dataset according to the requirements of the DETR model, divide all the data into a training set, a validation set, and a test set according to a ratio, and generate the training set, validation set, and test set json annotation files required for training the model according to the xml file containing the annotation information of the picture data.
3. The safety helmet detection method based on noise cancellation training and dimensionality reduction attention mechanism according to claim 1, characterized in that: Step S5 is specifically as follows: Step S51: Obtain the optimal values of the hyperparameters according to the preset experiment, and tuning the hyperparameters will make the training model of the lightweight DETR object detection network reach the optimal; Step S52: Set the number of training iterations to N, set the data picture reading batch size batch-sizes, and obtain the safety helmet detection model after training; Step S53: Use the safety helmet detection model to perform detection tasks, and the input data pictures are processed by the feature extraction network and the encoder and decoder to obtain the prediction results; Step S54: Draw the obtained prediction results in the original image and output the final result image.
Citation Information
Patent Citations
Image processing method, intelligent terminal and storage medium
CN114882226A
Light-weight safety helmet detection method based on sub-feature fusion
CN115100495A