Safety helmet wearing identification method and system based on multi-scale feature fusion

Through the combination of the panoramic feature pyramid network, the two-layer target enhancement mechanism and multi-scale detection head, the problems of insufficient fusion of multi-scale feature and background interference in the hard helmet wear recognition are solved, and high-precision and stable hard helmet wear status recognition are achieved.

CN120339592AActive Publication Date: 2025-07-18BEIJING ZHONGKE JINCAI TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510489708.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing safety helmet wear recognition methods have low recognition accuracy and poor robustness under complex backgrounds or different angles, insufficient fusion of multi-scale features, serious background interference, and inability to effectively focus the target area.

Method used

The panoramic feature pyramid network is used to fusion of multi-scale features, combining the two-layer target enhancement mechanism and multi-scale detection head, and the regression loss function is optimized to improve detection accuracy and stability through the weighted operation of the spatial and channel attention mechanism.

Benefits of technology

It significantly improves the accuracy and robustness of the wearable status of the safety helmet, and can accurately detect safety helmets of different sizes and angles in complex environments, reduces the rate of missed detection and false detection, and improves the applicability and safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339592A_ABST
    Figure CN120339592A_ABST
Patent Text Reader

Abstract

The invention discloses a safety helmet wearing identification method and system based on multi-scale feature fusion. The method comprises the steps that a target detection model fuses feature maps of different scales through a panoramic feature pyramid network; the target detection model adopts a double-layer target enhancement mechanism to perform weighting operation on the fused feature map so as to focus attention on a target area; multi-scale detection heads are introduced into the target detection model, and the multiple detection heads process targets of different sizes through anchor frames of different scales and feature maps; and constructing a total regression loss function for the target detection model, and adjusting parameters by using an optimization method in a model training process to minimize the total regression loss function so as to obtain a trained target detection model. The method has the advantages that the detection capability of a multi-scale target is effectively improved, the concentration on a target area is enhanced, and misrecognition and missed recognition under a complex background are reduced, so that higher accuracy and robustness are obtained in the recognition of the wearing state of the safety helmet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly relates to a safety helmet wearing recognition method and system based on multi-scale feature fusion. Background Art

[0002] In the field of safety helmet wearing recognition, single-scale or simple multi-scale methods are usually adopted for object detection, such as the traditional Feature Pyramid Network (FPN) and Path Aggregation Network (PANet). Although these methods can handle objects of different scales, they often perform poorly in object detection in complex backgrounds or at different angles, resulting in serious false detection and missed detection. The existing detection methods mainly have the following problems:

[0003] 1. Insufficient multi-scale feature fusion: The existing methods adopt simple feature fusion strategies (such as feature summation or direct splicing), which cannot fully explore the relationship between different scales, resulting in low recognition accuracy when dealing with objects with large size differences.

[0004] 2. Background interference problem: The traditional object detection methods have poor suppression effect on complex backgrounds, cannot effectively ignore the interference of irrelevant backgrounds, and are prone to introducing noise, reducing the recognition accuracy.

[0005] 3. Insufficient focus on the target area: In multi-object scenarios, the existing methods often cannot accurately focus on the target area, are easily interfered by other objects, and affect the recognition performance. Summary of the Invention

[0006] The purpose of the present invention is to provide a safety helmet wearing recognition method and system based on multi-scale feature fusion, so as to solve the foregoing problems existing in the prior art.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0008] A safety helmet wearing recognition method based on multi-scale feature fusion includes the following steps,

[0009] S1. Multi-scale feature fusion: The object detection model performs splicing and weighting operations on feature maps of different scales through a panoramic feature pyramid network to fuse the feature maps of different scales;

[0010] S2. Double-layer object enhancement: The object detection model adopts a double-layer object enhancement mechanism constructed based on the spatial attention mechanism and the channel attention mechanism to perform a weighting operation on the fused feature map to concentrate the attention on the target area;

[0011] S3. Introduction of multi-scale detection heads: Multi-scale detection heads are introduced into the object detection model. Multiple detection heads process objects of different sizes through anchor boxes and feature maps of different scales to optimize the accuracy of object detection;

[0012] S4. Optimization of regression loss function: The position regression loss function and the class regression loss function are used to construct the total regression loss function for the object detection model. During the training of the object detection model, the parameters are continuously adjusted using optimization methods to minimize the total regression loss function, thereby obtaining a trained object detection model.

[0013] Preferably, the panoramic feature pyramid network realizes the fusion of multi-scale information by splicing and weighted summation of feature maps of multiple scales. The calculation formula is

[0014] PFPN i = concat(W i ·F i , ∑(a j ·F j ))

[0015] where PFPN i is the result of multi-scale information fusion; F i and F j are the feature maps of the i-th layer and the j-th layer respectively; W i is the weighting matrix of the i-th layer feature map, used to perform a weighting operation on the i-th layer feature map; a j is the weighting coefficient of the j-th layer feature map; ∑(a i ·F j ) is the weighted summation of all scale feature maps F j ; concat is the operation of splicing multiple feature maps along the channel dimension;

[0016] Preferably, a loss function is constructed for the panoramic feature pyramid network, and the panoramic feature pyramid network is trained by backpropagation based on the loss function, thereby continuously adjusting the weighting matrix W i of the feature map and the weighting coefficient a j of the feature map.

[0017] Preferably, step S2 specifically includes the following contents

[0018] S21. The spatial attention mechanism weights each spatial position in the feature map according to the importance of different positions in the image. The calculation formula is

[0019] Attention spatial = σ(Conv(F))

[0020] where Attention spatialis the spatial weighting coefficient obtained through the spatial attention mechanism; F is the input feature map; Conv(F) is the spatial weighting coefficient extracted from the feature map through the convolution operation; σ is the activation function used to map the spatial weighting coefficient extracted from the feature map through the convolution operation to the interval [0, 1];

[0021] S22. The channel attention mechanism weights each channel of the feature map to emphasize the importance of different channels in a specific task; the calculation formula is

[0022] Attention channel =σ(FC(f))

[0023] Among them, Attention channel is the channel weighting coefficient obtained through the channel attention mechanism; FC(F) is the weighting coefficient of each channel extracted from the feature map through the fully connected layer; σ is the activation function used to map the weighting coefficient of each channel extracted from the feature map through the fully connected layer to the interval [0, 1];

[0024] S23. After the spatial attention and the channel attention respectively weight the feature map, the weighted results of the two are fused to enhance the network's ability to focus on the target area; the calculation formula is

[0025] Attention fused =Attention spatial ·Attention channel ·F

[0026] Among them, Attention fused is the final fusion result.

[0027] Preferably, a loss function is constructed for the double-layer target enhancement mechanism, and the double-layer target enhancement mechanism is trained by backpropagation based on the loss function, so as to continuously adjust the spatial weighting coefficient Attention spatial and the channel weighting coefficient Attention channel .

[0028] Preferably, the multi-scale detection head optimizes the accuracy of the target by adding anchor boxes of different scales in the network and processing them with different feature maps for each scale; the calculation formula is

[0029] Detection output =∑(Anchor t ·Feature_map t )

[0030] Among them, Detection outputThe processing result of the multi-scale detection head; Anchor i is the anchor box at the t-th scale; Feature_map t is the feature map at the t-th scale.

[0031] Preferably, the total regression loss function is constructed by weighting the position regression loss function and the class regression loss function, and the calculation formula is

[0032] L total = λ1·L bbox + λ2·L class

[0033] L bbox = ∑(1 - IOU(predict, ground truth ))

[0034]

[0035] L class = -∑(y·log(p) + (1 - y)·log(1 - p))

[0036] where L total is the total regression loss function; L bbox and L class are the position regression loss function and the class regression loss function respectively; λ1 and λ2 are the weight coefficients of the position regression loss function and the class regression loss function respectively; IOU is the intersection over union used to measure the overlap degree between the predicted box and the ground truth box; predict is the predicted box; ground truth is the ground truth box; Area_of Intersection is the area of the intersection of the predicted box and the ground truth box; Area_of Union is the area of the union of the predicted box and the ground truth box; y is the true class label, which is 1 when the target belongs to a certain class and 0 when the target does not belong to that class; p is the predicted probability of the object detection model for the target belonging to a certain class.

[0037] Preferably, after step S4, it further includes

[0038] S5. Safety helmet wearing recognition: Use the trained object detection model to recognize the safety helmet wearing situation and output the safety helmet wearing recognition result.

[0039] Preferably, before step S1, it further includes

[0040] S0. Dataset preparation: Manually annotate the safety helmet wearing status in the collected images to construct a safety helmet wearing status dataset.

[0041] The object of the present invention also lies in providing a safety helmet wearing recognition system based on multi-scale feature fusion, which can implement the method described above. The recognition system includes:

[0042] Multi-scale feature fusion module: The object detection model performs splicing and weighting operations on feature maps of different scales through a panoramic feature pyramid network to fuse feature maps of different scales.

[0043] Double-layer object enhancement module: The object detection model uses a double-layer object enhancement mechanism constructed based on spatial attention mechanism and channel attention mechanism to perform weighting operations on the fused feature maps, so as to focus attention on the target area.

[0044] Multi-scale detection head introduction module: A multi-scale detection head is introduced into the object detection model, and multiple detection heads process targets of different sizes through anchor boxes and feature maps of different scales to optimize the accuracy of object detection.

[0045] Regression loss function optimization module: A total regression loss function is constructed for the object detection model by using a position regression loss function and a class regression loss function. During the training process of the object detection model, parameters are continuously adjusted by using an optimization method to minimize the total regression loss function, so as to obtain a trained object detection model.

[0046] The beneficial effects of the present invention are as follows: 1. By introducing the optimization of the panoramic feature pyramid network (PFPN), double-layer object enhancement mechanism (BGA) and multi-scale detection head, the present invention solves the problems of low recognition accuracy and poor robustness of the safety helmet wearing state in complex backgrounds and different angles in the prior art, thereby improving the accuracy and stability of object detection. 2. Through the panoramic feature pyramid network of PFPN, the present invention can better fuse features of different scales, reduce information loss, and significantly improve the recognition ability for targets of different sizes, thereby improving the accuracy of the safety helmet wearing state, especially for targets with large size changes. 3. The double-layer object enhancement mechanism (BGA) combines spatial and channel attention, effectively improving the focusing ability of the model on the target area, reducing the interference of background noise, and significantly improving the detection accuracy in complex environments. 4. By introducing a multi-scale detection head, the model can better adapt to targets of different scales, optimize the anchor box scale, further improve the detection accuracy for safety helmets of different sizes, and reduce the incidence of missed detection and false detection. 5. In a dynamically changing working environment, the present invention can stably handle a variety of complex scenarios, including different shooting angles, background changes, and lighting conditions, thereby improving the robustness and applicability of the overall recognition system. 6. The present invention can greatly improve the recognition accuracy and efficiency of safety helmet wearing in fields such as construction sites and industrial production, reduce the need for human intervention, and further ensure the safety of workers through high-precision recognition. Description of the Drawings

[0047] Figure 1 It is a flowchart of the recognition method in the embodiments of the present invention. Detailed implementation manners

[0048] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] In this embodiment, a safety helmet wearing recognition method based on multi-scale feature fusion is provided, aiming to improve the recognition accuracy and robustness of the safety helmet wearing state in complex backgrounds and at different angles. By introducing an improved object detection framework and combining the panoramic feature pyramid network (PFPN), the double-layer object enhancement mechanism (BGA), and the optimization of the multi-scale detection head, the object recognition ability is significantly improved. First, through the panoramic feature pyramid network (PFPN), splicing and weighting operations are performed on multi-level feature maps to effectively fuse features of different scales and enhance the perception ability of multi-scale objects. Then, using the double-layer object enhancement mechanism (BGA), spatial and channel attention mechanisms are introduced to effectively focus on the target area and ignore background noise, further improving the model's focus and recognition accuracy for the target. Next, the multi-scale detection head is introduced to optimize the network's detection ability for safety helmets of different sizes. Finally, the optimized regression loss function combines the position loss and the class loss to further improve the target positioning accuracy and classification accuracy. The comprehensive technical solution of the present invention shows a significant improvement in the safety helmet wearing recognition task, especially in object detection in complex environments and at different angles, and has high practicality and promotion value. As Figure 1 shown, the method of the present invention mainly includes the following parts

[0050] I. Dataset preparation

[0051] Manually annotate the safety helmet wearing state in the collected images to construct a safety helmet wearing state dataset.

[0052] II. Multi-scale feature fusion

[0053] The object detection model performs splicing and weighting operations on feature maps of different scales through the panoramic feature pyramid network to fuse feature maps of different scales.

[0054] The Panoramic Feature Pyramid Network (PFPN) is a method to enhance multi-scale feature fusion. It improves the recognition ability of deep learning models for multi-scale targets through inter-layer splicing operations and dense feature connections. This network is particularly suitable for object detection tasks, especially when dealing with targets of different scales and complex backgrounds, and can effectively improve the performance of the model. PFPN mainly enhances the network's perception ability for targets of different sizes by weighted summation and splicing of feature maps at different scales.

[0055] 2.1. Feature Maps and Scales

[0056] The core idea of PFPN is to enhance the detection ability for targets of different sizes through the fusion of multi-scale feature maps. The scale differences of targets in images are relatively large, which requires the model to be able to simultaneously focus on the features of small and large targets. In traditional Convolutional Neural Networks (CNNs), as the number of network layers increases, the spatial resolution of the feature maps gradually decreases, resulting in a weakened detection ability for small targets. PFPN can effectively solve this problem by fusing feature maps of multiple scales. Suppose there are n feature maps of different scales, and the i-th layer feature map Fi corresponds to a feature map of one scale. For each layer of feature map, a weighting coefficient Wi is introduced to weight the feature map to ensure that feature maps of different layers can be adjusted according to their importance during fusion.

[0057] 2.2. Feature Map Fusion and Weighting

[0058] The core operations of PFPN are the concatenation (concat) and weighted summation (Σ) of feature maps. Specifically, PFPN realizes the fusion of multi-scale information by concatenating and weighted summing feature maps of multiple scales. The formula can be expressed as:

[0059] PFPN i = concat(W i ·F i , ∑(a j ·F j ))

[0060] where PFPN i is the result of multi-scale information fusion; F i and F j are the feature maps of the i-th layer and the j-th layer respectively, usually high-dimensional feature matrices generated by convolutional layers; the dimension of F i is (Wi, Hi, Ci), where Wi and Hi represent the width and height of the feature map respectively, and Ci represents the number of channels of the feature map; F jThe dimension of each layer of feature maps is also (Wj, Hj, Cj), but feature maps of different scales may have different spatial resolutions and numbers of channels. W i is the weighting matrix of the i-th layer of feature maps, which is used to perform a weighting operation on the feature maps of the i-th layer. W i has a dimension of (Ci, Ci'), that is, it converts the number of channels of the feature maps to ensure that the weighted feature maps have the same number of channels when concatenated with feature maps of other scales. a j is the weighting coefficient of the j-th layer of feature maps, indicating the importance of this layer of feature maps in feature fusion. a j The value of is usually adjusted according to the depth of the layer, the quality of the features, and the contribution to object recognition. The value of aj can be a scalar or a vector, which is used to adjust the feature maps of each layer. ∑(a i ·F j ) is to perform a weighted sum on all-scale feature maps F j . By summing up the weighted feature maps of each layer, useful information from different scales can be aggregated, thus avoiding information loss caused by scale differences. concat is an operation to concatenate multiple feature maps along the channel dimension. The concatenated feature maps contain feature information from different scales, and this information is used for subsequent object detection tasks.

[0061] Through this operation of concatenation and weighted summation, PFPN can fuse multi-scale information, significantly enhancing the network's perception ability for objects of different sizes. Especially when facing objects of different sizes and complex backgrounds, the model can more accurately perform object localization and classification.

[0062] 2.3. Optimization and Training

[0063] To better learn the feature map weighting coefficients W i and a j , the PFPN network usually adopts an end-to-end training method. The loss function of the network combines common loss terms in object detection, such as classification loss and regression loss. The classification loss is used to measure the prediction accuracy of the model for object categories, and the regression loss is used to measure the prediction accuracy of the model for object positions (such as bounding boxes).

[0064] For example, the loss function L can be expressed as:

[0065] L = L classification + L regression

[0066] where L classification and L regression are the classification error for calculating object categories and the regression error for calculating object positions (bounding boxes), respectively.

[0067] Through backpropagation, the model continuously adjusts the feature weighting coefficient W i and a j , enabling the network to more effectively fuse multi-scale features, thereby improving the overall detection performance.

[0068] 2.4. Practical Applications

[0069] The multi-scale feature fusion ability of PFPN performs excellently in many practical applications, especially in object detection tasks. For example, in the fields of industrial safety monitoring, intelligent transportation, medical image analysis, etc., PFPN can effectively identify targets of different sizes, such as traffic signs, human bodies, vehicles, etc. In the scenario of safety helmet wearing detection, PFPN can help the model identify safety helmets at different angles and scales. Whether it is a small target at close range or a large target at a long distance, PFPN can effectively process and accurately identify them.

[0070] The Panoramic Feature Pyramid Network (PFPN) plays a key role in multi-scale feature fusion in the method of the present invention. By splicing and weighting feature maps at different scales, PFPN can effectively fuse feature information from different levels, improving the network's ability to process targets of different sizes and different backgrounds. In safety helmet wearing recognition, different shooting angles and image sizes may lead to scale differences in the target, and PFPN precisely enhances the network's robustness to these scale differences through dense feature connections, ensuring that the model can accurately detect the safety helmet wearing status in various scenarios. PFPN provides high-quality feature inputs for the subsequent Bilateral Goal Attention Mechanism (BGA) and multi-scale detection heads. The role of PFPN is to enhance the feature expression ability, while BGA and multi-scale detection heads use these enhanced features to further optimize the detection accuracy and regression ability. PFPN provides the fusion of multi-scale information for the entire model and is the basis for the model to work better in complex backgrounds.

[0071] III. Bilateral Goal Attention

[0072] The object detection model uses a bilateral goal attention mechanism constructed based on the spatial attention mechanism and the channel attention mechanism to perform a weighting operation on the fused feature map to focus the attention on the target area.

[0073] The Bilateral Goal Attention Mechanism (BGA) is a weighting operation method that combines the spatial and channel attention mechanisms, aiming to enhance the network's ability to focus on important target areas. By introducing two levels of weighting strategies, spatial attention and channel attention, this mechanism helps the network to more effectively focus on the target area while suppressing the interference of irrelevant backgrounds when processing complex images, thereby improving the target recognition accuracy of the model.

[0074] 3.1. Spatial Attention Mechanism

[0075] The core idea of the spatial attention mechanism is to weight each spatial position in the feature map according to the importance of different positions in the image. Spatial attention aims to enhance the network's attention to the target region and reduce the interference of background information on the model's judgment. It is expressed by the formula:

[0076] Attention spatial = σ(Conv(F))

[0077] Among them, Attention spatial is the spatial weighting coefficient obtained through the spatial attention mechanism, with the dimension of (W, H, 1). F is the input feature map, with the dimension of (W, H, C), where W and H are the width and height of the feature map respectively, and C is the number of channels. Conv(F) is the spatial weighting coefficient extracted from the feature map F through the convolution operation. The convolution operation captures the distribution pattern of spatial information in the image by learning filters (kernels). The output after convolution is a spatial weighting coefficient map, with the dimension of (W, H, 1), where each element represents the importance of that spatial position. σ is the activation function, which is used to map the convolution output to the interval [0, 1], representing the attention weight of each spatial position. In this way, higher values will give higher attention weights to important regions, and lower values will suppress the influence of background regions.

[0078] The spatial attention mechanism weights each position in the spatial feature map, enabling the network to focus on the target region in the image and improving the recognition ability of the target position.

[0079] 3.2. Channel Attention Mechanism

[0080] The goal of the channel attention mechanism is to weight each channel of the feature map, thereby emphasizing the importance of different channels in a specific task. Through channel weighting, the model can selectively focus on the features that are most valuable for target recognition according to the information contribution of different channels. It is expressed by the formula:

[0081] Attention channel = σ(FC(F))

[0082] Among them, Attention channelis the channel weighting coefficient obtained through the channel attention mechanism, with a dimension of (1, 1, C). FC(F) processes the input feature map F through a fully connected layer to obtain the weighting coefficient for each channel. The fully connected layer usually first performs global average pooling on the input feature map to obtain the global features of each channel, and then calculates the attention coefficient for each channel through the fully connected layer. σ is the activation function used to map the channel weighting coefficient to the interval [0, 1]. For each channel, a larger value indicates a higher importance of the channel, and a smaller value indicates a smaller contribution of the channel.

[0083] The channel attention mechanism weights the features of different channels, enabling the network to enhance the attention to the key channel features according to the importance of the channels and suppress the interference of irrelevant channel information.

[0084] 3.3. Double-layer target enhancement mechanism

[0085] In the BGA mechanism, after the spatial and channel attention respectively weight the feature map, the weighted results of the two are finally fused to further enhance the network's focusing ability on the target area. Specifically, the fused feature map is obtained by element-wise multiplying the weighting coefficients of the spatial attention and the channel attention with the original feature map. The fused formula is expressed as:

[0086] Attention fused = attention spatial · Attention channel · F

[0087] where, Attention fused is the final fusion result.

[0088] Through this element-wise multiplication operation, the weighted information in the space and channels is integrated into the feature map, so that the features of each spatial position and each channel are weighted accordingly, thereby improving the model's recognition ability for the target area.

[0089] 3.4. Optimization and implementation of the double-layer target enhancement mechanism

[0090] In practical implementation, the BGA mechanism is often used in combination with a Convolutional Neural Network (CNN) and embedded into the network as a module for enhancing features. During specific implementation, spatial attention and channel attention can be learned through independent convolutional layers and fully connected layers, and the fused feature maps can be used for subsequent object detection tasks. The training process of BGA generally adopts an end-to-end training method and is optimized using various loss functions such as Cross-Entropy Loss and Regression Loss. Through backpropagation, the model can adaptively adjust the weighting coefficients of space and channels, continuously improving the recognition ability of the target area.

[0091] 3.5. Practical Applications

[0092] BGA has demonstrated excellent performance in many practical applications, especially in object detection tasks under complex backgrounds. For example, in tasks such as safety helmet wearing recognition, traffic sign recognition, and human pose estimation, BGA can effectively focus on the target area, suppress background interference, and improve the accuracy of object detection. In specific applications, BGA can accurately identify targets under dynamic lighting, different angles, and complex backgrounds. Even when the target is small or partially occluded in the image, it can still be effectively detected. In a safety helmet wearing detection system, BGA can help the network focus on the human head area through the spatial attention mechanism, while the channel attention can strengthen the key features related to the safety helmet wearing state, thus significantly improving the detection accuracy and robustness.

[0093] By introducing the Bilayer Goal Attention (BGA) mechanism, the model can focus more on the target area and suppress the interference of irrelevant information, thereby improving the accuracy and robustness of target recognition, especially suitable for object detection tasks in complex scenarios.

[0094] The Bilayer Goal Attention (BGA) mechanism combines spatial and channel attention mechanisms. It highlights the target area through weighted operations on the feature maps, ignoring the background and irrelevant parts. Through the spatial attention mechanism, the model can concentrate more attention on the location where the target is located; through the channel attention mechanism, the model can further optimize the feature responses of each channel, enhancing the model's focus ability on key areas (such as safety helmets). BGA acts directly on the feature maps output by PFPN and further strengthens the feature expression of the safety helmet wearing area through weighted operations. Its role is to improve the network's focus on the target area and ensure that the model can accurately identify targets under complex backgrounds. Through this kind of weighting, BGA improves the effect of the Panoptic Feature Pyramid Network, so it forms an effective synergistic effect with PFPN.

[0095] IV. Introduction of Multi-Scale Detection Heads

[0096] The Multi-scale Detection Head is a method that optimizes the detection ability of a neural network for objects of different sizes by introducing anchor boxes of multiple scales. In object detection tasks, especially when detecting objects of different sizes, a single-scale detection head often cannot effectively handle the large range of target size differences. By introducing multiple detection heads in the network, with each detection head responsible for processing targets of different scales, the detection performance of the network in complex scenarios can be significantly improved, especially the detection accuracy of small or large objects in complex backgrounds.

[0097] 4.1. Working Principle of the Multi-scale Detection Head

[0098] The multi-scale detection head optimizes the accuracy of object detection by adding anchor boxes of different scales in the network and processing them using different feature maps for each scale. An anchor box is a rectangular box used to represent objects of different sizes, and a feature map is the feature information at different levels extracted from the original image by a convolutional neural network (CNN). In traditional object detection methods, usually one detection head is used to process objects of different sizes. However, due to the large differences in target sizes, it is often difficult for a single-scale anchor box to accurately match all targets. The introduction of the multi-scale detection head solves this problem. Each detection head specifically processes targets of different sizes through anchor boxes and feature maps of different scales. The formula is as follows:

[0099] Detection output =∑(Anchor t ·Feature_map t )

[0100] Where Detection output is the processing result of the multi-scale detection head; Anchor i is the anchor box of the t-th scale. An anchor box is a predefined rectangular box used to match possible target objects in the image. The size and aspect ratio of each anchor box are usually adjusted according to the actual size and distribution of the target object. For different scales, the size and position of the anchor box will also be different. Feature_map t is the feature map of the t-th scale. The feature map of each scale is extracted through a convolutional network and contains feature information at different levels. Deep feature maps usually contain high-level information of larger targets, while shallow feature maps focus more on low-level information of smaller targets. Σ represents the weighted sum of anchor boxes and feature maps of all scales. This operation enables the network to make a comprehensive judgment based on the outputs of anchor boxes and feature maps of different scales, thereby improving the accuracy of object detection.

[0101] 4.2 Advantages of the multi-scale detection head

[0102] Improve the detection accuracy of small objects: Since the detection head for each scale focuses on targets within a specific size range, it can avoid the accuracy problem of traditional single-scale methods for small object detection. Especially in complex backgrounds, small objects are often easily masked by background noise. Using a multi-scale detection head can perform reinforcement learning through features of multiple scales to improve the recognition accuracy of small objects.

[0103] Enhance the detection ability of large objects: By introducing larger-scale anchor boxes for large objects, the network can better adapt to the localization and regression tasks of large objects. Anchor boxes of different scales can help the network capture more detailed features of large objects and improve the accuracy when detecting large objects.

[0104] Enhance the robustness of the model: The multi-scale detection head synthesizes feature maps of different scales through weighted combination, making the model more robust in the face of various scenarios. Whether it is large objects, small objects, or targets at different distances and angles, the network can accurately identify them based on features of different scales.

[0105] 4.3 Implementation of the multi-scale detection head

[0106] In actual implementation, each multi-scale detection head is connected to certain convolutional layers in the network. These convolutional layers are responsible for extracting the feature maps of the input image and assigning appropriate weights to the anchor boxes of each scale. In object detection frameworks such as YOLO, the deep feature maps of the network are often used to detect large objects, while the shallow feature maps are suitable for detecting small objects. The size and aspect ratio of the anchor boxes for each scale are optimized based on prior experience and the target sizes in the training set. During the training process of the network, the loss function for each scale is calculated separately and weighted according to the detection tasks of each scale. Finally, the loss functions of all scales are aggregated for the optimization of the overall network. This enables the network to be independently optimized for each scale and improve the recognition accuracy of targets of different sizes. Specifically, in the training stage, an image dataset labeled with target bounding boxes is first required. In each image, the network selects appropriate anchor boxes based on the overlap degree between the anchor boxes and the targets (usually measured by IoU - Intersection over Union). If the IoU between the anchor box and the target exceeds a preset threshold, it is considered that the anchor box has successfully detected the target. The detection head for each scale is responsible for processing targets within the corresponding scale range.

[0107] 4.4 Practical applications

[0108] In practical applications, the introduction of a multi-scale detection head is particularly suitable for scenarios where the target sizes vary significantly. For example, in helmet-wearing detection, the network needs to detect helmets in images taken from different distances and angles. Due to the large variations in the size and angle of helmets, a single-scale detection head is difficult to handle the diverse target sizes. After introducing a multi-scale detection head, the network can process small-sized helmets (e.g., helmets worn by people in the distance) and large-sized helmets (e.g., helmets captured in close-up) separately, thus greatly improving the detection accuracy and robustness. In addition, traffic sign recognition is also a typical application scenario, where traffic signs of different sizes need to be efficiently recognized through a multi-scale detection head. In this application, the system can automatically adjust the anchor boxes of different scales to quickly detect traffic signs from far to near, ensuring the accuracy and real-time nature of the detection process.

[0109] By introducing a multi-scale detection head, the network can perform object detection at different scales, thereby significantly improving the detection ability for multi-sized objects. Especially in complex scenarios, anchor boxes of multiple scales can effectively compensate for the deficiencies of traditional single-scale detection methods, improving the detection accuracy, especially for small objects or large objects in complex backgrounds. This method is widely applicable to object detection tasks such as helmet-wearing detection and traffic sign recognition, and can improve the robustness and practicality of the system.

[0110] The multi-scale detection head optimizes the network's recognition ability for various-sized objects by performing detections at different scales. In this invention, in view of the size differences of helmets, a fourth detection head is added to optimize the model's detection of objects at different scales. This detection head not only enhances the network's ability to process objects of different sizes, but also further improves the detection accuracy of the helmet-wearing state through the fusion of multi-scale features. The multi-scale detection head complements the PFPN and BGA mechanisms. PFPN provides rich multi-scale features, BGA further strengthens the expression of these features, and the multi-scale detection head optimizes the detection accuracy and regression ability of the model by performing object detection at multiple scales. Through this design, the model can more accurately identify the helmet-wearing state when dealing with multi-scale objects.

[0111] V. Optimization of the regression loss function

[0112] In the object detection task, the regression loss function is an important tool for measuring the difference between the positions and classes predicted by the model and the true labels. To improve the accurate regression of the object detection model for the object position and class, it is necessary to optimize the regression loss function to enhance the detection accuracy. In the problem of safety helmet wearing detection, accurately locating the position of the object (safety helmet) and judging the wearing state of the object is very crucial. Therefore, it is necessary to comprehensively consider the position regression loss and the class regression loss in the regression loss function.

[0113] 5.1. Total Regression Loss Function

[0114] To enable the model to better regress the position and class of the object, the position regression loss (L bbox ) and the class regression loss (L class ) are combined into a total regression loss function L total . The expression of this loss function is:

[0115] L total = λ1·L bbox + λ2·L class

[0116] where L total is the total loss function, which includes the position loss and the class loss. By optimizing L total , the network can simultaneously learn how to accurately locate the object position and judge the object class (such as the wearing state of the safety helmet).

[0117] λ1 and λ2 are the weight coefficients of the loss function, used to balance the influence of the position loss and the class loss. Usually, the position loss and the class loss will have different influences in actual training. Therefore, it is necessary to adjust λ1 and λ2 to ensure a reasonable ratio of the two in the total loss and avoid one loss being too large and affecting the overall optimization.

[0118] L bbox is the position regression loss, used to measure the gap between the bounding box predicted by the model and the true bounding box. Usually, the intersection over union (IoU) is used to measure the overlapping degree of the predicted box and the true box.

[0119] L class is the class regression loss, used to measure the accuracy of the model's prediction of the object class (such as the wearing state of the safety helmet), usually calculated using the cross-entropy loss.

[0120] 5.2. Position Regression Loss

[0121] The goal of the location regression loss is to accurately regress the bounding box of the target object as precisely as possible, that is, the overlap degree between the predicted box and the ground truth box. To calculate the location loss, the Intersection over Union (IoU) is usually used to measure the overlap degree between the predicted box and the ground truth box. The definition of IoU is as follows:

[0122]

[0123] where, IOU(predict, ground truth ) is the Intersection over Union used to measure the overlap degree between the predicted box and the ground truth box. predict is the predicted box, that is, the bounding box predicted by the model. ground truth is the ground truth box, that is, the actual annotated bounding box. Area_of Intersection is the area of the intersection between the predicted box and the ground truth box. Area_of Union is the area of the union between the predicted box and the ground truth box.

[0124] The location regression loss L bbox measures the gap between the predicted box and the ground truth box by calculating the difference in IoU. The formula is

[0125] L bbox = ∑(1 - IOU(predict, ground truth ))

[0126] where, the optimization goal of L bbox is to reduce the gap between the predicted box and the ground truth box, making the IoU value as close to 1 as possible. This loss function represents the optimization goal of location regression: improving the localization accuracy by reducing the IoU difference between the predicted box and the ground truth box. The model will minimize this loss function through optimization methods such as gradient descent to improve the location prediction of the target.

[0127] 5.3. Classification Regression Loss

[0128] The classification regression loss is used to measure the accuracy of the model's prediction of the target class (such as "wearing a safety helmet" or "not wearing a safety helmet"). In the classification task, the most commonly used loss function is the Cross-Entropy Loss, and its formula is:

[0129] L class = -∑(y·log(p) + (1 - y)·log(1 - p))

[0130] where, L total is the total regression loss function; L bbox and L classThey are the location regression loss function and the class regression loss function respectively; λ1 and λ2 are the weight coefficients of the location regression loss function and the class regression loss function respectively;;;;

[0131] Among them, y is the true class label, usually a binary value (0 or 1), where 1 indicates that the target belongs to a certain class (such as wearing a safety helmet), and 0 indicates that it does not belong to this class (not wearing a safety helmet). p is the predicted probability that the target detection model predicts that the target belongs to a certain class, usually the probability value output by the softmax or sigmoid function. L class is the class regression loss function, that is, the cross-entropy loss, which is used to measure the gap between the predicted probability and the true label.

[0132] The goal of the cross-entropy loss is to maximize the consistency between the predicted probability and the true label. When the true label y is 1, L class tends to minimize log(p), and when y is 0, L class tends to minimize log(1 - p). By optimizing L class , the network can better learn the features of the target class and improve the classification accuracy.

[0133] 5.4. Loss Function Optimization

[0134] By simultaneously optimizing the location regression loss (L bbox ) and the class regression loss (L class ), the model can not only accurately locate the position of the target, but also effectively judge the class of the target. The total loss function L total balances the importance of the location loss and the class loss in the overall optimization by adjusting the weight coefficients of λ1 and λ2. During the training process, the network will continuously adjust the parameters to minimize L total , thereby improving the accuracy of safety helmet wearing detection. In practical applications, the optimization of the regression loss function is crucial. For example, in security monitoring, accurately detecting personnel wearing safety helmets is crucial for ensuring construction site safety. By optimizing the regression loss function, the model can efficiently handle various complex factors such as illumination and angle changes, improving the reliability and accuracy of the safety helmet wearing recognition system. In addition, the optimized loss function can also effectively reduce the occurrence of false detections and missed detections, improving the overall performance of the system. Through such optimization of the regression loss function, the model can not only accurately locate the bounding box of the target, but also accurately classify whether the target is wearing a safety helmet, making the system more adaptable and robust in practical applications.

[0135] The optimization of the regression loss function mainly adjusts the weights of the location regression loss and the class regression loss in the loss function to accurately regress the location and class of the target. This part is crucial for the precise positioning and classification of the helmet wearing status. By optimizing the loss function, the model can better perform object positioning and classification in complex scenarios, reducing the occurrence of false detections and missed detections. The optimization of the regression loss function is the final optimization step, which depends on the previous feature extraction and object enhancement mechanisms (PFPN, BGA, and multi-scale detection heads). By optimizing the regression loss function, the model can more accurately predict the location and wearing status of the safety helmet, thereby improving the overall detection effect. The optimization of the regression loss function is closely related to the previous steps, and finally improves the detection performance and accuracy of the network through this step.

[0136] VI. Helmet Wearing Recognition

[0137] Use the trained object detection model to recognize the helmet wearing situation and output the helmet wearing recognition result.

[0138] In this embodiment, a helmet wearing recognition system based on multi-scale feature fusion is also provided. The recognition system can implement the above method. The recognition system includes:

[0139] (1) Multi-scale feature fusion module: The object detection model performs splicing and weighting operations on feature maps of different scales through the panoramic feature pyramid network to fuse feature maps of different scales;

[0140] (2) Double-layer object enhancement module: The object detection model uses a double-layer object enhancement mechanism constructed based on the spatial attention mechanism and the channel attention mechanism to perform a weighting operation on the fused feature map to focus the attention on the target area;

[0141] (3) Multi-scale detection head introduction module: Introduce a multi-scale detection head into the object detection model. Multiple detection heads process objects of different sizes through anchor boxes and feature maps of different scales to optimize the accuracy of object detection;

[0142] (4) Regression loss function optimization module: Use the location regression loss function and the class regression loss function to construct a total regression loss function for the object detection model, and continuously adjust the parameters using an optimization method during the training of the object detection model to minimize the total regression loss function, thereby obtaining a trained object detection model.

[0143] In this embodiment, PFPN solves the problem of recognizing targets of different scales through multi-scale feature fusion. Compared with using the features of a single level alone, PFPN can significantly improve the network's processing ability for targets of different sizes. BGA enables the network to focus on key regions through spatial and channel weighting operations, reduces background interference, and improves the recognition accuracy of the model for targets. It can not only enhance the attention to targets but also avoid excessive attention to irrelevant backgrounds, which greatly improves the recognition effect of the helmet-wearing status in complex backgrounds. Introducing a multi-scale detection head further improves the network's adaptability to targets of different sizes. It enables the network to not only process helmets of different sizes but also optimize the scales of anchor boxes, making the detection more accurate. By optimizing the regression loss function, the target can be accurately located and its wearing status can be judged, enabling the model to have higher accuracy in target localization and classification tasks.

[0144] After the combination of the above four core steps, they interact with each other to improve the robustness and accuracy of the entire model. The final recognition effect far exceeds the sum of single features. For example, PFPN provides rich multi-scale features, while BGA and the multi-scale detection head further optimize the feature expression and multi-scale adaptation ability. The optimization of the regression loss function ensures the accuracy of the model during the detection process.

[0145] The method of the present invention can effectively adapt to scenarios of different sizes, different angles, and complex backgrounds. The combination of PFPN and BGA enables the model to improve the recognition ability of the helmet-wearing status in a changing environment. Through multi-scale feature fusion, target enhancement, and regression loss optimization, the model can still maintain a high-precision detection effect in a complex environment, and reduce background interference and positioning errors at different angles. By optimizing the regression loss function, the model can not only accurately locate the bounding box of the target but also correctly judge whether the target is wearing a helmet, improving the accuracy and reliability in practical applications.

[0146] In summary, the various technical steps of the entire method of the present invention support and complement each other, greatly improving the performance of the model in the helmet-wearing recognition task, and demonstrating excellent technical innovation and application prospects.

[0147] By adopting the above technical solutions disclosed in the present invention, the following beneficial effects are obtained:

[0148] The present invention provides a safety helmet wearing recognition method and system based on multi-scale feature fusion. By introducing the Panoramic Feature Pyramid Network (PFPN), the Double-layer Goal Enhancement Mechanism (BGA), and the optimization of the multi-scale detection head, the present invention solves the problems of low recognition accuracy and poor robustness of the safety helmet wearing state in complex backgrounds and at different angles in the prior art, thereby improving the accuracy and stability of object detection. 2. Through the panoramic feature pyramid network of PFPN, the present invention can better fuse features of different scales, reduce information loss, and significantly improve the recognition ability for objects of different sizes, thus improving the accuracy of the safety helmet wearing state, especially for objects with large size variations. The Double-layer Goal Enhancement Mechanism (BGA) combines spatial and channel attention, effectively improving the model's focusing ability on the target area, reducing the interference of background noise, and significantly improving the detection accuracy in complex environments. By introducing the multi-scale detection head, the model can better adapt to objects of different scales, optimize the anchor box scale, further improve the detection accuracy for safety helmets of different sizes, and reduce the occurrence rate of missed detections and false detections. In a dynamically changing working environment, the present invention can stably handle a variety of complex scenarios, including different shooting angles, background changes, and lighting conditions, thereby improving the robustness and applicability of the overall recognition system. The present invention can greatly improve the accuracy and efficiency of safety helmet wearing recognition in fields such as construction sites and industrial production, reduce the need for human intervention, and further ensure the safety of workers through high-precision recognition.

[0149] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also fall within the protection scope of the present invention.

Claims

1. A safety helmet wearing recognition method based on multi-scale feature fusion, characterized in that: It includes the following steps: S1. Multi-scale feature fusion: The object detection model performs splicing and weighting operations on feature maps of different scales through a panoramic feature pyramid network to fuse feature maps of different scales; S2. Double-layer object enhancement: The object detection model uses a double-layer object enhancement mechanism constructed based on a spatial attention mechanism and a channel attention mechanism to perform a weighting operation on the fused feature map to focus attention on the target area; S3. Introduction of multi-scale detection heads: Multi-scale detection heads are introduced into the object detection model, and multiple detection heads process targets of different sizes through anchor boxes and feature maps of different scales to optimize the accuracy of object detection; S4. Optimization of the regression loss function: A total regression loss function is constructed for the object detection model using a position regression loss function and a class regression loss function. During the training of the object detection model, parameters are continuously adjusted using an optimization method to minimize the total regression loss function, thereby obtaining a trained object detection model.

2. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 1, wherein: The panoramic feature pyramid network realizes the fusion of multi-scale information by splicing and weighted summation of feature maps of multiple scales. The calculation formula is PFPN i = concat(W i ·F i , ∑(a j ·F j )) Among them, PFPN i is the multi-scale information fusion result; F i and F j are the feature maps of the i-th layer and the j-th layer respectively; W i is the weighting matrix of the feature map of the i-th layer, which is used to perform a weighting operation on the feature map of the i-th layer; a j is the weighting coefficient of the feature map of the j-th layer; ∑(a i ·F j ) is to perform a weighted sum on all scale feature maps F j ; aconcat is to perform a splicing operation on multiple feature maps along the channel dimension.

3. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 2, characterized in that: Construct a loss function for the panoramic feature pyramid network, and train the panoramic feature pyramid network through backpropagation based on the loss function, so as to continuously adjust the weighted matrix W of the feature map i and the weighted coefficient a of the feature map j .

4. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 1, wherein: Step S2 specifically includes the following content: S21. The spatial attention mechanism weights each spatial position in the feature map according to the importance of different positions in the image. The calculation formula is Attention spatial = σ(Conv(F)) Among them, Attention spatial is the spatial weighting coefficient obtained through the spatial attention mechanism; F is the input feature map; Conv(F) is the spatial weighting coefficient extracted from the feature map through the convolution operation; σ is the activation function, which is used to map the spatial weighting coefficient extracted from the feature map through the convolution operation to the interval [0, 1]; S22. The channel attention mechanism weights each channel of the feature map to emphasize the importance of different channels in a specific task. The calculation formula is Attention channel = σ(FC(F)) Among them, Attention channel is the channel weighting coefficient obtained through the channel attention mechanism; FC(F) is the weighting coefficient of each channel extracted from the feature map through the fully connected layer; σ is the activation function, which is used to map the weighting coefficient of each channel extracted from the feature map through the fully connected layer to the interval [0, 1]; S23. After the spatial attention and channel attention respectively weight the feature map, the weighted results of the two are fused to enhance the network's focusing ability on the target area. The calculation formula is Attention fused =Attention spatial ·Attention channel ·F Among them, Attention fused is the final fusion result.

5. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 4, characterized in that: Construct a loss function for the double-layer target enhancement mechanism, and train the double-layer target enhancement mechanism based on the loss function through backpropagation, so as to continuously adjust the spatial weighting coefficient Attention spatial and the channel weighting coefficient Attention channel .

6. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 1, wherein: The multi-scale detection head optimizes the accuracy of the target by adding anchor boxes of different scales to the network and using different feature maps for processing at each scale. The calculation formula is Detection output = ∑(Anchor t · Feature_map t ) Among them, Detection output is the processing result of the multi-scale detection head; Anchor i is the anchor box at the t-th scale; Feature_map t is the feature map at the t-th scale.

7. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 1, wherein: The total regression loss function is constructed by weighting the position regression loss function and the class regression loss function. The calculation formula is L total = λ1·L bbox + λ2·L class L class = -∑(y·log(p) + (1 - y)·log(1 - p)) Among them, L total is the total regression loss function; L bbox and L class are the location regression loss function and the class regression loss function respectively; λ1 and λ2 are the weight coefficients of the location regression loss function and the class regression loss function respectively; IOU is the intersection over union used to measure the overlap degree between the predicted bounding box and the ground truth bounding box; predict is the predicted bounding box; ground truth is the ground truth bounding box; Area_of Intersection is the area of the intersection of the predicted bounding box and the ground truth bounding box; Area_of Union is the area of the union of the predicted bounding box and the ground truth bounding box; y is the true class label, which is 1 when the target belongs to a certain class and 0 when the target does not belong to that class; p is the predicted probability that the object detection model predicts that the target belongs to a certain class.

8. The method for identifying the wearing of a safety helmet based on multi-scale feature fusion according to claim 1, wherein: After step S4, it also includes S5. Safety helmet wearing recognition: Use the trained object detection model to recognize the wearing situation of safety helmets and output the recognition result of safety helmet wearing.

9. The method for recognizing safety helmet wearing based on multi-scale feature fusion according to claim 1, wherein: Before step S1, it also includes S0. Dataset preparation: Manually annotate the wearing status of safety helmets in the collected images to construct a dataset of safety helmet wearing status.

10. A safety helmet wearing recognition system based on multi-scale feature fusion, characterized in that: The recognition system can implement the method described in any one of claims 1 to 9 above. The recognition system includes Multi-scale feature fusion module: The object detection model performs splicing and weighting operations on feature maps of different scales through a panoramic feature pyramid network to fuse feature maps of different scales; Double-layer object enhancement module: The object detection model uses a double-layer object enhancement mechanism constructed based on a spatial attention mechanism and a channel attention mechanism to perform a weighting operation on the fused feature map to focus attention on the target area; Introduction of multi-scale detection heads module: Multi-scale detection heads are introduced into the object detection model, and multiple detection heads process targets of different sizes through anchor boxes and feature maps of different scales to optimize the accuracy of object detection; Regression loss function optimization module: Construct a total regression loss function for the object detection model using the location regression loss function and the class regression loss function, and continuously adjust the parameters using an optimization method during the training of the object detection model to minimize the total regression loss function, thereby obtaining a trained object detection model.

Citation Information

Patent Citations

  • Safety helmet wearing detection method and device based on single-model prediction

    CN111723786A

  • Safety helmet wearing detection method in complex scene

    CN114419659A

  • Helmet wearing detection method based on implicit expression

    CN114463676A

  • Safety helmet wearing detection method based on global attention

    CN114463677A

  • Safety helmet wearing detection method based on improved YOLOV5 model

    CN115546614A