Safety helmet feature extraction system and method based on dynamic convolution and wavelet transform

By introducing the WEDConv module based on dynamic convolution and wavelet transform and the improved loss function in the YOLO architecture, the problem of insufficient feature extraction capabilities and insufficient difficult sample optimization capabilities in complex construction scenarios is solved, and more efficient safety helmet detection performance and robustness are achieved.

CN120125837APending Publication Date: 2025-06-10CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145893.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing YOLO architecture has insufficient feature extraction capabilities in complex construction scenarios, the Ciou loss function has insufficient optimization capabilities for difficult samples, and the Focal Loss classification loss function has limited effect when dealing with category imbalance and difficult-to-classify samples.

Method used

A safety helmet feature extraction system based on dynamic convolution and wavelet transform is designed. The WEDConv module combines wavelet transform and dynamic convolution to enhance the multi-scale capability and robustness of feature extraction, and the Focal-CIoU and Varifocal Loss loss functions are introduced to improve the model's optimization ability for difficult samples.

Benefits of technology

It significantly improves the accuracy and real-timeness of the safety helmet inspection tasks in complex construction site environments, improves the model's detection ability of small targets, occlusion targets and complex backgrounds, and enhances the attention to difficult-to-classify samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125837A_ABST
    Figure CN120125837A_ABST
Patent Text Reader

Abstract

The invention discloses a safety helmet feature extraction system and method based on dynamic convolution and wavelet transform, and relates to the field of image target detection. The safety helmet feature extraction system takes a YOLOv10 network structure as a basic network, and comprises a BackBone backbone network, the BackBone backbone network comprises a WEDConv module, and the WEDConv module comprises a wavelet transform module, a dynamic convolution module and a standard convolution module; the dynamic convolution module comprises a channel attention module, a filter attention module, a convolution kernel attention module, a space attention module and a self-adaptive temperature adjusting module; a Neck neck network; and the Decouled Head is used for decoupling the detection head. The safety helmet feature extraction method based on the system comprises the following steps of data preprocessing, feature extraction, feature fusion, information decoding and prediction. According to the method, the multi-scale feature extraction capability of wavelet transform and the efficient detection framework of YOLO are combined, so that the performance of a safety helmet detection task can be remarkably optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image target detection, and particularly relates to a safety helmet feature extraction system and method based on dynamic convolution and wavelet transform. Background Technique

[0002] In the field of computer vision, target detection is an important branch, which aims to classify, locate and perform target box regression on targets simultaneously. In real-world scenarios, safety helmet detection has wide applications, such as construction site safety management, industrial production supervision, etc. However, in the detection of safety helmet wearing at construction sites, problems such as diverse shapes, complex color distributions of target objects, and high similarity between the target and the background often exist, which makes the performance of existing target detection algorithms in complex scenarios less than ideal.

[0003] In recent years, the main frameworks of target detection algorithms can be divided into two categories: Transformer-based detection algorithms (such as the DETR series) and convolutional neural network-based detection algorithms (such as the YOLO series). The YOLO series of algorithms has achieved rapid development in the past few years. From the initial YOLOv1 to the current YOLOv11, in just 9 years, it has achieved a leap from a single detection framework to a multi-scale decoupled detection head, greatly improving the detection performance. During this process, the YOLO series of algorithms compared the advantages and disadvantages of anchor box and anchor-free detection methods, discussed the application of non-maximum suppression (NMS), and the advantages of multi-scale detection.

[0004] YOLOv10 and YOLOv11, as the two major mainstream algorithms in the current YOLO series, although they have achieved excellent performance on mainstream datasets such as COCO and continued the advantages of YOLO in real-time target detection, they also face some common deficiencies. First, in complex background or small target detection scenarios, missed detection or false detection may still occur, and the ability to recognize occluded targets needs to be further improved. Second, both algorithms still have relatively high requirements for hardware computing power during actual deployment, which is not conducive to large-scale applications on resource-constrained edge devices. In addition, although YOLOv10 cancels NMS in the inference stage and adopts a "dual label assignment strategy" to improve detection accuracy and inference speed, and YOLOv11 has further improved in network structure and feature fusion, in the process of balancing model lightweight and high precision, both still need more in-depth optimization and improvement to meet wider application requirements. Specifically, the following problems exist:

[0005] 1) The existing YOLO architecture has insufficient feature extraction ability in complex construction scenarios;

[0006] Although the YOLO series of models perform excellently in terms of lightweight and real-time performance, their feature extraction capabilities are still limited in complex scenarios such as construction sites. For example, safety helmets may interfere with the surrounding environment due to color, texture, or background complexity, making it difficult for the model to accurately capture target details and key features. This is mainly because the basic network (backbone) adopted by YOLO has limited support for multi-scale feature fusion and cannot fully utilize depth features and spatial features, thus restricting its robustness and detection performance in complex scenarios to a certain extent.

[0007] 2) Existing Ciou loss functions have defects in dealing with the diversity of target shapes;

[0008] Although Ciou (Complete IoU) takes into account the geometric relationship of detection boxes to a certain extent, it still has limitations in dealing with small targets, targets with large aspect ratio differences, and detection tasks with complex backgrounds. Specifically, when dealing with difficult samples such as blurred, partially occluded, or small-sized targets, Ciou has insufficient optimization ability, often resulting in these samples having lower weights during training and being difficult to receive sufficient attention, thus affecting the overall detection accuracy. In addition, the current loss function lacks an appropriate weighting strategy for such difficult samples, resulting in the model's insufficient ability to distinguish and recognize them during the learning process, thereby affecting the recall rate.

[0009] 3) Existing Focal Loss classification loss functions have limited effectiveness in dealing with class imbalance and difficult-to-classify samples;

[0010] In environments with complex backgrounds such as complex construction scenarios, Focal Loss is still insufficient in enhancing the model's attention to key features, easily leading to confusion between targets and the environment and affecting the accuracy of detection. At the same time, Focal Loss has limited effectiveness in dealing with class imbalance and difficult-to-classify samples such as occlusion or similar colors, and cannot apply sufficient training weights to such samples, thus reducing the overall detection performance and robustness. Therefore, there is an urgent need to improve the classification loss function to enhance the model's optimization ability for difficult samples and feature extraction effect, providing more stable support for applications in complex environments.

[0011] Currently, dynamic convolution, as an emerging convolution module, has received extensive attention. By dynamically adjusting parameters such as the weights of convolution kernels and the importance weights of each channel, it enhances the model's adaptability to diverse scenarios and its ability to capture target details. These characteristics make dynamic convolution a widely used feature extraction tool in modern object detection models.

[0012] In addition, as an effective signal processing method, wavelet transform can perform multi-resolution analysis on images at different scales and frequencies. Through wavelet transform, the time-frequency characteristics of the image can be extracted, enhancing the model's ability to distinguish complex backgrounds and subtle targets. Wavelet transform can effectively reduce noise interference during the feature extraction process and improve the robustness of target detection. Combining the multi-scale feature extraction ability of wavelet transform can further enhance the performance of the detection algorithm when dealing with targets of diverse shapes and complex colors.

[0013] The network module design of the present invention based on dynamic convolution and wavelet transform can combine the multi-scale feature extraction ability of wavelet transform and the efficient detection framework of YOLO, significantly optimizing the performance of the safety helmet detection task. This method can not only adapt to the complex construction site environment but also meet the real-time requirements while maintaining high accuracy, providing an innovative and practical solution for construction safety monitoring. Summary of the Invention

[0014] The purpose of the present invention is to provide a safety helmet feature extraction system and method based on dynamic convolution and wavelet transform to solve the problems such as insufficient feature extraction ability of the backbone network of the target detection algorithm of the existing YOLO architecture, insufficient optimization ability of the existing Ciou loss function for difficult samples, and limited effect of the existing Focal Loss classification loss function in dealing with class imbalance and difficult-to-classify samples.

[0015] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0016] In the first aspect of the present invention, a safety helmet feature extraction system based on dynamic convolution and wavelet transform is proposed. This system uses the YOLOv10 network structure as the basic network and includes:

[0017] The BackBone backbone network includes a WEDConv module. The WEDConv module includes a wavelet transform module, a dynamic convolution module, and a standard convolution module. The dynamic convolution module includes a channel attention, a filter attention, a convolution kernel attention, a spatial attention, and an adaptive temperature adjustment module. The BackBone backbone network is used for image feature extraction.

[0018] The Neck neck network is mainly composed of a feature pyramid and is used for image feature fusion during the upsampling and downsampling processes of the feature information extracted by the backbone network.

[0019] The Decoupled Head decoupled detection head is used for information decoding and prediction.

[0020] Preferably, the processing process of the WEDConv module is as follows:

[0021] First, assume that the input feature map is F in , and input F in into the wavelet transform module. After passing through

[0022] the high-pass filter and low-pass filter corresponding to the Haar wavelet, it is decomposed into four sub-bands: LL, HL, LH, and HH;

[0023] Subsequently, the four sub-bands are restored to the original size through transposed convolution;

[0024] Then, the four sub-bands are integrated into a comprehensive feature F combined through a splicing operation. Input F combined into the dynamic convolution module. The feature map passes through the channel attention, filter attention, convolution kernel attention, spatial attention, and adaptive temperature adjustment module in sequence to obtain the adjusted feature map F atten ;

[0025] Finally, input F atten into the standard convolution module for feature extraction, and add the extracted features to the original input feature map through residual connection to obtain the final output F output ;

[0026] The process of the WEDConv module is represented as follows:

[0027] LL, LH, HL, HH = Haar(F in )

[0028] f combined = Concat(LL, LH, HL, HH)

[0029] F atten = DynamicMoudle(F combined )

[0030] F output = Conv(F atten ) + Residual(F in )

[0031] Among them, LL represents low-frequency - low-frequency, which is used to capture the overall contour and low-frequency information; LH represents low-frequency - high-frequency, which is used to capture the edges and details in the horizontal direction; HL represents high-frequency - low-frequency, which is used to capture the edges and details in the vertical direction; HH represents high-frequency - high-frequency, which is used to capture the details in the diagonal direction; Concat represents the splicing operation; DynamicMoudle represents dynamic convolution; Conv represents the convolution operation.

[0032] Further, the process of decomposing into four sub-bands is specifically as follows:

[0033] The Haar wavelet corresponds to the following two filters in the discrete wavelet transform:

[0034] Low-pass filter:

[0035] High-pass filter:

[0036] Represent the image as X[i, k], and perform the following decomposition through the Haar wavelet:

[0037] Apply the Haar discrete wavelet transform to each row of the image to obtain the approximation coefficient A[i, k] and the horizontal detail coefficient H[i, k]:

[0038] A[i, k] = h(X[i, 2k] + X[i, 2k + 1])

[0039] H[i, k] = g(X[i, 2k] - X[i, 2k + 1])

[0040] Apply the Haar discrete wavelet transform to each column of A[i, k] and H[i, k] again to obtain four sub-bands LL, HL, LH, and HH.

[0041] Preferably, the Decoupled Head decoupled detection head decodes the feature information, and each piece of feature information obtains the localization information and the classification information; the decoding process is represented as follows:

[0042] Bbox = Conv(Conv(Conv2d(Feature map)))

[0043] Cls = Conv(Conv(Conv2d(Feature map)))

[0044] Wherein, Bbox and Cls represent the localization information and the classification information of the decoupled detection head; Conv2d represents a common two-dimensional convolution; Feature map represents the feature information.

[0045] Preferably, during the training phase of the safety helmet feature extraction system, the focal complete intersection over union loss is used to calculate the localization loss for the localization information obtained by the decoupled detection head, and at the same time, the variable focal loss is used to calculate the classification loss for the classification information obtained by the decoupled detection head; according to the calculated loss value, the network parameters of the safety helmet feature extraction system are adjusted through backpropagation gradient calculation for optimization.

[0046] Furthermore, the localization loss is calculated using the focal complete intersection over union loss, specifically as follows:

[0047]

[0048] Focal = (IoU) γ

[0049] L Focalciou = L ciou - Focal

[0050] Among them, B p , B gt are the predicted bounding box and the ground truth bounding box, ρ(b, b gt ) is the distance between the centers of the predicted bounding box and the ground truth bounding box, and c is the diagonal length of the minimum enclosing matrix of the predicted bounding box and the ground truth bounding box; γ represents the non - linear weight for adjusting the Focal term; L CIoU represents the CIoU loss; L Focalciou represents the Focal Complete Intersection over Union loss.

[0051] Furthermore, the classification loss is calculated using the variable - focus loss, specifically as follows:

[0052]

[0053] Among them, C represents the total number of categories; y c represents the true label of this sample for the C - th category; p c represents the probability predicted by the model for the C - th category; α c represents the weighting coefficient for the C - th category, used to balance class imbalance or the importance of different categories; γ is the modulation factor, used to control the attention to difficult samples and easy - to - classify samples.

[0054] The second aspect of the present invention proposes a safety helmet feature extraction method based on dynamic convolution and wavelet transform, including the following steps:

[0055] S1. Data pre - processing; perform image enhancement on the input image, process it to a unified size, and perform normalization operations;

[0056] S2. Feature extraction; input the pre - processed image into the backbone network, map the semantic information of the RGB three channels to a high - dimensional space, and generate feature maps of three scales;

[0057] S3. Feature fusion; first perform up - sampling and then down - sampling on the feature information extracted by the backbone, and synthesize the feature maps of the three scales into one feature information by splicing during the sampling process;

[0058] S4. Information decoding; form the fused feature information into a feature sequence, input the feature sequence into the decoupled detection head for decoding, and each feature information generates localization information and classification information;

[0059] S5. Safety helmet prediction; obtain the classification result and the localization result by parsing the localization information and the classification information.

[0060] Preferably, three-scale feature maps are generated in S2, specifically as follows:

[0061] After image preprocessing, it is fixed at 640*640. The preprocessed image undergoes a downsampling process from 640*640 to 20*20 in the backbone network, and then an upsampling process from 20*20 to 80*80; during this process, the feature data of 160*160 is input into the WEDConv module, and then the output value of the WEDConv module is concatenated with the feature value of 80*80; three-scale feature maps are generated.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] (1) The key technology in the present invention is the design of the WEDConv convolutional module; on the one hand, this module captures multi-scale features of different subbands through Haar wavelet decomposition, and on the other hand, it adaptively strengthens key information through methods such as channel attention, filter attention, convolutional kernel attention, and spatial attention, and then cooperates with adaptive temperature adjustment to improve the flexibility of feature selection, so as to achieve more comprehensive, accurate feature expression and robustness while maintaining high-efficiency operation.

[0064] (2) The key technology in the present invention is the design of the positioning loss function; in order to strengthen the learning ability for difficult-to-optimize samples (such as small targets and occluded targets), Focal-CIoU is invented after improving CIoU. It adds a Focal term to CIoU, and assigns a higher loss weight to difficult samples with a smaller IoU through dynamic weight adjustment. The Focal term is based on IoU and is non-linearly adjusted according to the sample difficulty, enabling the model to more effectively optimize the sample features in complex scenarios. This design fully combines the consistency of IoU, the distance between the center of the bounding box, and the aspect ratio, and at the same time focuses on difficult-to-optimize target samples, thereby improving the convergence speed and detection accuracy of the model.

[0065] (3) The key technology in the present invention is the design of the variable-focus classification loss function; in order to simultaneously consider the imbalance problem between classification confidence and target localization accuracy, the present invention introduces the Varifocal Loss. This loss is an improvement on the traditional Focal Loss. During model training, Focal Loss only weights the difference between the predicted probability and the true label, ignoring the influence of localization accuracy (such as IoU information) in the evaluation of sample weights. Varifocal Loss incorporates localization information such as IoU into the calculation of the classification loss, so that difficult-to-classify and accurately localized samples receive higher attention, thereby significantly improving the overall detection performance. Description of the Drawings

[0066] Figure 1 This is the structural block diagram of the safety helmet feature extraction system based on dynamic convolution and wavelet transform in the present invention;

[0067] Figure 2 This is the structural block diagram of the WEDConv module in the present invention. Specific implementation manners

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] Embodiment 1:

[0070] The safety helmet feature extraction method based on dynamic convolution and wavelet transform can be divided into four stages: data preprocessing, image feature extraction, image feature fusion, information decoding and prediction. In the data preprocessing stage, the input image is enhanced by methods such as random cropping, random rotation, and random occlusion, and then the image size is fixed to a unified size to facilitate the model to extract features. In the feature extraction stage, the preprocessed image is input into the backbone network, and the semantic information of the RGB three channels is mapped to a high-dimensional space, and then feature maps of multiple scales are obtained. In the feature fusion stage, the image feature information extracted from the backbone network is fused through upsampling and downsampling of the feature pyramid. In the information decoding stage, the multiple fused feature information is formed into a feature sequence and then input into the detection head for decoding. Finally, each feature information will generate a positioning information and a classification information, and the classification and positioning results of the model can be obtained by translating this information. The model of n size is as Figure 1 shown. The models of different sizes only have different modules that make up the network, and the overall process remains the same.

[0071] The safety helmet feature extraction method based on dynamic convolution and wavelet transform is as follows:

[0072] Step 1. Data preprocessing.

[0073] First, image enhancement is performed. In order to handle the input of different image sizes, facilitate inputting the image into the network for feature extraction and prediction, and enhance the robustness of the model, image enhancement is required. The methods include but are not limited to random cropping, blurring, randomly adding occlusion, translation and flipping, etc. These enhancement methods can effectively improve the robustness of the model.

[0074] Then, images of different sizes are fixed to a unified size. To prevent image distortion, it is generally chosen to scale the image according to the original ratio, and then smear the empty pixels.

[0075] Finally, the size of the image to be input is fixed to 640×640, and then each pixel is normalized. This operation can fix the pixels of each channel within a certain range (such as between 0 and 1 or between -1 and 1). Without changing their size relationship when changing the numerical size, this method can effectively improve the computing efficiency of the computer.

[0076] Step 2: Feature extraction of the BackBone backbone network for images.

[0077] The preprocessed image is input into the backbone network to map the semantic information of the RGB three channels into a high-dimensional space. Then, feature maps of three scales are generated. These feature maps come from different levels of the network structure, so they contain semantic information at different levels, facilitating the network to identify target objects of different sizes.

[0078] Specifically, after the preprocessed image is input into the backbone network, the overall process is a downsampling process from 640*640 to 20*20, and then an upsampling process from 20*20 to 80*80. During this upsampling process, feature data larger than 160*160 in size is not obtained. Therefore, based on the original YOLOV10 network structure, the feature data of 160*160 is input into the WEDConv module, and the output value of WEDConv is concatenated with the feature value of 80*80. The network structure of the WEDConv module is as Figure 2 shown.

[0079] The WEDConv (Wavelet-Enhanced Dynamic Convolution) module is an advanced feature extraction module based on the idea of dynamic convolution, designed specifically to improve the object detection performance in complex environments. The WEDConv module consists of a wavelet transform module, a dynamic convolution module, and a normal convolution operation. The module first decomposes the input feature map into four sub-bands (LL, LH, HL, HH) through the discrete wavelet transform (DWT) to achieve multi-resolution and multi-frequency feature extraction, thus effectively capturing the time-frequency information of the image. Subsequently, these sub-bands are concatenated together and input into the dynamic convolution module integrated with various attention mechanisms, which includes channel attention, filter attention, convolutional kernel attention, and spatial attention. By dynamically adjusting the convolutional kernel weights, the module significantly enhances the model's adaptability to diverse scenarios and subtle targets. In addition, the adaptive temperature adjustment mechanism further optimizes the response of the attention module to ensure the stability and accuracy of feature extraction. The entire WEDConv module also adopts residual connections to effectively prevent gradient vanishing, promoting the training efficiency and model performance of deep networks. By comprehensively applying the ideas of wavelet transform and dynamic convolution, the WEDConv module achieves efficient and robust feature extraction, significantly improving the accuracy and real-time performance of the hard hat detection task in complex construction site environments.

[0080] The main working principle of the WEDConv module is as follows:

[0081] Assume the input feature map is F in , and the WEDConv module performs the following processing steps in sequence:

[0082] First, the input feature F in passes through the high-pass and low-pass filters corresponding to the Haar wavelet and is decomposed into four sub-bands: LL (low frequency - low frequency, capturing the overall contour and low-frequency information), LH (low frequency - high frequency, capturing the horizontal edges and details), HL (high frequency - low frequency, capturing the vertical edges and details), and HH (high frequency - high frequency, capturing the diagonal details). Specifically as follows:

[0083] The scaling function and wavelet function corresponding to the Haar wavelet are defined as follows:

[0084] Scaling function:

[0085] Wavelet function:

[0086] Based on the above functions, the Haar wavelet corresponds to the following two filters in the discrete wavelet transform:

[0087] Low-pass filter (scaling filter):

[0088] High-pass filter (wavelet filter):

[0089] For a two-dimensional image, without considering the number of channels, if the image is represented as X[i, k], the following decomposition can be performed using the Haar wavelet:

[0090] First, apply the Haar discrete wavelet transform to each row of the image to obtain the approximation coefficient A[i, k] and the horizontal detail coefficient H[i, k]:

[0091] A[i, k] = h(X[i, 2k] + X[i, 2k + 1])

[0092] H[i, k] = g(X[i, 2k] - X[i, 2k + 1])

[0093] Then, apply the Haar discrete wavelet transform to each column of A[i, k] and H[i, k] again to obtain four sub-bands LL, HL, LH, HH.

[0094] Subsequently, restore the four sub-bands to the original size through transposed convolution.

[0095] Then, integrate the four sub-bands into a comprehensive feature F through a concatenation operation combined , and then immediately input F combined into the dynamic convolution module. The feature map passes through the channel attention, filter attention, convolution kernel attention, spatial attention, and adaptive temperature adjustment module in sequence to obtain the adjusted feature map F atten .

[0096] Finally, input F atten into the standard convolution module Conv for feature extraction, and add the extracted features to the original input through a residual connection to obtain the final output F output .

[0097] The WEDConv process is represented as follows: (Fin, LL, LH, HL, HH, ).

[0098] LL, LH, HL, HH = Haar(F in )

[0099] F commbined = Concat(LL, LH, HL, HH)

[0100] F atten = DynamicMoudle(F combined )

[0101] Foutput = Conv(F atten ) + Residual(F in )

[0102] Throughout the backbone network, considering the trade-off between the change of model parameters and the feature extraction ability, only one WEDConv module is added. The obtained feature map contains rich potential semantic information, which can fully extract the potential information in the high-dimensional space.

[0103] Step 3: Feature fusion of the Neck network for images.

[0104] After the image passes through the backbone network, the low-dimensional semantic information is extended to the high-dimensional space. The Neck network is mainly composed of a feature pyramid, which fully fuses the feature information during the upsampling and downsampling processes. The aim is to enhance the feature expression ability and adaptability, so as to better capture the semantic information and spatial relationship of the target object.

[0105] Specifically, in the way of a feature pyramid, the feature information extracted by the backbone is first upsampled and then downsampled. During the sampling process, the P3, P4, and P5 feature maps extracted by the backbone network are synthesized into a new tensor without loss by splicing, and then input into the module for channel adjustment and information fusion.

[0106] Step 4: Classification and localization of the Decoupled Head decoupled detection head.

[0107] The Decoupled Head decoupled detection head decodes and predicts information. During the feature fusion stage, many feature maps are generated, and then three of them are selected as prediction samples. These three feature maps have different widths and heights. The larger the feature map, the more effective it is for detecting small targets. Because the larger the feature map, the smaller the area corresponding to each element in the original image, and vice versa. Then the three feature maps are combined to perform object classification and localization.

[0108] Specifically, detection tensors of different sizes are input into the decoupled detection head. After processing, each detection tensor will obtain a classification tensor and a localization tensor, and the final classification and localization results can be obtained by analyzing these two tensors.

[0109] The specific decoding process can be expressed as follows:

[0110] Bbox = Conv(Conv(Conv2d(Feature map)))

[0111] Cls = Conv(Conv(Conv2d(Feature map)))

[0112] Among them, Bbox and Cls represent the localization and classification results of the decoupled head. Conv2d represents a common two-dimensional convolution.

[0113] The above is the entire process of the model inference stage. In the training stage of YOLOvl0, two label assignment strategies, one-to-one and one-to-many, are used, and the balance between the two requires the use of a compensation directional consistency matching metric. In the inference stage, only the one-to-one label assignment strategy is used. When the model is in the training state, the network calculates the localization loss using the Focal Complete IoU (FocalCIoU) based on the tensors obtained from the detection head, and at the same time calculates the classification loss using the Varifocal Loss. The gradients of each layer of the model are calculated based on the obtained losses, and then the model parameters are appropriately adjusted.

[0114] When the model is in the training stage, it is necessary to calculate the classification loss and localization loss of the feature map of each input detection head separately, and then combine them to obtain the total loss of the model. It should be noted that YOLOv10 trains both label assignment strategies simultaneously in the training stage, and the total loss is obtained by integrating the losses of the two label assignment strategies. After the model obtains the final loss value, the corresponding network parameters are adjusted through backpropagation gradient calculation to achieve the purpose of learning.

[0115] Specifically, the localization loss is calculated using the Focal Complete IoU (Focal-CIoU).

[0116] In order to simultaneously solve the problems of target localization accuracy and uneven sample difficulty, the present invention proposes the Focal Complete IoU loss.

[0117] This loss improves the CIoU (Complete IoU). During model training, CIoU only considers the relative shape and position information of the bounding boxes, ignoring the influence of the attributes of the bounding boxes themselves on the evaluation results. CIoU can be expressed as:

[0118]

[0119]

[0120] where, B p , B gt are the predicted box and the ground truth box, ρ(b, b gt ) is the distance between the centers of the predicted box and the ground truth box, and c is the diagonal length of the minimum circumscribed matrix of the predicted box and the ground truth box.

[0121] As can be seen from the formula, the weights for each sample are the same, and it is impossible to dynamically adjust the weights according to the difficulty of the samples. During the optimization process of the model, it is impossible to specifically focus on optimizing the difficult samples in scenarios with small IoU or complex scenes, resulting in limitations in detection accuracy. Small targets usually have a small bounding box area. Even if the center point distance and aspect ratio are optimized, their IoU will still be very small. Since the loss of CIoU mainly depends on the change of IoU, the contribution of small targets to the overall loss can be almost ignored. Therefore, based on the above loss function, a new improved scheme is proposed: FocalCIoU, and its formula can be expressed as:

[0122] Focal = (IoU)γ

[0123] L FocalCiou = L Ciou - Focal

[0124] Among them, γ represents the non - linear weight that adjusts the Focal term. Usually γ > 0. When γ > 1, more attention is paid to samples with small IoU (samples that are difficult to optimize), and the weight increases. When γ ->, the Focal term degenerates into the standard CIoU, without additional enhancement of attention to difficult samples. The larger the value of γ, the more significant the impact of the Focal term on samples with small IoU, strengthening the model's learning of samples that are difficult to optimize. A smaller γ value maintains a higher weight for samples that are easy to optimize, contributing to the stable optimization of the overall model.

[0125] Specifically, the classification loss is calculated through the Varifocal Loss.

[0126] In order to simultaneously take into account the imbalance between classification confidence and target localization accuracy, the present invention introduces the Varifocal Loss.

[0127] This loss is an improvement based on the traditional Focal Loss. During model training, Focal Loss only weights the difference between the predicted probability and the true label, ignoring the impact of localization accuracy (such as IoU information) in the evaluation of sample weights. Varifocal Loss incorporates localization information such as IoU into the calculation scope of the classification loss, enabling samples that are difficult to classify and have accurate localization to receive more attention, thereby significantly improving the overall detection performance. Varifocal Loss can be expressed as:

[0128]

[0129] Among them, C represents the total number of categories, y c represents the true label of this sample in the C - th category, usually using one - hot encoding (y c ∈{0, 1}, and )。p c Represents the probability of the model's prediction for the C-th class, output by Sigmoid. α c Represents the weighting coefficient for the C-th class, used to balance class imbalance or the importance of different classes. γ is a modulation factor that controls the attention to hard samples and easy-to-classify samples. γ usually takes values such as 1, 1.5, 2, etc. The larger the value, the more emphasis is placed on hard samples. In the implementation of general Focal Loss, the weights of positive samples are mostly written as 1, and the "localization accuracy" of positive samples themselves is not utilized.

[0130] Varifocal Loss, on the other hand, makes the weight of positive samples = y c and sets y c to IoU (or a score related to IoU) during the preprocessing stage of training data. In this way, the classification loss automatically "sees" the localization quality of this positive sample. The more accurate the localization, the higher the loss amplification factor; if the localization is inaccurate (low IoU value), the proportion of this positive sample in the classification loss is correspondingly reduced. The purpose of "incorporating IoU into the classification loss" is achieved, thus better solving the problems of target localization accuracy and sample difficulty imbalance simultaneously.

[0131] Based on the above invention content, the following evaluation of technical effects is proposed:

[0132] The present invention has been experimentally evaluated on the personal-made SHDD (Safety Helmet Detection Dataset) dataset. This dataset contains 8,722 high-resolution real construction site images, where the training set contains 6,977 images and the test set contains 1,745 images. This dataset includes three categories: safety helmets, people wearing safety helmets, and people not wearing safety helmets.

[0133] The baseline model YOLOv10 has model versions with multiple parameters. Since the use scenario requires the model to achieve real-time detection effects, the present invention selects the model with a parameter scale of n as the baseline model in this case. The comparison test results with the baseline model on the SHDD dataset are shown in Table 1:

[0134] Table 1 Comparison test results on the SHDD dataset

[0135] Model Parameters GFLOP(G) P(%) R(%) mAP50(%) YOLOv8-n 3006623 8.1 80.9 78.2 79.9 YOLOv10-n 2708210 8.2 83.2 74.2 77.6 YOLOv11-n 2582737 6.3 80.7 75.3 79.6 Ours-n 3010465 9.9 87.7 75.6 80.8

[0136] At the n size, compared with the baseline model, the accuracy rate, average recall rate, and average accuracy rate of fifty retrieval results of the present invention increase by 4.5, 1.4 percentage points, and 3.2 percentage points respectively.

[0137] As described above, it is only used to help understand the method of the present invention and its core concept. However, the protection scope of the present invention is not limited thereto. For those of ordinary skill in the art in the technical scope disclosed by the present invention, any equivalent substitution or change made according to the technical solution of the present invention and its inventive concept should be covered within the protection scope of the present invention. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A helmet feature extraction system based on dynamic convolution and wavelet transform, characterized in that: The basic network is based on the YOLOv10 network structure, including: BackBone backbone network, including WEDConv module, the WEDConv module includes wavelet transform module, dynamic convolution module and standard convolution module; the dynamic convolution module includes channel attention, filter attention, convolution kernel attention, space attention and adaptive temperature adjustment module; the BackBone backbone network is used for image feature extraction; Neck network, mainly composed of feature pyramids, is used to fuse image features of feature information extracted by the backbone network during upsampling and downsampling; Decoupled Head is a decoupled detection head used for information decoding and prediction.

2. The helmet feature extraction system based on dynamic convolution and wavelet transform according to claim 1 is characterized in that: The processing of the WEDConv module is as follows: First, assume that the input feature map is F in , F in The input wavelet transform module is decomposed into four sub-bands LL, HL, LH, and HH after passing through the high-pass filter and low-pass filter corresponding to the Haar wavelet; Subsequently, the four subbands are restored to their original sizes through transposed convolution; Then, the four sub-bands are integrated into a comprehensive feature F through splicing operation combined , F combined The input is a dynamic convolution module, and the feature map is sequentially passed through the channel attention, filter attention, convolution kernel attention, spatial attention and adaptive temperature adjustment modules to obtain the adjusted feature map F. atten ; Finally, F atten The standard convolution module is input for feature extraction, and the extracted features are added to the original input feature map through residual connection to obtain the final output F output ; The process of the WEDConv module is shown as follows: LL,LH,HL,HH=Hair(F in ) F commbined =Concat(LL,LH,HL,HH) F atten =DynamicMoudle(F combined ) F output =Conv(F atten )+Residual(F in ) Among them, LL stands for low frequency-low frequency, which is used to capture the overall contour and low-frequency information; LH stands for low frequency-high frequency, which is used to capture the edges and details in the horizontal direction; HL stands for high frequency-low frequency, which is used to capture the edges and details in the vertical direction; HH stands for high frequency-high frequency, which is used to capture the details in the diagonal direction; Concat represents splicing operation; DynamicMoudle represents dynamic convolution; Conv represents convolution operation.

3. The helmet feature extraction system based on dynamic convolution and wavelet transform according to claim 2 is characterized in that: The process of decomposing into four sub-bands is as follows: The Haar wavelet corresponds to the following two filters in discrete wavelet transform: Low pass filter: High Pass Filter: The image is represented as X[i, k] and decomposed by Haar wavelet as follows: Applying the Haar discrete wavelet transform to each row of the image yields the approximation coefficient A[i, k] and the horizontal detail coefficient H[i, k]: A[i,k]=h(X[i,2k]+X[i,2k+1]) H[i,k]=g(X[i,2k]-X[i,2k+1]) The Haar discrete wavelet transform is applied again to each column of A[i, k] and H[i, k] to obtain four sub-bands LL, HL, LH, and HH.

4. The helmet feature extraction system based on dynamic convolution and wavelet transform according to claim 3 is characterized in that: The Decoupled Head decouples the detection head to decode the feature information, and each feature information obtains positioning information and classification information; the decoding process is represented as follows: Bbox=Conv(Conv(Conv2d(Feature map))) Cls=Conv(Conv(Conv2d(Feature map))) Among them, Bbox and Cls represent the positioning information and classification information of the decoupled detection head; Conv2d represents an ordinary two-dimensional convolution; Feature map represents feature information.

5. The helmet feature extraction system based on dynamic convolution and wavelet transform according to claim 4 is characterized in that: During the training phase, the helmet feature extraction system uses the focal complete intersection-over-union loss to calculate the positioning loss for the positioning information obtained by the decoupling detection head, and uses the variable focus loss to calculate the classification loss for the classification information obtained by the decoupling detection head. Based on the calculated loss values, the network parameters of the helmet feature extraction system are adjusted and optimized through back-propagation gradient calculation.

6. The helmet feature extraction system based on dynamic convolution and wavelet transform according to claim 5 is characterized in that: The positioning loss is calculated using the focal complete intersection-over-union loss, as follows: Focal=(I0U) γ THE FocalCiou =L Ciou -Focal Among them, B p , B gt is the predicted box and the real box, ρ(b, b gt ) is the center distance between the predicted box and the real box, c is the diagonal length of the minimum circumscribed matrix between the predicted box and the real box; γ represents the nonlinear weight for adjusting the Focal term; L CIoU represents CIoU loss; L Focalciou Indicates the focal complete intersection loss.

7. The helmet feature extraction system based on dynamic convolution and wavelet transform according to claim 5 is characterized in that: The classification loss is calculated using variable focus loss, as follows: Where C represents the total number of categories; y c represents the true label of the sample in category C; p c Represents the probability of the model predicting the Cth class; α represents the weighting coefficient of the Cth class, which is used to balance the class imbalance or the importance of different classes; γ is a modulation factor used to control the attention paid to difficult samples and easy-to-classify samples.

8. A method for extracting safety helmet features based on dynamic convolution and wavelet transform using the safety helmet feature extraction system according to any one of claims 1 to 7, characterized in that: The following steps are involved: S1. Data preprocessing: perform image enhancement on the input image, process it into a uniform size, and perform normalization operations; S2, feature extraction; The preprocessed image is input into the backbone network, the semantic information of the three RGB channels is mapped to a high-dimensional space, and feature maps of three scales are generated; S3, feature fusion: the feature information extracted from the backbone is first upsampled and then downsampled. During the sampling process, the feature maps of the three scales are spliced ​​to form one feature information; S4, information decoding: the fused feature information is combined into a feature sequence, the feature sequence is decoded, and each feature information generates positioning information and classification information; S5, helmet prediction; The classification result and the positioning result are obtained by parsing the positioning information and the classification information.

9. The method for extracting safety helmet features based on dynamic convolution and wavelet transform according to claim 8, characterized in that: In S2, feature maps of three scales are generated, as follows: After image preprocessing, the image is fixed to 640*640. The preprocessed image undergoes a downsampling process from 640*640 to 20*20 in the backbone network, and then an upsampling process from 20*20 to 80*80. In this process, the 160*160 feature data is input into the WEDConv module, and then the output value of the WEDConv module is concatenated with the 80*80 feature value to generate feature maps of three scales.