Computer vision-based construction site safety helmet wearing detection system and method

Through computer vision methods of multimodal data fusion and multi-scale feature extraction, the challenge of safety helmet detection in complex environments on construction sites is solved, and efficient and accurate safety helmet wear detection is achieved, suitable for resource-constrained edge devices.

CN119091219BActive Publication Date: 2025-09-02CHINA RAILWAY CONSTRUCTION ENGINEERING GROUP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411244233.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-09-02
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

The existing computer vision safety helmet detection technology faces challenges such as large changes in lighting, serious occlusion, messy background, frequent overlapping scenes for multiple people, and poor time consistency of the detection results in construction sites. It is difficult to meet real-time requirements while ensuring the detection accuracy.

Method used

Multimodal data fusion preprocessing, multi-scale feature extraction and fusion, depth-aware density clustering RPN algorithm, multi-stage safety helmet detection and classification method are adopted, combined with multi-stage post-processing and result optimization, and the detection efficiency and accuracy are improved through technologies such as multi-modal sensor array synchronous acquisition, deep adaptive ICP algorithm, multi-scale spatiotemporal DTW algorithm, etc.

Benefits of technology

It significantly improves the efficiency and accuracy of safety helmet detection in complex construction site environments, and can identify the wearing status of the safety helmet in real time and accurately, making it suitable for resource-constrained edge equipment deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091219B_ABST
    Figure CN119091219B_ABST
Patent Text Reader

Abstract

The present invention discloses a computer vision-based construction site hardhat wearing detection system and method, comprising the following steps: obtaining an aligned and enhanced multimodal dataset, processing it using a multi-scale feature extraction and fusion method to obtain a fused feature set; obtaining the fused feature set and the enhanced depth image, processing them using a depth-aware density clustering (RPN) algorithm to obtain a candidate region set; then processing them using a multi-stage hardhat detection and classification method to obtain preliminary hardhat detection and classification results, forming a classification result set; obtaining the obtained candidate region set and classification result set, as well as the aligned and enhanced multimodal dataset, and processing them using a multi-stage post-processing and result optimization method to obtain the final hardhat detection and classification results. This method significantly improves detection efficiency and has good performance in practical scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to building safety monitoring technology, and in particular to a method and system for detecting safety helmets through computer vision. Background Art

[0002] Construction site safety management is crucial for ensuring the safety of construction workers and the smooth progress of projects. Properly wearing a hard hat is one of the most fundamental and crucial safety measures. However, due to the complex environment and high personnel mobility of construction sites, traditional manual inspections are often inefficient and have limited coverage. Therefore, developing an automated hard hat detection system based on computer vision is of great practical significance. Such a system not only monitors site safety conditions around the clock and in all directions, but also promptly detects and warns of potential safety hazards, significantly improving the efficiency and accuracy of safety management. Furthermore, through long-term data accumulation and analysis, such a system can provide data support for the formulation and optimization of construction site safety management strategies, thereby improving the safety level of the entire construction industry at a macro level.

[0003] Currently, research on hardhat detection based on computer vision has made considerable progress. Mainstream methods focus on two main technical approaches: object detection and image classification. In object detection, researchers have experimented with methods ranging from the traditional HOG+SVM approach to more recent deep learning methods such as Faster R-CNN, YOLO, and SSD. These methods are able to effectively locate and identify hardhats under ideal conditions. In image classification, research focuses on extracting more discriminative features and designing more efficient classifiers. With the recent development of deep learning technology, methods based on convolutional neural networks (CNNs), such as VGG and ResNet, have demonstrated superior performance in hardhat classification tasks. Furthermore, some researchers have attempted to combine object detection with image classification, proposing end-to-end methods for hardhat detection and wearing status classification, further improving the practicality of the system.

[0004] Although existing research has achieved certain results, it still faces many challenges in practical applications. For example, the construction site environment is complex and changeable, with large changes in lighting, severe occlusion, and cluttered backgrounds, which poses a huge challenge to traditional computer vision algorithms. Safety helmets are small in size and similar in shape, and are easily confused with other objects in long-distance or high-altitude work scenarios, resulting in a decrease in detection accuracy. Multiple overlapping scenes frequently occur on construction sites, and existing methods are not effective in detecting safety helmets in dense crowds. Due to the dynamic nature of construction sites, the temporal consistency of detection results is also an issue that needs to be addressed urgently. In terms of feature extraction, how to effectively fuse multimodal information such as color, texture, shape, and timing to improve the robustness of detection remains an open problem.

[0005] In practical applications, how to design lightweight models to meet real-time requirements while ensuring detection accuracy also needs to be solved, which requires research and development and innovation. Summary of the Invention

[0006] Purpose of the invention:

[0007] According to one aspect of the present application, a method for detecting the wearing of a safety helmet at a construction site based on computer vision is provided, comprising the following steps:

[0008] Step S1: collecting or acquiring multimodal raw data of a construction site, and processing the data using a multimodal data fusion preprocessing method to obtain an aligned and enhanced multimodal dataset;

[0009] Step S2: obtaining an aligned and enhanced multimodal dataset, processing it using a multi-scale feature extraction and fusion method to obtain a fused feature set;

[0010] Step S3: Obtain the fused feature set and the enhanced depth image, process them using the depth-aware density clustering RPN algorithm to obtain a candidate region set; then use a multi-stage helmet detection and classification method to obtain preliminary helmet detection and classification results to form a classification result set;

[0011] Step S4: Obtain the obtained candidate region set and classification result set, as well as the aligned and enhanced multimodal dataset, and process them using a multi-stage post-processing and result optimization method to obtain the final helmet detection and classification results.

[0012] According to another aspect of the present application, a computer vision-based construction site safety helmet wearing detection system is provided, comprising:

[0013] at least one processor; and,

[0014] a memory communicatively connected to at least one of the processors; wherein,

[0015] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the computer vision construction site safety helmet wearing detection method described in any of the above technical solutions.

[0016] The beneficial effect is that the detection efficiency is greatly improved and the effect is good in actual scenes. The specific technical effects will be described in detail below in conjunction with the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of the present invention.

[0018] Figure 2It is a flow chart of step S1 of the present invention.

[0019] Figure 3 It is a flow chart of step S2 of the present invention.

[0020] Figure 4 It is a flow chart of step S3 of the present invention.

[0021] Figure 5 It is a flow chart of step S4 of the present invention. DETAILED DESCRIPTION

[0022] In this embodiment, the specific process is as follows:

[0023] Step S1: Acquire multimodal raw data from the construction site and process it using a multimodal data fusion preprocessing method to obtain an aligned and enhanced multimodal dataset. Specifically, first, acquire multimodal raw data, including high-resolution RGB images, depth images, and thermal images. Then, perform spatiotemporal alignment on these raw data to obtain an aligned multimodal dataset. Next, generate a multiscale image pyramid for the aligned RGB images and perform illumination normalization to obtain an illumination-normalized multiscale image pyramid. Finally, enhance the depth image to obtain an enhanced depth image and depth edge map.

[0024] Step S11: Real-time scene data from the construction site is acquired and processed using a multimodal sensor array synchronous acquisition method to obtain a raw multimodal dataset. Specifically, a multimodal sensor array, including a high-resolution RGB camera, a depth camera, and a thermal imaging camera, is used to synchronously acquire data from the real-time construction site scene. The high-resolution RGB camera utilizes stacked CMOS sensor technology, the depth camera employs hybrid depth sensing technology combining structured light and time-of-flight (ToF), and the thermal imaging camera utilizes uncooled vanadium oxide (VOx) microbolometer technology. This improves the sensitivity and resolution of thermal imaging. This multimodal synchronous acquisition method captures comprehensive scene information, providing a rich data foundation for subsequent hardhat inspections.

[0025] Step S12: Obtain the original multimodal dataset and process it using the deep adaptive ICP algorithm to obtain a spatially aligned multimodal dataset. Then, process the spatially aligned multimodal dataset using the multiscale spatiotemporal DTW algorithm to obtain a spatiotemporally aligned multimodal dataset. Specifically, the deep adaptive ICP algorithm is first used for spatial alignment. This algorithm introduces an adaptive weighting mechanism to dynamically adjust the weights of point correspondences based on the reliability of depth information. Then, the multiscale spatiotemporal DTW algorithm is applied for temporal alignment. This algorithm introduces a multiscale analysis strategy to perform alignment at different time scales.

[0026] By aligning data at different time scales, the sampling rate differences and time delays between multimodal data are effectively addressed. This improved multimodal data alignment method significantly improves the consistency between different sensor data, laying the foundation for subsequent feature extraction and fusion.

[0027] Step S13: Obtain the aligned RGB image and process it using the edge-enhanced multi-directional Laplacian pyramid algorithm to generate a multi-scale image pyramid. Specifically, this algorithm introduces an adaptive edge-enhancement filter in each pyramid layer, dynamically adjusting the filter parameters based on the local image structure. Furthermore, a multi-directional decomposition technique based on the wavelet transform is introduced to generate an image pyramid that contains multi-scale information and multi-directional features.

[0028] The adaptive edge enhancement filter dynamically adjusts its parameters based on the local image structure, effectively suppressing noise while preserving image detail. Furthermore, a multi-directional decomposition technique based on wavelet transforms is introduced, allowing the generated image pyramid to incorporate not only multi-scale information but also multi-directional features, thereby better capturing the shape and texture characteristics of the helmet. This improved multi-scale image pyramid generation method significantly improves the effectiveness of subsequent feature extraction, especially for detecting helmets of different scales and orientations.

[0029] Step S14: Obtain a multi-scale image pyramid and aligned thermal imaging data, and process them using the thermal imaging-guided CLAHE algorithm to obtain a multi-scale image pyramid after illumination normalization. Specifically, the thermal imaging data is first used to create an illumination intensity map, derived through a nonlinear mapping relationship between the thermal imaging data and the visible light image. Then, the thermal imaging-guided CLAHE algorithm is applied to each scale image. This algorithm incorporates a local contrast adaptation mechanism based on the thermal imaging data, dynamically adjusting the degree of contrast enhancement based on the thermal characteristics of different regions. Finally, the generated illumination intensity map is used to perform pixel-level brightness adjustment, using a nonlinear mapping function to convert the thermal imaging data into a brightness adjustment factor. This conversion from thermal imaging data to brightness adjustment factors ensures that details in low-light areas are enhanced while maintaining the overall visual quality of the image. This adaptive illumination normalization method effectively handles images under complex lighting conditions and significantly improves the robustness of subsequent helmet detection. Morphological operations are combined to optimize edge continuity and closure. This significantly improves the quality and integrity of depth information, providing reliable data support for subsequent 3D-based helmet detection.

[0030] Step S15: Acquire the aligned depth image and process it using a depth-adaptive bilateral filtering algorithm to obtain an enhanced depth image. Then, process the enhanced depth image using a depth completion algorithm to obtain a filled depth image. Finally, process the filled depth image using a multi-scale depth gradient algorithm to obtain a depth edge map. Specifically, first, apply a depth-adaptive bilateral filtering algorithm to remove noise from the depth map. This algorithm introduces an adaptive kernel function based on the local depth distribution. Then, use a depth completion algorithm to fill holes in the depth map. This algorithm combines an interpolation method based on the diffusion equation with learning-based depth prior knowledge. Finally, generate a depth edge map. Use a multi-scale depth gradient operator to extract depth edges, and combine morphological operations to process the edges.

[0031] Step S2: Obtain the aligned and enhanced multimodal dataset obtained in step S1 and process it using a multi-scale feature extraction and fusion method to obtain a fused feature set. Specifically, color and texture features are first extracted from the illumination-normalized multi-scale image pyramid; shape features are then extracted from the enhanced depth image and depth edge map; and temporal features are then extracted using multiple consecutive image frames. Finally, all extracted features are fused to obtain a comprehensive fused feature set.

[0032] Step S21: Obtain the illumination-normalized multi-scale image pyramid obtained in step S14 and process it using an adaptive multi-space color descriptor method to obtain a multi-scale color feature set. Specifically, first, the discriminative scores of multiple color spaces are calculated for each image scale, and the color space with the highest score is selected. Then, the local color moments and color covariance matrix are calculated in the selected color space. Finally, the features at different scales are integrated using a weighted average method to obtain the final multi-scale color feature set.

[0033] In the embodiments, adaptive color space transformation technology is used to dynamically select the optimal color space (such as RGB, HSV, or Lab) based on the image content. Within the selected color space, an improved local color descriptor extraction algorithm is applied. This algorithm combines color moments and color covariance matrices to more comprehensively capture local color distribution characteristics. A multi-scale pyramid structure is used to construct a scale-space representation of color features. This improved color feature extraction method effectively captures the color characteristics of the helmet and is highly robust to lighting changes and occlusion.

[0034] Step S22: Obtain the illumination-normalized multi-scale image pyramid obtained in step S14 and process it using the scale-invariant and direction-aware LBP algorithm to obtain a multi-scale texture feature set. Specifically, the sampling radius and number of sampling points of the LBP operator are first dynamically determined for each scale of the image; then, a rotation-invariant LBP code is calculated; then, a wavelet transform is applied to perform multi-scale and multi-directional analysis; finally, an LBP histogram is constructed and features at different scales are fused using an adaptive weighting method to obtain a multi-scale texture feature set.

[0035] In this embodiment, a modified Local Binary Pattern (LBP) algorithm is used to extract texture features. This improved algorithm incorporates rotational invariance and scale adaptability. By dynamically adjusting the sampling radius and number of sampling points of the LBP operator, adaptive description of textures of varying scales is achieved. Furthermore, an LBP encoding scheme based on the principal direction is introduced to ensure rotational invariance of the features. Furthermore, by combining wavelet transform technology, multi-scale and multi-directional texture analysis is achieved, further enhancing the discriminative power of texture features. This improved texture feature extraction method effectively captures subtle texture differences on the helmet surface, improving detection accuracy.

[0036] Step S23: Obtain the enhanced depth image and depth edge map obtained in step S15 and process them using a 3D-aware Hough transform and curvature analysis method to obtain a shape feature set. Specifically, candidate circles are first detected using adaptive edge extraction and a multi-scale Hough transform. Then, the principal curvature is calculated and a curvature descriptor is constructed. Next, the 3D surface normal is calculated using local plane fitting. Finally, the circularity metric, curvature features, and normal information are integrated to construct a comprehensive shape descriptor, resulting in a shape feature set.

[0037] In this example, an improved Hough transform algorithm is used to detect circular and elliptical shapes. This improved algorithm incorporates an adaptive voting threshold and a multi-scale analysis strategy, enhancing the detection of partially occluded and deformed helmets. Curvature features are then extracted, and local curvature is calculated using an improved differential geometry method that effectively handles noise and discontinuities in depth maps. Finally, the 3D surface normal is calculated using an improved principal component analysis (PCA) method, incorporating local depth distribution information to improve the accuracy of normal estimation. This improved shape feature extraction method comprehensively describes the 3D geometric characteristics of helmets, providing an important basis for subsequent detection and classification.

[0038] Step S24: Obtain the continuous multi-frame illumination-normalized multi-scale image pyramid obtained in step S14 and the enhanced depth image obtained in step S15, and process them using the depth-enhanced optical flow STIP algorithm to obtain a temporal feature set. Specifically, feature point detection and matching are first performed, followed by estimation of sparse-to-dense optical flow and a depth consistency check. Next, a spatiotemporal volume is constructed and adaptive scale selection is performed to detect spatiotemporal points of interest. HOG, HOF, and MBH features are then extracted. Finally, a modified Fisher Vector encoding method is used to obtain a temporal feature set.

[0039] In this embodiment, an improved sparse-dense optical flow estimation algorithm is used, combining feature point tracking and variational optical flow methods to accurately capture motion information in a scene. Then, an improved spatiotemporal interest point (STIP) detection algorithm is applied, incorporating an adaptive scale selection mechanism and a multimodal fusion strategy to detect and describe significant motion patterns in the spatiotemporal domain. Furthermore, a recurrent neural network (RNN)-based temporal modeling technique is introduced to capture long-term temporal dependencies. This temporal feature extraction method effectively exploits the temporal correlation between consecutive frames, improving the ability to identify helmet wearing status in dynamic scenes.

[0040] Step S25: The multi-scale color feature set, multi-scale texture feature set, shape feature set, and temporal feature set obtained in steps S21 to S24 are obtained and processed using a GCN-guided adaptive feature fusion network method to obtain a fused feature set. Specifically, PCA dimensionality reduction and feature selection based on the Fisher discriminant criterion are first performed on each feature set; then, feature importance learning is performed using an improved AdaBoost algorithm; then, a feature relationship graph is constructed and a GCN layer is designed to model feature relationships; finally, an adaptive weight fusion strategy is used to obtain the final fused feature set.

[0041] Principal component analysis (PCA) is applied to reduce the dimensionality of each feature subset to reduce redundant information. A feature selection algorithm based on the Fisher discriminant criterion is then used to select the most discriminative feature combinations. Finally, an adaptive weight fusion strategy is employed to dynamically adjust the fusion weights based on the importance of different features in the current scenario. This weight adjustment is based on the discriminative ability and stability of the features, and is learned and updated online using an improved AdaBoost algorithm. Furthermore, feature relationship modeling based on a graph convolutional network (GCN) is introduced to capture the potential correlations between different features. This improved feature fusion method fully leverages the complementarity of multiple features to generate a highly discriminative fused feature set, providing a powerful feature representation for subsequent helmet detection and classification.

[0042] Step S3: Take the fused feature set obtained in step S2 and the enhanced depth image obtained in step S1 and process them using a multi-stage hardhat detection and classification method to obtain preliminary hardhat detection and classification results. Specifically, candidate regions are first generated based on the fused feature set and depth information. Then, hardhat-specific features are extracted for each candidate region. Finally, these features are used to classify the hardhat, obtaining the classification result and confidence score for each candidate region.

[0043] Step S31: Obtain the fused feature set from step S2 and the enhanced depth image from step S1, and process them using the depth-aware density clustering RPN algorithm to obtain a set of candidate regions. Specifically, the sliding window size is dynamically adjusted based on the depth information to generate initial candidate regions on the fused feature map. Then, the objectness score of each candidate window is calculated. Next, a depth-based soft non-maximum suppression algorithm is applied to screen candidate regions. Finally, a density clustering algorithm is used to further refine the candidate regions to obtain the final set of candidate regions.

[0044] An improved multi-scale sliding window algorithm is applied, dynamically adjusting the window size based on the depth information in the depth image I_depth_enhanced to ensure that the helmet is adequately covered at all distances. The sliding window operates on the fused feature set F_fused, leveraging the spatial information of the feature map to improve efficiency. An objectness score is then calculated for each window position, based on the local distribution characteristics and depth consistency of the fused features. Next, a modified non-maximum suppression (NMS) algorithm is used to screen candidate regions. This algorithm incorporates a depth-based soft NMS strategy, allowing multiple candidate boxes to be retained in overlapping regions with significant depth differences. Finally, a density clustering algorithm is applied to further refine the candidate regions, merging similar candidate regions and removing isolated false positives. This improved region proposal generation method effectively reduces the computational burden of subsequent processing while maintaining high recall.

[0045] Step S32: Obtain the candidate region set from step S31 and the fused feature set from step S2, and process them using a multimodal pyramid hardhat feature extractor to obtain a hardhat feature set. Specifically, multi-scale spatial pyramid pooling is first performed on each candidate region; then, improved HOG features and color distribution features are extracted; then, a 3D shape descriptor is constructed based on depth information; finally, a hardhat feature vector for each candidate region is obtained through feature selection and combination, forming the hardhat feature set.

[0046] For each candidate region r_i, the corresponding local features are extracted from the fused feature set F_fused. Then, an improved multi-scale spatial pyramid pooling operation is applied to unify the features of candidate regions of different sizes into a fixed dimension. Next, the following three feature extractors are used:

[0047] 1) An improved Histogram of Oriented Gradients (HOG) feature extractor introduces an adaptive unit partitioning strategy to dynamically adjust HOG units based on the shape and size of the candidate region;

[0048] 2) An enhanced color distribution feature extractor that combines local color histograms and global color moments to capture the color characteristics of the helmet;

[0049] 3) An innovative 3D shape descriptor that utilizes the depth information in step S1 to construct a local 3D point cloud and extract shape features based on 3D curvature and normal vector distribution.

[0050] Finally, these three features are optimized and combined using a feature selection algorithm to obtain the helmet feature vector f_h_i for each candidate region r_i. The feature vectors of all candidate regions form the helmet feature set F_helmet = {f_h_1, f_h_2, ..., f_h_m}. This helmet-specific feature extraction method comprehensively captures the helmet's appearance, color, and 3D structure, providing high-quality feature representation for subsequent classification tasks.

[0051] Step S33: Obtain the hardhat feature set from step S32 and process it using the knowledge distillation attention hardhat classification network method to obtain a classification result set. Specifically, a lightweight network structure is first designed that includes depthwise separable convolution and channel attention mechanisms. Then, knowledge distillation techniques are used to transfer knowledge from the pre-trained complex model. Next, an improved Focal Loss is used for training. Finally, an ensemble learning strategy is used to fuse the predictions of multiple models to obtain the classification result and confidence level for each candidate region.

[0052] A lightweight convolutional neural network architecture employs depthwise separable convolutions and a channel-wise attention mechanism, significantly reducing computational complexity while maintaining high accuracy. The network input is the helmet feature vector f_h_i, and the output is a probability distribution over three categories: worn, not worn, and incorrectly worn. Knowledge distillation is then introduced, using a complex model pre-trained on a large-scale dataset to guide the training of the lightweight network, improving its generalization. An improved focal loss is then applied as the loss function to better handle class imbalance. During the inference phase, an ensemble learning strategy is used to weightedly combine the predictions of multiple trained lightweight models, further improving classification accuracy and robustness. Finally, for each candidate region r_i, a classification result c_i is obtained, including the class label and corresponding confidence score. The classification results for all candidate regions form the classification result set C = {c_1, c_2, ..., c_m}. This innovative classification approach significantly reduces computational requirements while maintaining high accuracy, making it suitable for deployment on resource-constrained edge devices.

[0053] Step S4: Take the candidate region set and classification result set obtained in step S3, as well as the enhanced depth image and original RGB image obtained in step S1, and process them using a multi-stage post-processing and result optimization method to obtain the final helmet detection and classification results. Specifically, the method first processes the scene with multiple people overlapping; then optimizes the results using temporal information; then enhances the small target detection results; and finally, generates a visualization result and a detection report.

[0054] Step S41: Obtain the candidate region set from step S3, the classification result set, and the enhanced depth image from step S1, and process them using the depth layered graph cut overlap parsing algorithm to obtain an optimized detection result. Specifically, an improved region growing algorithm is first used for depth layering; then, an adaptive weighted non-maximum suppression algorithm is applied; then, a graph cut algorithm is used to finely segment the overlapping regions; finally, the overlapping regions are re-evaluated and occlusion processed to obtain the optimized detection result.

[0055] The depth image I_depth_enhanced is used to stratify overlapping regions. Using an improved region growing algorithm, starting from the center of each candidate region, regions are expanded based on depth similarity until a significant depth discontinuity is encountered. Then, for each stratified region, an adaptively weighted non-maximum suppression (NMS) algorithm is applied, which takes into account the confidence, overlap, and depth consistency of the candidate boxes. Next, a graph cut algorithm is used to finely segment the overlapping regions into multiple subregions. For each subregion, the classification result is re-evaluated, employing a multimodal fusion strategy based on local features and global context. Finally, an improved occlusion handling algorithm is applied, which leverages partially visible helmet features and prior knowledge of human pose to estimate the state of the helmet in occluded areas. After these processes, the optimized detection results D_opt = {d_1, d_2, ..., d_k} are obtained, where each d_i contains the optimized bounding box, classification label, and confidence. This improved multi-person overlap handling method effectively solves the helmet detection problem in dense scenes, improving detection accuracy and robustness.

[0056] Step S42: Obtain the optimized detection results from step S41 and the historical detection results of multiple consecutive frames, and process them using the spatiotemporal graph Kalman filter (CRF) optimization algorithm to obtain a temporally consistent detection result. Specifically, the improved Kalman filter is first used to track the detection results; then, a spatiotemporal graph model is constructed; then, a graph convolutional network is applied to learn temporal dependencies; then, a conditional random field is used for temporal smoothing; finally, a temporal voting mechanism is used to obtain a temporally consistent detection result.

[0057] Each detection result is tracked using an improved Kalman filter. This improved algorithm incorporates an adaptive noise estimation mechanism to better handle the uncertainty of helmet motion. A spatiotemporal graph model is then constructed, representing detection results from multiple consecutive frames as nodes in the graph, with edges representing the correspondence between frames. A graph convolutional network (GCN) is then applied to learn the spatiotemporal dependencies between nodes, capturing long-term temporal patterns. Furthermore, a temporal voting mechanism is implemented that considers the consistency of detection results in historical frames and the confidence level of the current frame. For sudden changes in detection results, a conditional random field (CRF) model is used for smoothing to ensure temporal consistency. Finally, based on the optimized temporal information, the classification label and confidence level of each detection result are adjusted to obtain a temporally consistent detection result D_temp = {d_t1, d_t2, ..., d_tk}. This temporal consistency optimization method effectively reduces single-frame false detections and missed detections, improving the stability and reliability of detection results.

[0058] Step S43: Take the time-aligned detection results from step S42 and the multi-scale image pyramid from step S1 and process them using an attention-guided super-resolution small object enhancement network to obtain enhanced detection results. Specifically, small objects are first identified and multi-scale features are extracted. Then, an improved super-resolution reconstruction network is used to improve the image quality of small objects. Next, attention-guided feature enhancement is applied. Finally, the small objects are reclassified and the results are integrated to obtain enhanced detection results.

[0059] Small targets in D_temp are identified, defined as detections whose bounding box area is smaller than a predetermined threshold. For each small target, the optimal scale layer is located in the multi-scale image pyramid P_norm. Next, an improved super-resolution reconstruction algorithm is applied to magnify the small target area. This algorithm combines a deep learning-based super-resolution network with an image prior-based reconstruction method to effectively restore the details of the small target. During the reconstruction process, an attention mechanism is introduced to focus on the key feature areas of the helmet. The feature extraction and classification methods in steps S32 and S33 are reapplied to the reconstructed small target image to obtain a more accurate classification result. Finally, the updated small target detection result is integrated with the original result to obtain the enhanced detection result D_final = {d_f1, d_f2, ..., d_fn}. This small target enhancement method can significantly improve the detection of helmets in long-distance or high-altitude work scenarios and increase the overall detection recall rate.

[0060] Step S44: Take the enhanced detection results from step S43 and the original RGB image from step S1 and process them using the intelligent interactive security analysis visualization system method to obtain a visualization result image and detection report. Specifically, high-level image rendering is first performed, including bounding box drawing, label rendering, and heat map generation. Then, statistical information is calculated, including category statistics, time series analysis, and spatial distribution analysis. Next, natural language generation technology is used to generate the detection report. Finally, an interactive visualization component is created to obtain the final visualization result and detection report.

[0061] The detection results are plotted on the original RGB image I_rgb, using bounding boxes of different colors to represent different helmet wearing states (e.g., green for correct wearing, red for not wearing, and yellow for incorrect wearing). A confidence label and ID are then added next to each bounding box. Next, a heat map is generated to show the spatial distribution of helmet wearing conditions. For the detection report, statistics are first calculated for each category, including the correct wearing rate, non-wearing rate, and incorrect wearing rate. Time series data is then analyzed to generate a trend chart of helmet wearing conditions. Next, high-risk areas are identified, i.e., areas with a high proportion of non-wearing or incorrect wearing. Finally, natural language generation technology is used to convert statistical data and analysis results into an easy-to-understand text report. This visualization and report generation method can intuitively display detection results and provide valuable safety management insights, helping construction site managers quickly identify and address safety hazards.

[0062] In another embodiment of the present application, the process of the deep adaptive ICP algorithm is as follows:

[0063] The RGB image point cloud and depth image point cloud are acquired and processed using the depth-adaptive ICP algorithm to obtain a spatially aligned point cloud and transformation matrix. Specifically, the FPFH features are first extracted from the input RGB point cloud and depth point cloud, while the confidence of the depth information is used as a weighting factor for feature calculation. Then, a nearest neighbor search is performed using a kd-tree to establish initial point correspondences, and a depth consistency check is introduced to remove corresponding point pairs with excessive depth differences. Next, an iterative optimization is performed, including calculating the weights of corresponding point pairs based on depth value reliability and feature similarity, solving the transformation matrix using the weighted SVD method, applying the transformation and updating the correspondences, and using a dynamic search radius and adaptive threshold for convergence checks. Finally, local optimization is performed on high-curvature areas, and Gaussian process regression is used to smooth the transformation to obtain the final transformation matrix and spatially aligned point cloud.

[0064] The specific data processing process is as follows:

[0065] Get the RGB image point cloud P_rgb and the depth image point cloud P_depth; initialize the transformation matrix T to the identity matrix.

[0066] Extract FPFH (Fast Point Feature Histograms) features for P_rgb and P_depth respectively. Use an improved FPFH algorithm and add the confidence level of depth information as a weighting factor in feature calculation.

[0067] Initial correspondence establishment: Use the kd-tree to perform nearest neighbor search to establish the initial point correspondence. Introduce depth consistency check to eliminate corresponding point pairs with large depth differences.

[0068] The iterative optimization process is as follows:

[0069] For each iteration:

[0070] a) Calculate the weights of corresponding point pairs:

[0071] Reliability σ based on depth value d Calculate the weight w with feature similarity s: w = exp(-||p rgb -p depth || 2 / (2σ d 2 )) s; depth reliability σ d Estimation by local depth variance: σ d = 1 / (1 + var(local_depth));

[0072] b) Weighted least squares solution:

[0073] Use the weighted SVD (singular value decomposition) method to solve the transformation matrix T;

[0074] Minimize the weighted error function: E = Σ w i ||Tp_rgb_i - p_depth_i|| 2 ;

[0075] c) Apply transformation: Apply the obtained transformation matrix T to P_rgb;

[0076] d) Update correspondences: Use the updated P_rgb to find the nearest neighbor again. Apply a dynamic search radius, adaptively adjusting it based on the average registration error of the current iteration.

[0077] e) Convergence check:

[0078] Calculate the average registration error of the current iteration and use the adaptive threshold to judge convergence: threshold = max(base_threshold, median_error scale_factor);

[0079] If the average error is less than the threshold or the maximum number of iterations is reached, the iteration is stopped.

[0080] Apply a local optimization strategy to perform additional fine alignment on high curvature areas and use Gaussian ProcessRegression to smooth the transformation and reduce overfitting. Output the final transformation matrix T and the aligned point cloud

[0081] By introducing an adaptive weighting mechanism and reliable estimation of depth information, the registration accuracy and robustness are significantly improved in complex scenes such as construction sites. In particular, better alignment results are achieved when dealing with areas with occlusions, reflective surfaces, or unreliable depth sensors.

[0082] In another embodiment of the present application, the multi-scale spatiotemporal DTW algorithm has the following specific process:

[0083] RGB image sequences and depth image sequences are acquired and processed using a multi-scale spatiotemporal DTW algorithm to obtain the time-aligned sequences and timestamp mapping functions. Specifically, the input RGB and depth image sequences are first decomposed at multiple scales to generate an image pyramid structure. Then, an improved spatiotemporal interest point detector is used to extract temporal features, and the feature representation is enhanced by combining depth information. Next, multi-scale DTW calculations are performed from coarse to fine scales, including initializing the DTW matrix, introducing adaptive bandwidth constraints to calculate the cost matrix, using an improved step size mode and adaptive slope constraints for dynamic programming, and projecting the coarse-scale alignment path to a finer scale as the initial path. Then, sub-frame interpolation technology is used at the finest scale for precise alignment, and a Kalman filter is applied to smooth the final alignment path. Finally, a timestamp mapping relationship is established based on the alignment path, and piecewise linear interpolation is used to process the mapping of non-integer frames, resulting in the time-aligned sequences and mapping functions.

[0084] The specific data processing process is as follows:

[0085] Get the RGB image sequence X = {x_1, ..., x_n} and the depth image sequence Y = {y_1, ..., y_m}.

[0086] Perform multi-scale decomposition on the input sequence and generate pyramid structures X_pyramid and Y_pyramid.

[0087] Temporal features are extracted for each frame in X and Y. An improved spatiotemporal interest point (STIP) detector is used to enhance feature representation by combining depth information.

[0088] Multi-scale DTW calculation:

[0089] At each scale level l from coarse to fine, perform:

[0090] a) Initialize the DTW matrix D_l and the path matrix P_l.

[0091] b) Calculate the cost matrix C_l:

[0092] Introduce adaptive bandwidth constraint: w_l = base_w (2^l), where base_w is the base bandwidth.

[0093] Compute the cost C_l(i,j) = dist(X_l[i], Y_l[j]) + penalty(i,j) within bandwidth w_l.

[0094] penalty(i,j) is a penalty term based on depth consistency.

[0095] c) Dynamic programming to fill the DTW matrix:

[0096] Use the modified step pattern: {(1,1), (1,2), (2,1), (1,3), (3,1)}.

[0097] Apply an adaptive slope constraint: max_slope_l = base_slope (0.5^l).

[0098] d) Backtrack to obtain the optimal path P_l.

[0099] e) Project P_l to the next finer scale as the initial path.

[0100] At the finest scale, subframe interpolation techniques are used for precise alignment.

[0101] Apply Kalman filtering to smooth the final alignment path and reduce jitter.

[0102] Based on the final alignment path, a timestamp mapping relationship between X and Y is established.

[0103] Mapping of non-integer frames is handled using piecewise linear interpolation.

[0104] Outputs the time-aligned sequences X' and Y', and the timestamp mapping function between them.

[0105] By introducing multi-scale analysis and adaptive constraints, the efficiency and accuracy of processing multimodal data with different sampling rates and time delays are significantly improved. In particular, better alignment results can be achieved when processing long sequences or real-time streaming data.

[0106] In another embodiment of the present application, the edge enhancement multi-directional Laplacian pyramid algorithm has the following process:

[0107] The original image is processed using an edge-enhanced multi-directional Laplacian pyramid algorithm to produce a multi-scale image pyramid. Specifically, a Gaussian pyramid is first generated and Gaussian filtering is performed using an adaptive kernel size based on local variance. Next, Laplacian differences are calculated and an adaptive edge enhancement filter is applied, where the enhancement factor is dynamically adjusted based on the local edge strength. The Laplacian differences at each layer are then decomposed in a multi-directional manner using a discrete wavelet transform, and features in the main directions are enhanced using a directional lifting wavelet. Finally, the enhanced image is reconstructed layer by layer, starting from the top layer, to produce a multi-scale image pyramid that contains the original resolution image and directionally enhanced details.

[0108] The specific data processing process is as follows:

[0109] Get the original image I_0, for each layer l (l = 0, 1, ..., n-1):

[0110] a) Apply adaptive Gaussian filtering:

[0111] Use adaptive kernel size based on local variance: σ_l(x,y) = base_σ (1 + α var(local_patch)));

[0112] G_l+1 = downsample(gaussianFilter(G_l, σ_l));

[0113] Laplacian pyramid generation:

[0114] For each layer l (l = 0, 1, ..., n-1):

[0115] a) Calculate the Laplace difference:

[0116] L_l = G_l - upsample(G_l+1);

[0117] b) Apply an adaptive edge enhancement filter:

[0118] Calculate local edge strength: E_l = sobelFilter(G_l);

[0119] Adaptive enhancement factor: α_l(x,y) = base_α (1 + β E_l(x,y));

[0120] L_l_enhanced = L_l + α_l L_l;

[0121] Multi-directional decomposition based on wavelet transform:

[0122] For each layer of Laplace difference L_l_enhanced:

[0123] a) Apply discrete wavelet transform (DWT): [LL, LH, HL, HH] = DWT(L_l_enhanced);

[0124] b) Directional enhancement of LH, HL, and HH sub-bands:

[0125] Use directional lifting wavelet to enhance the four main directions of 0°, 45°, 90°, and 135°;

[0126] c) L_l_directional after reconstruction and enhancement by inverse discrete wavelet transform (IDWT).

[0127] Pyramid reconstruction:

[0128] Starting from the top layer, reconstruct the enhanced image layer by layer, I_reconstructed = G_n + sum(upsample(L_l_directional)) for l = n-1 to 0.

[0129] Output multi-scale image pyramid P = {I_0, I_1, ..., I_n}, where each layer contains the original resolution image and directionally enhanced details.

[0130] By introducing adaptive edge enhancement and multi-directional decomposition, the quality of multi-scale representation is significantly improved, especially in terms of preserving edge details and capturing directional features. This provides richer and more discriminative feature representations for subsequent hardhat detection, especially when dealing with hardhats of different scales and orientations.

[0131] In another embodiment of the present application, thermal imaging guides the CLAHE algorithm, and the specific process is as follows:

[0132] Images and corresponding thermal imaging data from a multi-scale image pyramid are acquired and processed using a thermal imaging-guided CLAHE algorithm to obtain an illumination-normalized image. Specifically, the thermal imaging data is first converted into illumination intensity estimates using a nonlinear mapping function and smoothed using guided filtering. The CLAHE block size is then adaptively determined based on the gradient information of the illumination intensity map. Next, a local histogram is calculated and an adaptive clipping threshold based on local illumination intensity is applied. Finally, a local enhancement factor is calculated based on illumination intensity for adaptive contrast enhancement. Finally, pixel-level brightness adjustment is performed based on the illumination intensity map to obtain an illumination-normalized image.

[0133] The specific data processing process is as follows:

[0134] Obtain the image I_i and the corresponding thermal imaging data I_thermal in the multi-scale image pyramid P.

[0135] a) Apply a nonlinear mapping function f to convert the thermal imaging data into light intensity estimates:

[0136] M_light = f(I_thermal), where f is the sigmoid-based adaptive mapping function;

[0137] b) Use guided filtering to smooth M_light and preserve edge information;

[0138] Based on the gradient information of M_light, the block size of CLAHE is adaptively determined; smaller blocks are used in areas with drastic lighting changes, and larger blocks are used instead;

[0139] Local histogram calculation:

[0140] For each chunk:

[0141] a) Calculate the local histogram.

[0142] b) Apply adaptive clipping threshold based on M_light:

[0143] clip_limit = base_clip (1 + γ mean(M_light_local)).

[0144] c) Reallocate the portion that exceeds clip_limit.

[0145] Contrast Adaptive Enhancement:

[0146] For each pixel (x,y):

[0147] a) Calculate the local enhancement factor based on M_light:

[0148] α(x,y) = base_α (1 - δ M_light(x,y)).

[0149] b) Apply adaptive enhancement: I_enhanced(x,y) = I_original(x,y) + α(x,y) (I_clahe(x,y) - I_original(x,y)).

[0150] Perform pixel-level brightness adjustment based on M_light: I_final(x,y) = I_enhanced(x,y) (1 + ε(1 - M_light(x,y))).

[0151] Output the illumination normalized image I_norm_i.

[0152] By introducing an adaptive mechanism guided by thermal imaging data, the algorithm can better handle image enhancement issues under complex lighting conditions. It effectively improves visibility in low-light areas while preserving image details, providing higher-quality input data for subsequent helmet detection tasks.

[0153] In another embodiment of the present application, the depth adaptive bilateral filtering algorithm is specifically as follows:

[0154] A depth image is acquired and processed using a depth-adaptive bilateral filtering algorithm to obtain a denoised depth image and depth edge map. Specifically, the local depth mean and standard deviation are first calculated for each pixel to estimate local depth reliability. Then, an adaptive kernel function is designed, in which the parameters of the depth value kernel function are dynamically adjusted based on the local depth reliability. Next, an improved bilateral filter is applied using adaptive weights for filtering. Finally, a depth edge map is calculated. Finally, edge-preserving enhancement is performed using edge-guided filtering to obtain a denoised depth image and depth edge map.

[0155] The specific data processing process is as follows:

[0156] Get the depth image I_depth and perform statistical analysis on the local depth:

[0157] For each pixel (x,y):

[0158] a) Calculate the local depth mean μ_d and standard deviation σ_d

[0159] b) Estimate local depth reliability: r(x,y) = 1 / (1 + k σ_d / μ_d)

[0160] Adaptive kernel function design:

[0161] a) Spatial kernel function: G_s(x,y) = exp(-(x^2+y^2) / (2σ_s^2))

[0162] b) Depth kernel function: G_r(d) = exp(-d^2 / (2σ_r(x,y)^2))

[0163] Among them, σ_r(x,y) = base_σ_r (1 + λ (1-r(x,y)))

[0164] Improved bilateral filtering:

[0165] For each pixel (x,y):

[0166] a) Define the local window W(x,y)

[0167] b) Calculate adaptive weights:

[0168] w(i,j) = G_s(ix, jy) G_r(I_depth(i,j) - I_depth(x,y)) r(i,j)

[0169] c) Apply filtering:

[0170] I_filtered(x,y) = Σ(i,j)∈W(x,y) w(i,j) I_depth(i,j) / Σ(i,j)∈W(x,y) w(i,j)

[0171] Edge Preserving Enhancement:

[0172] Calculate the depth edge map: E_depth = sobelFilter(I_filtered)

[0173] Apply edge-directed filtering:

[0174] I_enhanced = guidedFilter(I_filtered, E_depth)

[0175] Output the depth image I_depth_filtered after noise removal and the depth edge map E_depth.

[0176] By introducing an adaptive kernel function based on the local depth distribution, we can better preserve depth edges while effectively removing noise. This achieves particularly good filtering results in areas prone to noise from depth sensors, such as object edges or reflective surfaces. This provides more reliable input data for subsequent depth-based helmet detection.

[0177] In another embodiment of the present application, the adaptive multi-space color descriptor process is specifically as follows:

[0178] A multi-scale image pyramid is obtained after illumination normalization and processed using an adaptive multi-space color descriptor method to obtain a multi-scale color feature set. Specifically, the discriminative scores of multiple color spaces (such as RGB, HSV, and Lab) are first calculated for each image scale, and the color space with the highest score is selected. Then, in the selected color space, local color moments and color covariance matrices are calculated to construct an enhanced color descriptor. Finally, features from different scales are integrated using a weighted average method to obtain the final multi-scale color feature set.

[0179] The specific data processing process is as follows:

[0180] Get the illumination normalized multi-scale image pyramid P_norm = {I_norm_0, I_norm_1, ...,I_norm_n}.

[0181] Adaptive color space selection:

[0182] For each scale image I_norm_i:

[0183] a) Calculate the discriminative scores of multiple color spaces:

[0184] S_j = Σ(inter_class_variance_j) / Σ(intra_class_variance_j).

[0185] Among them, j represents different color spaces (such as RGB, HSV, Lab, etc.),

[0186] inter_class_variance_j and intra_class_variance_j represent the between-class variance and within-class variance, respectively.

[0187] b) Select the color space with the highest score:

[0188] best_space = argmax_j(S_j)

[0189] Improved local color descriptor extraction:

[0190] For each image I_color in the chosen color space:

[0191] a) Calculate the local color moment:

[0192] M_k = (1 / N) Σ(x,y)∈W (I_color(x,y) - μ)^k.

[0193] Where W is the local window, N is the number of pixels in the window, μ is the average color in the window, and k is the order of the moment (usually 1, 2, or 3).

[0194] b) Calculate the color covariance matrix:

[0195] C = (1 / N) Σ(x,y)∈W (I_color(x,y) - μ)(I_color(x,y) - μ)^T.

[0196] c) Constructing enhanced color descriptors:

[0197] F_color = [M_1, M_2, M_3, eigenvalues(C), eigenvectors(C)].

[0198] Multi-scale feature integration:

[0199] Use weighted averaging to integrate features of different scales:

[0200] F_color_final = Σ(w_i F_color_i) / Σ(w_i).

[0201] Among them, w_i is the weight of the i-th scale, which can be dynamically adjusted according to the importance of the scale.

[0202] Output multi-scale color feature set F_color = {F_color_0, F_color_1, ..., F_color_n}.

[0203] By adaptively selecting the most discriminative color space and combining it with high-order color statistics, the color features of the helmet can be captured more comprehensively, improving the distinguishing ability and robustness of the features.

[0204] In another embodiment of the present application, the scale-invariant direction-aware LBP algorithm has the following process:

[0205] A multi-scale image pyramid is obtained after illumination normalization and processed using the scale-invariant and direction-aware LBP algorithm to obtain a multi-scale texture feature set. Specifically, the sampling radius and number of sampling points of the LBP operator are first dynamically determined for each scale of the image. Then, a rotation-invariant LBP code is calculated. Next, a wavelet transform is applied for multi-scale and multi-directional analysis, and a directional LBP is calculated for each sub-band. Finally, an LBP histogram is constructed and the histograms of multiple sub-bands are concatenated. Finally, an adaptive weighting method is used to fuse features at different scales to obtain a multi-scale texture feature set.

[0206] The specific data processing process is as follows:

[0207] Get the illumination normalized multi-scale image pyramid P_norm = {I_norm_0, I_norm_1, ...,I_norm_n}.

[0208] Adaptive LBP operator design:

[0209] For each scale image I_norm_i:

[0210] a) Dynamically determine the sampling radius R and the number of sampling points P:

[0211] R = base_R scale_factor_i

[0212] P = base_P scale_factor_i

[0213] Among them, base_R and base_P are basic parameters, and scale_factor_i is the scaling factor of the current scale

[0214] Rotation-invariant LBP encoding:

[0215] For each pixel (x,y) in the image:

[0216] a) Calculate the basic LBP code:

[0217] LBP_P,R = Σ(p=0 to P-1) s(g_p - g_c) 2^p

[0218] Among them, g_c is the gray value of the center pixel, g_p is the gray value of the neighboring pixel,

[0219] s(x) = 1 if x ≥ 0, else 0.

[0220] b) Find the minimum value to obtain the rotation-invariant LBP code:

[0221] LBP_ri_P,R = min(ROR(LBP_P,R, i) | i = 0, 1, ..., P-1).

[0222] Here, ROR(x, i) means rotating x right by i bits.

[0223] Multi-scale and multi-directional analysis:

[0224] a) Apply wavelet transform to perform multi-scale decomposition:

[0225] [LL, LH, HL, HH] = DWT(I_norm_i)

[0226] b) Calculate the directional LBP on each subband:

[0227] LBP_dir_k = LBP_ri_P,R(subband_k), k ∈ {LL, LH, HL, HH}.

[0228] Feature vector construction:

[0229] a) Calculate the LBP histogram of each subband:

[0230] H_k = histogram(LBP_dir_k), k ∈ {LL, LH, HL, HH}.

[0231] b) Concatenate histograms of multiple subbands:

[0232] F_texture_i = [H_LL, H_LH, H_HL, H_HH].

[0233] Use adaptive weighting method to fuse features of different scales:

[0234] F_texture_final = Σ(α_i F_texture_i) / Σ(α_i).

[0235] Among them, α_i is the weight of the i-th scale, which can be dynamically adjusted through feature importance analysis.

[0236] Output multi-scale texture feature set F_texture = {F_texture_0, F_texture_1, ..., F_texture_n}.

[0237] By introducing rotation invariance and scale adaptability and combining multi-scale and multi-directional analysis of wavelet transform, the subtle texture differences on the surface of the helmet can be captured more effectively, and the discriminative ability and scale invariance of the features are improved.

[0238] In another embodiment of the present application, the 3D perception Hough transform and curvature analysis process is specifically as follows:

[0239] The enhanced depth image and depth edge map are acquired and processed using a 3D-aware Hough transform and curvature analysis method to obtain a shape feature set. Specifically, adaptive edge extraction and multi-scale Hough transform are first performed to detect candidate circles. Then, the principal curvatures are calculated and a curvature descriptor is constructed, including mean curvature, Gaussian curvature, curvature intensity, and shape index. Next, the 3D surface normal is calculated through local plane fitting and normal smoothing is performed. Finally, the circularity metric, curvature features, and normal information are integrated to construct a comprehensive shape descriptor, resulting in a shape feature set.

[0240] Get the enhanced depth image I_depth_enhanced and depth edge map E_depth.

[0241] Hough transform detection:

[0242] a) Adaptive edge extraction:

[0243] E_adaptive = adaptiveThreshold(E_depth, window_size, C)

[0244] Among them, window_size is the adaptive window size, C is the constant offset

[0245] b) Multi-scale Hough transform:

[0246] For each scale s:

[0247] H_s = HoughTransform(E_adaptive, scale=s)

[0248] Peaks_s = findPeaks(H_s, threshold=T_s)

[0249] Among them, H_s is the Hough transform result, T_s is the scale-related threshold

[0250] c) Candidate circle detection:

[0251] Circles = detectCircles(Peaks_s, min_radius, max_radius)

[0252] Curvature feature extraction:

[0253] a) Calculate the principal curvatures:

[0254] For each point p:

[0255] [k1, k2] = principalCurvatures(I_depth_enhanced, p)

[0256] Where k1 and k2 are the maximum and minimum principal curvatures

[0257] b) Construct curvature descriptor:

[0258] C_feature = [H, K, S, C]

[0259] Where H = (k1 + k2) / 2 is the mean curvature, K = k1 k2 is the Gaussian curvature,

[0260] S = sqrt(k1^2 + k2^2) is the curvature strength, C = (k1 - k2) / (k1 + k2) is the shape index

[0261] 3D surface normal calculation:

[0262] a) Local plane fitting:

[0263] For each point p and its neighborhood N(p):

[0264] [n_x, n_y, n_z] = fitPlane(N(p))

[0265] Among them, [n_x, n_y, n_z] is the normal vector of the fitting plane

[0266] b) Normal vector smoothing:

[0267] n_smooth = gaussianFilter(n, σ)

[0268] Where σ is the standard deviation of the Gaussian filter

[0269] Shape feature integration:

[0270] a) Circularity Measurement:

[0271] Circularity = 4π Area / Perimeter^2

[0272] b) Constructing a comprehensive shape descriptor:

[0273] F_shape = [Circularity, C_feature, n_smooth]

[0274] Output shape feature set F_shape.

[0275] By combining multi-scale Hough transform, accurate curvature calculation and stable normal vector estimation, the 3D geometric characteristics of the helmet can be fully described, providing strong feature support for subsequent detection and classification.

[0276] In another embodiment of the present application, the depth enhanced optical flow STIP algorithm has the following process:

[0277] The illumination-normalized multi-scale image pyramid and enhanced depth image of multiple consecutive frames are acquired and processed using the depth-enhanced optical flow (STIP) algorithm to obtain a temporal feature set. Specifically, feature point detection and matching are performed first, followed by estimation of sparse-to-dense optical flow and a depth consistency check. Next, a spatiotemporal volume is constructed and adaptively scaled to detect spatiotemporal points of interest. HOG, HOF, and MBH features are then extracted for each detected STIP point. Finally, a modified Fisher Vector encoding method is used to obtain the temporal feature set.

[0278] The specific data processing process is as follows:

[0279] Get the illumination normalized multi-scale image pyramid P_norm and enhanced depth image I_depth_enhanced of multiple consecutive frames

[0280] Improved optical flow estimation:

[0281] For two consecutive frames t and t+1:

[0282] a) Feature point detection and matching:

[0283] Keypoints_t = FAST(I_norm_t);

[0284] Keypoints_t+1 = FAST(I_norm_t+1);

[0285] Matches = KLT_tracker(Keypoints_t, I_norm_t, I_norm_t+1);

[0286] b) Sparse to dense optical flow estimation:

[0287] Flow_sparse = calcOpticalFlowPyrLK(I_norm_t, I_norm_t+1, Keypoints_t);

[0288] Flow_dense = interpolateFlow(Flow_sparse, I_norm_t);

[0289] c) Deep consistency check:

[0290] Flow_refined = refineFlow(Flow_dense, I_depth_t, I_depth_t+1);

[0291] Among them, the refineFlow function adjusts the optical flow estimation based on depth consistency;

[0292] Improved Spatio-Temporal Interest Point (STIP) Detection:

[0293] a) Constructing space-time volume:

[0294] V = stack(I_norm_t-n, ..., I_norm_t, ..., I_norm_t+n)

[0295] b) Adaptive scale selection:

[0296] For each spatial position (x,y) and time point t:

[0297] σ^2 = argmax_s (s^2 |det(μ(x,y,t;s))|)

[0298] Among them, μ is the space-time second-order moment matrix, s is the scale parameter;

[0299] c) STIP response function calculation:

[0300] R = det(μ) - k trace(μ)^3

[0301] Where k is an empirical constant, usually 0.04-0.06

[0302] d) Non-maximum suppression:

[0303] STIP = NMS(R, window_size_space, window_size_time);

[0304] Timing feature description:

[0305] For each detected STIP point:

[0306] a) Extract HOG features:

[0307] HOG = computeHOG(V, STIP, σ_spatial);

[0308] b) Extract HOF features:

[0309] HOF = computeHOF(Flow_refined, STIP, σ_temporal);

[0310] c) Construct MBH (Motion Boundary Histogram) features:

[0311] MBH_x = computeHOG(dFlow_x / dx, STIP, σ_mbh);

[0312] MBH_y = computeHOG(dFlow_y / dy, STIP, σ_mbh);

[0313] Among them, dFlow_x / dx and dFlow_y / dy are the spatial derivatives of optical flow.

[0314] Temporal feature encoding:

[0315] Use the improved Fisher Vector encoding method:

[0316] F_temporal = FisherVector([HOG, HOF, MBH_x, MBH_y], GMM).

[0317] Among them, GMM is a Gaussian mixture model, which is used for feature encoding.

[0318] Output temporal feature set F_temporal.

[0319] By combining robust optical flow estimation and adaptive STIP detection, the motion characteristics of the helmet in the temporal dimension can be effectively captured, improving the ability to recognize the wearing status of the helmet in dynamic scenes.

[0320] In another embodiment of the present application, the GCN guides the adaptive feature fusion network, and the process is specifically as follows:

[0321] Multi-scale color, texture, shape, and temporal feature sets are acquired and processed using a GCN-guided adaptive feature fusion network to produce a fused feature set. Specifically, each feature set is first subjected to PCA dimensionality reduction and feature selection based on the Fisher discriminant criterion. Feature importance is then learned using an improved AdaBoost algorithm. A feature relationship graph is then constructed and a GCN layer is designed to model feature relationships. The GCN is then trained to minimize task-related loss and regularization terms. Finally, an adaptive weight fusion strategy is employed to combine feature importance scores with the GCN output to produce the final fused feature set.

[0322] The specific data processing process is as follows:

[0323] Get the multi-scale color feature set F_color, the multi-scale texture feature set F_texture, the shape feature set F_shape, and the temporal feature set F_temporal

[0324] Feature dimensionality reduction and selection:

[0325] a) Apply principal component analysis (PCA) for preliminary dimensionality reduction:

[0326] F_i_reduced = PCA(F_i, n_components_i)

[0327] Among them, F_i represents each feature set, and n_components_i is the number of principal components retained

[0328] b) Feature selection based on Fisher discriminant criterion:

[0329] J(f) = (μ_1 - μ_2)^2 / (σ_1^2 + σ_2^2)

[0330] F_i_selected = selectTopK(F_i_reduced, J, K_i)

[0331] Among them, μ_1, μ_2 are the means of the two types of samples on feature f, σ_1^2, σ_2^2 are the variances, and K_i is the number of selected features

[0332] Feature Importance Assessment:

[0333] Use the improved AdaBoost algorithm for feature importance learning:

[0334] Initialize sample weights: w_t(i) = 1 / N, i = 1, ..., N

[0335] For each round t = 1, ..., T:

[0336] a) Train the weak classifier h_t(x)

[0337] b) Calculate error rate: ε_t = Σ(w_t(i) I(h_t(x_i) ≠ y_i)) / Σ(w_t(i))

[0338] c) Calculate the classifier weight: α_t = 0.5 log((1 - ε_t) / ε_t)

[0339] d) Update sample weights: w_t+1(i) = w_t(i) exp(-α_t y_i h_t(x_i)) / Z_t

[0340] Among them, Z_t is the normalization factor

[0341] Calculate feature importance score: S(f) = Σ(α_t I(f used in h_t))

[0342] Graph Convolutional Network (GCN) feature relationship modeling:

[0343] a) Constructing feature relationship graph:

[0344] G = (V, E), where V is the node set (features) and E is the edge set (feature relationships)

[0345] A_ij = similarity(F_i, F_j), construct the adjacency matrix A

[0346] b) Design the GCN layer:

[0347] H^(l+1) = σ(D^(-1 / 2) AD^(-1 / 2) H^(l) W^(l))

[0348] Where D is the degree matrix, H^(l) is the feature representation of the lth layer, W^(l) is the learnable weight matrix, and σ is the activation function

[0349] c) Training GCN:

[0350] Minimize the loss function: L = L_task + λ L_reg

[0351] Among them, L_task is the task-related loss (such as classification cross entropy), and L_reg is the regularization term

[0352] Adaptive weight fusion:

[0353] a) Calculate the initial weight:

[0354] w_i = softmax(S(F_i_selected))

[0355] b) Apply GCN output for weight adjustment:

[0356] w_i_adjusted = w_i (1 + β GCN_output_i)

[0357] Among them, β is the adjustment factor

[0358] c) Feature fusion:

[0359] F_fused = Σ(w_i_adjusted F_i_selected) / Σ(w_i_adjusted)

[0360] Output fused feature set F_fused.

[0361] By introducing feature importance learning and graph convolutional networks, we can adaptively adjust the weights of different features and capture the potential relationships between features. The technical effects are as follows:

[0362] 1. Use the improved AdaBoost algorithm to evaluate feature importance, taking into account the contribution of features in the classification task.

[0363] 2. The introduction of graph convolutional networks to model the complex relationships between features improves the effectiveness of feature fusion.

[0364] 3. Adopting an adaptive weight fusion strategy, combining feature importance scores and GCN output, a more flexible and robust feature fusion is achieved.

[0365] By fully leveraging the complementarity of multiple features, a highly discriminative fusion feature set is generated, providing a powerful feature representation for subsequent helmet detection and classification tasks. This is especially true when dealing with complex construction site scenarios, as it can better adapt to different environmental conditions and changes in detection targets.

[0366] In another embodiment of the present application, the depth-aware density clustering RPN algorithm has the following specific process:

[0367] The fused feature set and enhanced depth image are obtained and processed using the depth-aware density clustering (RPN) algorithm to obtain a set of candidate regions. Specifically, the sliding window size is first dynamically adjusted based on the depth information to generate initial candidate regions on the fused feature map. Then, the objectness score of each candidate window is calculated, taking into account feature distribution characteristics, depth consistency, and edge density. Next, a depth-based soft non-maximum suppression algorithm is applied to screen candidate regions. Finally, a distance matrix is ​​constructed and density clustering is performed using the DBSCAN algorithm. Similar candidate regions are merged to obtain the final set of candidate regions.

[0368] The specific data processing process is as follows:

[0369] Get the fused feature set F_fused and the enhanced depth image I_depth_enhanced

[0370] Multi-scale sliding window:

[0371] a) Dynamically adjust the window size based on depth information:

[0372] W(x, y) = base_W (f_depth(x, y) / f_ref)

[0373] Among them, W(x, y) is the window size at position (x, y), base_W is the base window size,

[0374] f_depth(x, y) is the depth value at (x, y), and f_ref is the reference depth value

[0375] b) Sliding window on the fused feature map:

[0376] R_initial = slidingWindow(F_fused, W, stride)

[0377] Among them, stride is the sliding step size, which can be dynamically adjusted according to computing resources

[0378] Objectness score calculation:

[0379] For each candidate window r in R_initial:

[0380] a) Calculate characteristic distribution characteristics:

[0381] μ_r = mean(F_fused(r))

[0382] Σ_r = cov(F_fused(r))

[0383] b) Calculate depth consistency:

[0384] D_r = 1 - (std(I_depth_enhanced(r)) / mean(I_depth_enhanced(r)))

[0385] c) Calculate the Objectness score:

[0386] S_r = w1 det(Σ_r) + w2 D_r + w3 edge_density(r)

[0387] Among them, w1, w2, w3 are weight coefficients, and edge_density(r) is the edge density within the window

[0388] Improved non-maximum suppression (NMS):

[0389] a) Depth-based soft NMS:

[0390] For each pair of overlapping windows (r_i, r_j):

[0391] S_i = S_i exp(-overlap(r_i, r_j)^2 / σ^2) if depth_diff(r_i, r_j) <threshold

[0392] Among them, overlap() calculates IoU, depth_diff() calculates the average depth difference, and σ is the parameter that controls the attenuation speed

[0393] b) Keep the K candidate regions with the highest scores:

[0394] R_nms = topK(R_initial, S, K)

[0395] Density clustering optimization:

[0396] a) Construct a distance matrix:

[0397] D_ij = w_s spatial_dist(r_i, r_j) + w_f feature_dist(F_fused(r_i), F_fused(r_j))

[0398] Among them, spatial_dist() calculates spatial distance, feature_dist() calculates feature distance, w_s and w_f are weights

[0399] b) Apply the DBSCAN algorithm:

[0400] Clusters = DBSCAN(R_nms, D, eps, min_samples)

[0401] Among them, eps is the neighborhood radius, min_samples is the minimum number of samples to form a core point

[0402] c) Merge similar candidate regions:

[0403] For each cluster C_k, merge the candidate regions:

[0404] r_merged = merge(C_k), the merge() function can be a weighted average or select the region with the highest score

[0405] Output candidate region set R = {r_1, r_2, ..., r_m}.

[0406] By combining depth information and density clustering, we can more accurately locate areas that may contain helmets while effectively reducing redundant candidate areas.

[0407] 1. Use depth information to dynamically adjust the sliding window size to adapt to targets at different distances.

[0408] 2. Introduce an objectness scoring mechanism based on feature distribution and depth consistency.

[0409] 3. Adopt a depth-based soft NMS strategy to retain candidate regions that may be incorrectly suppressed by traditional NMS.

[0410] 4. The density clustering algorithm is used to further optimize the candidate regions, thereby improving the quality and diversity of the candidate regions.

[0411] In another embodiment of the present application, a multimodal pyramid helmet feature extractor is constructed as follows:

[0412] A set of candidate regions and a fused feature set are obtained and processed using a multimodal pyramid hardhat feature extractor to obtain a hardhat feature set. Specifically, multi-scale spatial pyramid pooling is first performed on each candidate region. Then, improved HOG features are extracted using an adaptive unit partitioning strategy. Next, color distribution features are calculated, including local color histograms and global color moments. A 3D shape descriptor, including FPFH features and 3D curvature, is constructed based on depth information. Finally, feature selection is performed using PCA dimensionality reduction and random forest feature importance evaluation to obtain a hardhat feature vector for each candidate region, forming the hardhat feature set.

[0413] The specific data processing process is as follows:

[0414] Get the candidate region set R and the fused feature set F_fused

[0415] Multi-scale spatial pyramid pooling:

[0416] For each candidate region r_i in R:

[0417] a) Constructing a spatial pyramid:

[0418] levels = [1x1, 2x2, 4x4]

[0419] b) Adaptive Pooling:

[0420] F_pyramid = []

[0421] For each level in levels:

[0422] F_level = adaptivePooling(F_fused(r_i), level)

[0423] F_pyramid.append(F_level)

[0424] Among them, the adaptivePooling() function adaptively performs pooling operations based on the input size and target level

[0425] Improved Histogram of Oriented Gradients (HOG) feature extraction:

[0426] a) Adaptive cell partitioning:

[0427] cell_size = base_size sqrt(area(r_i) / base_area)

[0428] b) Calculate the gradient direction and magnitude:

[0429] [magnitude, orientation] = computeGradient(I(r_i))

[0430] c) Construct a direction histogram:

[0431] For each cell:

[0432] hist = weightedHistogram(magnitude, orientation, n_bins)

[0433] Among them, weightedHistogram() uses bilinear interpolation for voting

[0434] d) Block Normalization:

[0435] For each block (2x2 cells):

[0436] block_feature = normalize(concatenate(hist_cells))

[0437] e) Construct HOG features:

[0438] F_hog = concatenate(block_features)

[0439] Color distribution feature extraction:

[0440] a) Calculate the local color histogram:

[0441] hist_local = colorHistogram(I(r_i), n_bins_color)

[0442] b) Calculate the global color moment:

[0443] μ_color = mean(I(r_i))

[0444] Σ_color = cov(I(r_i))

[0445] c) Constructing color features:

[0446] F_color = [hist_local, μ_color, eigenvalues(Σ_color)]

[0447] 3D shape descriptor extraction:

[0448] a) Constructing a local 3D point cloud:

[0449] P_3d = depthTo3D(I_depth_enhanced(r_i))

[0450] b) Calculate the local surface normal:

[0451] N = estimateNormals(P_3d, k_neighbors)

[0452] c) Extract FPFH (Fast Point Feature Histograms) features:

[0453] F_fpfh = computeFPFH(P_3d, N, radius)

[0454] d) Calculate 3D curvature:

[0455] [k1, k2] = computePrincipalCurvatures(P_3d, N)

[0456] K = k1 k2 / / Gaussian curvature

[0457] H = (k1 + k2) / 2 / / mean curvature

[0458] e) Construct 3D shape features:

[0459] F_3d = [F_fpfh, K, H]

[0460] Feature selection and combination:

[0461] a) Apply principal component analysis (PCA) to reduce dimensionality:

[0462] F_reduced = PCA([F_pyramid, F_hog, F_color, F_3d], n_components)

[0463] b) Feature importance assessment:

[0464] importance = randomForestFeatureImportance(F_reduced, labels)

[0465] c) Select Top-K features:

[0466] F_selected = selectTopK(F_reduced, importance, K)

[0467] Output helmet feature set F_helmet = {f_h_1, f_h_2, ..., f_h_m}

[0468] By combining multi-scale spatial pyramid pooling, improved HOG features and 3D shape descriptors, the appearance, color and 3D structural features of the helmet can be fully captured. Specifically,

[0469] 1. Use multi-scale spatial pyramid pooling to adapt to candidate regions of different sizes.

[0470] 2. Improved HOG feature extraction and introduction of adaptive unit division strategy.

[0471] 3. Combine local color histogram and global color moment to describe color features more comprehensively.

[0472] 4. Introduce a 3D shape descriptor based on depth information to enhance the representation ability of the 3D structure of the helmet.

[0473] 5. Use feature selection strategies to retain the most discriminative feature combinations.

[0474] In another embodiment of the present application, the knowledge distillation attention helmet classification network is specifically processed as follows:

[0475] A hardhat feature set is obtained and processed using a knowledge distillation attention hardhat classification network method to obtain a set of classification results. Specifically, a lightweight network structure is first designed, which includes depthwise separable convolution and channel attention mechanisms. Then, a teacher network is pre-trained and knowledge distillation techniques are used to transfer knowledge from the complex model to the lightweight network. Next, an improved Focal Loss is used for training to address the class imbalance problem. Finally, multiple model variants are trained using different initialization and data augmentation strategies. Finally, the prediction results of multiple models are fused through an ensemble learning strategy to obtain the classification result and confidence level for each candidate region.

[0476] The specific data processing process is as follows:

[0477] Get the helmet feature set F_helmet = {f_h_1, f_h_2, ..., f_h_m}.

[0478] Lightweight network structure design:

[0479] a) Depthwise Separable Convolution Block:

[0480] DWConv(in_c, kernel_size, stride) → BN → ReLU

[0481] PWConv(in_c, out_c) → BN → ReLU

[0482] b) Channel Attention Module:

[0483] F_squeeze = GlobalAvgPool(F_in)

[0484] F_excitation = σ(W2 ReLU(W1 F_squeeze))

[0485] F_out = F_in F_excitation

[0486] Among them, σ is the sigmoid activation function, W1 and W2 are learnable weight matrices

[0487] c) Network Architecture:

[0488] Input → DWSConv1 → CA1 → DWSConv2 → CA2 → ... → GAP → FC →Output

[0489] Among them, DWSConv is a depth-separable convolution block, CA is a channel attention module, GAP is a global average pooling, and FC is a fully connected layer.

[0490] Knowledge Distillation:

[0491] a) Teacher network pre-training:

[0492] T = trainTeacherNetwork(F_helmet, labels)

[0493] b) Calculation of distillation loss:

[0494] L_KD = α T^2 KL(softmax(z_s / T), softmax(z_t / T)) + (1-α) CE(y_pred, y_true)

[0495] Where T is the temperature parameter, z_s and z_t are the logits of the student network and the teacher network respectively, α is the balance factor, KL is the KL divergence, and CE is the cross entropy

[0496] Improved Focal Loss:

[0497] L_focal = -α_t (1-p_t)^γ log(p_t)

[0498] Among them, p_t is the predicted probability, α_t is the category balance factor, and γ is the focusing parameter

[0499] Network training:

[0500] a) Initialization:

[0501] Initialize network parameters using He initialization method

[0502] b) Optimizer selection:

[0503] Use Adam optimizer with learning rate decay strategy

[0504] c) Training loop:

[0505] For each epoch:

[0506] For each batch:

[0507] 1. Forward Propagation

[0508] 2. Calculate the loss: L = L_focal + λ L_KD

[0509] 3. Backpropagation

[0510] 4. Update parameters

[0511] Ensemble learning strategy:

[0512] a) Training multiple model variants:

[0513] Train N models using different initializations, data augmentation strategies, and hyperparameter settings

[0514] b) Model fusion:

[0515] For each candidate region r_i:

[0516] p_ensemble(r_i) = Σ(w_j p_j(r_i)) / Σ(w_j)

[0517] Among them, p_j(ri) is the predicted probability of the jth model for ri, and w_j is the model weight

[0518] Output classification result set C = {c_1, c_2, ..., c_m}

[0519] Each c_i contains the category label (worn, not worn, incorrectly worn) and the corresponding confidence

[0520] By combining depthwise separable convolution, channel attention mechanism, and knowledge distillation technology, the computational complexity is significantly reduced while maintaining high accuracy. Specifically,

[0521] 1. Use depth-wise separable convolution and channel attention mechanisms to improve model efficiency and feature representation capabilities.

[0522] 2. Introducing knowledge distillation technology to use pre-trained complex models to guide the learning of lightweight networks.

[0523] 3. Use improved Focal Loss to better handle the problem of category imbalance.

[0524] 4. Use ensemble learning strategies to improve classification by fusing multiple model variants

[0525] In another embodiment of the present application, the process of the deep layered graph cut overlap parsing algorithm is as follows:

[0526] A set of candidate regions, a set of classification results, and an enhanced depth image are obtained and processed using a depth-layered graph cut overlap parsing algorithm to obtain optimized detection results. Specifically, an improved region growing algorithm is first used for depth stratification and depth discontinuities are detected. Then, an adaptive weighted non-maximum suppression algorithm is applied, taking into account classification confidence, depth consistency, and size rationality. A graph model is then constructed and a graph cut algorithm is used to finely segment the overlapping regions. Local features and global context are then extracted to re-evaluate the overlapping regions. Finally, the occlusion ratio is estimated and the occlusion state is inferred based on partially visible features to obtain optimized detection results.

[0527] The specific data processing process is as follows:

[0528] Get the candidate region set R, the classification result set C, and the enhanced depth image I_depth_enhanced

[0529] Deep layering processing:

[0530] a) Improved region growing algorithm:

[0531] For each candidate region r_i in R:

[0532] seed = centroid(r_i)

[0533] region = regionGrow(I_depth_enhanced, seed, threshold_depth)

[0534] Among them, the regionGrow() function expands the region based on depth similarity, and threshold_depth is the depth threshold

[0535] b) Depth discontinuity detection:

[0536] edges = cannyEdgeDetection(I_depth_enhanced)

[0537] discontinuities = findSignificantEdges(edges, threshold_edge)

[0538] Adaptive weighted non-maximum suppression (NMS):

[0539] a) Calculate the candidate box weight:

[0540] w_i = conf_i depth_consistency_i size_penalty_i

[0541] Among them, conf_i is the classification confidence, depth_consistency_i is the depth consistency score,

[0542] size_penalty_i is the size penalty term (preferring candidate boxes of reasonable size)

[0543] b) Adaptive NMS:

[0544] For each pair of overlapping candidate boxes (r_i, r_j):

[0545] If IoU(r_i, r_j) > threshold_iou and |depth(r_i) - depth(r_j)| <threshold_depth:

[0546] r_suppress = argmin(w_i, w_j)

[0547] R = R - {r_suppress}

[0548] Among them, IoU() calculates the intersection-over-union ratio, and depth() calculates the average depth

[0549] Graph cut algorithm fine segmentation:

[0550] a) Build a graph model:

[0551] G = (V, E), where V is the set of pixels and E is the edge between pixels

[0552] b) Define the energy function:

[0553] E(L) = Σ(D_p(L_p)) + Σ(V_pq(L_p, L_q))

[0554] Where L is the label configuration, D_p is the data term, and V_pq is the smoothing term

[0555] c) Set data items:

[0556] D_p(L_p) = -log(P(I_p | L_p))

[0557] Where P(I_p | L_p) is the probability that pixel p belongs to label L_p, based on color and depth features

[0558] d) Set the smoothing item:

[0559] V_pq(L_p, L_q) = exp(-β ||I_p - I_q||^2) [L_p ≠ L_q]

[0560] Among them, β controls the smoothing strength, [·] is the indicative function

[0561] e) Solve the minimum cut:

[0562] L = argmin(E(L)), solved using the maximum flow algorithm

[0563] Overlapping area reassessment:

[0564] For each overlapping region:

[0565] a) Extract local features:

[0566] F_local = extractFeatures(region, I_depth_enhanced)

[0567] b) Consider the global context:

[0568] F_context = extractContextFeatures(region, R, I_depth_enhanced)

[0569] c) Multimodal fusion classification:

[0570] score = classifyRegion([F_local, F_context])

[0571] Occlusion processing:

[0572] a) Estimated occlusion ratio:

[0573] occlusion_ratio = estimateOcclusion(r_i, discontinuities)

[0574] b) Inference based on partially visible features:

[0575] visible_features = extractVisibleFeatures(r_i, I_depth_enhanced)

[0576] full_prediction = inferFullState(visible_features, occlusion_ratio)

[0577] Output optimized detection results D_opt = {d_1, d_2, ..., d_k}

[0578] Each d_i contains the optimized bounding box, classification label and confidence.

[0579] By combining depth information, adaptive NMS and graph cut algorithm, it can effectively handle complex overlapping scenes. Specifically,

[0580] 1. Use an improved region growing algorithm for deep layering to more accurately segment overlapping objects.

[0581] 2. Introduce adaptive weighted NMS to consider depth consistency and size rationality.

[0582] 3. Apply graph cut algorithm to perform fine segmentation and improve the segmentation accuracy of overlapping areas.

[0583] 4. Re-evaluate overlapping regions by considering local features and global context.

[0584] 5. Introduce an occlusion inference mechanism based on partially visible features.

[0585] In another embodiment of the present application, the spatiotemporal graph Kalman CRF optimization algorithm is specifically implemented as follows:

[0586] The optimized detection results and historical detection results of multiple consecutive frames are obtained and processed using the spatiotemporal graph Kalman filter (CRF) optimization algorithm to obtain temporally consistent detection results. Specifically, the improved Kalman filter is first used to track the detection results, and an adaptive noise estimation mechanism is introduced. Then, a spatiotemporal graph model is constructed, defining node and edge features. Next, a graph convolutional network is designed and trained, and a temporal attention mechanism is introduced to learn long-term dependencies. Then, a conditional random field is applied for temporal smoothing, and an energy function is designed and solved using mean field inference. Finally, a temporal voting mechanism is used to fuse the historical consistency score and the current frame confidence to obtain temporally consistent detection results.

[0587] The specific data processing process is as follows:

[0588] Get the optimized detection result D_opt and historical detection results of multiple consecutive frames.

[0589] Improved Kalman filter tracking:

[0590] a) State space model:

[0591] x_k = F_k x_{k-1} + w_k

[0592] z_k = H_k x_k + v_k

[0593] Among them, x_k is the state vector (position, speed, size), z_k is the observation vector,

[0594] F_k is the state transfer matrix, H_k is the observation matrix, w_k and v_k are process noise and observation noise

[0595] b) Adaptive noise estimation:

[0596] Q_k = adaptNoiseEstimation(x_{k-1}, z_k)

[0597] Among them, the adaptNoiseEstimation() function dynamically adjusts the process noise covariance based on historical status and current observations

[0598] c) Prediction step:

[0599] x_k^- = F_k x_{k-1}

[0600] P_k^- = F_k P_{k-1} F_k^T + Q_k

[0601] d) Update steps:

[0602] K_k = P_k^- H_k^T (H_k P_k^- H_k^T + R_k)^(-1)

[0603] x_k = x_k^- + K_k (z_k - H_k x_k^-)

[0604] P_k = (I - K_k H_k) P_k^-

[0605] Where K_k is the Kalman gain and P_k is the state estimation covariance matrix

[0606] Construction of spatiotemporal graph model:

[0607] a) Build graph structure:

[0608] G = (V, E), where V is the detection result node and E is the spatiotemporal edge

[0609] b) Define node characteristics:

[0610] f_v = [bbox, conf, class, appearance_feature]

[0611] c) Define edge features:

[0612] f_e = [IoU, feature_similarity, temporal_distance]

[0613] Graph Convolutional Network (GCN) Temporal Dependency Learning:

[0614] a) GCN layer design:

[0615] H^(l+1) = σ(D^(-1 / 2) AD^(-1 / 2) H^(l) W^(l))

[0616] Among them, A is the adjacency matrix, D is the degree matrix, H^(l) is the node representation of the lth layer,

[0617] W^(l) is the learnable weight matrix, σ is the activation function

[0618] b) Temporal Attention Mechanism:

[0619] α_ij = softmax(a^T [Wh_i || Wh_j])

[0620] h_i' = σ(Σ(α_ij Wh_j))

[0621] Among them, a is the learnable attention vector, h_i and h_j are node representations.

[0622] Conditional Random Field (CRF) smoothing:

[0623] a) Define the energy function:

[0624] E(X) = Σ(ψ_u(x_i)) + Σ(ψ_p(x_i, x_j))

[0625] Among them, ψ_u is the unary potential energy and ψ_p is the pairwise potential energy.

[0626] b) Design of unary potential energy:

[0627] ψ_u(x_i) = -log(P(x_i)), P(x_i) is the probability of GCN output.

[0628] c) Design paired potential:

[0629] ψ_p(x_i, x_j) = μ(x_i, x_j) k(f_i, f_j)

[0630] Among them, μ is the label compatibility function and k is the feature similarity kernel function.

[0631] d) Find the optimal label configuration:

[0632] X = argmin(E(X)), solving using the mean-field extrapolation algorithm.

[0633] Sequential voting mechanism:

[0634] a) Calculate the historical consistency score:

[0635] s_consistency = weightedVote(predictions_history).

[0636] b) Fusion of current frame confidence:

[0637] s_final = λ s_current + (1-λ) s_consistency.

[0638] Among them, λ is the balance factor, which is dynamically adjusted according to the timing stability.

[0639] Output the detection results with consistent timing D_temp = {d_t1, d_t2, ..., d_tk}.

[0640] Each d_ti contains the time-optimized bounding box, classification label, and confidence.

[0641] By combining the improved Kalman filter, spatiotemporal graph model and conditional random field, the temporal consistency and stability of the detection results can be effectively improved. Specifically,

[0642] 1. An improved Kalman filter with adaptive noise estimation improves the adaptability to complex motion patterns.

[0643] 2. Introduce spatiotemporal graph models and GCN to capture long-term temporal dependencies.

[0644] 3. Apply conditional random fields to perform time series smoothing to ensure the consistency of detection results.

[0645] 4. Design a sequential voting mechanism to balance current frame information and historical consistency.

[0646] In another embodiment of the present application, attention guides the super-resolution small target enhancement network, and the specific process is as follows:

[0647] The method obtains temporally consistent detection results and a multi-scale image pyramid, and processes them using an attention-guided super-resolution small object enhancement network to obtain enhanced detection results. Specifically, small objects are first identified and multi-scale features are extracted. Then, a generator network containing residual dense blocks and a channel-wise attention mechanism is used for super-resolution reconstruction. Next, spatial and channel-wise attention modules are applied for feature enhancement. The enhanced features, multi-scale features, and the reconstructed image are then combined to reclassify small objects. Finally, the small object detection results are updated and integrated with the original results to obtain the enhanced detection results.

[0648] The specific data processing process is as follows:

[0649] Obtain time-consistent detection results D_temp and multi-scale image pyramid P_norm

[0650] Small target recognition:

[0651] a) Define the criteria for judging small goals:

[0652] is_small(d_i) = (area(d_i) < threshold_area) && (conf(d_i) <threshold_conf)

[0653] b) Screening small goals:

[0654] D_small = {d_i | is_small(d_i) for d_i in D_temp}

[0655] Multi-scale feature extraction:

[0656] For each small target d_i in D_small:

[0657] a) Optimal scale for positioning:

[0658] s_best = argmax_s(response(d_i, P_norm[s]))

[0659] Among them, the response() function calculates the response intensity of the target at a specific scale

[0660] b) Extracting multi-scale features:

[0661] F_multi = [extract(d_i, P_norm[s]) for s in range(s_best-1, s_best+2)]

[0662] Improved super-resolution reconstruction:

[0663] a) Design the generator network G:

[0664] Using residual dense blocks and channel attention mechanism

[0665] G(x) = F(x) + x, where F(x) is the residual learning function

[0666] b) Design the discriminator network D:

[0667] Use the PatchGAN structure to determine the authenticity of local image blocks

[0668] c) Define the loss function:

[0669] L_G = L_content + λ_adv L_adv + λ_perceptual L_perceptual

[0670] Among them, L_content is the content loss, L_adv is the adversarial loss, and L_perceptual is the perceptual loss

[0671] d) Super-resolution reconstruction:

[0672] I_sr = G(I_small)

[0673] Attention-guided feature enhancement:

[0674] a) Spatial Attention Module:

[0675] M_s = σ(f([AvgPool(F), MaxPool(F)]))

[0676] Among them, f is the convolution operation and σ is the sigmoid activation function

[0677] b) Channel Attention Module:

[0678] M_c = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))

[0679] Among them, MLP is a multi-layer perceptron

[0680] c) Feature Enhancement:

[0681] F_enhanced = F M_s M_c

[0682] Small target reclassification:

[0683] a) Build enhancement features:

[0684] F_final = [F_enhanced, F_multi, flatten(I_sr)]

[0685] b) Apply the classifier:

[0686] score_new = classifier(F_final)

[0687] Among them, the classifier can be the lightweight neural network trained in step S33

[0688] Results integration:

[0689] a) Update small target detection results:

[0690] -7. Results integration:

[0691] a) Update small target detection results:

[0692] D_small_updated = {updateDetection(d_i, score_new_i) for d_i in D_small}

[0693] The updateDetection() function updates the detection results with the new classification scores and possible bounding box adjustments.

[0694] b) Merge updated small target results:

[0695] D_final = (D_temp - D_small) ∪ D_small_updated

[0696] Output enhanced detection result D_final = {d_f1, d_f2, ..., d_fn}

[0697] where each d_fi contains the possibly augmented and reclassified bounding box, classification label, and confidence.

[0698] By combining super-resolution reconstruction and attention-guided feature enhancement, the detection of helmets in long-distance or high-altitude working scenarios can be significantly improved.

[0699] 1. Use multi-scale feature extraction to fully utilize the information in the image pyramid.

[0700] 2. Introduce an improved super-resolution reconstruction network to improve the image quality of small objects.

[0701] 3. Design an attention-guided feature enhancement mechanism to highlight key features.

[0702] 4. Combining original multi-scale features and enhanced features for reclassification to improve the accuracy of small target detection.

[0703] In another embodiment of the present application, the intelligent interactive security analysis visualization system has the following specific processes:

[0704] The enhanced detection results and original RGB images are obtained and processed using an intelligent interactive security analysis and visualization system approach to produce a visualization result image and detection report. Specifically, advanced image rendering is first performed, including adaptive bounding box drawing, optimized label position rendering, and transparency heatmap generation. Next, statistical information is calculated, including category statistics, time series analysis, and spatial distribution analysis. Key information is then extracted and a structured detection report is generated using natural language generation techniques. Finally, interactive visualization components are created, including a timeline slider, category filters, and hotspot area interaction features, resulting in the final visualization result and detection report.

[0705] The specific data processing process is as follows:

[0706] Get the enhanced detection result D_final and the original RGB image I_rgb

[0707] Advanced Image Rendering:

[0708] a) Bounding box drawing:

[0709] For each detection result d_i in D_final:

[0710] color = getColorByClass(d_i.class)

[0711] drawBox(I_rgb, d_i.bbox, color, thickness=f(d_i.conf))

[0712] The getColorByClass() function returns the color based on the category, and the f() function adjusts the line thickness based on the confidence level.

[0713] b) Label rendering:

[0714] For each detection result d_i in D_final:

[0715] label = f"{d_i.class}: {d_i.conf:.2f}"

[0716] position = optimizeLabelPosition(d_i.bbox, I_rgb)

[0717] drawLabel(I_rgb, label, position, color, font=adaptive_font(d_i.bbox))

[0718] Among them, the optimizeLabelPosition() function optimizes the label position to avoid overlap, and the adaptive_font() function adjusts the font according to the bounding box size.

[0719] c) Transparency heatmap:

[0720] heatmap = generateHeatmap(D_final, I_rgb.shape)

[0721] I_visual = blendImages(I_rgb, heatmap, α=0.4)

[0722] Among them, the generateHeatmap() function generates a heat map representing the detection density, and the blendImages() function performs image fusion.

[0723] Statistics calculation:

[0724] a) Category statistics:

[0725] class_counts = countByClass(D_final)

[0726] class_ratios = {c: count / len(D_final) for c, count in class_counts.items()}

[0727] b) Time Series Analysis:

[0728] trend = analyzeTrend(history_detections + D_final)

[0729] Among them, the analyzeTrend() function analyzes the time trend of the test results

[0730] c) Spatial distribution analysis:

[0731] spatial_density = computeSpatialDensity(D_final, I_rgb.shape)

[0732] high_risk_areas = identifyHighRiskAreas(spatial_density, threshold_risk)

[0733] Natural language report generation:

[0734] a) Report template design:

[0735] template = loadTemplate("safety_report_template")

[0736] b) Key information extraction:

[0737] key_points = extractKeyPoints(class_ratios, trend, high_risk_areas)

[0738] c) Natural Language Generation:

[0739] report_text = generateNLG(template, key_points)

[0740] Among them, the generateNLG() function uses the template and key information to generate a natural language report

[0741] Interactive visualization components:

[0742] a) Timeline Slider:

[0743] slider = createTimeSlider(history_detections + D_final)

[0744] b) Category filter:

[0745] filters = createClassFilters(class_counts.keys())

[0746] c) Hotspot area interaction:

[0747] hotspots = createInteractiveHotspots(high_risk_areas)

[0748] Output:

[0749] a) Visualization result image:

[0750] I_visual (rendered image containing bounding boxes, labels, and heatmaps)

[0751] b) Test report:

[0752] Report (structured report containing statistical information, trend analysis, and natural language description)

[0753] c) Interactive components:

[0754] InteractiveComponents (a component set that includes timeline sliders, category filters, and hotspot area interactions)

[0755] The technical advantages are:

[0756] 1. Use adaptive rendering technology to dynamically adjust visual elements based on the characteristics of the detection results.

[0757] 2. Introduce a transparency heat map to visually display the spatial distribution of helmet wearing conditions.

[0758] 3. Combine time series analysis and spatial distribution analysis to comprehensively assess safety status.

[0759] 4. Use natural language generation technology to automatically generate understandable text reports.

[0760] 5. Design interactive visualization components to support in-depth analysis and exploration of detection results.

[0761] It should be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present invention will not further describe various possible combinations.

Claims

1. A construction site safety helmet wearing detection method based on computer vision, characterized in that: include: Step S1: collecting or acquiring multimodal raw data of a construction site, and processing the data using a multimodal data fusion preprocessing method to obtain an aligned and enhanced multimodal dataset; Step S2: obtaining an aligned and enhanced multimodal dataset, processing it using a multi-scale feature extraction and fusion method to obtain a fused feature set; Step S3: Obtain the fused feature set and the enhanced depth image, process them using the depth-aware density clustering RPN algorithm to obtain a candidate region set; then use a multi-stage helmet detection and classification method to obtain preliminary helmet detection and classification results to form a classification result set; Step S4: Obtain the obtained candidate region set and classification result set, as well as the aligned and enhanced multimodal dataset, and process them using a multi-stage post-processing and result optimization method to obtain the final helmet detection and classification results; The step S1 is specifically as follows: Step S11: Acquire real-time scene data of the construction site and process it using a multimodal sensor array synchronous acquisition method to obtain an original multimodal dataset; the multimodal sensor array includes an RGB camera, a depth camera, and a thermal imaging camera; the original multimodal dataset includes an RGB image, a depth image, and a thermal imaging image; Step S12: obtaining the original multimodal dataset, processing it with the deep adaptive ICP algorithm, and obtaining a spatially aligned multimodal dataset; processing the spatially aligned multimodal dataset with the multiscale spatiotemporal DTW algorithm, and obtaining a spatiotemporally aligned multimodal dataset; Step S13: Read the aligned RGB image from the spatiotemporally aligned multimodal dataset, and process it using an edge-enhanced multi-directional Laplacian pyramid algorithm to obtain a multi-scale image pyramid; Step S14: obtaining a multi-scale image pyramid and aligned thermal imaging data, and processing them using a thermal imaging-guided CLAHE algorithm to obtain a multi-scale image pyramid after illumination normalization; Step S15: Acquire the aligned depth image, process it with a depth adaptive bilateral filtering algorithm to obtain an enhanced depth image; process the enhanced depth image with a depth completion algorithm to obtain a filled depth image; process the filled depth image with a multi-scale depth gradient algorithm to obtain a depth edge map; Step S14 is specifically as follows: Obtain images and corresponding thermal imaging data in a multi-scale image pyramid; Thermal imaging data are converted into light intensity estimates using a nonlinear mapping function and smoothed using guided filtering; Adaptively determine the block size of CLAHE based on the gradient information of the illumination intensity map; Compute local histograms and apply adaptive clipping thresholds based on local illumination intensity; Calculate the local enhancement factor according to the light intensity and perform contrast adaptive enhancement; Pixel-level brightness adjustment is performed based on the light intensity map to obtain a light-normalized image.

2. The computer vision-based construction site safety helmet wearing detection method according to claim 1, characterized in that: The step S2 is specifically as follows: Step S21: obtaining a multi-scale image pyramid after illumination normalization, and processing it using an adaptive multi-space color descriptor method to obtain a multi-scale color feature set; Step S22: obtaining a multi-scale image pyramid after illumination normalization, and processing it using a scale-invariant direction-aware LBP algorithm to obtain a multi-scale texture feature set; Step S23: Obtain the enhanced depth image and depth edge map, and process them using 3D-aware Hough transform and curvature analysis methods to obtain a shape feature set; Step S24: obtaining a continuous multi-frame illumination normalized multi-scale image pyramid and an enhanced depth image, and processing them using a depth enhanced optical flow STIP algorithm to obtain a temporal feature set; Step S25: Based on the multi-scale color feature set, the multi-scale texture feature set, the shape feature set and the temporal feature set, a GCN-guided adaptive feature fusion network method is used for processing to obtain a fused feature set.

3. The computer vision-based construction site safety helmet wearing detection method according to claim 1, characterized in that: The step S3 is specifically as follows: Step S31: Obtain the fused feature set and the enhanced depth image, and process them using the depth-aware density clustering RPN algorithm to obtain a candidate region set; Step S32: obtaining a candidate region set and a fused feature set, and processing them using a multimodal pyramid helmet feature extractor method to obtain a helmet feature set; Step S33: Obtain a safety helmet feature set, process it using the knowledge distillation attention safety helmet classification network method, and obtain a classification result set.

4. The method for detecting safety helmet wearing at a construction site based on computer vision according to claim 1, wherein: The step S4 is specifically as follows: Step S41: Obtain a candidate region set, a classification result set, and an enhanced depth image, and process them using a depth layered graph cut overlapping parsing algorithm to obtain an optimized detection result; Step S42: Obtain the optimized detection results and the historical detection results of multiple consecutive frames, and process them using the spatiotemporal graph Kalman CRF optimization algorithm to obtain detection results with consistent time sequence; Step S43: Obtain the detection results and multi-scale image pyramid with consistent time sequence, and process them using an attention-guided super-resolution small target enhancement network method to obtain an enhanced detection result; Step S44: Obtain the enhanced detection result and the original RGB image, and process them using an intelligent interactive security analysis visualization system method to obtain a visualization result image and a detection report.

5. The method for detecting safety helmet wearing at a construction site based on computer vision according to claim 1, wherein: The deep adaptive ICP algorithm is used for processing, specifically: Step S121: extract FPFH features from the input RGB point cloud and depth point cloud, and use the confidence of the depth information as a weight factor for feature calculation; Step S122: Use the kd tree to perform nearest neighbor search, establish the initial point correspondence, and introduce a depth consistency check to remove corresponding point pairs with too large depth differences; Step S123: performing iterative optimization, including calculating the weights of corresponding point pairs based on depth value reliability and feature similarity, solving the transformation matrix using the weighted SVD method, applying the transformation and updating the correspondence, and performing convergence checks using a dynamic search radius and an adaptive threshold; Step S124: Locally optimize the high curvature area and use Gaussian process regression to smooth the transformation to obtain the final transformation matrix and spatially aligned point cloud.

6. The computer vision-based construction site safety helmet wearing detection method according to claim 1, characterized in that: The multi-scale spatiotemporal DTW algorithm is used to process the spatially aligned multimodal dataset, specifically: Step S125: performing multi-scale decomposition on the input RGB and depth image sequences to generate an image pyramid structure; Step S126: extracting temporal features using an improved spatiotemporal interest point detector and enhancing feature representation by combining depth information; Step S127: Perform multi-scale DTW calculation from coarse to fine, including initializing the DTW matrix, introducing adaptive bandwidth constraints to calculate the cost matrix, using an improved step size mode and adaptive slope constraints for dynamic programming, and projecting the coarse-scale alignment path to a finer scale as the initial path; Step S128: Use subframe interpolation technology to perform precise alignment at the finest scale, and apply Kalman filtering to smooth the final alignment path; Step S129: Establish a timestamp mapping relationship according to the alignment path, use piecewise linear interpolation to process the mapping of non-integer frames, obtain the time-aligned sequence and mapping function, and then obtain the spatiotemporally aligned multimodal dataset.

7. The method for detecting safety helmet wearing at a construction site based on computer vision according to claim 1, wherein: The edge enhancement multi-directional Laplace pyramid algorithm is used for processing, specifically: Step S131: First, generate a Gaussian pyramid and perform Gaussian filtering using an adaptive kernel size based on local variance; Step S132: Calculate the Laplace difference and apply an adaptive edge enhancement filter, where the enhancement factor is dynamically adjusted according to the local edge strength; Step S133: performing multi-directional decomposition based on discrete wavelet transform on the Laplace difference of each layer, and using directional lifting wavelet to enhance the features of the main direction; Step S134 : reconstructing the enhanced image layer by layer starting from the top layer to obtain a multi-scale image pyramid including the original resolution image and the directionally enhanced details.

8. The method for detecting safety helmet wearing at a construction site based on computer vision according to claim 1, wherein: The deep adaptive bilateral filtering algorithm is used for processing, specifically: Step S151: Calculate the local depth mean and standard deviation for each pixel to estimate the local depth reliability; Step S152: constructing an adaptive kernel function, wherein the parameters of the depth value kernel function are dynamically adjusted according to the local depth reliability; Step S153: Apply improved bilateral filtering and use adaptive weights to perform filtering operations; calculate a depth edge map; Step S154: Apply edge-guided filtering to perform edge-preserving enhancement to obtain a depth image and a depth edge map after noise removal.

9. A computer vision construction site safety helmet wearing detection system, characterized by: include: at least one processor; as well as, a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the computer vision construction site safety helmet wearing detection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • RGB-D multi-modal feature fusion 3D target detection method

    CN113408584A

  • Outdoor photovoltaic field operation safety management and control method based on AI vision

    CN118277947A