Improved YOLOv11s safety helmet wearing detection model and optimization method thereof
By integrating RGB with near-infrared image features, optical flow estimation and graph convolution networks, the problems of traditional safety helmet detection methods in lighting changes and hardware adaptation are solved, and high-precision and rapid detection in extreme scenarios are achieved, suitable for industrial monitoring and autonomous driving.
Patent Information
- Application Number
- CN202510299960.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional safety helmet detection methods rely on single visible light sensor data, and the characteristic representation dimension is limited. In severe light changes or target reflective scenes, texture details are easily lost and missed detection rates are increased. The detection framework based on single-frame images ignores timing context information, resulting in a lack of continuity in the prediction of fast moving target trajectory and trajectory fracture misjudgment; the model training process does not consider cross-construction environment differences, resulting in a decrease in detection accuracy of newly deployed scenarios; the static network architecture cannot adapt to heterogeneous hardware computing power, and inference delay and memory usage are unstable during edge devices deployment.
The multimodal fusion module is used to fuse RGB and near-infrared image features, and generate multimodal feature maps through dynamic weight allocation and spatial transformation networks; the spatiotemporal analysis module uses optical flow estimation network and gated cycle units to model the target trajectory, and adjusts the confidence threshold with motion sensitivity factors; the domain adaptation module generates harsh environment simulation images through CycleGAN, and adjusts the weights based on the domain-perceived focus loss function; the topology optimization module encodes spatial topological relationships through graph convolution networks to eliminate logical conflicts; the dynamic architecture module optimizes model parameters through differentiable channel pruning and hardware simulator.
In extreme scenarios, improve detection accuracy, reduce false detection rates, enhance model generalization capabilities, optimize the inference speed and memory usage of edge devices, and meet the needs of real-time industrial monitoring.
Smart Images

Figure CN120356237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision target detection, and specifically to a safety helmet wearing detection model for improving YOLOv11s and its optimization method. Background Art
[0002] This model belongs to the technical field of target detection in computer vision, and its core focuses on real-time target recognition and positioning. The target detection technology automatically extracts image features through a deep learning model to complete the dual tasks of target classification and bounding box regression, involving key technical links such as feature pyramid construction, multi-scale prediction, and non-maximum suppression. It is widely used in scenarios such as industrial monitoring, autonomous driving, and security systems. The safety helmet wearing detection model for improving YOLOv11s is a dedicated detection system optimized based on the YOLO architecture, aiming to improve the recognition accuracy and real-time performance of the safety helmet wearing state in complex construction site scenarios.
[0003] Traditional safety helmet detection methods rely on single visible light sensor data, with limited feature representation dimensions, and are prone to texture detail loss in scenes with drastic lighting changes or target reflection, resulting in an increased missed detection rate. The detection framework based on single-frame images ignores temporal context information, lacks continuous modeling for the trajectory prediction of fast-moving targets, and produces trajectory break misjudgments when there is a short-term occlusion. In the model training process, fixed-scene datasets are mostly used, without considering cross-construction site environmental differences, and the feature distribution shift leads to a decrease in detection accuracy in newly deployed scenarios. The lack of spatial topological relationship modeling makes it difficult for the system to distinguish non-safety helmet yellow objects with similar appearances, resulting in false detections in abnormal situations where the safety helmet is removed from the wearing position. The static network architecture cannot adapt to the differences in heterogeneous hardware computing power, and the model structure needs to be manually adjusted when deploying on edge devices, with the inference delay fluctuating range and the memory peak occupancy limiting the long-term operation stability of the device. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a safety helmet wearing detection model for improving YOLOv11s and its optimization method, which solves the problems that traditional safety helmet detection methods rely on single visible light sensor data, have limited feature representation dimensions, are prone to texture detail loss in scenes with drastic lighting changes or target reflection, resulting in an increased missed detection rate. The detection framework based on single-frame images ignores temporal context information, lacks continuous modeling for the trajectory prediction of fast-moving targets, and produces trajectory break misjudgments when there is a short-term occlusion.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A safety helmet wearing detection model for improving YOLOv11s, characterized by including the following modules: multi-modal fusion module, spatio-temporal analysis module, domain adaptation module, topological optimization module, dynamic architecture module;
[0006] The multi-modal fusion module, based on the input RGB and near-infrared images, uses channel-separated convolution to achieve feature decoupling, fuses visible light and thermal radiation features through a dynamic weight allocation algorithm, and performs affine transformation alignment on multi-scale features using a spatial transformation network to generate a multi-modal feature map;
[0007] The multi-modal fusion module includes a data decoupling sub-module, a dynamic fusion sub-module, and a spatial alignment sub-module;
[0008] The spatio-temporal analysis module, based on the multi-modal feature map, uses an optical flow estimation network to extract the inter-frame displacement vector, constructs a target trajectory prediction model through a gated recurrent unit, and dynamically adjusts the classification confidence threshold in combination with a motion sensitivity factor to generate a spatio-temporal feature vector;
[0009] The spatio-temporal analysis module includes an optical flow extraction sub-module, a trajectory prediction sub-module, and a motion optimization sub-module;
[0010] The domain adaptation module, based on the spatio-temporal feature vector, uses a gradient reversal layer to achieve cross-domain feature alignment, generates simulated images of harsh environments through CycleGAN, and dynamically adjusts the weights in combination with a domain-aware focal loss function to generate a domain-invariant feature set;
[0011] The domain adaptation module includes an adversarial alignment sub-module, a data augmentation sub-module, and a loss optimization sub-module;
[0012] The topology optimization module, based on the domain-invariant feature set, uses a graph convolutional network to construct a spatial topology relationship encoder, learns the geometric constraints of the target through a self-attention mechanism, and eliminates spatial logical conflicts using a topology consistency loss to generate topology-optimized features;
[0013] The topology optimization module includes a graph structure modeling sub-module, a relationship encoding sub-module, and a logic correction sub-module;
[0014] The dynamic architecture module, based on the topology-optimized features, uses a differentiable channel pruning strategy to dynamically select and retain channels, searches for the optimal operator combination through reinforcement learning, and evaluates the deployment efficiency under multi-objective constraints in combination with a hardware simulator to generate optimized model parameters;
[0015] The dynamic architecture module includes a pruning decision sub-module, an architecture search sub-module, and a hardware adaptation sub-module.
[0016] Preferably, the data decoupling sub-module, based on the input RGB and near-infrared images, uses a channel-separated convolution algorithm to decouple spectral features, extracts the texture details of the visible light channel and the thermal radiation features of the near-infrared channel respectively, and generates a visible light feature map and a near-infrared feature map;
[0017] The dynamic fusion sub-module, based on the visible light feature map and the near-infrared feature map, calculates the channel weights of the feature map through a dynamic weight allocation algorithm, fuses the bimodal features and generates preliminary fusion features, and generates bimodal fusion features;
[0018] The spatial alignment sub-module, based on the bimodal fusion features, uses a spatial transformation network to perform an affine transformation on the multi-scale feature layers, eliminates the spatial misalignment problem caused by the perspective difference, and generates a multi-modal feature map.
[0019] Preferably, the optical flow extraction sub-module, based on the multi-modal feature map, uses an optical flow estimation network to calculate the displacement vector field of the target between consecutive frames, captures the trajectory change trend of the moving target, and generates an inter-frame optical flow field;
[0020] The trajectory prediction sub-module, based on the inter-frame optical flow field, models the target motion trajectory through a gated recurrent unit, predicts the displacement path of the target within the next 3 frames, and generates a target trajectory sequence;
[0021] The motion optimization sub-module, based on the target trajectory sequence, dynamically adjusts the classification confidence threshold in combination with the motion sensitivity factor, suppresses the false detection caused by fast movement, and generates a spatio-temporal feature vector.
[0022] Preferably, the adversarial alignment sub-module, based on the spatio-temporal feature vector, adversarially trains the domain classifier through a gradient reversal layer, aligns the feature distributions of the source domain and the target domain, and generates domain-aligned features;
[0023] The data augmentation sub-module, based on the domain-aligned features, uses CycleGAN to generate simulated images of harsh environments such as rain, fog, and dust, expands the domain diversity of the training data, and generates cross-domain augmented data;
[0024] The loss optimization sub-module, based on the cross-domain augmented data, designs a domain-aware focal loss function to dynamically balance the weights of cross-domain samples, optimizes the generalization ability of the model for unknown domains, and generates a domain-invariant feature set.
[0025] Preferably, the graph structure modeling sub-module, based on the domain-invariant feature set, models the detection box and the human key points as graph nodes, constructs a spatial topological relationship graph, and generates a spatial topological graph;
[0026] The relationship encoding sub-module, based on the spatial topological graph, uses a graph convolutional network to encode the geometric constraint relationships between nodes, learns the spatial position correlation between the safety helmet and the head, and generates topological encoded features;
[0027] The logic correction sub-module, based on the topological encoded features, eliminates the spatial logic conflicts through a topological consistency loss function, corrects the isolated false detection targets, and generates topologically optimized features.
[0028] Preferably, the pruning decision sub-module calculates the channel importance score using a differentiable channel pruning strategy based on the topology optimization features, identifies redundant feature channels, and generates a channel pruning mask;
[0029] The architecture search sub-module searches for the optimal combination of convolutional operators through reinforcement learning based on the channel pruning mask and generates an efficient operator architecture;
[0030] The hardware adaptation sub-module evaluates the inference latency and memory occupancy by combining with a hardware simulator based on the efficient operator architecture, generates the final deployment parameters, and generates optimized model parameters.
[0031] An optimization method for a safety helmet wearing detection model that improves YOLOv11s, characterized by including the following steps:
[0032] S1: Multi-modal feature fusion and spatio-temporal modeling. Based on the input RGB and near-infrared images, channel-separated convolution is used to decouple spectral features, and visible light texture and near-infrared thermal radiation features are extracted respectively. The dual-modal features are fused through a dynamic weight allocation algorithm. The spatial transformation network is used to align multi-scale features to generate a multi-modal feature map. Based on this feature map, an optical flow estimation network is constructed to calculate the displacement vector between consecutive frames. Combining with a gated recurrent unit to predict the target trajectory sequence, and dynamically adjusting the confidence threshold through a motion sensitivity factor to generate spatio-temporal optimization features;
[0033] S2: Cross-domain adaptation and topology logic correction. Based on the spatio-temporal optimization features, a gradient reversal layer is used to adversarially train a domain classifier to align the feature distributions of different construction site scenarios. A cycle generative adversarial network is used to generate simulated data for harsh environments such as rain, fog, and dust. The domain-aware focal loss function is used to dynamically balance the weights of cross-domain samples to generate a domain-invariant feature set. Based on this feature set, a graph convolutional network is constructed, and the detection box and human key points are modeled as graph nodes. The self-attention mechanism is used to encode the spatial topology relationship, and the topology consistency loss function is used to correct logical conflicts to generate topology-constrained features;
[0034] S3: Dynamic architecture compression and hardware adaptation. Based on the topology-constrained features, a differentiable channel pruning strategy is introduced to calculate the channel importance score, generate a channel pruning mask, search for the optimal combination of convolutional operators through learning, construct a multi-objective optimization function to jointly constrain the accuracy and efficiency metrics, and combine with a hardware simulator to real-time evaluate the inference latency and memory occupancy of different architectures on edge devices, generate the final deployment parameters, and generate a hardware-optimized model.
[0035] Preferably, S1: Multimodal feature fusion and spatio-temporal modeling. Based on the input RGB and near-infrared images, channel-separated convolution is used to decouple spectral features, extracting visible light texture and near-infrared thermal radiation features respectively. The dual-modal features are fused through a dynamic weight allocation algorithm. A spatial transformation network is used to align multi-scale features to generate a multimodal feature map. Based on this feature map, an optical flow estimation network is constructed to calculate the displacement vector between consecutive frames. Combining with a gated recurrent unit to predict the target trajectory sequence, and the confidence threshold is dynamically adjusted by a motion-sensitive factor to generate spatio-temporal optimized features, which includes the following steps:
[0036] S101: Based on the input RGB and near-infrared images, the channel-separated convolution algorithm is used to decouple spectral features, extracting the texture features of the visible light channel and the thermal radiation intensity map of the near-infrared channel respectively, generating a visible light feature map and a near-infrared feature map;
[0037] S102: Based on the visible light feature map and the near-infrared feature map, the response weights of each channel are calculated through a dynamic weight allocation algorithm. When fusing the dual-modal features, a non-local attention mechanism is introduced to suppress the illumination interference noise, generating dual-modal fusion features;
[0038] S103: Based on the dual-modal fusion features, a spatial transformation network is used to perform an affine transformation on the feature layer, and the geometric distortion of multi-view imaging is compensated through a thin plate spline interpolation algorithm, generating spatially aligned features;
[0039] S104: Based on the spatially aligned features, an optical flow estimation network is constructed to extract the displacement vector field of 5 consecutive frames. Combining with a gated recurrent unit to model the motion trajectory, and the target displacement prediction error is corrected through Kalman filtering, generating spatio-temporal optimized features.
[0040] Preferably, S2: Cross-domain adaptation and topological logic correction. Based on the spatio-temporal optimized features, a gradient reversal layer is used to adversarially train a domain classifier to align the feature distributions of different construction site scenarios. A cycle generative adversarial network is used to generate simulated data of harsh environments such as rain, fog, and dust. The domain-aware focal loss function is used to dynamically balance the weights of cross-domain samples, generating a domain-invariant feature set. Based on this feature set, a graph convolutional network is constructed, modeling the detection box and human key points as graph nodes, using a self-attention mechanism to encode the spatial topological relationship, and correcting the logical conflict through a topological consistency loss function, generating topological constraint features, which includes the following steps;
[0041] S201: Based on the spatio-temporal optimized features, a gradient reversal layer is used to adversarially train a domain classifier, and the maximum mean discrepancy loss is used to force the alignment of the feature distributions of the source domain and the target domain, generating domain-aligned features;
[0042] S202: Based on the domain-aligned features, a cycle generative adversarial network is used to generate simulated images of rain, fog, and dust environments, and adaptive instance normalization is used to enhance the cross-domain data diversity, generating cross-domain enhanced data;
[0043] S203: Based on cross - domain enhanced data, design a domain - aware focal loss function to dynamically adjust the sample weights, implement gradient amplification optimization for low - confidence cross - domain samples, and generate a domain - invariant feature set;
[0044] S204: Based on the domain - invariant feature set, construct a graph convolutional network to encode the spatial topological relationship between detection boxes and human key points, and strengthen the geometric constraints of the safety helmet on the head through a multi - head self - attention mechanism to generate topological constraint features.
[0045] Preferably, S3: Dynamic architecture compression and hardware adaptation. Based on the topological constraint features, introduce a differentiable channel pruning strategy to calculate the channel importance scores, generate a channel pruning mask, learn to search for the optimal combination of convolutional operators, construct a multi - objective optimization function to jointly constrain the accuracy and efficiency metrics, and combine with a hardware simulator to evaluate the inference latency and memory occupancy of different architectures on edge devices in real - time, generate the final deployment parameters, and generate a hardware - optimized model, including the following steps;
[0046] S301: Based on the topological constraint features, adopt a differentiable channel pruning strategy to calculate the channel importance scores, dynamically identify redundant channels through Gibbs sampling, and generate a channel pruning mask;
[0047] S302: Based on the channel pruning mask, construct a reinforcement learning search space to define the convolutional operator combination strategy, optimize the accuracy - speed trade - off coefficient through the Q - learning algorithm, and generate an efficient operator architecture;
[0048] S303: Based on the efficient operator architecture, deploy a hardware simulator to monitor the GPU and CUDA core utilization in real - time, optimize the operator execution timing through dynamic voltage and frequency adjustment, and generate hardware - aware parameters;
[0049] S304: Based on the hardware - aware parameters, adopt a multi - objective evolutionary algorithm to search for the optimal deployment plan, jointly optimize the model parameter quantity, inference latency, and peak memory occupancy, and generate a hardware - optimized model.
[0050] The present invention provides a safety helmet wearing detection model for improving YOLOv11s and its optimization method. It has the following beneficial effects:
[0051] The present invention enhances the complementarity of target texture and thermal radiation features under complex lighting conditions by fusing visible light and near-infrared spectral features and implementing dynamic weight allocation, solves the problem of feature distortion of single-modal data in strong backlight or low-illumination scenarios, constructs a temporal motion model based on optical flow estimation and gated recurrent units to capture the target displacement law between consecutive frames, dynamically corrects the classification threshold through a motion sensitivity factor to reduce misjudgments of trajectory breaks caused by fast movement or short-term occlusion, introduces a gradient reversal layer adversarial training mechanism to align cross-domain feature distributions, combines a cyclic generative adversarial network to synthesize data in harsh environments, improves the generalization ability of the model in unlabeled construction site scenarios, uses a graph convolutional network to encode the spatial topological relationship between detected targets and human key points, strengthens the logical constraint of the safety helmet wearing position through a self-attention mechanism to eliminate misdetection of spatial misalignment of the safety helmet on the head caused by perspective distortion or partial occlusion, deploys a differentiable channel pruning and hardware-aware architecture search strategy to dynamically optimize the model calculation path, realizes an improvement in inference speed and compression of memory occupancy on a Jetson Nano edge device. The above design improves the average accuracy of safety helmet detection in extreme scenarios such as night, rain, fog, and dust, reduces the false detection rate, and reduces the parameter scale, meeting the requirements of industrial real-time monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a system block diagram of the present invention;
[0053] Figure 2 is a schematic diagram of the main steps of the present invention;
[0054] Figure 3 is a detailed schematic diagram of S1 of the present invention;
[0055] Figure 4 is a detailed schematic diagram of S2 of the present invention;
[0056] Figure 5 is a detailed schematic diagram of S3 of the present invention;
[0057] Figure 6 is a schematic diagram of the user usage process. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment:
[0060] Such as Figures 1-6As shown in the figure, an embodiment of the present invention provides: a safety helmet wearing detection model that improves YOLOv11s, which is characterized by including the following modules: a multimodal fusion module, a spatio-temporal analysis module, a domain adaptation module, a topology optimization module, and a dynamic architecture module;
[0061] The multimodal fusion module, based on the input RGB and near-infrared images, uses channel-separated convolution to achieve feature decoupling, fuses visible light and thermal radiation features through a dynamic weight allocation algorithm, and performs affine transformation alignment on multi-scale features using a spatial transformation network to generate a multimodal feature map;
[0062] The multimodal fusion module includes a data decoupling sub-module, a dynamic fusion sub-module, and a spatial alignment sub-module;
[0063] The spatio-temporal analysis module, based on the multimodal feature map, uses an optical flow estimation network to extract the inter-frame displacement vector, constructs a target trajectory prediction model through a gated recurrent unit, and dynamically adjusts the classification confidence threshold in combination with a motion sensitivity factor to generate a spatio-temporal feature vector;
[0064] The spatio-temporal analysis module includes an optical flow extraction sub-module, a trajectory prediction sub-module, and a motion optimization sub-module;
[0065] The domain adaptation module, based on the spatio-temporal feature vector, uses a gradient reversal layer to achieve cross-domain feature alignment, generates simulated images of harsh environments through CycleGAN, and dynamically adjusts the weights in combination with a domain-aware focal loss function to generate a domain-invariant feature set;
[0066] The domain adaptation module includes an adversarial alignment sub-module, a data augmentation sub-module, and a loss optimization sub-module;
[0067] The topology optimization module, based on the domain-invariant feature set, uses a graph convolutional network to construct a spatial topology relationship encoder, learns the geometric constraints of the target through a self-attention mechanism, and eliminates spatial logical conflicts using a topology consistency loss to generate topology-optimized features;
[0068] The topology optimization module includes a graph structure modeling sub-module, a relationship encoding sub-module, and a logic correction sub-module;
[0069] The dynamic architecture module, based on the topology-optimized features, uses a differentiable channel pruning strategy to dynamically select and retain channels, searches for the optimal operator combination through reinforcement learning, and evaluates the deployment efficiency under multi-objective constraints in combination with a hardware simulator to generate optimized model parameters;
[0070] The dynamic architecture module includes a pruning decision sub-module, an architecture search sub-module, and a hardware adaptation sub-module.
[0071] The data decoupling sub-module, based on the input RGB and near-infrared images, uses the channel separation convolution algorithm to decouple spectral features, extracts the texture details of the visible light channel and the thermal radiation features of the near-infrared channel respectively, and generates a visible light feature map and a near-infrared feature map;
[0072] The dynamic fusion sub-module, based on the visible light feature map and the near-infrared feature map, calculates the channel weights of the feature map through the dynamic weight allocation algorithm, fuses the bimodal features and generates preliminary fusion features, and generates bimodal fusion features;
[0073] The spatial alignment sub-module, based on the bimodal fusion features, uses the spatial transformation network to perform affine transformation on the multi-scale feature layers, eliminates the spatial misalignment problem caused by the perspective difference, and generates a multi-modal feature map.
[0074] The data decoupling sub-module: Based on the input RGB and near-infrared images, it adopts a hybrid architecture of grouped convolution and depthwise separable convolution to separate the HOG texture features of the visible light channel and the local binary pattern radiation features of the near-infrared channel, and generates a visible light feature map and a near-infrared feature map.
[0075] The dynamic fusion sub-module: Based on the bimodal feature map, through the dynamic weight allocation formula calculate the channel weights (μ v , μ n is the channel mean, is the variance, is the gradient magnitude), and introduce a non-local attention mechanism to strengthen the complementary feature response during fusion to generate bimodal fusion features.
[0076] The spatial alignment sub-module: Based on the bimodal fusion features, it adopts the thin plate spline interpolation algorithm of the spatial transformation network, corrects the multi-view geometric distortion through the affine matrix, eliminates the spatial misalignment of the feature map, and generates a multi-modal feature map.
[0077] The dynamic fusion sub-module depends on the visible light and near-infrared features of the data decoupling sub-module. The spatial alignment sub-module receives the output features of the dynamic fusion sub-module for geometric correction, and finally outputs a multi-modal feature map for the spatio-temporal analysis module to use.
[0078] The optical flow extraction sub-module: Based on the multi-modal feature map, it adopts an improved FlowNet 2.0 architecture, calculates the displacement vector field of 5 consecutive frames through the pyramid loss function, and generates an inter-frame optical flow field.
[0079] The trajectory prediction sub-module: Based on the optical flow field, it constructs a temporal prediction model of the gated recurrent unit. The input sequence is the target coordinates and velocities of the previous 5 frames, and the displacement paths of the next 3 frames are output. The trajectory noise is corrected by combining the Kalman filter to generate a target trajectory sequence.
[0080] Motion Optimization Sub-module: Based on the trajectory sequence, design a motion sensitivity factor α = 1 - e -β·∥Δv∥2 (Δv is the rate of change of velocity, β = 0.05), dynamically lower the classification confidence threshold of fast-moving targets to 80% of the original value, and generate spatio-temporal feature vectors.
[0081] The optical flow extraction sub-module processes the multi-modal feature map to generate an optical flow field, the trajectory prediction sub-module outputs a trajectory sequence based on the optical flow field, and the motion optimization sub-module integrates the trajectory data to adjust the confidence and outputs spatio-temporal feature vectors for the domain adaptation module to use.
[0082] Adversarial Alignment Sub-module: Based on the spatio-temporal feature vectors, deploy a gradient reversal layer, and minimize the maximum mean discrepancy between the source domain and the target domain through adversarial training of the domain classifier to generate domain-aligned features.
[0083] Data Augmentation Sub-module: Based on the domain-aligned features, adopt the residual block generator of CycleGAN to synthesize images of harsh environments such as rain, fog, and dust, expand the training data to 300% of the original scale, and generate cross-domain augmented data.
[0084] Loss Optimization Sub-module: Based on the cross-domain augmented data, define a domain-aware focal loss function (γ d is the cross-domain weight, θ = 2), dynamically amplify the gradients of cross-domain difficult samples, and generate a domain-invariant feature set.
[0085] The adversarial alignment sub-module receives spatio-temporal feature vectors for domain alignment, the data augmentation sub-module generates cross-domain data, and the loss optimization sub-module outputs a domain-invariant feature set for the topology optimization module to use.
[0086] Graph Structure Modeling Sub-module: Based on the domain-invariant feature set, model the center coordinates of the detection box and the coordinates of 17 key points of the human body as graph nodes, construct a spatial topology relationship graph, and generate a spatial topology graph.
[0087] Relationship Encoding Sub-module: Based on the spatial topology graph, adopt the multi-head self-attention mechanism of the graph convolutional network to learn the Euclidean distance constraint (d max = 20 pixels) between the safety helmet and the head node, and generate topology-encoded features.
[0088] Logic Correction Sub-module: Based on the topology-encoded features, design a topology consistency loss function L topo = ∑max(0, d - d max ), eliminate misdetected targets with excessive distances, and generate topology-optimized features.
[0089] The graph structure modeling sub-module processes the domain-invariant feature set to construct a topology graph, the relationship encoding sub-module extracts geometric constraints, and the logic correction sub-module outputs topology-optimized features for the dynamic architecture module to use.
[0090] Pruning decision submodule: Based on topology optimization features, differentiable channel pruning is adopted, channel retention probability is calculated through Gumbel-Softmax sampling, and channel pruning mask is generated.
[0091] Architecture search submodule: Based on pruning mask, it builds a proximal policy optimization algorithm for reinforcement learning, searches for the optimal combination of standard convolution and depthwise separable convolution, and generates an efficient operator architecture.
[0092] Hardware adapter module: Based on the efficient operator architecture, deploy the hardware simulator, evaluate the inference latency and memory peak of the Jetson Nano platform, and generate optimized model parameters.
[0093] The pruning decision submodule compresses the topology optimization features, the architecture search submodule optimizes the operator combination, and the hardware adaptation submodule outputs the final optimized model parameters to complete the end-to-end deployment.
[0094] The optical flow extraction submodule uses an optical flow estimation network to calculate the displacement vector field of the target between consecutive frames based on the multimodal feature map, captures the trajectory change trend of the moving target, and generates the optical flow field between frames;
[0095] The trajectory prediction submodule, based on the inter-frame optical flow field, models the target motion trajectory through the gated recurrent unit, predicts the displacement path of the target in the next three frames, and generates the target trajectory sequence;
[0096] The motion optimization submodule dynamically adjusts the classification confidence threshold based on the target trajectory sequence and the motion sensitivity factor to suppress false detection caused by rapid movement and generate a spatiotemporal feature vector.
[0097] Based on the multimodal feature map, the improved FlowNet 2.0 architecture is used to calculate the displacement vector field of 5 consecutive frames through the pyramid loss function. The specific process includes: Feature pyramid construction: downsampling the input multimodal feature map at 3 levels to extract multi-scale spatial gradient features;
[0098] Rough optical flow estimation: At the 64×64 resolution layer, the features of adjacent frames are matched by correlation volume calculation to generate the initial optical flow field;
[0099] Optical flow refinement: Upsample to the original resolution step by step, combine convolutional gated recurrent units to iteratively correct the displacement error, and finally output the inter-frame optical flow field, including horizontal and vertical displacement components.
[0100] Based on the inter-frame optical flow field, the temporal motion law is modeled through a bidirectional gated recurrent unit:
[0101] Trajectory initialization: The target center coordinates (x t ,y t) Starting from t-4 , extract the optical flow vectors of the first 5 frames {Δx t-4},..., {Δx t , Δy t}} to form the input sequence;
[0102] Temporal modeling: Set the dimension of the Bi-GRU hidden layer to 64. The forward and backward propagations respectively learn the historical and future trends of the target motion, and output the hidden state h t ;
[0103] Displacement prediction: The fully connected layer maps h t to the coordinate increments of the next 3 frames {Δx t+1 , Δy t+1},..., {Δx t+3 , Δy t+3}}. Smooth the trajectory noise through Kalman filtering with the observation noise covariance (R = 0.1) to generate the target trajectory sequence.
[0104] Based on the target trajectory sequence, design a dynamic threshold mechanism driven by a motion sensitivity factor:
[0105] Motion state quantization: Calculate the displacement standard deviation σ d of the target within 10 consecutive frames. Set the threshold σ th = 2.5 pixels / frame. When σ d > σ th , it is determined to be moving fast;
[0106] Threshold adjustment rule: The classification confidence threshold T c is dynamically decreased (α = 0.3 is the adjustment coefficient). For example, when σ = 4.0, T d ′ = 0.8T c ; c False detection suppression: Eliminate the detection boxes with confidence lower than T
[0107] ′. The remaining boxes are further filtered by weighted non-maximum suppression to generate the spatio-temporal feature vector. c ′. The adversarial alignment sub-module, based on the spatio-temporal feature vector, adversarially trains the domain classifier through the gradient reversal layer to align the feature distributions of the source domain and the target domain, and generates the domain-aligned features;
[0108] The data augmentation sub-module, based on the domain-aligned features, uses CycleGAN to generate simulated images of harsh environments such as rain, fog, and dust to expand the domain diversity of the training data and generate cross-domain augmented data;
[0109]
[0110] The loss optimization sub-module designs a domain-aware focal loss function based on cross-domain enhanced data to dynamically balance the weights of cross-domain samples, optimize the generalization ability of the model for the unknown domain, and generate a domain-invariant feature set.
[0111] Based on the spatio-temporal feature vectors, adopting the domain adversarial neural network architecture, aligning the feature distributions through the gradient reversal layer, inputting the spatio-temporal features into the fully connected layer, mapping them to the shared latent space to generate domain-invariant latent features, the domain classifier receives the latent features, calculates the domain classification loss, and at the same time the feature extractor maximizes the classifier error through the gradient reversal, measures the feature distance between the source domain and the target domain by the maximum mean discrepancy, and stops the adversarial training when the MMD value is lower than the threshold of 0.05 to generate domain-aligned features.
[0112] The graph structure modeling sub-module models the detection boxes and human key points as graph nodes based on the domain-invariant feature set, constructs a spatial topology relationship graph, and generates a spatial topology graph;
[0113] The relationship encoding sub-module encodes the geometric constraint relationships between nodes based on the spatial topology graph, learns the spatial position correlation between the safety helmet and the head, and generates topological encoding features;
[0114] The logic correction sub-module eliminates spatial logic conflicts through the topological consistency loss function based on the topological encoding features, corrects isolated misdetected targets, and generates topologically optimized features.
[0115] Graph structure modeling → Relationship encoding: The spatial topology graph is used as the input of the relationship encoding sub-module, and the geometric constraint features are extracted through the GCN;
[0116] Relationship encoding → Logic correction: The topological encoding features are input into the logic correction sub-module to trigger spatial logic verification and misdetection filtering;
[0117] The dimension of the topologically optimized features output by the logic correction sub-module is exactly the same as the module definition, ensuring seamless connection of downstream modules.
[0118] The pruning decision sub-module calculates the channel importance scores using the differentiable channel pruning strategy based on the topologically optimized features, identifies redundant feature channels, and generates a channel pruning mask;
[0119] The architecture search sub-module searches for the optimal combination of convolutional operators through reinforcement learning based on the channel pruning mask, and generates an efficient operator architecture;
[0120] The hardware adaptation sub-module combines the hardware simulator to evaluate the inference latency and memory occupancy based on the efficient operator architecture, generates the final deployment parameters, and generates optimized model parameters.
[0121] Pruning decision → Architecture search: The channel pruning mask constrains the action space of the architecture search, and only the channels corresponding to the non-zero masks are retained for the search;
[0122] Architecture Search → Hardware Adaptation: Efficient operator architecture is used as input configuration for the hardware simulator to ensure that deployment parameters strictly match the target hardware;
[0123] The optimized model parameters output by the hardware adapter module are consistent with the module definition and can be directly deployed to the edge device.
[0124] First, the user clicks the Register button, enters the user name information, password information, confirms the password information, and then clicks the Register button to successfully register the user. Then enter the username and password information after registration, click the Login button, and log in to the main page of the helmet detection system. Select the Image button to upload the image to be detected, and then click the Image Detection button and click the Result Display button to detect whether the construction workers in the uploaded image are wearing helmets. Click the Image Detection Result Export button to export the detection results. Click the End Image Detection button to end this image detection. When performing the image detection task, the system first sends it to the image detection module of the helmet wearing detection system. In this module, the system focuses on extracting the area of the person's head in the image to judge the helmet wearing status and display the final detection results. Select the Video button to upload the video to be detected, and then click the Video Detection button to detect whether the construction workers in the uploaded video are wearing helmets. When the helmet wearing detection system performs the video detection task, the video stream is used as input by default. When the system is running, users can obtain video streams by uploading local video files. The system then decomposes the video into continuous frames and sends these frames one by one to the video detection module of the helmet wearing detection system for processing. Click the Pause Detection button to pause this video detection. Click the Video Detection Result Export button to export the video detection results. Click the End Video Detection button to end this video detection. By opening the camera button and then clicking the Camera Detection button, you can detect in real time whether the construction workers are wearing helmets. In the real-time camera monitoring task of the helmet wearing detection system, the video stream is also used as the monitoring method by default. Compared with video detection, real-time camera monitoring only includes the function of camera capture, and the rest of the processing flow remains highly consistent. During the operation of the system, the video stream can be captured instantly with the help of an external camera, and then the video is frame-decomposed, and the decomposed frames are sent to the helmet wearing detection model for analysis. And the results of helmet wearing are recorded and marked. The location information, target quantity information, and confidence information of the helmet are displayed in the real-time screen. Click the End Camera button to end this real-time camera detection.
[0125] An optimization method for a safety helmet wearing detection model that improves YOLOv11s, characterized by including the following steps:
[0126] S1: Multi-modal feature fusion and spatio-temporal modeling. Based on the input RGB and near-infrared images, channel-separated convolution is used to decouple spectral features, and visible light texture and near-infrared thermal radiation features are extracted respectively. The dual-modal features are fused through a dynamic weight allocation algorithm. A spatial transformation network is used to align multi-scale features to generate a multi-modal feature map. Based on this feature map, an optical flow estimation network is constructed to calculate the displacement vector between consecutive frames. Combining with a gated recurrent unit to predict the target trajectory sequence, and the confidence threshold is dynamically adjusted through a motion sensitivity factor to generate spatio-temporal optimized features;
[0127] S2: Cross-domain adaptation and topological logic correction. Based on the spatio-temporal optimized features, a gradient reversal layer is used to adversarially train a domain classifier to align the feature distributions of different construction site scenarios. A cycle generative adversarial network is used to generate simulation data for harsh environments such as rain, fog, and dust. The domain-aware focal loss function is used to dynamically balance the weights of cross-domain samples to generate a domain-invariant feature set. Based on this feature set, a graph convolutional network is constructed. The detection box and human key points are modeled as graph nodes, and a self-attention mechanism is used to encode the spatial topological relationship. The topological consistency loss function is used to correct logical conflicts to generate topological constraint features;
[0128] S3: Dynamic architecture compression and hardware adaptation. Based on the topological constraint features, a differentiable channel pruning strategy is introduced to calculate the channel importance score, generate a channel pruning mask, and learn to search for the optimal combination of convolutional operators. A multi-objective optimization function is constructed to jointly constrain the accuracy and efficiency metrics. Combining with a hardware simulator to evaluate the inference latency and memory occupancy of different architectures on edge devices in real time to generate the final deployment parameters and generate a hardware-optimized model.
[0129] The spatio-temporal optimized features of S1 are used as the input basis for S2 to achieve full-link optimization from data fusion to temporal analysis
[0130] The topological constraint features output by S2 inherit the spatio-temporal and cross-domain optimization results and provide high-precision features for the architecture compression of S3
[0131] S3 is based on the geometric logic constraint features of S2 to complete the final deployment of model lightweight and hardware adaptation
[0132] This optimization method realizes three technical goals of improving detection accuracy, enhancing cross-domain generalization ability, and optimizing edge inference speed through three-level progressive processing.
[0133] S1: Multimodal Feature Fusion and Spatiotemporal Modeling. Based on the input RGB and near-infrared images, use channel-separated convolution to decouple spectral features, extract visible light texture and near-infrared thermal radiation features respectively, fuse the bimodal features through a dynamic weight allocation algorithm, use a spatial transformation network to align multi-scale features to generate a multimodal feature map, build an optical flow estimation network based on this feature map, calculate the displacement vector between consecutive frames, and combine a gated recurrent unit to predict the target trajectory sequence. Dynamically adjust the confidence threshold through a motion-sensitive factor to generate spatiotemporal optimized features, including the following steps:
[0134] S101: Based on the input RGB and near-infrared images, use the channel-separated convolution algorithm to decouple spectral features, extract the texture features of the visible light channel and the thermal radiation intensity map of the near-infrared channel respectively, and generate a visible light feature map and a near-infrared feature map;
[0135] S102: Based on the visible light feature map and the near-infrared feature map, calculate the response weights of each channel through a dynamic weight allocation algorithm, introduce a non-local attention mechanism to suppress light interference noise when fusing bimodal features, and generate bimodal fusion features;
[0136] S103: Based on the bimodal fusion features, use a spatial transformation network to perform an affine transformation on the feature layer, and compensate for the geometric distortion of multi-view imaging through a thin plate spline interpolation algorithm to generate spatially aligned features;
[0137] S104: Based on the spatially aligned features, build an optical flow estimation network to extract the displacement vector field of 5 consecutive frames, combine a gated recurrent unit to model the motion trajectory, and correct the target displacement prediction error through Kalman filtering to generate spatiotemporal optimized features.
[0138] S101→S102: The visible light feature map and the near-infrared feature map are used as the input of the dynamic weight allocation, and the weight value determines the fusion ratio;
[0139] S102→S103: The bimodal fusion features are input into the spatial transformation network, and the affine matrix parameters depend on the spatial distribution of the fusion features;
[0140] S103→S104: The spatially aligned features are used as the input of the optical flow estimation network to ensure that the displacement field calculation is not interfered by perspective distortion.
[0141] S2: Cross-Domain Adaptation and Topological Logic Correction. Based on spatio-temporal optimized features, use the gradient reversal layer to adversarially train the domain classifier to align the feature distributions of different construction site scenarios. Use the cycle generative adversarial network to generate simulated data of harsh environments such as rain, fog, and dust. Dynamically balance the cross-domain sample weights through the domain-aware focal loss function to generate a domain-invariant feature set. Based on this feature set, construct a graph convolutional network, model the detection boxes and human key points as graph nodes, use the self-attention mechanism to encode the spatial topological relationship, and correct logical conflicts through the topological consistency loss function to generate topological constraint features, including the following steps;
[0142] S201: Based on spatio-temporal optimized features, use the gradient reversal layer to adversarially train the domain classifier, and force the source domain and target domain feature distributions to align through the maximum mean discrepancy loss to generate domain-aligned features;
[0143] S202: Based on the domain-aligned features, use the cycle generative adversarial network to generate simulated images of rain, fog, and dust environments, and use adaptive instance normalization to enhance cross-domain data diversity to generate cross-domain enhanced data;
[0144] S203: Based on the cross-domain enhanced data, design a domain-aware focal loss function to dynamically adjust the sample weights, perform gradient amplification optimization on low-confidence cross-domain samples, and generate a domain-invariant feature set;
[0145] S204: Based on the domain-invariant feature set, construct a graph convolutional network to encode the spatial topological relationship between the detection boxes and human key points, and strengthen the geometric constraints of the safety helmet on the head through the multi-head self-attention mechanism to generate topological constraint features.
[0146] S201→S202: The domain-aligned features are used as the content input of CycleGAN to ensure that the generated images retain the semantic information of the safety helmet;
[0147] S202→S203: The cross-domain enhanced data is weighted by the domain-aware loss to strengthen the generalization ability of the model to the synthetic data;
[0148] S203→S204: The domain-invariant feature set is input into the graph convolutional network, and the node features rely on cross-domain invariance constraints.
[0149] S3: Dynamic Architecture Compression and Hardware Adaptation. Based on the topological constraint features, introduce a differentiable channel pruning strategy to calculate the channel importance scores, generate a channel pruning mask, learn to search for the optimal combination of convolutional operators, construct a multi-objective optimization function to jointly constrain the accuracy and efficiency metrics, combine with a hardware simulator to evaluate the inference latency and memory occupancy of different architectures on edge devices in real time, and generate the final deployment parameters to generate a hardware-optimized model, including the following steps;
[0150] S301: Based on the topological constraint features, adopt a differentiable channel pruning strategy to calculate the channel importance scores, dynamically identify redundant channels through Gibbs sampling, and generate a channel pruning mask;
[0151] S302: Based on the channel pruning mask, construct a reinforcement learning search space to define the convolution operator combination strategy, optimize the accuracy-speed trade-off coefficient through the Q-learning algorithm, and generate an efficient operator architecture;
[0152] S303: Based on the efficient operator architecture, deploy a hardware simulator to monitor the GPU and CUDA core utilization in real time, optimize the operator execution timing through dynamic voltage and frequency scaling, and generate hardware-aware parameters;
[0153] S304: Based on the hardware-aware parameters, adopt a multi-objective evolutionary algorithm to search for the optimal deployment plan, jointly optimize the model parameter quantity, inference latency, and peak memory occupancy, and generate a hardware-optimized model.
[0154] S301→S302: The number of channels for architecture search is constrained by the channel pruning mask, and only the channels with a mask of 1 are retained for search;
[0155] S302→S303: The efficient operator architecture is input into the hardware simulator, and the DVFS parameters are dynamically adjusted based on the architecture calculation load;
[0156] S303→S304: The hardware-aware parameters are used as the fitness evaluation basis for the evolutionary algorithm to guide multi-objective optimization.
[0157] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An improved safety helmet wearing detection model for YOLOv11 s, characterized in that It includes the following modules: multimodal fusion module, spatio-temporal analysis module, domain adaptation module, topology optimization module, and dynamic architecture module; The multimodal fusion module, based on the input RGB and near-infrared images, uses channel-separated convolution to achieve feature decoupling, fuses visible light and thermal radiation features through a dynamic weight allocation algorithm, and performs affine transformation alignment on multi-scale features using a spatial transformation network to generate a multimodal feature map; The multimodal fusion module includes a data decoupling sub-module, a dynamic fusion sub-module, and a spatial alignment sub-module; The spatio-temporal analysis module, based on the multimodal feature map, uses an optical flow estimation network to extract the inter-frame displacement vector, constructs a target trajectory prediction model through a gated recurrent unit, and dynamically adjusts the classification confidence threshold in combination with a motion sensitivity factor to generate a spatio-temporal feature vector; The spatio-temporal analysis module includes an optical flow extraction sub-module, a trajectory prediction sub-module, and a motion optimization sub-module; The domain adaptation module, based on the spatio-temporal feature vector, uses a gradient reversal layer to achieve cross-domain feature alignment, generates simulated images of harsh environments through CycleGAN, and dynamically adjusts the weights in combination with a domain-aware focal loss function to generate a domain-invariant feature set; The domain adaptation module includes an adversarial alignment sub-module, a data augmentation sub-module, and a loss optimization sub-module; The topology optimization module, based on the domain-invariant feature set, uses a graph convolutional network to construct a spatial topology relationship encoder, learns the geometric constraints of the target through a self-attention mechanism, and eliminates spatial logical conflicts using a topology consistency loss to generate topology-optimized features; The topology optimization module includes a graph structure modeling sub-module, a relationship encoding sub-module, and a logic correction sub-module; The dynamic architecture module, based on the topology-optimized features, uses a differentiable channel pruning strategy to dynamically select and retain channels, searches for the optimal operator combination through reinforcement learning, and evaluates the deployment efficiency under multi-objective constraints in combination with a hardware simulator to generate optimized model parameters; The dynamic architecture module includes a pruning decision sub-module, an architecture search sub-module, and a hardware adaptation sub-module.
2. An improved helmet-wearing detection model for YOLOv11s according to claim 1, characterized in that: The data decoupling sub-module, based on the input RGB and near-infrared images, uses a channel-separated convolution algorithm to decouple spectral features, extracts the texture details of the visible light channel and the thermal radiation features of the near-infrared channel respectively, and generates a visible light feature map and a near-infrared feature map; The dynamic fusion sub-module, based on the visible light feature map and the near-infrared feature map, calculates the channel weights of the feature map through a dynamic weight allocation algorithm, fuses the bimodal features and generates a preliminary fusion feature, and generates a bimodal fusion feature; The spatial alignment sub-module, based on the bimodal fusion feature, uses a spatial transformation network to perform affine transformation on the multi-scale feature layers, eliminates the spatial misalignment problem caused by perspective differences, and generates a multimodal feature map.
3. An improved helmet-wearing detection model for YOLOv11 s according to claim 1, characterized in that: The optical flow extraction sub-module, based on the multimodal feature map, uses an optical flow estimation network to calculate the displacement vector field of the target between consecutive frames, captures the trajectory change trend of the moving target, and generates an inter-frame optical flow field; The trajectory prediction sub-module, based on the inter-frame optical flow field, models the target motion trajectory through a gated recurrent unit, predicts the displacement path of the target within the next 3 frames, and generates a target trajectory sequence; The motion optimization sub-module, based on the target trajectory sequence, dynamically adjusts the classification confidence threshold in combination with the motion sensitivity factor to suppress false detections caused by rapid movement and generates spatio-temporal feature vectors.
4. An improved helmet-wearing detection model for YOLOv11s according to claim 1, characterized in that: The adversarial alignment sub-module, based on the spatio-temporal feature vectors, adversarially trains the domain classifier through the gradient reversal layer to align the feature distributions of the source domain and the target domain and generates domain-aligned features. The data augmentation sub-module, based on the domain-aligned features, uses CycleGAN to generate simulated images of harsh environments such as rain, fog, and dust to expand the domain diversity of the training data and generates cross-domain augmented data. The loss optimization sub-module, based on the cross-domain augmented data, designs a domain-aware focal loss function to dynamically balance the weights of cross-domain samples, optimizes the generalization ability of the model for unknown domains, and generates a domain-invariant feature set.
5. An improved helmet-wearing detection model for YOLOv11s according to claim 1, characterized in that: The graph structure modeling sub-module, based on the domain-invariant feature set, models the detection boxes and human key points as graph nodes, constructs a spatial topology graph, and generates a spatial topology graph. The relationship encoding sub-module, based on the spatial topology graph, uses a graph convolutional network to encode the geometric constraint relationships between nodes, learns the spatial position correlation between the safety helmet and the head, and generates topology-encoded features. The logic correction sub-module, based on the topology-encoded features, eliminates spatial logic conflicts through the topology consistency loss function, corrects isolated false detection targets, and generates topology-optimized features.
6. An improved helmet-wearing detection model for YOLOv11s according to claim 1, wherein Based on: The pruning decision sub-module, based on the topology-optimized features, uses a differentiable channel pruning strategy to calculate the channel importance scores, identifies redundant feature channels, and generates a channel pruning mask. The architecture search sub-module, based on the channel pruning mask, searches for the optimal combination of convolutional operators through reinforcement learning and generates an efficient operator architecture. The hardware adaptation sub-module, based on the efficient operator architecture, combines a hardware simulator to evaluate the inference latency and memory occupancy, generates the final deployment parameters, and generates optimized model parameters.
7. An optimization method for a safety helmet wearing detection model that improves YOLOv11 s, characterized in that Including the following steps: S1: Multi-modal feature fusion and spatio-temporal modeling. Based on the input RGB and near-infrared images, use channel-separated convolution to decouple spectral features, extract visible light texture and near-infrared thermal radiation features respectively, fuse the bimodal features through the dynamic weight allocation algorithm, use the spatial transformation network to align multi-scale features to generate a multi-modal feature map, construct an optical flow estimation network based on this feature map, calculate the displacement vector between consecutive frames, combine the gated recurrent unit to predict the target trajectory sequence, and dynamically adjust the confidence threshold through the motion sensitivity factor to generate spatio-temporal optimized features. S2: Cross-domain adaptation and topology logic correction. Based on the spatio-temporal optimized features, use the gradient reversal layer to adversarially train the domain classifier to align the feature distributions of different construction site scenarios, use the cycle generative adversarial network to generate simulated data of harsh environments such as rain, fog, and dust, dynamically balance the weights of cross-domain samples through the domain-aware focal loss function, generate a domain-invariant feature set, construct a graph convolutional network based on this feature set, model the detection boxes and human key points as graph nodes, use the self-attention mechanism to encode the spatial topology relationship, and correct the logic conflicts through the topology consistency loss function to generate topology-constrained features. S3: Dynamic architecture compression and hardware adaptation. Based on topological constraint features, introduce a differentiable channel pruning strategy to calculate channel importance scores, generate a channel pruning mask, learn to search for the optimal combination of convolution operators, construct a multi-objective optimization function to jointly constrain accuracy and efficiency metrics, combine with a hardware simulator to evaluate the inference latency and memory occupancy of different architectures on edge devices in real time, generate final deployment parameters, and generate a hardware optimization model.
8. An optimization method for a safety helmet wearing detection model that improves YOLOv11 s according to claim 7, characterized in that: S1: Multi-modal feature fusion and spatio-temporal modeling. Based on the input RGB and near-infrared images, use channel-separated convolution to decouple spectral features, extract visible light texture and near-infrared thermal radiation features respectively, fuse the bimodal features through a dynamic weight allocation algorithm, use a spatial transformation network to align multi-scale features to generate a multi-modal feature map, construct an optical flow estimation network based on this feature map, calculate the displacement vector between consecutive frames, combine with a gated recurrent unit to predict the target trajectory sequence, and dynamically adjust the confidence threshold through a motion-sensitive factor to generate spatio-temporal optimized features, including the following steps: S101: Based on the input RGB and near-infrared images, use the channel-separated convolution algorithm to decouple spectral features, extract the texture features of the visible light channel and the thermal radiation intensity map of the near-infrared channel respectively, and generate a visible light feature map and a near-infrared feature map; S102: Based on the visible light feature map and the near-infrared feature map, calculate the response weights of each channel through a dynamic weight allocation algorithm, introduce a non-local attention mechanism to suppress light interference noise when fusing bimodal features, and generate bimodal fusion features; S103: Based on the bimodal fusion features, use a spatial transformation network to perform an affine transformation on the feature layer, and compensate for the geometric distortion of multi-view imaging through a thin plate spline interpolation algorithm to generate spatially aligned features; S104: Based on the spatially aligned features, construct an optical flow estimation network to extract the displacement vector field of 5 consecutive frames, combine with a gated recurrent unit to model the motion trajectory, and correct the target displacement prediction error through Kalman filtering to generate spatio-temporal optimized features.
9. An optimized method for a safety helmet wearing detection model that improves YOLOv11 s, characterized in that: S2: Cross-domain adaptation and topological logic correction. Based on the spatio-temporal optimized features, use a gradient reversal layer to adversarially train a domain classifier to align the feature distributions of different construction site scenarios, use a cycle generative adversarial network to generate simulation data of harsh environments such as rain, fog, and dust, dynamically balance the cross-domain sample weights through a domain-aware focal loss function, generate a domain-invariant feature set, construct a graph convolutional network based on this feature set, model the detection box and human key points as graph nodes, use a self-attention mechanism to encode the spatial topological relationship, and correct the logical conflict through a topological consistency loss function to generate topological constraint features, including the following steps; S201: Based on the spatio-temporal optimized features, use a gradient reversal layer to adversarially train a domain classifier, and force the source domain and target domain feature distributions to align through a maximum mean discrepancy loss to generate domain-aligned features; S202: Based on the domain-aligned features, use a cycle generative adversarial network to generate simulation images of rain, fog, and dust environments, and use adaptive instance normalization to enhance cross-domain data diversity to generate cross-domain enhanced data; S203: Based on cross-domain enhanced data, design a domain-aware focal loss function to dynamically adjust the sample weights, implement gradient amplification optimization for low-confidence cross-domain samples, and generate a domain-invariant feature set; S204: Based on the domain-invariant feature set, construct a graph convolutional network to encode the spatial topological relationship between the detection box and human key points, and strengthen the geometric constraints of the safety helmet on the head through the multi-head self-attention mechanism to generate topological constraint features.
10. An optimization method for a safety helmet wearing detection model that improves YOLOv11 s according to claim 7, characterized in that: S3: Dynamic architecture compression and hardware adaptation. Based on the topological constraint features, introduce a differentiable channel pruning strategy to calculate the channel importance score, generate a channel pruning mask, learn to search for the optimal combination of convolution operators, construct a multi-objective optimization function to jointly constrain the accuracy and efficiency metrics, and combine with a hardware simulator to real-time evaluate the inference latency and memory occupancy of different architectures on edge devices, generate the final deployment parameters, and generate a hardware optimization model, including the following steps; S301: Based on the topological constraint features, use a differentiable channel pruning strategy to calculate the channel importance score, and dynamically identify redundant channels through Gibbs sampling to generate a channel pruning mask; S302: Based on the channel pruning mask, construct a reinforcement learning search space to define the convolution operator combination strategy, optimize the accuracy-speed trade-off coefficient through the Q-learning algorithm, and generate an efficient operator architecture; S303: Based on the efficient operator architecture, deploy a hardware simulator to real-time monitor the GPU and CUDA core utilization rates, and optimize the operator execution timing through dynamic voltage and frequency adjustment to generate hardware-aware parameters; S304: Based on the hardware-aware parameters, use a multi-objective evolutionary algorithm to search for the optimal deployment plan, jointly optimize the model parameter quantity, inference latency, and peak memory occupancy, and generate a hardware optimization model.
Citation Information
Cited By
Video stream real-time target detection and tracking system based on deep learning
CN120564107A
A Real-Time Target Detection and Tracking System for Video Streams Based on Deep Learning
CN120564107B
Visible light and infrared fusion-based river no-fishing ship monitoring method and system
CN120894696A
Feature fusion-based aviation food truck docking identification method and system
CN120910581A
A method and system for docking and identification of airline food trucks based on feature fusion
CN120910581B