Intelligent safety helmet identification monitoring method and system based on deep learning
By using deep learning technology to identify the standardization of safety helmet wearing and the rationality of worker behavior, a knowledge graph and temporal attention network are constructed, which solves the problems of inaccurate identification and insufficient prevention capabilities of existing safety helmet monitoring systems, and realizes high-precision safety helmet wearing monitoring and preventive intervention.
Patent Information
- Application Number
- CN202511442843.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-03-06
AI Technical Summary
Existing safety helmet monitoring systems cannot identify whether safety helmets are worn correctly, and have difficulty distinguishing between improper wearing such as loose wearing or not fastening the buckle. They also cannot accurately distinguish between reasonable temporary removal of safety helmets and illegal removal of helmets by workers, and lack the ability to predict workers' potential illegal intentions. This results in poor safety supervision, high false alarm rate, and weak prevention capabilities.
A deep learning-based intelligent safety helmet identification and monitoring method is adopted. Fine-grained human posture estimation technology is used to identify key points of the worker's head. Combined with the head and safety helmet relative position modeling algorithm, the wearing standard is evaluated. A knowledge graph of construction site scene is constructed to analyze the rationality of behavior. Temporal attention network is used to learn the precursor patterns of violations. Based on the behavioral economics model, risk perception and decision-making tendency are estimated. A violation risk scoring system is constructed to implement preventive intervention measures.
It has achieved accurate identification of seven types of non-standard wearing status, improved the accuracy of wearing standardization assessment, reduced the missed detection rate, and can accurately understand the characteristics of the work environment and job type, identify violations in advance, reduce false alarm rate and provide personalized preventive intervention, thereby improving the level of safety management.
Smart Images

Figure CN121617022A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of construction site safety monitoring technology, and more specifically, to a deep learning-based intelligent safety helmet recognition and monitoring method and system. Background Technology
[0002] In environments such as construction sites, mines, and factories, the proper wearing of safety helmets is a crucial measure to ensure the personal safety of workers. With the development of computer vision technology, image recognition-based safety helmet monitoring systems have been applied in industrial safety management.
[0003] Current safety helmet identification and monitoring technologies mainly employ traditional target detection algorithms or simple deep learning models, which can only detect whether workers are wearing safety helmets. These technologies have the following significant technical shortcomings: First, existing systems can only detect whether a safety helmet is worn, but cannot identify improper wearing conditions such as loose fits or unfastened buckles, leading to situations that appear compliant but actually pose safety hazards. Second, existing systems struggle to distinguish between workers temporarily removing their helmets for necessary operations (such as wiping sweat) and unauthorized helmet removal, resulting in numerous false alarms or missed alarms. Finally, traditional safety helmet monitoring systems passively respond to past violations, lacking the ability to anticipate potential violations and thus failing to provide preventative intervention.
[0004] These technical issues have resulted in existing safety helmet monitoring systems having shortcomings such as poor safety supervision, high false alarm rates, and weak prevention capabilities, failing to meet the high standards of industrial safety production. There is an urgent need for an intelligent safety helmet monitoring method that can accurately identify the compliance of safety helmet wearing, intelligently understand the rationality of workers' behavior, and have the ability to predict violations proactively. Summary of the Invention
[0005] This invention provides a deep learning-based intelligent safety helmet identification and monitoring method and system, which solves the technical problems of poor safety supervision effect, high false alarm rate and weak prevention capability of existing safety helmet monitoring systems in related technologies.
[0006] This invention provides a deep learning-based intelligent safety helmet recognition and monitoring method, comprising:
[0007] Fine-grained human pose estimation technology is used to identify key points on the worker's head, and a modeling algorithm for the relative position of the head and safety helmet is used to evaluate the compliance of the safety helmet wearing.
[0008] Based on the assessment results of the standardization of safety helmet wearing, a knowledge graph of construction site scenarios is constructed, which includes a three-layer relationship model of safety regulations for different types of work, and the rationality of workers' safety helmet operation behavior is analyzed.
[0009] Based on the results of the safety helmet wearing standardization assessment and the results of the behavioral rationality analysis, the characteristics of the violation intention were extracted from the workers' micro-behavioral sequences, and the precursor patterns of the violation were learned through a temporal attention network.
[0010] Based on the comprehensive assessment results of helmet wearing standards, the analysis results of behavioral rationality, and the patterns of early signs of violations, a violation risk scoring system was constructed and preventive intervention measures were implemented by estimating workers' risk perception and decision-making tendencies using a behavioral economics model.
[0011] Furthermore, the fine-grained human pose estimation technique includes:
[0012] A deep convolutional neural network with multi-scale feature fusion is used to process input surveillance video frames to identify and locate the coordinates of key points on the human body.
[0013] The algorithm for modeling the relative position of the head helmet includes:
[0014] The detected head key points are used to construct a 3D head model, the safety helmet detection box is converted into a cylindrical model in 3D space, and the spatial relationship parameters between the head model and the safety helmet model are calculated.
[0015] Based on the aforementioned spatial relationship parameters, a wearing standardization scoring function is constructed to quantitatively evaluate the degree of standardization in helmet wearing.
[0016] Furthermore, the spatial relationship parameters include the distance between the center point of the head and the center point of the helmet, the angle between the head tilt angle and the main axis of the helmet, the distance between the top of the head and the bottom of the helmet, and the buckle status.
[0017] Furthermore, the node types of the construction site scene knowledge graph include job type nodes, operation nodes, location nodes, and safety standard nodes; the edge types include job type operation edges, operation location edges, and operation safety standard edges.
[0018] The steps for analyzing the rationality of workers' hat-removal behavior based on the knowledge graph include:
[0019] Calculate the probability of a worker’s possible behavioral intentions given a sequence of behaviors, location information, and time information;
[0020] The rationality of the worker's hat-removal behavior is assessed based on the calculated probability distribution of intent.
[0021] Based on the reasonableness score, the worker's hat-removal behavior is classified into reasonable behavior, suspicious behavior, or violation behavior.
[0022] Furthermore, the features of the violation intent extracted from the worker's micro-behavioral sequences include:
[0023] Head posture change features are extracted by tracking the motion trajectory of key points on the head.
[0024] Body language characteristics, analyzing workers' body postures and movements;
[0025] Social interaction features were used to extract the interaction patterns between workers and their colleagues.
[0026] Spatiotemporal trajectory characteristics were analyzed to determine the movement trajectories and dwell time patterns of workers in different areas of the construction site.
[0027] Furthermore, the temporal attention network includes a feature embedding layer, a temporal encoding layer, a multi-head self-attention model, and a violation intent prediction layer. The temporal encoding layer uses a bidirectional LSTM network to encode the behavior sequence and capture temporal dependencies. The multi-head self-attention model calculates the attention weights at each time step in the behavior sequence through multi-head self-attention.
[0028] Furthermore, the behavioral economics model includes:
[0029] Prospect theory models are used to assess workers' subjective perceptions of risk and their decision-making preferences.
[0030] A decision framing effect model was used to analyze the differences in workers' decision preferences under different frames.
[0031] The social influence factor model considers the impact of peer behavior and group norms on individual decision-making.
[0032] Furthermore, the violation risk scoring system also considers the following factors:
[0033] Environmental factor characteristic vectors, including temperature, humidity, noise level, and light conditions;
[0034] The work status feature vector includes work duration, work intensity, rest interval, and task complexity;
[0035] Physiological state feature vectors, including heart rate, respiratory rate, and body temperature.
[0036] Furthermore, the preventive intervention measures include:
[0037] Real-time safety alerts, using personalized alert methods based on workers' risk perception characteristics;
[0038] Incentive systems, based on workers' decision-making preferences, provide positive incentives or negative warnings;
[0039] Safety knowledge delivery: Customized safety information is delivered to address workers' biases in risk perception.
[0040] Adjust the work environment by adjusting work environment parameters or arranging appropriate rest based on environmental impact factor analysis;
[0041] Management intervention: For high-risk workers, notify on-site management personnel for direct intervention.
[0042] A deep learning-based intelligent safety helmet recognition and monitoring system, used to execute the aforementioned deep learning-based intelligent safety helmet recognition and monitoring method, includes:
[0043] A fine-grained human pose estimation module is used to identify key points on the worker's head and, combined with a modeling algorithm for the relative position of the head and safety helmet, to evaluate the compliance of the safety helmet wearing.
[0044] The construction site scenario knowledge graph module is used to construct a knowledge graph containing a three-layer relationship model of work operation safety regulations, and to analyze the rationality of workers' safety helmet operation behavior;
[0045] The micro-behavior analysis module is used to extract features of the intent to violate regulations from the worker's micro-behavior sequence and learn the precursor patterns of the violation through a temporal attention network;
[0046] The risk assessment and intervention module is used to estimate workers' risk perception and decision-making tendencies based on behavioral economics models, construct a violation risk scoring system, and implement preventive intervention measures.
[0047] The beneficial effects of this invention are as follows: This invention can identify 7 different non-standard wearing states, improves the accuracy of wearing standardization assessment, reduces the missed detection rate, and improves the level of precision in safety supervision.
[0048] Through the construction site scenario knowledge graph and the behavior rationality reasoning engine, the system can accurately understand contextual information such as the work environment, job characteristics, and operational necessity, thereby reducing the false alarm rate while maintaining a high detection rate.
[0049] Through microbehavior analysis and temporal attention networks, the system can identify workers' intentions to violate regulations in advance, improving the accuracy of prediction and providing a sufficient time window for preventive intervention.
[0050] Based on behavioral economics models and environmental factor analysis, the system can implement personalized preventive intervention measures for different worker characteristics and environmental conditions, thereby reducing the incidence of violations and improving the level of construction site safety management. Attached Figure Description
[0051] Figure 1 This is a flowchart of a deep learning-based intelligent safety helmet recognition and monitoring method in this invention;
[0052] Figure 2 This is a flowchart of the sub-steps of step 1 of the present invention;
[0053] Figure 3 This is a flowchart of the sub-steps of step 2 of the present invention;
[0054] Figure 4 This is a flowchart of the sub-steps of step 3 of the present invention;
[0055] Figure 5 This is a flowchart of a sub-step in step 4 of the present invention. Detailed Implementation
[0056] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0057] At least one embodiment of the present invention discloses a deep learning-based intelligent safety helmet recognition and monitoring method, such as... Figures 1 to 5 As shown, it includes the following steps:
[0058] Step 1: Use fine-grained human pose estimation technology to identify key points on the worker's head, and combine this with a head safety helmet relative position modeling algorithm to evaluate the compliance of the safety helmet wearing.
[0059] Sub-step 1.1, Human body key point detection;
[0060] A deep convolutional neural network model employing multi-scale feature fusion is used to process input surveillance video frames, identifying and locating the coordinates of key points on the human body. This neural network model includes a feature extraction backbone network, a multi-scale feature fusion module, and a key point regression module.
[0061] The feature extraction backbone network uses an improved ResNet structure, which reduces the gradient vanishing problem through residual connections and improves feature extraction capabilities. The multi-scale feature fusion module uses FPN (Feature Pyramid Network) technology, defined as follows:
[0062] ;
[0063] in, Indicates the first Layer fusion feature map Indicates the first Layer fusion feature map Indicates an upsampling operation. This represents the convolution operation. express Convolution operation, Indicates the backbone network Feature maps output by the layer.
[0064] The keypoint regression module predicts 17 standard keypoints of the human body using heatmaps, particularly 5 keypoints in the head region (both eyes, both ears, and the nose). Each keypoint... The heatmap is calculated as follows:
[0065] ;
[0066] in Indicate key points In coordinates The thermal response value at that location, and These represent the x and y coordinates in the image, respectively. and These are the key points The x and y coordinates of the true coordinates, It is the standard deviation parameter that controls the peak width in the heatmap. This represents an exponential function.
[0067] The output is a set of coordinates for 17 key points of each detected human body:
[0068] ;
[0069] in Represents the set of coordinates of key points on the human body. , , These represent the coordinates of the 1st, 2nd, and 17th keypoints, respectively. The 17 keypoints include the 5 keypoints of the head (both eyes, both ears, and nose) as well as other keypoints of the body (such as shoulders, elbows, wrists, hips, knees, and ankles).
[0070] In its implementation, this deep convolutional neural network employs a multi-stage cascaded structure, including an initial feature extraction stage and multiple refinement stages. The initial stage uses convolutional layers with a stride of 2 to downsample the input image, generating a low-resolution feature map. Subsequent refinement stages progressively increase the feature map resolution through transposed convolutions and fuse features at different scales using cross-stage connections. Network training utilizes the mean squared error loss function of keypoint heatmaps, and the optimization method employs stochastic gradient descent with momentum. The initial learning rate is set to 0.001 and dynamically adjusted using a cosine annealing strategy.
[0071] In construction site safety helmet monitoring scenarios, this network has been specifically optimized for common worker postures (such as bending over, looking up, and turning around). By increasing the proportion of training samples for specific postures, the detection accuracy of key head points under complex postures has been improved. The system can maintain a high accuracy rate in detecting key head points under actual working conditions such as partial worker occlusion, changes in lighting, and varying angles, providing reliable basic data for subsequent safety helmet wearing compliance assessments.
[0072] Sub-step 1.2, safety helmet detection and positioning;
[0073] A helmet detection model is built based on the improved YOLOv5 object detection algorithm. It processes the input video frames, detects helmets in the scene, and returns their position and type information. This model introduces an attention calculation module and an improved feature extraction network on the basis of standard YOLOv5 to improve detection accuracy and speed.
[0074] The model outputs the detection bounding box of the safety helmet:
[0075] ;
[0076] in This represents the set of safety helmet detection boxes. and These represent the coordinates of the top-left and bottom-right corners of the detection box, respectively. Indicates the detection confidence level. This indicates the category of safety helmets (different colors or types of safety helmets).
[0077] The improved YOLOv5 model implemented in this embodiment includes the following aspects:
[0078] First, depthwise separable convolution is introduced into the backbone network, which decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution, reducing the number of parameters and computational complexity.
[0079] Secondly, it integrates the channel attention calculation module and the spatial attention calculation module, and enhances the expressive power of safety helmet-related features by learning the importance weights of feature channels and spatial positions respectively;
[0080] Specifically, the channel attention module Defined as:
[0081] ;
[0082] in This indicates that the channel attention module is applied to the feature map. The processing results This is the input feature map for helmet detection. and These represent global average pooling and global max pooling, respectively. This represents a multilayer perceptron. This represents the Sigmoid activation function.
[0083] Furthermore, considering the varying sizes of safety helmets in construction site environments, the feature pyramid structure was improved, and an adaptive feature alignment module was added. Deformable convolution adjusts the receptive field size to better adapt to safety helmet targets of different proportions. During model training, hybrid precision training techniques and a progressive learning rate adjustment strategy were employed, and FocalLoss was introduced to replace the standard cross-entropy loss, better addressing the class imbalance problem.
[0084] In practical applications, the model has been specifically optimized for complex scenarios commonly encountered in construction site environments (such as occlusion, changes in lighting, and dense crowds), generating training samples that include these scenarios through data augmentation. A hierarchical classification system has been established for seven different colors of safety helmets (red, yellow, blue, white, orange, green, and gray) and three main types (standard, with communication equipment, and with protective face shield), enabling the system to simultaneously identify the presence, location, color, and type of safety helmets, providing richer information for subsequent analysis.
[0085] Sub-step 1.3: Modeling the relative position of the head and helmet;
[0086] The head key points detected in sub-step 1.1 and the helmet information detected in sub-step 1.2 are correlated to construct a three-dimensional spatial positional relationship model between the head and the helmet. This model is implemented through the following steps:
[0087] A 3D head model is constructed based on five key points (both eyes, both ears, and nose). The Perspective-n-Point (PnP) algorithm is used to estimate the 3D pose of the head from the 2D key points. The solution equation is expressed as:
[0088] ;
[0089] in A homogeneous representation of image coordinates. These are the 2D coordinates of the key points. These are the corresponding 3D coordinates. It is the camera intrinsic parameter matrix. and These are the rotation matrix and the translation vector, respectively. This represents the extrinsic parameter matrix of the camera. It is a scale factor. This indicates the transpose operation, and 1 indicates the additional dimension of the homogeneous coordinates;
[0090] By solving this system of equations, the position and orientation of the head in 3D space can be determined.
[0091] The helmet detection frame is converted into a cylindrical model in 3D space, and the spatial relationship parameters between the head model and the helmet model are calculated, including:
[0092] Distance between the center point of the head and the center point of the helmet ;
[0093] The angle between the head tilt angle and the main axis of the helmet ;
[0094] Distance between the top of the head and the bottom of the helmet .
[0095] Sub-step 1.4, wearing standardization scoring algorithm;
[0096] Based on the head-helmet spatial relationship model established in sub-step 1.3, a helmet-wearing compliance scoring algorithm is constructed. Scoring function. Defined as:
[0097] ;
[0098] in This indicates the score for proper helmet wearing. , , and These represent the weighting coefficients for the distance between the center point of the head and the center point of the helmet, the angle between the head tilt angle and the main axis of the helmet, the distance between the top of the head and the bottom of the helmet, and the buckle status, respectively. , , and These are the scoring functions for the corresponding parameters; This indicates the buckle status, and a specific pattern recognition algorithm is used to determine whether the helmet buckle is properly fastened.
[0099] Rating results The range is [0, 1], with higher values indicating more proper wearing; the system categorizes the wearing status into three levels based on the scoring results:
[0100] "Wearing it properly" );
[0101] "Minor irregularities" );
[0102] "Seriously irregular" ).
[0103] Through the above sub-steps, the output of this step is a safety helmet wearing compliance score and level classification for each detected worker, providing basic data for subsequent steps. Specific output includes:
[0104] Worker IDs and their corresponding keypoint coordinates;
[0105] The results of the safety helmet inspection include location, type, and confidence level;
[0106] Head-helmet spatial relationship parameters ( , , and state);
[0107] Wearing Standardization Score And the level of compliance with the wearing standards ("standard wearing", "minor non-compliance" or "serious non-compliance").
[0108] Step 2: Based on the assessment results of the standardization of safety helmet wearing, construct a knowledge graph of construction site scenarios, including a three-layer relationship model of safety regulations for different types of work, and analyze the rationality of workers' safety helmet operation behavior;
[0109] This step, based on the output of step 1, uses the worker's safety helmet wearing status information and combines it with a human behavior recognition algorithm to analyze the rationality of the worker's safety helmet operation behavior in order to distinguish between normal operation and violation; specifically, it includes the following sub-steps:
[0110] Sub-step 2.1: Construct a knowledge graph of the construction site scenario;
[0111] A construction site scenario knowledge graph is constructed, which includes a three-layer relationship model of job type, operation, and safety regulations. This knowledge graph comprises the following main components:
[0112] Node types include:
[0113] The job type node represents different types of jobs, such as welder, electrician, scaffolder, etc.
[0114] ;
[0115] Represents the set of job type nodes. , , They represent the 1st, 2nd, and 3rd respectively. Each work type node Indicates the number of nodes for a particular job type;
[0116] Operation nodes represent various operations that a worker may perform, such as welding, handling, climbing, etc.
[0117] ;
[0118] Represents the set of operation nodes. , , They represent the 1st, 2nd, and 3rd respectively. Each operation node Indicates the number of nodes being operated on;
[0119] Location nodes represent different areas or locations on the construction site:
[0120] ;
[0121] Represents a set of location nodes. , , They represent the 1st, 2nd, and 3rd respectively. Location nodes, Indicates the number of location nodes;
[0122] Safety specification nodes indicate various regulatory requirements related to wearing safety helmets:
[0123] ;
[0124] Represents a set of security specification nodes. , , They represent the 1st, 2nd, and 3rd respectively. Each security standard node This indicates the number of nodes in the security specification.
[0125] Edge type:
[0126] Job type - Operation side This indicates the operations typically performed by a particular job type;
[0127] Operation - Position Edge This indicates where a certain operation is typically performed;
[0128] Operation - Safety Guidelines This indicates the safety regulations applicable to a certain operation.
[0129] Edge weight Indicates the strength or applicability of the relationship, ranging from [0, 1].
[0130] Knowledge graphs in triplet form Storage, in Indicates the head entity. Indicates a relationship. Represents the tail entity; the atlas can be formally represented as:
[0131] ;
[0132] in This represents the entire knowledge graph. Represents a set of nodes. Represents the set of edges;
[0133] ;
[0134] ;
[0135] in The union of sets is represented by the union of sets. , , and These represent the sets of job type nodes, operation nodes, location nodes, and safety specification nodes, respectively. , and These represent the sets of edge pairs for job type-operation, operation-location, and operation-safety regulations, respectively.
[0136] In this embodiment, the construction site scene knowledge graph is constructed in the following way:
[0137] First, based on construction industry standards and safety regulations, 15 common trade nodes (including concrete workers, steelworkers, carpenters, electricians, welders, scaffolders, etc.), 42 basic operation nodes, 25 construction site location nodes, and 30 safety specification nodes are predefined.
[0138] Secondly, the basic relationships and weights between various nodes were initialized using expert knowledge and historical security management data.
[0139] The knowledge graph is stored using an attribute graph model and managed and queried using the graph database Neo4j. To improve the completeness and accuracy of the knowledge graph, the system implements an automated expansion mechanism: it automatically discovers implicit relationships through a combination of rule-based reasoning and statistical analysis; simultaneously, it dynamically adjusts relationship weights using reinforcement learning methods, based on feedback from on-site monitoring data, enabling the graph to adapt to the specific conditions of different construction sites.
[0140] In practical applications, this knowledge graph has been developed into specific versions for different types of construction projects (such as high-rise buildings, bridges, and tunnels). Each version contains the job-operation-safety regulation relationships unique to that type of project. Specifically, for special circumstances involving the use of safety helmets, the graph clearly defines 14 specific operational scenarios that allow for the temporary removal of safety helmets (such as short breaks in designated rest areas or specific inspections in enclosed spaces), along with their time limits and preconditions, providing a knowledge foundation for subsequent rationale analysis of behavior. The graph supports multi-angle queries, such as querying applicable safety regulations given job type and location, or determining safety helmet usage requirements given operation and time, greatly improving the system's ability to understand complex construction site scenarios.
[0141] Sub-step 2.2, worker behavior sequence encoding;
[0142] Human behavior in consecutive video frames is identified and encoded to generate worker behavior sequence vectors; the specific implementation steps are as follows:
[0143] Using 3D convolutional neural networks to extract spatiotemporal features, given continuous... Frame video sequence:
[0144] ;
[0145] in Represents a video sequence. , , Representing the 1st frame, the 2nd frame, and the 3rd frame respectively. Frame video sequence, This indicates the total number of frames in the video sequence.
[0146] Extracting spatiotemporal features using a 3D-CNN model:
[0147] ;
[0148] in Indicates spatiotemporal characteristics, This represents a 3D-CNN model with output feature dimensions of . .
[0149] Long Short-Term Memory (LSTM) networks are used to model temporal features and capture long-term dependencies in behavioral sequences.
[0150] ;
[0151] in Is it LSTM at time step The hidden state, It is a time step Features This represents the Long Short-Term Memory (LSTM) network.
[0152] Classify behavior categories using MLP (Multilayer Perceptron) and the Softmax function:
[0153] ;
[0154] in It is the probability distribution of behavior categories. It is the hidden state of the last time step. This represents the Softmax function. This represents a multilayer perceptron.
[0155] The output is a sequence of worker behaviors:
[0156] ;
[0157] in Represents a sequence of worker behaviors. , , They represent the 1st, 2nd, and 3rd respectively. Each identified behavior type This indicates the total number of behavior types, which include helmet-related actions such as removing the helmet, adjusting the helmet, and wiping sweat.
[0158] Sub-step 2.3: Reasoning about the rationality of behavior based on knowledge graph;
[0159] Based on the constructed construction site scene knowledge graph and the identified worker behavior sequences, a graph inference engine is built to analyze the rationality of workers removing their hats. The inference process mainly includes the following steps:
[0160] Behavioral intention probability calculation, given a sequence of worker behaviors Location information and time information Calculate the worker's possible behavioral intentions The probability of:
[0161] ;
[0162] in Represents the probability distribution of behavioral intentions. This represents a sequence of worker behaviors observed under given intent, location, and time conditions. The probability, This represents the prior probability of an intention given a location and time. It represents the marginal probability of a sequence of actions given a location and time condition.
[0163] The behavioral rationality score assesses the rationality of the worker's act of removing his hat based on the calculated probability distribution of his intention.
[0164] ;
[0165] in Indicates the reasonableness score of the behavior. This indicates a normal set of intentions (such as adjusting a helmet, briefly removing the helmet to wipe sweat, etc.). This refers to a collection of actions that indicate a violation of regulations (such as removing a helmet for an extended period of time or deliberately not wearing one). express The index of an element in a set express The index of an element in a set This indicates a summation operation.
[0166] Reasonableness assessment, based on reasonableness score. The worker's act of removing his hat is categorized as follows:
[0167] "Reasonable conduct": ;
[0168] "Suspicious behavior":
[0169] “Violation”: ;
[0170] in, and These represent the threshold parameters for reasonable and illegal behaviors, respectively.
[0171] Through the above sub-steps, the output of this step is a reasonableness score and classification result for the worker's safety helmet operation behavior, used to distinguish between normal operation and violation; the specific output includes:
[0172] Construction Site Scene Knowledge Graph This includes job type, operation, location, and safety specification nodes and their relationships;
[0173] Worker Behavior Sequence It records operations related to the safety helmet;
[0174] Behavioral Intent Probability Distribution ;
[0175] Behavioral rationality score ;
[0176] The results of the behavior reasonableness classification (“reasonable behavior”, “suspicious behavior”, or “violation”).
[0177] Step 3: Based on the results of the safety helmet wearing standardization assessment and the results of the behavior rationality analysis, extract the characteristics of the violation intention from the worker's micro-behavioral sequence, and learn the precursor patterns of the violation through a temporal attention network;
[0178] This step utilizes the outputs of steps 1 and 2, including information on workers' helmet wearing status and the results of behavioral rationality analysis. By analyzing the workers' micro-behavioral sequences, it identifies and learns the precursor patterns of violations, enabling the prediction of intentions to violate regulations. Specifically, it includes the following sub-steps:
[0179] Sub-step 3.1, Micro-behavioral feature extraction;
[0180] Extracting workers' micro-behavioral features from surveillance videos, including but not limited to the following categories:
[0181] Head posture change features are extracted by tracking the movement trajectory of key points on the head, such as frequent head turning, looking left and right, etc. Head posture features are defined as follows:
[0182] ;
[0183] in Indicates head posture characteristics. , and Representing time steps The nose yaw angle, pitch angle, and roll angle, Indicates a time step. This represents the total number of time steps.
[0184] Body language characteristics are analyzed by examining workers' postures and movements, such as frequently approaching blind spots or hidden corners, or creating obstructive relationships with colleagues. Body language characteristics are defined as follows:
[0185] ;
[0186] in Indicates body language characteristics, Indicates time step The set of key point locations and These represent the corresponding velocity and acceleration characteristics, respectively.
[0187] Social interaction features extract the interaction patterns between workers and their colleagues, such as specific gestures, dialogues, or collaborative behaviors. Social interaction features are defined as follows:
[0188] ;
[0189] in Indicates characteristics of social interaction. Indicates workers and workers Interaction feature descriptors between them Represents a group of workers. Indicates workers and workers They are not the same person.
[0190] Spatiotemporal trajectory features are analyzed to determine the movement trajectories and dwell time patterns of workers in different areas of the construction site. The spatiotemporal trajectory features are defined as follows:
[0191] ;
[0192] in Represents spatiotemporal trajectory characteristics, Indicates time step Spatial location coordinates, The semantic label indicating the location (such as danger zone, rest area, etc.).
[0193] The above features are fused through multimodal feature fusion to obtain the final micro-behavioral feature vector:
[0194] ;
[0195] in Represents the feature vector of micro-behavior. , , and These represent head posture change features, body language features, social interaction features, and spatiotemporal trajectory features, respectively.
[0196] Sub-step 3.2: Construction of the temporal attention network;
[0197] A temporal attention network is constructed to model the historical behavioral sequences of workers and learn the precursor patterns of violations. The network structure includes the following components:
[0198] The feature embedding layer takes the micro-behavioral feature vector extracted in sub-step 3.1 and embeds it into the feature embedding layer. Mapping to the latent space yields the embedded representation:
[0199] ;
[0200] in Represents the feature vector of micro-behavior Embedded representation, and These are the weight matrix and bias vector of the embedding layer, respectively.
[0201] The temporal coding layer uses a bidirectional LSTM network to encode behavioral sequences and capture temporal dependencies.
[0202] ;
[0203] ;
[0204] ;
[0205] in, and These are the forward and backward LSTMs at time steps. The hidden state, It is the result of their splicing. Indicates time step Micro-behavioral feature vectors Embedded representation, and These are forward and backward LSTMs, respectively. Indicates a time step. This represents the total number of time steps. This represents the concatenation operator. , Representing time steps and time step The hidden state.
[0206] Multi-head self-attention model: This model calculates the attention weights at each time step in the action sequence.
[0207] ;
[0208] ;
[0209] ;
[0210] in , , , , , , , , , , , , These represent the query matrix, key matrix, value matrix, hidden state matrix, query transformation matrix, key transformation matrix, value transformation matrix, output transformation matrix, and the first... The output of each attention head, the number of attention heads, the total number of time steps, and the dimension of the key vector. This represents the Softmax function. This represents the concatenation operator. This represents the transpose operator. This indicates the attention calculation method. This represents a multi-head attention calculation method.
[0211] The layer predicts the probability of illegal intent based on attention-weighted sequence representation:
[0212] ;
[0213] in Indicates at time step The probability of a violation of intent. The weights are calculated using an attention-based computation method. It is a prediction function that includes a fully connected layer and a sigmoid activation function. Indicates time step Micro-behavioral feature vectors The embedded representation is weighted and summed. Indicates time step Micro-behavioral feature vectors Embedded representation.
[0214] The temporal attention network in this embodiment is specifically implemented as follows: The LSTM unit uses a gated version, namely GRU (Gated Recurrent Unit), which controls the information flow through reset and update gates, reducing the gradient vanishing problem and improving the ability to model long sequences. The feature embedding layer uses 256-dimensional hidden states, the LSTM hidden layer has a dimension of 128, and the number of attention heads is 8, with each head having a dimension of 16. To enhance the network's generalization ability, a Dropout layer (with a dropout rate of 0.2) and layer normalization operations are added to reduce the risk of overfitting.
[0215] The network training employs a two-stage strategy: the first stage uses supervised learning, training the network to recognize known violation patterns using labeled violation data; the second stage combines contrastive learning, constructing positive and negative sample pairs to enhance the network's ability to distinguish similar violation patterns. The loss function combines binary cross-entropy loss and triplet loss, and the optimization algorithm uses the Adam optimizer with an initial learning rate of 0.0005, dynamically adjusted using an exponential decay strategy.
[0216] In construction site safety monitoring applications, this network can process behavioral sequences of up to 120 seconds, identifying various micro-behavioral patterns of workers before removing their safety helmets. The system is specifically optimized for seven common warning signs of violations on construction sites (such as frequently raising hands to touch safety helmets, repeatedly glancing at monitoring equipment, whispering to colleagues before moving quickly, etc.), and can distinguish subtle differences between these behaviors and normal work actions. By analyzing the temporal characteristics of historical violation cases, the system has built a behavioral pattern library, enabling the network to quickly adapt to and identify potential violations in new construction site environments.
[0217] Sub-step 3.3, learning and detecting early signs of violations;
[0218] Based on the constructed temporal attention network, the system learns and detects precursor patterns of violations. The specific steps are as follows:
[0219] Model training: A temporal attention network is trained using a labeled action sequence dataset.
[0220] ;
[0221] in Represents the training dataset. It is the first The behavioral sequence of a sample These are the corresponding tags (0 indicates normal behavior, 1 indicates violation). Represents the total number of samples. Indicates from time step 1 to The complete sequence;
[0222] The loss function is defined as:
[0223] ;
[0224] in Represents the total loss function. It is a regularization term used to prevent the model from overfitting. It is the regularization coefficient, which controls the strength of regularization. Indicates the sample index. Represents the total number of samples. Represents the logarithmic function. Represents a given sequence of actions Predicted probability of intent to violate regulations under given conditions. Indicates the first The true labels of each sample (0 represents normal behavior, 1 represents illegal behavior). This represents the summation over all samples; the first part of the loss function is the binary cross-entropy loss, which measures the difference between the predicted probability and the true label, and the second part is the regularization term, which controls the model complexity.
[0225] Precursor pattern extraction involves analyzing the attention distribution of a trained model on a sequence of violations to extract key precursor patterns.
[0226] ;
[0227] in This represents the set of extracted early warning patterns of violations. This is the attention weight threshold; only behaviors with an attention weight exceeding this threshold are considered part of the violation precursor pattern. Indicates time step , behavior Indicates time step Attention weights Indicates the time step index. This indicates the total length of the sequence.
[0228] Real-time violation intent detection, analyzing worker behavior sequences in online surveillance videos. Using the trained model, calculate the probability of a violation of intent:
[0229] ;
[0230] in, This represents the probability of predicting a violation intent given a sequence of online behaviors. This represents the probability of the final intent to violate the rules.
[0231] When the probability of illegal intent Exceeding the preset threshold At that time, the system identifies the intent to violate the rules and issues a corresponding warning.
[0232] Through the above sub-steps, the output of this step is the probability of the worker's intent to violate regulations and the identification results of early warning patterns of violations, providing a basis for subsequent risk assessment and intervention measures. Specific outputs include:
[0233] Worker's micro-behavioral feature vector ;
[0234] Temporal attention network computes behavioral sequence attention weights ;
[0235] Probability of illegal intent ;
[0236] Extracted pattern of early signs of violation .
[0237] Step 4: Based on the comprehensive assessment results of helmet wearing compliance, the analysis results of behavioral rationality, and the warning patterns of violations, estimate workers' risk perception and decision-making tendencies using a behavioral economics model, construct a violation risk scoring system, and implement preventive intervention measures.
[0238] This step comprehensively utilizes the outputs of the previous three steps, including the helmet-wearing compliance score (Step 1), the behavioral rationality analysis results (Step 2), and the probability of violation intent (Step 3). By constructing a behavioral economics model and combining it with environmental factor analysis, it comprehensively assesses the risk of worker violations and implements preventative intervention measures. Specifically, it includes the following sub-steps:
[0239] Sub-step 4.1, Risk perception and decision-making model construction;
[0240] A worker risk perception and decision-making model is constructed based on behavioral economics theory to assess workers' risk perception and decision-making tendencies in different situations. The specific implementation method is as follows:
[0241] The prospect theory model uses the prospect theory framework to assess workers' subjective perception of risk and decision preferences.
[0242] In this model, the worker's subjective value function and probability weight function are defined as follows:
[0243] ;
[0244] ;
[0245] in Indicates the worker's perception of gain or loss Subjective value assessment Indicates the worker's perception of objective probability Subjective weight conversion, Indicates gain or loss. Represents objective probability. and It is a risk preference parameter. It is the loss aversion coefficient. These are probability weight parameters;
[0246] These parameters are estimated individually by analyzing workers' historical behavioral data and questionnaire results.
[0247] The decision framing effect model analyzes the differences in workers' decision preferences under different framing (e.g., emphasizing safety benefits or emphasizing violation penalties); the differences in decision preferences are calculated as follows:
[0248] ;
[0249] in This indicates the degree of preference difference caused by the decision-making framing effect. This represents the probability of choice within the payoff framework. This represents the choice probability within the loss framework.
[0250] Social influence factor, which considers the impact of peer behavior and group norms on individual decision-making; social influence factor is defined as:
[0251] ;
[0252] in Indicates the overall intensity of social impact. and These represent the number of companions who complied with the rules and the number who violated them, respectively. This indicates the strength of an authority figure's influence. , and These represent the weighting coefficients for the influence of peer compliance, peer violation, and authority, respectively.
[0253] Sub-step 4.2, Analysis of environmental factors and working conditions;
[0254] Analyze the environmental factors and working conditions that influence worker safety behavior, and construct a multi-dimensional influencing factor model:
[0255] ;
[0256] in Represents the feature vector of environmental factors. , , , These represent temperature, humidity, noise level, and lighting conditions, respectively.
[0257] ;
[0258] in Represents the feature vector of working state. , , , These represent work duration, work intensity, rest interval, and task complexity, respectively.
[0259] Collect workers' physiological indicators using wearable devices:
[0260] ;
[0261] in Represents a feature vector of physiological state. , , These represent heart rate, respiratory rate, and body temperature, respectively.
[0262] The environment-work status integrated impact model inputs the aforementioned feature vectors into a multilayer perceptron (MLP) to calculate the combined impact factor of environment and work status. :
[0263] ;
[0264] in This represents the combined influencing factors of environment and working conditions; MLP stands for Multilayer Perceptron Neural Network Model. , , These represent the feature vectors of environmental factors, working state, and physiological state, respectively.
[0265] The range of the combined influence factor of environment and working conditions is [0, 1]. The larger the value, the more likely the current environment and working conditions are to lead to worker violations.
[0266] Sub-step 4.3: Construction of the violation risk scoring system;
[0267] Based on the outputs of the first two sub-steps, a comprehensive violation risk scoring system is constructed to assess the level of worker violation risk in real time:
[0268] Violation Risk Scoring Function Defined as:
[0269] ;
[0270] in This represents the worker's overall risk score for violations. It is the probability of illegal intent output in step 3. It is the subjective expected value of the risk perception model. Indicates gain or loss. Represents objective probability. It is a social impact factor. It is a comprehensive influencing factor of environment and working conditions. , , and These represent the weighting coefficients of the probability of intent to violate regulations, subjective expected value, social impact factor, and the comprehensive impact factor of the environment and working conditions, respectively.
[0271] Risk levels are classified based on the calculated violation risk score:
[0272] "Low risk": ;
[0273] "Medium risk": ;
[0274] "High risk": ;
[0275] in, and These represent the risk threshold parameters for low and high risk, respectively.
[0276] The risk warning method triggers a corresponding level of risk warning when the calculated risk score exceeds a specific threshold.
[0277] Low risk: Regular safety reminders;
[0278] Medium risk: Close monitoring and targeted safety alerts;
[0279] High risk: Immediate intervention and notification to management.
[0280] Sub-step 4.4: Implementation of personalized preventive intervention measures;
[0281] Based on risk scoring results and individual worker characteristics, develop and implement personalized preventative intervention measures:
[0282] An intervention strategy selection algorithm, based on reinforcement learning, selects the optimal intervention strategy for workers with different risk characteristics. Strategy selection function. Defined as:
[0283] ;
[0284] in Indicates the state Next-choice intervention strategy The probability distribution or decision rule, It indicates the worker's current status (including risk score, individual characteristics, etc.). Indicate possible intervention strategies, Indicates the state The following strategy The expected utility Indicates the choice to make The biggest strategy This function evaluates the expected effects of different intervention strategies and selects the most suitable personalized intervention for the current worker's condition.
[0285] Intervention types and content: The system supports the following intervention types:
[0286] Real-time safety alerts: Personalized alert methods (visual, auditory, or tactile) are used based on workers' risk perception characteristics.
[0287] Incentive system: Design positive incentives or negative warnings based on workers' decision-making preferences;
[0288] Safety knowledge delivery: Delivering customized safety knowledge to address workers' risk perception biases;
[0289] Work environment adjustment: Based on the analysis of environmental impact factors, adjust the work environment parameters or arrange appropriate rest periods;
[0290] Management intervention: For high-risk workers, notify on-site management personnel for direct intervention.
[0291] Intervention effectiveness evaluation and optimization: A / B testing was used to evaluate the effectiveness of different intervention strategies, and the intervention strategies were continuously optimized using a multi-armed bandit algorithm.
[0292] ;
[0293] in Representation strategy The upper bound of confidence, It is a strategy The estimated utility It is the current time step. It is a strategy Number of selections, These are exploration parameters. This represents the natural logarithm function.
[0294] Through the above sub-steps, the output of this step is a worker's violation risk score, risk level classification, and personalized preventive intervention measures, realizing a shift from a passive response to a proactive prevention-based safety management model. Specific outputs include:
[0295] Workers' risk perception and decision-making model parameters ( , , , wait);
[0296] Comprehensive influencing factors of environment and working conditions ;
[0297] Violation Risk Score ;
[0298] Risk level classification results ("low risk", "medium risk" or "high risk");
[0299] Personalized intervention strategies And its execution results.
[0300] These outputs together constitute the final result of this method, realizing full-process management of smart safety helmet monitoring, from wearing compliance detection, behavior rationality analysis, prediction of violation intent to risk assessment and intervention, forming a complete closed-loop system.
[0301] A deep learning-based intelligent safety helmet recognition and monitoring system, used to execute the aforementioned deep learning-based intelligent safety helmet recognition and monitoring method, includes:
[0302] A fine-grained human pose estimation module is used to identify key points on the worker's head and, combined with a modeling algorithm for the relative position of the head and safety helmet, to evaluate the compliance of the safety helmet wearing.
[0303] The construction site scenario knowledge graph module is used to construct a knowledge graph containing a three-layer relationship model of work operation safety regulations, and to analyze the rationality of workers' safety helmet operation behavior;
[0304] The micro-behavior analysis module is used to extract features of the intent to violate regulations from the worker's micro-behavior sequence and learn the precursor patterns of the violation through a temporal attention network;
[0305] The risk assessment and intervention module is used to estimate workers' risk perception and decision-making tendencies based on behavioral economics models, construct a violation risk scoring system, and implement preventive intervention measures.
[0306] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A deep learning-based intelligent safety helmet recognition monitoring method, characterized in that, The method comprises the following steps: Identify worker head key points using fine-grained human pose estimation technology, and evaluate safety helmet wearing standardization by combining head safety helmet relative position modeling algorithm; Based on the safety helmet wearing standardization evaluation results, construct a construction site scene knowledge graph, including a three-layer relationship model of operation safety specifications for different types of work, to analyze the rationality of workers' safety helmet operation behavior; According to the safety helmet wearing standardization evaluation results and the behavior rationality analysis results, extract the rule violation intention features from the worker's micro-behavior sequence, and learn the precursor pattern of the rule violation behavior through the time sequence attention network; Integrate the safety helmet wearing standardization evaluation results, the behavior rationality analysis results and the rule violation behavior precursor pattern, estimate the worker's risk perception and decision-making tendency based on the behavioral economics model, construct a rule violation risk scoring system and implement preventive intervention measures. 2.The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The fine-grained human pose estimation technology comprises: Using a deep convolutional neural network with multi-scale feature fusion to process the input monitoring video frames, identifying and locating the key point coordinates of the human body; The head safety helmet relative position modeling algorithm comprises: The detected head key points are constructed into a head 3D model, and the safety helmet detection box is converted into a cylindrical model in 3D space, and the spatial relationship parameters between the head model and the safety helmet model are calculated; Based on the spatial relationship parameters, a wearing standardization scoring function is constructed to quantitatively evaluate the standardization degree of safety helmet wearing. 3.The intelligent safety helmet recognition monitoring method based on deep learning according to claim 2, characterized in that, The spatial relationship parameters include the distance between the head center point and the safety helmet center point, the included angle between the head tilt angle and the safety helmet main axis, the distance between the head top and the safety helmet bottom, and the buckle state.
4. The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The node types of the construction site scene knowledge graph include work type nodes, operation nodes, location nodes and safety specification nodes; Edge types include work operation edges, operation location edges and operation safety specification edges; Based on the knowledge graph, the worker's hat removal behavior rationality is analyzed, including: Calculate the worker's possible behavior intention probability under the condition of given behavior sequence, location information and time information; According to the calculated intention probability distribution, the rationality of the worker's hat removal behavior is evaluated; According to the rationality score, the worker's hat removal behavior is classified as reasonable behavior, suspicious behavior or rule violation behavior.
5. The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The rule violation intention features extracted from the worker's micro-behavior sequence include: Head posture change features, extracted by tracking the motion trajectory of the head key points; Body language features, analyzing the worker's body posture and action; Social interaction features, extracting the interaction mode of the worker with the surrounding colleagues; Space-time trajectory features, analyzing the worker's moving trajectory and staying time pattern in different areas of the construction site.
6. The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The time sequence attention network comprises a feature embedding layer, a time sequence encoding layer, a multi-head self-attention model and a rule violation intention prediction layer, the time sequence encoding layer uses a bidirectional LSTM network to encode the behavior sequence, capturing the time sequence dependency, and the multi-head self-attention model calculates the attention weight of each time step in the behavior sequence through multi-head self-attention.
7. The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The behavioral economics model comprises: Prospect theory model, to evaluate the worker's subjective perception of risk and decision-making preference; Decision framework effect model, to analyze the worker's decision-making preference difference under different frameworks; Social influence factor model, considering the influence of peer behavior and group norms on individual decision-making. 8.The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The violation risk scoring system also takes into account the following factors: Environmental factor feature vector, including temperature, humidity, noise level and lighting conditions; Working state feature vector, including working duration, working intensity, rest interval and task complexity; Physiological state feature vector, including heart rate, respiratory rate and body temperature. 9.The intelligent safety helmet recognition monitoring method based on deep learning according to claim 1, characterized in that, The preventive intervention measures include: Real-time safety reminders, using personalized reminder methods according to the worker's risk perception characteristics; Incentive system, providing positive incentives or negative warnings based on the worker's decision-making preferences; Safety knowledge push, pushing customized safety knowledge according to the worker's risk cognitive bias; Work environment adjustment, adjusting work environment parameters or arranging appropriate rest according to environmental influence factor analysis; Management intervention, notifying on-site managers to intervene directly for high-risk workers. 10.A deep learning-based intelligent safety helmet recognition monitoring system, characterized in that, An intelligent safety helmet recognition monitoring method based on deep learning for performing any one of claims 1-9, comprising: A fine-grained human pose estimation module for identifying worker head key points and evaluating safety helmet wearing specifications using head safety helmet relative position modeling algorithms; A construction site scene knowledge graph module for constructing a knowledge graph containing a three-layer relationship model of operation safety specifications for different types of work and analyzing the rationality of worker safety helmet operation behavior; A micro-behavior analysis module for extracting violation intention features from the worker's micro-behavior sequence and learning the precursor patterns of violation behavior through a time series attention network; A risk assessment and intervention module for estimating the worker's risk perception and decision-making tendency based on a behavioral economics model, constructing a violation risk scoring system and implementing preventive intervention measures.