Construction site safety behavior real-time early warning system based on multi-mode perception
By constructing a multi-layered technical architecture and multi-modal data fusion, the problems of inefficient data fusion and static early warning strategies in construction site safety monitoring systems have been solved, enabling comprehensive monitoring and accurate early warning of construction site safety conditions, and improving the system's real-time performance and adaptability.
Patent Information
- Application Number
- CN202511926496.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing construction site safety monitoring systems suffer from inefficient data fusion methods, weak edge computing capabilities, static early warning strategies, and insufficient model generalization capabilities, resulting in poor real-time performance and accuracy, making it difficult to adapt to the dynamic environment and diverse needs of construction sites.
A multi-layered technical architecture is constructed, including a multimodal perception layer, an edge computing layer, and a cloud service layer. Multimodal data fusion, real-time processing, and dynamic early warning technologies are adopted. Data is collected using visual sensors, millimeter-wave radar, wearable devices, and environmental sensors. Three-dimensional spatial information is processed in combination with BIM models. Data fusion is performed through an improved Transformer model and DS evidence theory. Reinforcement learning algorithms are used to dynamically adjust the early warning threshold, and federated learning is used to optimize the model.
It enables comprehensive monitoring and precise early warning of safety conditions at construction sites, improves the accuracy and timeliness of monitoring, reduces false alarms and missed alarms, supports rapid identification and adaptation to new behaviors, and enhances the system's real-time performance and generalization capabilities.
Smart Images

Figure CN121789380A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of construction safety management technology, and more particularly to a real-time early warning system for construction site safety behaviors based on multimodal perception. Background Technology
[0002] As a vital pillar of the national economy, the construction industry has always been a major concern regarding workplace safety. According to data from the National Bureau of Statistics, 593 construction accidents occurred nationwide in 2023, resulting in 687 deaths. The main types of accidents included falls from heights, being struck by objects, crane-related injuries, and collapses. These accidents not only caused casualties but also resulted in significant economic losses and social impact.
[0003] The construction site environment is complex, with numerous dynamically changing risk factors, mainly including: 1. Personnel behavior risks: worker violations (such as not wearing safety protective equipment, unauthorized changes to construction procedures), fatigue, etc. 2. Equipment operation risks: crane malfunctions, improper operation of construction equipment, electrical equipment leakage, etc. 3. Environmental condition risks: severe weather (such as strong winds, heavy rain), high dust levels, low light levels, confined space operations, etc. 4. Management and coordination risks: unclear division of construction areas, insufficient coordination of multiple trades working simultaneously, inadequate safety supervision, etc.
[0004] Traditional construction site safety monitoring mainly relies on the following methods: 1. Manual inspection: Safety management personnel regularly inspect the site to identify safety hazards and urge rectification. This method suffers from strong subjectivity, limited coverage, and poor real-time performance, making it difficult to respond promptly to sudden risks. 2. Single sensor monitoring: Such as deploying only cameras for video monitoring or installing a single type of sensor (e.g., temperature sensor, gas sensor). This method has the following drawbacks: Incomplete monitoring information: A single sensor can only acquire specific types of data and cannot comprehensively reflect the safety status of the construction site. For example, video monitoring is easily affected by factors such as lighting and obstruction, and the accuracy of recognition decreases in complex environments. Insufficient data fusion: Data from different types of sensors is not effectively integrated, making it impossible to form a comprehensive safety assessment system. Limited real-time processing capabilities: A large amount of data needs to be transmitted to the cloud for processing, resulting in response delays and failing to meet the needs of real-time early warning.
[0005] In recent years, with the development of sensor technology and artificial intelligence technology, some studies have attempted to introduce multimodal perception technology (such as fusing video and sensor data) for construction site safety monitoring. However, existing technologies still have the following problems: 1. Inefficient data fusion methods: Most studies adopt simple feature stitching or early fusion strategies, failing to fully explore the spatiotemporal correlation and complementary information between different modal data. For example, the time synchronization and spatial calibration problems of video images and millimeter-wave radar data have not been effectively solved, resulting in low fusion accuracy. 2. Weak edge computing capabilities: Complex deep learning models rely on cloud servers for computation, while edge devices (such as smart terminals deployed on-site) have limited computing resources, making it difficult to support real-time multimodal data processing and behavior recognition, resulting in high overall system latency. 3. Static early warning strategies: The early warning thresholds of existing systems are usually set based on fixed rules or historical experience, and cannot be adaptively adjusted according to the dynamic environment of the construction site (such as changes in lighting, equipment status, and personnel fatigue), which easily leads to false alarms or missed alarms. 4. Insufficient model generalization ability: Traditional deep learning models perform poorly in recognizing new scenarios or new types of safety behaviors, requiring a large amount of labeled data for retraining, and are difficult to quickly adapt to the diverse needs of construction sites. Summary of the Invention
[0006] The purpose of this invention is to provide a real-time early warning system for construction site safety behavior based on multimodal perception, which achieves efficient fusion, real-time processing, and dynamic early warning of multi-source data by constructing a multi-layered technical architecture, thereby significantly improving the safety monitoring of construction sites.
[0007] This invention is achieved through the following measures: A real-time early warning system for construction site safety behaviors based on multimodal perception, characterized in that: It includes a multimodal perception layer, an edge computing layer, and a cloud service layer; The multimodal perception layer includes visual sensors, millimeter-wave radar, wearable devices, and environmental sensors. The multimodal perception layer is used to collect multi-dimensional data from the construction site and combine it with BIM to obtain three-dimensional spatial information and hazardous area markings of the construction site. The multi-dimensional data includes at least video image data collected by a visual sensor, three-dimensional position, speed and direction of movement of personnel and equipment collected by millimeter-wave radar, physiological and motion data of workers collected by wearable devices, and environmental parameters collected by environmental sensors; wherein, the environmental parameters include humidity, dust, noise, light and gas concentration, etc., and the video image data includes personnel posture, equipment status and wearing of safety protective equipment. The edge computing layer is connected to the multimodal perception layer and is used for real-time processing and fusion of multimodal data. It includes at least a data preprocessing module, a multimodal fusion module, a real-time behavior recognition module, and a dynamic early warning engine module. The data preprocessing module is used to obtain spatiotemporally aligned multimodal feature data based on the multidimensional data output by the multimodal perception layer and the BIM model, performing time synchronization, spatial calibration, noise filtering, and feature extraction. The multimodal fusion module is used to perform cross-modal feature interaction and fusion on the spatiotemporally aligned multimodal feature data, generating a fused feature representation using a cross-attention mechanism and a dynamic weight allocation algorithm, and generating a fused decision result using DS evidence theory. The real-time behavior recognition module is used to output behavior recognition results of personnel safety behaviors based on the fused feature representation and fused decision results, combined with a spatiotemporal perception network and zero-shot learning technology. The dynamic early warning engine module is used to assess risk levels and adjust adaptive early warning thresholds based on the behavior recognition results and multidimensional risk factors, outputting the risk level and adaptive early warning threshold. The cloud service layer is connected to the edge computing layer and is used for global data management and model optimization. It includes at least a data storage and analysis module, a federated learning module, and a policy management module. The data storage and analysis module is used to store the behavior recognition results, early warning records, and historical multimodal data and generate a security situation report. The federated learning module is used to coordinate with each edge node to perform model training and global model updates and to send the updated model parameters to the edge computing layer. The policy management module is used to configure and optimize early warning strategies based on the risk level and the security situation report.
[0008] The invention also has the following specific features: Preferably, the visual sensor includes a 4K high-definition camera and a panoramic camera, the millimeter-wave radar is a 77GHz radar with a ranging accuracy of ±0.1 meters, and the wearable device includes a smart safety helmet and a smart bracelet that integrate heart rate, body temperature, and acceleration sensors.
[0009] Preferably, the data preprocessing module includes a time synchronization algorithm and a spatial calibration algorithm. The time synchronization error is less than 1ms, and the spatial calibration unifies the data of each modality to the global coordinate system of the BIM model.
[0010] Preferably, the multimodal fusion module adopts a "feature-level fusion + decision-level fusion" architecture. The feature-level fusion uses an improved Transformer model to perform cross-modal feature interaction, and the decision-level fusion uses DS evidence theory for decision fusion.
[0011] Preferably, the spatiotemporal awareness network consists of 3DCNN and LSTM, used to extract spatiotemporal features from video image data and identify behavior categories; Preferably, the zero-shot learning technique is used to support the generation of new behavior detection models through natural language descriptions, so as to expand the behavior recognition capability without a large amount of labeled data when new types of security behaviors emerge.
[0012] Preferably, the dynamic early warning engine module uses a reinforcement learning algorithm to dynamically adjust the early warning threshold. The state space includes environmental parameters, equipment status, personnel physiological status, and historical early warning records. The reward function is a weighted sum of early warning accuracy and response time.
[0013] Preferably, the federated learning module of the cloud service layer adopts horizontal federated learning technology, whereby edge nodes upload model gradients or parameter update information, and the cloud generates a global model through weighted average aggregation, and uses homomorphic encryption and differential privacy to protect data security.
[0014] Preferably, the BIM model is fused with real-time sensing data to generate a three-dimensional risk heat map, marking dangerous areas and displaying the real-time location of personnel and equipment; Preferably, the edge computing layer adopts the NVIDIA Jetson AGX Orin edge computing platform, which supports 5G / wired dual-link communication, realizes local real-time inference, and has a response time of less than 2 seconds.
[0015] Preferably, the implementation method of the system includes the following steps: S1. Collect multi-dimensional data through visual sensors, millimeter-wave radar, wearable devices and environmental sensors, and obtain the three-dimensional spatial information and hazardous area markings of the BIM model to form original multimodal perception data and BIM spatial benchmark information. S2. Using the original multimodal sensing data and BIM spatial reference information as input, perform time synchronization, spatial calibration, noise filtering and feature extraction to unify the modal data into the global coordinate system of the BIM model and obtain spatiotemporally aligned multimodal feature data. S3. Using spatiotemporally aligned multimodal feature data as input, a combination of feature-level fusion and decision-level fusion is adopted to achieve cross-modal feature interaction and fusion, and output fused feature representation and fusion decision results; S4. Using fusion feature representations or fusion decision results as input, based on spatiotemporal perception networks and combined with zero-shot learning techniques, identify personnel safety behavior categories and output behavior recognition results. S5. Using the behavior recognition results as input, the risk level is assessed by combining environmental parameters, equipment status, personnel physiological status and historical warning records. The warning threshold is dynamically adjusted using reinforcement learning algorithms, and the risk level and adaptive warning threshold are output. S6. Using risk level and adaptive warning threshold as input, trigger a graded warning response, and fuse the BIM model with real-time perception data to generate a three-dimensional risk heat map. At the same time, upload the behavior recognition results, warning records and related multimodal data to the cloud service layer for data storage and analysis. Through federated learning, realize the collaborative training of models of each edge node and the global model update, so as to send the updated model to the edge computing layer for subsequent real-time behavior recognition.
[0016] The graded early warning response includes a three-level early warning mechanism: Level 1 early warning alerts workers through the vibration and directional sound waves of smart safety helmets; Level 2 early warning triggers an audible and visual alarm and pushes a notification to management personnel; Level 3 early warning automatically shuts down equipment and activates the emergency plan.
[0017] The beneficial effects of this invention are as follows: 1. This invention collects multi-dimensional data through visual sensors, millimeter-wave radar, wearable devices, and environmental sensors, and combines this with BIM models to obtain three-dimensional spatial information and hazardous area markings. It adopts a "feature-level fusion + decision-level fusion" architecture, utilizing an improved Transformer model and DS evidence theory to fully explore the spatiotemporal correlation and complementary information between different modal data. This solves the problems of incomplete information and insufficient data fusion in traditional single-sensor monitoring, comprehensively reflects the safety status of the construction site, and improves monitoring accuracy.
[0018] 2. This invention employs a reinforcement learning algorithm. Based on environmental parameters, equipment status, personnel physiological status, and historical early warning records, the algorithm dynamically adjusts the early warning threshold using a weighted sum of early warning accuracy and response time as the reward function. This adapts to the dynamic environment of the construction site, reduces false alarms and missed alarms, and improves the accuracy and timeliness of early warnings.
[0019] 3. The zero-shot learning technology of the real-time behavior recognition module of this invention supports the generation of new behavior detection models through natural language descriptions. This allows for rapid adaptation to diverse new behavior recognition needs at construction sites without requiring a large amount of labeled data, thus improving the model's generalization ability. The federated learning module employs horizontal federated learning technology, where edge nodes upload model gradients or parameter update information, which is then aggregated in the cloud to generate a global model. Homomorphic encryption and differential privacy are used to protect data security, solving the problems of retraining traditional models in new scenarios and ensuring data security.
[0020] 4. The early warning response of this invention adopts a three-level mechanism. The first-level early warning reminds workers through the vibration of the smart safety helmet and directional sound waves. The second-level early warning triggers an audible and visual alarm and pushes a notification to the management personnel. The third-level early warning automatically shuts down the equipment and activates the emergency plan, so as to achieve precise response to different risk levels. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the system flow of the present invention. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0023] To clearly illustrate the technical features of this solution, the following detailed implementation method will be used to explain the solution.
[0024] Example 1 See Figure 1 A real-time early warning system for construction site safety behavior based on multimodal perception is proposed. It consists of a multimodal perception layer, an edge computing layer, and a cloud service layer. The layers cooperate with each other to realize real-time early warning and management of construction site safety behavior.
[0025] Multimodal sensing layer: This layer is responsible for collecting multi-dimensional data from the construction site, providing a rich source of information for subsequent data processing and analysis.
[0026] Visual sensors: These include 4K high-definition cameras and panoramic cameras. The 4K high-definition cameras are used to capture high-definition video image data, clearly extracting details such as personnel posture, equipment status, and the wearing of safety protective equipment; the panoramic cameras are used to acquire panoramic video images of the construction site, providing a wider field of view and ensuring no blind spots in monitoring.
[0027] Millimeter-wave radar: Employing a 77GHz radar, the ranging accuracy is ±0.1 meters. It can accurately collect the three-dimensional position, speed, and direction of personnel and equipment, unaffected by environmental factors such as lighting or obstruction, and can track the dynamics of personnel and equipment in real time in complex environments.
[0028] Wearable devices include smart safety helmets and smart bracelets that integrate heart rate, body temperature, and acceleration sensors. Smart safety helmets can collect workers' physiological data such as heart rate and body temperature, as well as head acceleration data, in real time to monitor workers' physiological state and movement. Smart bracelets can collect movement data from other parts of the worker's body, further enriching the monitoring of workers' movement status.
[0029] Environmental sensors: used to collect environmental parameters such as temperature, humidity, dust, noise, light, and gas concentration, to monitor the environmental conditions of the construction site in real time and provide data support for assessing environmental risks.
[0030] BIM model: Provides three-dimensional spatial information and hazardous area markings for the construction site, and digitally models information such as building structure and equipment layout of the construction site, providing a foundation for subsequent data fusion and spatial analysis.
[0031] Edge computing layer: It connects to the multimodal perception layer and is responsible for real-time processing and fusion of multimodal data to achieve identification of personnel safety behaviors and dynamic early warning.
[0032] Data preprocessing module Time synchronization algorithm: Time synchronization processing is performed on the data of each mode, with a time synchronization error of less than 1ms, to ensure that the data collected by different sensors are consistent in time, which facilitates subsequent data fusion and analysis.
[0033] Spatial calibration algorithm: Unifies the data of each modality into the global coordinate system of the BIM model, realizes the spatial alignment of data from different sensors, solves the spatial calibration problem of multimodal data, and improves the accuracy of data fusion.
[0034] Noise filtering and feature extraction: Appropriate algorithms are used to filter noise in each modality of data, remove interference factors from the data, and extract useful feature information to prepare for subsequent behavior recognition and data fusion.
[0035] Multimodal fusion module: adopts a "feature-level fusion + decision-level fusion" architecture.
[0036] Feature-level fusion: Utilizing an improved Transformer model for cross-modal feature interaction, and through self-attention and multi-head attention mechanisms, it captures the spatiotemporal correlation and complementary information between different modal data, achieving deep fusion of cross-modal features.
[0037] Decision-level fusion: Decision fusion is performed using DS evidence theory to integrate the identification results of different modalities and improve the accuracy and reliability of decision-making.
[0038] Real-time behavior recognition module: Spatiotemporal Awareness Network (STPN): Composed of 3DCNN and LSTM. 3DCNN is used to extract spatiotemporal features from video image data and capture the motion information of people and equipment in three-dimensional space; LSTM is used to process sequential data and capture feature changes in the time dimension, thereby realizing the recognition of personnel safety behaviors.
[0039] Zero-shot learning module: Supports the generation of new behavior detection models through natural language description. When new types of safety behaviors appear on the construction site, no large amount of labeled data is required. The corresponding detection model can be generated simply by describing the characteristics of the behavior in natural language, which can quickly adapt to the diverse needs of the construction site.
[0040] Dynamic early warning engine module: Employs reinforcement learning algorithms to dynamically adjust early warning thresholds.
[0041] State space: includes environmental parameters, equipment status, personnel physiological status, and historical early warning records, comprehensively reflecting the dynamic risk factors at the construction site.
[0042] The reward function is a weighted sum of early warning accuracy and response time. A reinforcement learning algorithm continuously optimizes the early warning threshold to improve the accuracy and timeliness of early warnings. Based on behavioral recognition results and multi-dimensional risk factors, a risk level assessment is performed, and the early warning threshold is adjusted according to the assessment results to achieve adaptive early warning.
[0043] Cloud service layer: It connects to the edge computing layer and is responsible for global data management and model optimization, thereby improving the overall performance and intelligence level of the system.
[0044] Data storage and analysis module: Stores historical data, including multimodal perception data, behavior recognition results, early warning records, etc., and analyzes this data to generate safety situation reports, providing decision support for managers and helping them understand the safety status and trends of the construction site.
[0045] Federated Learning Module: Employing horizontal federated learning technology, edge nodes upload model gradients or parameter update information, which are then aggregated in the cloud using weighted averaging to generate a global model. Homomorphic encryption and differential privacy are used to protect data security, ensuring that edge node data does not leave their local machines during model training, thus protecting data privacy. Through federated learning, collaborative model training and global model updates are achieved among edge nodes, improving the model's generalization ability and accuracy.
[0046] Strategy Management Module: Used to configure and optimize early warning strategies. Managers can set different early warning thresholds, early warning methods, and emergency response measures according to the actual conditions of the construction site, and optimize and adjust the early warning strategies based on actual operation to ensure the effectiveness and adaptability of the early warning system.
[0047] Tiered early warning and response mechanism: A three-tiered early warning mechanism is adopted, with different early warning response measures taken according to different risk levels.
[0048] Level 1 warning: When low-risk behavior is detected, the smart safety helmet vibrates and emits directional sound waves to alert the worker, promptly correcting the worker's unsafe behavior and preventing the risk from escalating.
[0049] Level 2 warning: When a medium-risk behavior is detected, an audible and visual alarm is triggered and a notification is sent to the management personnel. The management personnel can promptly understand the situation on site and take corresponding management measures, such as on-site inspection and ordering rectification.
[0050] Level 3 Early Warning: When high-risk behavior or an impending safety incident is detected, the equipment will be automatically shut down and the emergency plan will be activated to minimize accident losses and protect the lives and property of personnel.
[0051] Application of BIM model and real-time sensing data fusion By fusing BIM models with real-time sensing data, a 3D risk heat map is generated. Within the 3D space of the BIM model, hazardous areas are marked based on the real-time sensing data, and the real-time locations of personnel and equipment are displayed. Managers can intuitively understand the safety situation at the construction site through the 3D risk heat map, promptly identify potential safety hazards, and achieve visualized and intelligent management of the construction site.
[0052] Edge computing platform and communication support: The edge computing layer utilizes the NVIDIA Jetson AGX Orin edge computing platform, which boasts powerful computing capabilities and supports real-time inference of complex deep learning models locally. It also supports dual-link communication (5G / wired) to ensure stable and reliable data transmission, guaranteeing system operation even in poor network conditions and enabling real-time alerts and data transmission.
[0053] Example 2 See Figure 1 The implementation method of a real-time early warning system for construction site safety behavior based on multimodal perception includes: First, define the following symbols: The video image data acquired by the vision sensor is denoted as The three-dimensional position and motion information acquired by millimeter-wave radar is denoted as... Physiological and motor data collected by wearable devices are denoted as The environmental parameters collected by the environmental sensors are denoted as follows: The three-dimensional spatial information and hazardous area markings provided by the BIM model are denoted as... The features extracted from each modality by the data preprocessing module are denoted as follows: The fused feature representation output by the multimodal fusion module is denoted as . The probability vector of behavior categories output by the real-time behavior recognition module is denoted as... The corresponding behavior category is denoted as The risk level output by the dynamic early warning engine module is recorded as follows: The adaptive early warning threshold is denoted as The global model parameters obtained by aggregation from the cloud service layer are denoted as... The local model parameters of the edge nodes are denoted as .
[0054] The specific steps are as follows: S1. Collect multi-dimensional data from the construction site using visual sensors, millimeter-wave radar, wearable devices, and environmental sensors, and obtain the three-dimensional spatial information and hazardous area markings of the BIM model to form raw multimodal perception data. and BIM spatial reference information ; The preferred visual sensors include 4K high-definition cameras and panoramic cameras; the preferred millimeter-wave radar is 77GHz radar; the preferred wearable devices include smart safety helmets and smart bracelets that integrate heart rate, body temperature, and acceleration sensors; and the environmental sensors collect environmental parameters such as temperature, humidity, dust, noise, light, and gas concentration.
[0055] S2. Using the original multimodal sensing data and BIM spatial reference information as input, the data preprocessing module performs time synchronization, spatial calibration, noise filtering, and feature extraction to unify the modal data to the global coordinate system of the BIM model, obtaining spatiotemporally aligned multimodal feature data. Specifically: (1) Time synchronization Let the unified reference time axis for each mode in the edge computing layer be... The original timestamps for each modality are as follows: A correction function is established through a time synchronization algorithm, so that... ,in , This represents the time deviation estimate for the corresponding mode. The time synchronization objective function can be set as follows:
[0056] in For modality Key synchronization signals (e.g., video frame trigger markers, radar point cloud trigger markers, wearable device inertial change markers, environmental sampling trigger markers); For reference synchronization signal; where, To unify the reference timeline A discrete time interval is used to align and calculate errors for synchronization signals of different modes.
[0057] By using a time synchronization algorithm, the timestamps of any two modes can be synchronized. satisfy ,in Let Y be the timestamp after time synchronization. Both X and Y originate from... The two modes selected.
[0058] (2) Spatial calibration to the global coordinate system of the BIM model Let point A in the local coordinate system of the millimeter-wave radar be... The coordinates of key points of the personnel estimated by the visual sensor are: The wearable device estimates the location of people as follows: Constructing rigid body transformations from individual coordinate systems to the global coordinate system of the BIM model. , in , For rotation and scaling matrices, It is a translation vector. These are points mapped to the global coordinate system of the BIM model. A calibration error function is established using known structural feature points in the BIM model or pre-calibrated positioning benchmarks as constraints. , in These are the coordinates of the corresponding BIM model reference points. This spatial calibration algorithm unifies data from different sensors into the global coordinate system of the BIM model, improving the accuracy of subsequent multimodal fusion.
[0059] (3) Feature extraction: After noise filtering of each modality, a unified and fusionable representation is extracted:
[0060] in For modal Feature extraction networks or algorithms that allow the use of Spatial priors impose regional constraints on features (e.g., weighted sampling of personnel / equipment trajectories and posture features only near hazardous areas).
[0061] S3. Using spatiotemporally aligned multimodal feature data as input, a combination of feature-level fusion and decision-level fusion is employed to achieve cross-modal feature interaction and fusion, outputting a fused feature representation. And the integration of decision-making results, specifically: Feature-level fusion utilizes an improved Transformer model for cross-modal feature interaction, while decision-level fusion employs DS evidence theory for decision fusion. (1) Cross-attention in feature-level fusion: Map each modality feature to a query, key, and value vector: , in This is a trainable parameter matrix.
[0062] Taking the cross-attention of visual features to radar, wearable devices, and environmental features as an example, we define... , in , For the attention dimension.
[0063] The final cross-modal aggregation representation can be written as ,in Add linear mapping or gated fusion operators for splicing.
[0064] (2) Dynamic weight allocation To address the issue of modal reliability variations under different field conditions, dynamic weights based on environmental and quality indicators are introduced: ,in , The modal quality scoring function can be defined in the following fractional form:
[0065] Where occ is the visual occlusion estimation; lux is the illumination anomaly; clu is the radar point cloud clustering stability index; loss is the wearable device data packet loss rate; and noise is the environmental sensor noise intensity estimation. to These are positive weighting coefficients. Based on this, the weighted fusion representation is obtained.
[0066] Will and Further integration leads to the final result , in It is a linear mapping or gating unit.
[0067] (3) DS Evidence Theory of Decision-Level Fusion Let the basic probability assignment of each modality to the set of behavior categories be: The pairwise combination rule is as follows:
[0068] in This is a subset of behavior categories.
[0069] Then The final fusion decision is obtained by successively combining the remaining modes to improve robustness to complex environments. S4. Using the fusion feature representation or fusion decision result as input, based on the spatiotemporal perception network and combined with zero-shot learning technology, identify the category of personnel safety behavior and output the behavior identification result. Specifically, the behavior identification result includes the behavior category and the behavior category probability vector. Represented by fusion features As input, the real-time behavior recognition module outputs a behavior category probability vector based on a spatiotemporal awareness network (STPN). and behavioral categories And it utilizes a zero-shot learning module to support the rapid expansion of new behavior types; (1) Spatiotemporal Awareness Network (STPN) The Spatiotemporal Awareness Network (STPN) consists of 3DCNN and LSTM. 3DCNN extracts spatiotemporal representations from video image sequences and aligned radar motion features, while LSTM models sequence variations. This can be abstracted as...
[0070] in, It is an intermediate spatiotemporal feature; Sequence semantic features; Let be the classification mapping parameter matrix. The supervised training loss can be set as... , in One-hot tags for the true category, The current sample output by the real-time behavior recognition module belongs to the first... The predicted probability of each behavior category.
[0071] (2) The zero-shot learning module aims to "generate new behavior detection models through natural language descriptions" by constructing behavior text description embeddings. Embedded with vision-radar fusion Alignment learning: , in To fuse feature projectors, For text encoders. Employing a similarity-driven alignment loss: , in, Cosine similarity; Embed text descriptions for correct behavior. Embedding for negative sample text descriptions; This is the temperature coefficient.
[0072] When a new type of safety behavior emerges at the construction site, only a natural language description of the behavior needs to be input to generate a corresponding detection semantic prototype, thus achieving expansion without the need for a large amount of labeled data.
[0073] S5. Using the behavior recognition results as input, the risk level is assessed by combining environmental parameters, equipment status, personnel physiological status, and historical warning records. A reinforcement learning algorithm is then used to dynamically adjust the warning threshold, and the risk level is output. With adaptive warning threshold Specifically: (1) Multidimensional risk scoring: Constructing a continuous risk score: , in, For behavior categories With behavior category probability vector Behavioral risk component; Environmental risk component (by Dangerous area markings (Common constraints) As a component of equipment risk; Risk components of personnel's physiological state (by calculate); This is a risk component for historical early warning trends; to These are the weighting coefficients. The behavioral risk component is calculated using: , For category Prior coefficient of risk; This represents the index or specific category when traversing the set of behavior categories.
[0074] Environmental risk components can be expressed in fractional form: , Where dust is the normalized value of dust concentration; gas is the normalized value of gas concentration; noise is the normalized value of noise intensity; and lux is the normalized value of illumination conditions. to It is a positive coefficient. According to... and Relationship risk level assessment:
[0075] in The grading threshold can be dynamically updated by reinforcement learning.
[0076] (2) Reinforcement learning threshold adaptation: The threshold adjustment is modeled as a Markov decision process, defined as follows: state A state vector consisting of environmental parameters, equipment status, personnel physiological status, and historical early warning records; :right Increase, decrease, or maintenance; rewards This is the weighted sum of early warning accuracy and response time. Its reward function can be specified as:
[0077] Where Acc is the alert accuracy within a unit window; Delay is the average delay from the occurrence of the action to the triggering of the alert. Positive weights. Through policy iteration or value iteration, the following is achieved:
[0078] in, This serves as a discount factor to achieve a balance optimization between false alarms and missed alarms that dynamically change with the field situation.
[0079] S6, based on risk level With adaptive warning threshold As input, a tiered early warning response is triggered, and the BIM model is fused with real-time perception data to generate a 3D risk heat map. Simultaneously, behavior recognition results, early warning records, and related multimodal data are uploaded to the cloud service layer for data storage and analysis. Federated learning is used to achieve collaborative training of models at each edge node and global model updates, so that the updated model can be distributed to the edge computing layer for subsequent real-time behavior recognition. Specifically: (1) The graded early warning response includes a three-level early warning mechanism: when the risk level is a level one early warning, the worker is alerted by the vibration of the smart safety helmet and directional sound waves; when the risk level is a level two early warning, the sound and light alarm is triggered and a notification is sent to the management personnel; when the risk level is a level three early warning, the equipment is automatically shut down and the emergency plan is activated.
[0080] (2) Three-dimensional risk heat map: the risk level is divided into three-dimensional risk heat map. The location of personnel and equipment in the global coordinate system of the BIM model, and the overlay display of hazardous area labels, enable managers to intuitively obtain hazardous areas and real-time dynamics in the three-dimensional space of the BIM model.
[0081] (3) Global Model Update of Federated Learning The federated learning module of the cloud service layer adopts horizontal federated learning technology. Edge nodes upload model gradient or parameter update information, and the cloud generates a global model by weighted average aggregation. Homomorphic encryption and differential privacy can be used to protect data security.
[0082] Its aggregation form can be represented as: , in, The number of edge nodes participating in the training; For edge nodes The local sample size; For edge nodes Local model parameters; These are the aggregated global model parameters.
[0083] Issued By moving to the edge computing layer, the generalization ability of the real-time behavior recognition module and the dynamic early warning engine module in new scenarios can be improved.
[0084] (4) The edge computing layer preferably adopts the NVIDIA Jetson AGX Orin edge computing platform and supports 5G / wired dual-link communication to achieve local real-time inference and stable transmission.
[0085] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A real-time early warning system for construction site safety behaviors based on multimodal perception, characterized in that, It includes a multimodal perception layer, an edge computing layer, and a cloud service layer; The multimodal perception layer includes visual sensors, millimeter-wave radar, wearable devices, and environmental sensors. The multimodal perception layer is used to collect multi-dimensional data from the construction site and combine it with BIM to obtain three-dimensional spatial information and hazardous area markings of the construction site. The edge computing layer is connected to the multimodal perception layer and is used to process and fuse multimodal data in real time. It includes at least a data preprocessing module, a multimodal fusion module, a real-time behavior recognition module, and a dynamic early warning engine module. The cloud service layer is connected to the edge computing layer and is used for global data management and model optimization. It includes at least a data storage and analysis module, a federated learning module, and a policy management module.
2. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The multi-dimensional data includes at least video image data collected by a visual sensor, three-dimensional position, speed and direction of movement of personnel and equipment collected by millimeter-wave radar, physiological and motion data of workers collected by wearable devices, and environmental parameters collected by environmental sensors. The environmental parameters include humidity, dust, noise, light, and gas concentration, while the video image data includes personnel posture, equipment status, and the wearing of safety protective equipment. The visual sensors include a 4K high-definition camera and a panoramic camera; the millimeter-wave radar is a 77GHz radar with a ranging accuracy of ±0.1 meters; and the wearable devices include a smart safety helmet and a smart bracelet that integrate heart rate, body temperature, and acceleration sensors.
3. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The data preprocessing module is used to perform time synchronization, spatial calibration, noise filtering, and feature extraction on the multi-dimensional data output by the multi-modal perception layer and the BIM model to obtain spatiotemporally aligned multimodal feature data. The data preprocessing module includes a time synchronization algorithm and a spatial calibration algorithm. The time synchronization error is less than 1ms, and the spatial calibration unifies the data of each modality to the global coordinate system of the BIM model.
4. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The multimodal fusion module is used to perform cross-modal feature interaction and fusion on the spatiotemporally aligned multimodal feature data, and generates fused feature representations using a cross-attention mechanism and a dynamic weight allocation algorithm. And it combines DS evidence theory to generate fusion decision results; The multimodal fusion module adopts a "feature-level fusion + decision-level fusion" architecture. Feature-level fusion uses an improved Transformer model to perform cross-modal feature interaction, while decision-level fusion uses DS evidence theory for decision fusion.
5. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The real-time behavior recognition module is used to output the behavior recognition result of personnel safety behavior based on the fusion feature representation and fusion decision result, combined with the spatiotemporal awareness network and zero-shot learning technology. The spatiotemporal awareness network consists of 3DCNN and LSTM, and is used to extract spatiotemporal features from video image data and identify behavior categories. The zero-shot learning technique is used to support the generation of new behavior detection models through natural language descriptions, so as to expand behavior recognition capabilities without a large amount of labeled data when new types of security behaviors emerge.
6. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The dynamic early warning engine module is used to assess the risk level and adjust the adaptive early warning threshold based on the behavior recognition results and multi-dimensional risk factors, and output the risk level and the adaptive early warning threshold. The dynamic early warning engine module uses a reinforcement learning algorithm to dynamically adjust the early warning threshold. The state space includes environmental parameters, equipment status, personnel physiological status, and historical early warning records. The reward function is a weighted sum of early warning accuracy and response time.
7. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The data storage and analysis module is used to store the behavior recognition results, early warning records, and historical multimodal data, and generate a security situation report. The federated learning module is used to coordinate each edge node to perform model training and global model updates, and to send the updated model parameters to the edge computing layer. The federated learning module adopts horizontal federated learning technology, where edge nodes upload model gradients or parameter update information, and the cloud generates a global model through weighted average aggregation, and uses homomorphic encryption and differential privacy to protect data security. The strategy management module is used to configure and optimize early warning strategies based on the risk level and the security situation report.
8. The real-time early warning system for construction site safety behavior based on multimodal perception as described in claim 1, wherein the BIM model is fused with real-time perception data to generate a three-dimensional risk heat map, marking dangerous areas and displaying the real-time location of personnel and equipment; The edge computing layer adopts the NVIDIA Jetson AGX Orin edge computing platform, which supports 5G / wired dual-link communication, enables local real-time inference, and has a response time of less than 2 seconds.
9. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 1, characterized in that, The implementation method of the system includes the following steps: S1. Collect multi-dimensional data through visual sensors, millimeter-wave radar, wearable devices and environmental sensors, and obtain the three-dimensional spatial information and hazardous area markings of the BIM model to form original multimodal perception data and BIM spatial benchmark information. S2. Using the original multimodal sensing data and BIM spatial reference information as input, perform time synchronization, spatial calibration, noise filtering and feature extraction to unify the modal data into the global coordinate system of the BIM model and obtain spatiotemporally aligned multimodal feature data. S3. Using spatiotemporally aligned multimodal feature data as input, a combination of feature-level fusion and decision-level fusion is adopted to achieve cross-modal feature interaction and fusion, and output fused feature representation and fusion decision results; S4. Using fusion feature representations or fusion decision results as input, based on spatiotemporal perception networks and combined with zero-shot learning techniques, identify personnel safety behavior categories and output behavior recognition results. S5. Using the behavior recognition results as input, the risk level is assessed by combining environmental parameters, equipment status, personnel physiological status and historical warning records. The warning threshold is dynamically adjusted using reinforcement learning algorithms, and the risk level and adaptive warning threshold are output. S6. Using risk level and adaptive warning threshold as input, trigger a graded warning response, and fuse the BIM model with real-time perception data to generate a three-dimensional risk heat map. At the same time, upload the behavior recognition results, warning records and related multimodal data to the cloud service layer for data storage and analysis. Through federated learning, realize the collaborative training of models of each edge node and the global model update, so as to send the updated model to the edge computing layer for subsequent real-time behavior recognition.
10. The real-time early warning system for construction site safety behavior based on multimodal perception according to claim 9, characterized in that, The graded early warning response includes a three-level early warning mechanism: Level 1 early warning alerts workers through the vibration and directional sound waves of smart safety helmets; Level 2 early warning triggers an audible and visual alarm and pushes a notification to management personnel; Level 3 early warning automatically shuts down equipment and activates the emergency plan.