Construction site safety risk identification method, device, equipment and medium
By collecting and fusing data through a multimodal sensor network, the system can assess the risks of high-altitude operations in real time, proactively predict fall trajectories, and generate intervention commands. This solves the problems of accuracy and timeliness in risk assessment for high-altitude operations in existing technologies, and achieves a highly efficient transformation in safety protection.
Patent Information
- Application Number
- CN202511859233.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies cannot accurately assess dynamic risks in real time during high-altitude operations, making it difficult to achieve accurate early warning and pre-emptive intervention for fall accidents. Existing systems often rely on a single data source, resulting in high false alarm rates and poor robustness, and are unable to effectively prevent accidents.
By deploying fixed and mobile sensor nodes to synchronously collect multimodal sensing data, spatial registration is performed to generate dynamic motion state sequences and environmental point clouds. Combined with inertial data, positioning data, and wind speed and direction data, the data is input into a risk assessment model to calculate a comprehensive risk index. In critical situations, the system can proactively predict the fall trajectory and generate intervention commands.
It has enabled a shift from post-event alarms to pre-event prediction and early warning, significantly improving the accuracy and timeliness of risk identification, providing key decision-making basis for proactive protection, and enhancing the safety protection capabilities for high-altitude operations.
Smart Images

Figure CN121686352A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent safety monitoring and risk prevention technology, and in particular relates to a method, device, equipment and medium for identifying safety risks at construction sites. Background Technology
[0002] In high-altitude operations, especially construction work, falls remain the leading risk causing serious injuries and fatalities. Currently, the safety measures commonly relied upon in the industry mainly fall into two categories: passive physical protection and post-accident alarm systems. Passive protection, such as safety nets, guardrails, and safety belts, heavily relies on the completeness of their pre-installation and the standardized use by workers. They cannot adapt to dynamically changing work surfaces and essentially mitigate consequences after a hazard has occurred, failing to prevent falls altogether. In recent years, with the development of IoT technology, some intelligent safety systems based on wearable devices or positioning technology have emerged. For example, safety helmets or vests with integrated sensors can monitor a person's position or posture, issuing alarms when a fall or prolonged stillness is detected. However, most of these technological solutions still fall into the category of "post-accident response," their core logic being that alarms are triggered only after instability or a fall has occurred or even completed, thus missing the critical time window for preventing accidents. Furthermore, these methods often rely on a single type of data source, such as using only UWB for location determination or only IMU for attitude analysis. They lack comprehensive perception and collaborative analysis of the overall movement status of workers, the surrounding environment, and risk factors (such as strong winds, slippery ground, and heavy-load operations). This results in a high false alarm rate and poor robustness in the actual complex construction site environment. It is difficult to accurately distinguish between high-risk instability precursors and normal work actions, and it is even more impossible to achieve true risk prediction and pre-intervention.
[0003] Therefore, the existing technology system suffers from the dual defects of "lagging passive protection" and "one-sided intelligent early warning". There is an urgent need for a proactive safety identification method that can deeply integrate multi-dimensional information, accurately assess dynamic risks in real time, and provide effective early warning and even intervention decisions before danger occurs, so as to fundamentally improve the inherent safety level of high-altitude operations. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, device, equipment, and medium for identifying safety risks at construction sites to address the aforementioned technical problems.
[0005] Firstly, this application provides a method for identifying safety risks at construction sites, including:
[0006] S1. Multimodal perception data is collected synchronously by fixed sensing nodes deployed in the construction site work area and mobile sensing nodes worn by workers; among which, multimodal perception data includes environmental point cloud data collected by fixed stereo vision cameras and lidar, inertial data collected by inertial measurement units in mobile sensing nodes, positioning data collected by ultra-wideband positioning tags, and wind speed and direction data collected by environmental sensors.
[0007] S2. Spatial registration of multimodal sensing data to generate a temporal dynamic motion state sequence for each worker and a fused environmental point cloud of the work area;
[0008] S3. Based on the time-series dynamic motion state sequence, calculate the kinematic characteristics and posture stability characteristics of the workers; based on the fusion of environmental point cloud and wind speed and direction data, calculate the environmental interaction characteristics.
[0009] S4. Input the kinematic features, posture stability features, and environmental interaction features into the pre-trained risk assessment model and output a comprehensive risk index.
[0010] S5. Compare the comprehensive risk index with the preset risk threshold to obtain the comparison result; generate a risk identification decision signal based on the comparison result;
[0011] S6. When the risk identification decision signal indicates a critical risk, predict the person's fall trajectory based on the time-series dynamic motion state sequence, and generate an active intervention instruction including the target person's identification, suggested intervention type, and intervention parameters.
[0012] Secondly, this application also provides a construction site safety risk identification device for implementing the method described in the first aspect, the device comprising:
[0013] The multimodal perception data acquisition module is used to simultaneously acquire multimodal perception data through fixed sensing nodes deployed in the construction site work area and mobile sensing nodes worn by workers. The multimodal perception data includes environmental point cloud data acquired by fixed stereo vision cameras and lidar, inertial data acquired by inertial measurement units in mobile sensing nodes, positioning data acquired by ultra-wideband positioning tags, and wind speed and direction data acquired by environmental sensors.
[0014] The data registration module is used to spatially register multimodal sensing data, generate a temporal dynamic motion state sequence for each worker and a fused environmental point cloud of the work area;
[0015] The dynamic feature analysis module is used to calculate the kinematic characteristics and posture stability characteristics of workers based on time-series dynamic motion state sequences; and to calculate environmental interaction characteristics based on fused environmental point cloud and wind speed and direction data.
[0016] The risk assessment processing module is used to input kinematic features, posture stability features, and environmental interaction features into a pre-trained risk assessment model and output a comprehensive risk index.
[0017] The risk decision generation module is used to compare the comprehensive risk index with a preset risk threshold to obtain the comparison result; and to generate a risk identification decision signal based on the comparison result.
[0018] The proactive intervention decision module is used to predict the fall trajectory of a person based on a time-series dynamic motion state sequence when the risk identification decision signal indicates a critical risk, and to generate proactive intervention instructions including the target person's identification, suggested intervention type, and intervention parameters.
[0019] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a construction site safety risk identification method as described in the first aspect.
[0020] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a construction site safety risk identification method as described in the first aspect.
[0021] The aforementioned method, device, equipment, and medium for identifying safety risks at construction sites utilizes a multimodal sensor network to synchronously collect personnel movement, location, and environmental data. Through spatial registration and fusion, a dynamic digital twin accurately reflects the personnel's full-body posture and the surrounding environment is constructed. This twin then extracts multidimensional risk characteristics in real time, such as kinematics, posture stability, and environmental interaction, representing precursors to instability. These characteristics are input into an AI risk assessment model for predictive calculations. Ultimately, when a critical risk is identified, the model proactively predicts the fall trajectory and generates precise intervention commands. This represents a fundamental shift in the assessment of fall risks from post-event alarms to pre-event prediction and early warning, significantly improving the accuracy and timeliness of risk identification. It provides crucial decision-making support for activating active protection systems and effectively enhances the safety protection capabilities for high-altitude operations. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1A flowchart illustrating a method for identifying safety risks at construction sites provided by this invention;
[0024] Figure 2 This is a schematic diagram of the process of spatial registration of the multimodal sensing data in an optional embodiment of the present invention;
[0025] Figure 3 This is a structural schematic diagram of a construction site safety risk identification device provided by the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0027] refer to Figure 1 The application presents a flowchart illustrating a method for identifying safety risks at construction sites, which includes the following steps:
[0028] S1. Multimodal perception data is collected synchronously by fixed sensing nodes deployed in the construction site work area and mobile sensing nodes worn by workers. The multimodal perception data includes environmental point cloud data collected by fixed stereo vision cameras and lidar, inertial data collected by inertial measurement units in mobile sensing nodes, positioning data collected by ultra-wideband positioning tags, and wind speed and direction data collected by environmental sensors.
[0029] Specifically, fixed sensor nodes are distributed according to the spatial structure of the work area. Stereo vision cameras are used to ensure effective coverage and overlap of the field of view; their spatial positions are determined through calibration to achieve comprehensive capture of environmental visual information. LiDAR is used to collect 3D spatial point cloud data of the work area; its deployment location must avoid obstructions to ensure accurate perception of environmental features such as the work surface and obstacles. Environmental sensors are integrated into the fixed sensor nodes and are specifically used to capture environmental parameters affecting operational safety, providing environmental data support for risk assessment.
[0030] The mobile sensing node adopts a modular design and is integrated into the protective equipment worn by the workers, ensuring that it does not affect their mobility. The inertial measurement unit (IMU), as the core component, collects the inertial motion information of the workers. Its installation position must ensure consistency with the human body's motion state to avoid data distortion due to installation deviations. The ultra-wideband positioning tag works in conjunction with the IMU, constructing a positioning network through positioning base stations deployed within the work area to achieve accurate real-time tracking of the workers' locations. This tag contains unique identification information that can be bound to the worker's identity, facilitating the differentiation of motion data from different individuals. Data transmission between the mobile sensing node and the fixed sensing node is achieved via a wireless communication protocol, and an independent power supply module ensures continuous operation.
[0031] Data synchronization is achieved through a time synchronization mechanism, with all sensor nodes connected to a unified time calibration system to ensure consistency of various data across time dimensions. The collected multimodal sensing data is stored in a preset format, including core fields such as data type identifier, timestamp, sensor identifier, and raw data. Environmental point cloud data, inertial data, positioning data, and wind speed and direction data each employ storage formats adapted to their specific characteristics, and data encryption is performed to ensure security and privacy during data transmission and storage.
[0032] S2. Spatial registration of multimodal sensing data is performed to generate a temporal dynamic motion state sequence for each worker and a fused environmental point cloud of the work area.
[0033] Specifically, spatial registration of multimodal sensing data is a crucial step in eliminating differences in coordinate systems between different sensors and achieving data fusion. Its core objective is to unify various types of sensing data under the same world coordinate system, providing a consistent spatial benchmark for subsequent feature calculations and risk assessments. First, the coordinate system definition is clarified. The world coordinate system adopts the universal geodetic coordinate system, and the world coordinates of fixed sensor nodes are determined through positioning technology. Image data from stereo vision cameras needs to be converted to the camera coordinate system through intrinsic parameter calibration, and then converted to the world coordinate system through extrinsic parameter calibration. The radar coordinate system of lidar is directly converted to the world coordinate system through the spatial parameters of its deployment location. The positioning data of ultra-wideband positioning tags is based on the world coordinate system itself and requires no additional conversion. The body coordinate system of the inertial measurement unit is calibrated in real time using the position and attitude data of the ultra-wideband positioning tags to ensure the consistency of its motion data with the world coordinate system.
[0034] The spatial registration process is divided into two stages: sensor extrinsic parameter calibration and real-time data registration. Sensor extrinsic parameter calibration employs a hand-eye calibration method, using point cloud data acquired by the lidar as a reference. Data is collected at multiple feature points within the work area using a moving calibration tool. A calibration algorithm is then used to calculate the spatial transformation relationships between the sensors, including rotation matrices and translation vectors. The accuracy of the calibration results is verified through reprojection error analysis. Real-time data registration is based on a timestamp synchronization strategy. For each time step, the collected data from all sensors is extracted. A pre-set extrinsic parameter matrix is used to uniformly transform the environmental point cloud data, inertial data, and positioning data to the world coordinate system, forming a fused data frame.
[0035] The generation of time-series dynamic motion state sequences is based on individuals, indexed by the unique identifiers of ultra-wideband positioning tags. Inertial and positioning data of the same worker at consecutive time steps are concatenated chronologically to form a continuous time-series data sequence. During sequence generation, the raw data needs to be preprocessed to eliminate noise interference. The inertial data undergoes noise reduction using a Kalman filter algorithm. The core of this algorithm is to construct a filtering model based on the state equation and the observation equation. The state equation is... The observation equation is In the formula This represents the state vector composed of acceleration and angular velocity. The sensor's observations, Here is the state transition matrix. To control the input matrix, For the observation matrix, and The noise consists of process noise and observation noise, both of which follow a Gaussian distribution. This filtering algorithm can effectively suppress sensor noise and improve data accuracy. The positioning data is smoothed using a moving average filter to reduce the impact of positioning jitter on data quality.
[0036] The point cloud generation for the operational area employs an iterative nearest-point algorithm to register and fuse visual and LiDAR point clouds. First, both types of point cloud data undergo denoising to remove background noise points and outliers, improving registration accuracy. Then, using the LiDAR point cloud as the target point cloud and the visual point cloud as the source point cloud, the iteration parameters and convergence conditions of the iterative nearest-point algorithm are initialized. Point cloud registration is achieved by minimizing the Euclidean distance error between the source and target point clouds, with the objective function being: In the formula Represents any point in the source point cloud. Indicates the target point cloud and The nearest point is identified, and the error function is iteratively optimized to converge to a preset threshold, thus completing point cloud fusion. The fused environmental point cloud contains both geometric and visual information, accurately representing the environmental characteristics of the work area. It is stored step-by-step to form a temporal sequence of environmental point clouds, providing a data foundation for calculating environmental interaction features.
[0037] S3. Based on the time-series dynamic motion state sequence, calculate the kinematic characteristics and posture stability characteristics of the workers; based on the fusion of environmental point cloud and wind speed and direction data, calculate the environmental interaction characteristics.
[0038] Specifically, the calculation of kinematic features is based on a time-series dynamic motion state sequence, extracting quantitative indicators from core dimensions such as velocity, acceleration, and trajectory to reflect the worker's motion state. Velocity features are calculated using the central difference method. For positioning data at any time step in the time series, the velocity in each direction is calculated using the positional change and time interval between adjacent time steps. The resultant velocity is obtained by vector synthesis of velocities in each direction, expressed as follows: In the formula , , These represent the velocity components in the three spatial directions, respectively. This represents the resultant velocity. Acceleration is calculated using the change in velocity between adjacent time steps and the time interval, expressed as: In the formula This represents the resultant velocity at the previous time step. Indicates the time step interval. This represents acceleration. Jerk, as the rate of change of acceleration, is expressed as: In the formula This represents the acceleration at the previous time step. It indicates jerk.
[0039] The motion trajectory features include trajectory curvature and trajectory offset. Trajectory curvature is obtained by fitting positioning data from multiple consecutive time steps into a circular arc curve and calculating the reciprocal of the arc radius; the expression is: In the formula This represents the radius of the fitted circular arc. The curvature of the trajectory is represented by a value; a larger curvature indicates a greater degree of trajectory bending. The trajectory offset is calculated by determining the spatial distance between the current trajectory and the preset safe path, and is used to characterize the degree to which the worker deviates from the safe area. All kinematic features are calculated statistically within a sliding window to form a kinematic feature vector with fixed dimensions, providing a quantitative basis for subsequent risk assessment at the motion state level.
[0040] The calculation of attitude stability characteristics is based on inertial data collected by the inertial measurement unit and attitude calculation results. The core objective is to quantify the stability of the worker's body posture. Attitude calculation employs the quaternion method to avoid the gimbal lock problem inherent in Euler angles. The quaternion initialization is determined by the initial attitude of the sensors, and its update equation is:
[0041]
[0042] In the formula, Represents the quaternion of the current time step. This represents the quaternion of the previous time step. This represents quaternion multiplication. This represents the three-axis angular velocity vector acquired by the inertial measurement unit. , , These represent the angular velocity components of the three axes, respectively. The magnitude of the angular velocity vector. This represents the time step interval. The quaternion for each time step can be calculated using this update equation, and then converted into Euler angles (roll, pitch, and yaw) to visually reflect the operator's body posture.
[0043] Based on the calculated attitude angles, attitude stability characteristics are calculated, including the rate of change of attitude angles, attitude angle deviation, and attitude stability index. The rate of change of attitude angles is calculated by the difference between attitude angles at adjacent time steps and the time interval, reflecting the severity of attitude changes. The attitude angle deviation is calculated by the difference between the current attitude angle and the normal operating attitude angle, characterizing the degree of deviation between the current attitude and the safe attitude. The attitude stability index is calculated by integrating the rate of change of attitude angles, attitude angle deviation, and acceleration fluctuations in inertial data, using a weighted summation method, and is expressed as follows: In the formula This represents the attitude stability index. Statistical value representing the rate of change of attitude angle. The statistical value representing the attitude angle deviation. The standard deviation of acceleration, , , The weight coefficients for each feature are determined through calibration using experimental data.
[0044] The calculation of environmental interaction features is based on the fusion of environmental point cloud and wind speed and direction data. The core objective is to quantify the interaction between workers and their surrounding environment, as well as the impact of environmental factors on work safety. First, key environmental features of the work area are extracted from the fused environmental point cloud, including the slope of the work surface, edge contours, and obstacle distribution. A point cloud segmentation algorithm divides the environmental point cloud into different regions such as the work surface, obstacles, and background. Spatial interaction parameters, such as the distance between the worker and the edge of the work surface, and the distance to obstacles, are calculated. Based on wind speed and direction data, environmental impact parameters, such as the angle between the vector component of wind speed and the worker's movement direction, and the equivalent value of the force exerted by wind speed on the human body, are calculated. The environmental interaction features are obtained by fusing the above spatial interaction parameters and environmental impact parameters, forming a feature vector that comprehensively reflects the interaction state between the environment and the worker.
[0045] S4. Input the kinematic features, posture stability features, and environmental interaction features into the pre-trained risk assessment model and output a comprehensive risk index.
[0046] Specifically, the core function of the risk assessment model is to output a comprehensive risk index that characterizes the degree of safety risk to workers based on the input kinematic features, posture stability features, and environmental interaction features. The training process of the model requires the construction of a sample dataset containing different work scenarios and different risk states. The sample data is collected through simulated work scenarios and actual work scenarios, covering feature data under various states such as normal operation, low risk, medium risk, and high risk.
[0047] Before model training, sample data needs to be preprocessed, including data standardization and outlier removal, to improve training performance. Data standardization uses the Z-score standardization method to convert each feature data into a standard normal distribution with a mean of 0 and a variance of 1. The expression is: In the formula Represents the standardized eigenvalues. Represents the original feature values. This represents the mean of the feature. The standard deviation represents the characteristic.
[0048] The risk assessment model is constructed using a deep learning network, which consists of an input layer, hidden layers, and an output layer. The input layer has the same dimension as the feature vector, taking the vector concatenated from kinematic features, pose stability features, and environmental interaction features as input. The hidden layer consists of multiple fully connected layers and activation functions, which introduce non-linear mapping capabilities to improve the model's ability to fit complex features. The output layer uses a single neuron to output a comprehensive risk index ranging from 0 to 1, used to quantify the degree of risk.
[0049] During model training, cross-validation is used to optimize model parameters to avoid overfitting. The mean squared error loss function is used, with the expression: In the formula Indicates the loss value. Indicates the number of samples. This represents the true risk level label of the i-th sample (after normalization). Let represent the predicted risk index of the model for the i-th sample. The loss function is minimized using the gradient descent algorithm, and the model parameters are iteratively updated until the model converges.
[0050] During the model inference phase, the preprocessed feature vectors are input into the trained risk assessment model, and the comprehensive risk index is calculated through forward propagation, expressed as follows: In the formula This represents the overall risk index. This represents the activation function of the model. The weight matrix of the model is represented. This represents the input feature vector. This represents the bias vector of the model. This comprehensive risk index can fully reflect the current safety status of workers and provide a core basis for subsequent risk decisions.
[0051] S5. Compare the comprehensive risk index with the preset risk threshold to obtain the comparison result; generate a risk identification decision signal based on the comparison result.
[0052] Specifically, the generation of risk identification decision signals is based on a comparison between a comprehensive risk index and a preset risk threshold. The core principle is to accurately determine the risk level of operators and generate corresponding decision signals through reasonable threshold settings and comparison logic. The determination of preset risk thresholds requires consideration of the safety requirements of the actual work scenario, historical accident data, and expert experience. Statistical analysis methods are used for calibration to ensure that the thresholds can effectively distinguish different risk levels, avoiding both underreporting of high-risk events and minimizing the interference of false alarms on normal operations.
[0053] Risk thresholds include low-risk, medium-risk, and critical-risk thresholds, forming a three-tiered risk classification standard. The comprehensive risk index is compared with each of these thresholds: when the comprehensive risk index is below the low-risk threshold, it indicates that the workers are in a safe state, generating a safety decision signal; when the comprehensive risk index is between the low-risk and medium-risk thresholds, it indicates a potential safety risk, generating an early warning decision signal; when the comprehensive risk index is between the medium-risk and critical-risk thresholds, it indicates a high risk level, generating a warning decision signal; and when the comprehensive risk index is above the critical-risk threshold, it indicates an immediate safety hazard, generating a critical-risk decision signal.
[0054] The comparison process employs a continuous judgment logic, performing real-time evaluation of the comprehensive risk index at each time step to ensure the real-time nature of the decision signals. The decision signals are in the form of standardized digital signals, containing core information such as risk level identifiers, timestamps, and personnel identifiers, facilitating subsequent system analysis and response. Simultaneously, the decision signals are stored in a database to provide data support for subsequent risk tracing and model optimization.
[0055] S6. When the risk identification decision signal indicates a critical risk, predict the person's fall trajectory based on the time-series dynamic motion state sequence, and generate an active intervention instruction including the target person's identification, suggested intervention type, and intervention parameters.
[0056] Specifically, when the risk identification decision signal indicates a critical risk, an active intervention mechanism must be activated immediately. The core of this mechanism is to predict the personnel's fall trajectory and generate precise active intervention instructions to buy time to prevent or mitigate the consequences of the accident. Fall trajectory prediction is based on a time-series dynamic motion state sequence and is implemented using kinematic prediction algorithms. By analyzing the movement trends, posture changes, and environmental influencing factors of the workers, a fall motion model is constructed.
[0057] The fall motion model is based on classical kinematic equations, modified to incorporate the effects of environmental factors such as air resistance and wind speed. The expression is: In the formula This represents the position vector at the prediction time step. This represents the position vector at the current time step. Represents the velocity vector at the current time step. Represents the gravitational acceleration vector. This represents the environmental acceleration vector caused by wind speed and direction. This represents the acceleration vector corresponding to air resistance. This represents the prediction time step interval. The wind speed vector is converted into an equivalent acceleration acting on the human body by calculating wind speed and direction data and human body force model based on aerodynamic principles. Piecewise function modeling is used, which relates the magnitude of the current velocity vector to parameters such as air density and the frontal area of the human body. The expression is as follows: In the formula This indicates the total mass of the workers and their equipment. Indicates air density (collected with the assistance of environmental sensors or a preset standard value). This represents the windward area of the human body (based on the human body contour estimation from the posture calculation results). This indicates the air resistance coefficient (dynamically adjusted based on the operator's posture). The negative sign indicates that the air resistance is opposite to the direction of motion. Represents velocity vector The length of the module.
[0058] The trajectory prediction process employs an iterative calculation method. It uses the last time step of the time-series dynamic motion state sequence as the initial state, inputs it into the fall motion model, and calculates the position and velocity of the first predicted time step. This prediction result is then used as the initial state for the next iteration. Combined with real-time updates of environmental parameters (such as dynamic changes in wind speed and direction), the position vectors of subsequent prediction time steps are calculated sequentially to form a complete fall trajectory prediction sequence. The number of prediction time steps is determined based on the working height and safety response requirements, ensuring that the predicted trajectory covers the entire process from instability to potential contact with the ground or protective facilities, while allowing sufficient intervention response time. During the prediction process, collision detection is performed using fused environmental point cloud data. If the predicted trajectory collides with obstacles, edges of the working area, or other environmental features, subsequent prediction results are corrected promptly to ensure the accuracy and practicality of the trajectory prediction.
[0059] The generation of proactive intervention commands is based on the predicted fall trajectory, environmental characteristics of the work area, and the configuration of existing intervention equipment. It includes three core components: target personnel identification, suggested intervention type, and intervention parameters. The target personnel identification is directly linked to the unique identification information of the mobile sensor node. This identification is pre-bound to metadata such as the worker's identity information, work location, and equipment number. Through data association queries, complete personnel identity and work-related information can be obtained, ensuring that intervention commands can accurately locate specific at-risk personnel.
[0060] The determination of intervention types should be based on a combination of the characteristics of the predicted fall trajectory (such as the fall start location, fall speed, estimated fall time, and landing area) and the deployment of intervention equipment in the work area, using a rule-based reasoning and priority ranking mechanism. Common intervention types include emergency braking intervention, protective facility activation intervention, and warning and guidance intervention: Emergency braking intervention is suitable for scenarios where workers are on high-altitude platforms, scaffolding, or other structures with braking devices. It prevents workers from falling by triggering the braking mechanism of the wearable equipment or the work platform. Protective facility activation intervention is for scenarios where deployable protective nets, air cushions, or other equipment are deployed in the landing area. It activates the corresponding protective facilities based on the predicted landing coordinates to ensure effective protection along the fall path. Warning and guidance intervention is suitable for scenarios where there are evasive paths or temporary protective spaces. It issues directional guidance instructions to at-risk personnel through the sound and light warning modules of wearable equipment or the on-site broadcast system, while simultaneously issuing avoidance warnings to surrounding workers. The priority of intervention types is determined based on the urgency of the fall risk and the assessment of intervention effectiveness. Emergency braking intervention and protective facility activation intervention have higher priority than warning and guidance intervention, ensuring that intervention measures that can directly prevent or mitigate fall injuries are prioritized.
[0061] The calculation of intervention parameters must precisely match the recommended intervention type, and be quantitatively determined based on the predicted fall trajectory and environmental characteristics to ensure the effectiveness and safety of the intervention measures. For emergency braking intervention, the intervention parameters include braking trigger time, braking force threshold, and braking duration. The braking trigger time is determined by the initial stage time node of the predicted fall trajectory to ensure that braking can be initiated at the initial stage of personnel instability. The braking force threshold is calculated based on the worker's motion state and the load limit of the braking device to avoid secondary injuries caused by excessive braking force. For protective facility activation intervention, the intervention parameters include protective facility activation number, deployment time, and deployment range. The protective facility activation number is determined by spatial matching between the predicted landing point coordinates and the deployment location of the protective facility. The deployment time is calculated based on the difference between the predicted fall time and the facility deployment response time to ensure that the facility is deployed before the personnel reach the landing point. The deployment range is determined based on the possible deviation range of the predicted trajectory and the effective protection area of the facility, with a certain safety redundancy reserved. For warning and guidance intervention, the intervention parameters include warning frequency, guidance direction vector, and guidance distance. The warning frequency is set according to the urgency of the risk. The guidance direction vector is calculated from the spatial vector between the current location and the safe area. The guidance distance is the shortest path length from the current location to the safe area, ensuring that the guidance instructions are clear and accurate.
[0062] The generated proactive intervention commands adopt a standardized data format, including a command header, core data segment, and verification segment. The command header contains a command type identifier, a sending timestamp, and a sending node identifier. The core data segment contains the target personnel identifier, the suggested intervention type code, and the intervention parameter set. The verification segment uses the CRC32 checksum algorithm to ensure the integrity of the command transmission. The intervention commands are transmitted in real-time via a wireless communication network to the corresponding intervention execution equipment (such as braking devices, protective facility controllers, and wearable equipment warning modules), and simultaneously synchronized to the site safety monitoring center, facilitating real-time monitoring of the intervention execution status by management personnel. If the intervention command transmission fails or the execution equipment reports an anomaly, the system will automatically activate the backup transmission channel and redundant intervention plan to ensure reliable execution of intervention measures in critical and risky scenarios.
[0063] The aforementioned method for identifying safety risks at construction sites involves deploying a multimodal sensor network to synchronously collect data on personnel movement, location, and the environment. Through spatial registration and fusion, a dynamic digital twin is constructed that accurately reflects the personnel's full-body posture and the surrounding environment. From this twin, multidimensional risk features, such as kinematics, posture stability, and environmental interaction, representing signs of instability, are extracted in real time and input into an AI risk assessment model for predictive calculation. Ultimately, when a critical risk is identified, the model proactively predicts the fall trajectory and generates precise intervention commands. This represents a fundamental shift in the assessment of fall risks from post-event alarms to pre-event prediction and early warning, significantly improving the accuracy and timeliness of risk identification. It provides crucial decision-making support for activating active protection systems and effectively enhances the safety protection capabilities for high-altitude operations.
[0064] refer to Figure 2 In one optional embodiment, spatial registration is performed on the multimodal sensing data to generate a temporal dynamic motion state sequence for each worker and a fused environmental point cloud of the work area, including the following steps:
[0065] S11. Based on the pre-measured installation parameters of the fixed sensors, obtain the initial transformation matrix of each fixed sensor from its own coordinate system to the global coordinate system of the construction site.
[0066] Specifically, the core of this step is to obtain the initial transformation matrix of each fixed sensor from its own coordinate system to the global coordinate system of the construction site. This matrix is the fundamental prerequisite for achieving spatial unification of sensor data. First, the coordinate system definition is clarified. The global coordinate system of the construction site adopts a geodetic coordinate system, determined through satellite positioning technology combined with site benchmark control points. Its origin is selected from the geometric center of the construction site plane or a designated benchmark point. The coordinate axes are set according to the right-hand rule, where the Z-axis is perpendicular to the horizontal plane of the construction site and points towards the zenith, and the X-axis and Y-axis form an orthogonal coordinate system within the horizontal plane of the construction site. Fixed sensor installation parameters include key parameters such as the sensor's installation position coordinates, installation attitude angle, and installation height. These parameters are collected on-site using professional measuring equipment after sensor deployment to ensure that the parameters accurately reflect the actual installation state of the sensors in the global coordinate system of the construction site. The initial transformation matrix consists of a rotation matrix and a translation vector, expressed as follows: In the formula The initial rotation matrix represents the transition from the sensor's own coordinate system to the global coordinate system of the construction site. It is used to describe the attitude mapping relationship between the two coordinate systems and is calculated based on the conversion formula from Euler angles to the rotation matrix using the attitude angles (roll angle, pitch angle, yaw angle). This represents the initial translation vector, corresponding to the coordinates of the sensor's own coordinate system origin in the global coordinate system of the construction site, directly determined by the installation position coordinates; 0 is the zero vector, and 1 is the homogeneous coordinate identifier, ensuring that the transformation matrix meets the mathematical requirements of homogeneous coordinate transformation. The purpose of the initial transformation matrix is to achieve preliminary coordinate alignment of various fixed sensor data, providing a foundation for subsequent precise registration.
[0067] S12. Based on the initial transformation matrix, the stereo vision point cloud and lidar point cloud in the environmental point cloud data are transformed to the global coordinate system of the construction site; for the stereo vision point cloud and lidar point cloud transformed to the global coordinate system of the construction site, the iterative nearest point registration algorithm is executed in the spatially overlapping area to obtain the accurate transformation matrix of each fixed sensor.
[0068] Specifically, this step completes the point cloud coordinate transformation based on the initial transformation matrix, and optimizes it using an iterative nearest-point registration algorithm to obtain an accurate transformation matrix, thereby eliminating registration errors caused by initial installation errors and inherent sensor biases. First, using the initial transformation matrix obtained in S11, coordinate transformation operations are performed on the stereo vision point cloud acquired by the stereo vision camera and the lidar point cloud acquired by the lidar, transforming the two types of point clouds from their respective sensor coordinate systems to the global coordinate system of the construction site. The coordinate transformation process uses a homogeneous coordinate transformation method, for any point in the point cloud data... (Homogeneous coordinate representation), points transformed to the global coordinate system of the construction site The calculation formula is: In this formula, matrix multiplication is performed according to the homogeneous coordinate transformation rule to achieve synchronous transformation of the point's position and orientation. Subsequently, for the stereo vision point cloud and the LiDAR point cloud transformed to the global coordinate system of the construction site, the spatial overlap region is first determined by spatial range judgment. This region is determined based on the intersection of the coordinate ranges of the two types of point clouds in the global coordinate system of the construction site. Specifically, the maximum and minimum values of the x, y, and z coordinates of the stereo vision point cloud are compared with the corresponding coordinate ranges of the LiDAR point cloud, and the intersection is taken as the spatial overlap region. Within the spatial overlap region, an iterative nearest-point registration algorithm is executed. The core of this algorithm is to find the optimal transformation matrix, i.e., the accurate transformation matrix, that minimizes the spatial distance error between the two types of point clouds through iterative optimization.
[0069] The execution process of the iterative nearest neighbor registration algorithm is as follows: First, a set of sampling points is selected from the stereo vision point cloud (source point cloud). Then, the nearest neighbor of each sampling point is found in the lidar point cloud (target point cloud), constructing a set of point pairs. Based on this set of point pairs, the root mean square error between point clouds under the current transformation matrix is calculated. If the root mean square error is greater than a preset convergence threshold, a new transformation matrix is solved using the least squares method, and the coordinates of the source point cloud are updated. The above steps of sampling, finding nearest neighbors, calculating errors, and updating transformation matrices are repeated until the root mean square error is less than the convergence threshold or the maximum number of iterations is reached. The transformation matrix obtained at this time is the accurate transformation matrix for each fixed sensor. In the formula For an accurate rotation matrix, As a precise translation vector, compared to the initial transformation matrix, it can more accurately describe the mapping relationship between the sensor coordinate system and the global coordinate system of the construction site, thus improving the point cloud registration accuracy.
[0070] S13. Based on the precise transformation matrix, the environmental point cloud data collected by all fixed sensors are fused to generate a fused environmental point cloud.
[0071] Specifically, this step uses a precise transformation matrix to fuse all environmental point cloud data collected by fixed sensors, generating a fused environmental point cloud. First, for each fixed sensor's environmental point cloud data (including stereo vision and LiDAR point clouds), a coordinate transformation is performed using its corresponding precise transformation matrix to ensure all point cloud data are accurately mapped to the global coordinate system of the construction site. Then, point cloud fusion processing is performed, which includes three key steps: point cloud deduplication, point cloud completion, and point cloud optimization. Point cloud deduplication addresses duplicate points in spatially overlapping areas of point clouds collected by different sensors. By calculating the Euclidean distance between points, if the distance between two points is less than a preset deduplication threshold, it is considered a duplicate point. One of these points is retained (usually points from the LiDAR point cloud due to its higher geometric accuracy), or the coordinates of the duplicate points are weighted and averaged. The weighting coefficients are determined based on the sensor's measurement accuracy. Point cloud completion addresses missing areas in point cloud data from a single sensor by supplementing it with point cloud data from other sensors. For example, stereo vision point clouds may have missing data in poorly lit areas; by fusing LiDAR point cloud data for that area, complete point cloud coverage is achieved. Point cloud optimization removes outliers from the fused point cloud through statistical filtering. This filtering method calculates the average distance between each point and its neighbors; if the average distance exceeds a preset outlier threshold, the point is identified as an outlier and removed, thus improving the quality of the fused environmental point cloud. After the above fusion processing, the generated fused environmental point cloud not only includes visual information (such as color and texture) from the stereo vision point cloud but also incorporates high-precision geometric information from the LiDAR point cloud, enabling a comprehensive and accurate representation of the environmental characteristics of the work area, including the shape of the work surface, obstacle distribution, and edge contours.
[0072] S14. Perform Euclidean clustering on the fused environment point cloud to obtain independent human point cloud clusters corresponding to different workers, and assign a person identifier to each independent human point cloud cluster; input each independent human point cloud cluster into a pre-trained human pose estimation neural network, and output the visual joint coordinates of all joints of the corresponding person identifier in the global coordinate system of the construction site.
[0073] Specifically, this step obtains independent human point cloud clusters through Euclidean clustering and assigns personnel identifiers, and then outputs visual joint coordinates through a human pose estimation neural network. First, Euclidean clustering is performed on the fused environment point cloud. The core principle of this algorithm is to cluster and group points based on the Euclidean distance between points in the point cloud, dividing points that are closer into the same cluster and points that are farther apart into different clusters.
[0074] The specific execution process is as follows: An unlabeled point is selected from the fused environmental point cloud as a seed point. All points within the seed point's neighborhood whose distance is less than the clustering distance threshold are found and assigned to the current cluster. Then, a new unlabeled point is selected from the current cluster as a new seed point, and the process of finding neighboring points is repeated until no new points are added to the current cluster. Next, the next unlabeled point is selected as a new seed point, and a new clustering process begins until all points are labeled and assigned to a cluster. Through Euclidean clustering, the human point cloud in the fused environmental point cloud can be separated from the environmental point cloud (such as the work surface and obstacles), resulting in independent human point cloud clusters corresponding to different workers.
[0075] Each individual human point cloud cluster is assigned a personnel identifier. The assignment rule is based on the association between the spatial location of the point cloud cluster and the positioning data of the mobile sensor nodes. Specifically, the center coordinates of each individual human point cloud cluster are calculated, and spatial distance matching is performed with the positioning data of each mobile sensor node. The unique identifier of the nearest mobile sensor node is assigned to the human point cloud cluster, ensuring a one-to-one correspondence between the personnel identifier and the worker. Subsequently, each individual human point cloud cluster is input into a pre-trained human pose estimation neural network. This neural network is built on a deep learning framework, and the training process uses point cloud datasets containing different work scenarios and different human poses. The datasets cover various work postures such as standing, walking, bending over, and climbing. Each sample is labeled with the coordinate information of key human joints (such as head, neck, shoulder, elbow, wrist, hip, knee, and ankle).
[0076] The neural network takes point cloud data of individual human point cloud clusters as input (after standardization, the point cloud center is translated to the origin and scaled to a uniform scale), and outputs the visual joint coordinates of all joints of the corresponding personnel identifier in the global coordinate system of the construction site. Each joint coordinate is represented by a three-dimensional vector. It means that among them , , These are the x, y, and z axis coordinates of the joint in the global coordinate system of the construction site. These visual joint coordinates can accurately reflect the spatial position of each joint in the human body, providing joint-level constraint information for the subsequent construction of dynamic motion state sequences.
[0077] S15. Input the positioning data corresponding to each person identifier and the position data calculated based on the center of the independent human point cloud cluster into an extended Kalman filter for fusion filtering to obtain the fused global position and global velocity; input the inertial data corresponding to each person identifier into a complementary filter for attitude calculation to obtain the torso attitude quaternion.
[0078] Specifically, this step fuses positioning data and point cloud center location data using an extended Kalman filter, and calculates the torso pose quaternion using a complementary filter. For each personnel identifier, its corresponding positioning data (collected from an ultra-wideband positioning tag) and location data calculated based on the center of an independent human point cloud cluster are first obtained. The positioning data directly reflects the overall position information of the worker, while the center location data of the human point cloud cluster is obtained by calculating the average coordinates of all points in the independent human point cloud cluster, expressed as follows: In the formula Let M be the coordinates of the center of the point cloud cluster, and M be the number of points in an independent human point cloud cluster. Let be the coordinates of the i-th point in the point cloud cluster. These two types of position data are input into an extended Kalman filter (EPF) for fusion filtering. The EPF is suitable for state estimation of nonlinear systems. Its core principle is to describe the dynamic and observation processes of the system through state equations and observation equations, and to achieve optimal state estimation through iterative updates. The state equation of the EPF is defined as follows: In the formula The state vector at time k contains the global position. With global speed ; It is a nonlinear state transition function, constructed based on a kinematic model, describing the evolution of the state from time k-1 to time k; To control the input, set it to 0 here; The process noise follows a Gaussian distribution with a mean of 0 and a covariance matrix of Q. The observation equation is defined as follows: In the formula Let k be the observation vector at time k, containing ultra-wideband positioning data. Data on the center position of point cloud clusters ; For nonlinear observation functions, the state vector is mapped to the observation space; The observed noise follows a Gaussian distribution with a mean of 0 and a covariance matrix of R.
[0079] The extended Kalman filter (EPF) execution process includes two stages: prediction and update. In the prediction stage, the prior state estimate and prior covariance matrix at time k are calculated based on the state equation. In the update stage, the Kalman gain is calculated using the observation equation, and the prior state estimate is corrected by combining the observation vector to obtain the posterior state estimate (i.e., the fused global position and global velocity) and posterior covariance matrix at time k. This fusion filtering process fully combines the advantages of two types of position data (high update frequency of ultra-wideband positioning data and high accuracy of point cloud cluster center position data), resulting in more accurate and stable global position and global velocity. Simultaneously, the inertial data corresponding to each personnel identifier (three-axis acceleration and three-axis angular velocity collected by the inertial measurement unit) is input into a complementary filter for attitude calculation. The core principle of the complementary filter is to utilize the complementary characteristics of the accelerometer and gyroscope in the inertial measurement unit. The accelerometer can measure the gravitational acceleration component and can be used for long-term stable attitude estimation, but it is susceptible to motion acceleration interference; the gyroscope can measure angular velocity and can be used for short-term rapid attitude tracking, but it has drift errors. The complementary filter, by setting appropriate filter coefficients, performs a weighted fusion of the attitude calculated by the accelerometer and the attitude obtained by the gyroscope integration. The expression is as follows: In the formula For the torso attitude quaternion output by complementary filtering, The attitude quaternion is obtained by integrating the gyroscope angular velocity. The attitude quaternion is obtained based on the gravity component calculated from the accelerometer. The filter coefficients (ranging from 0 to 1) were determined through experimental calibration to balance the error characteristics of both. The final output is the torso pose quaternion. Satisfying the unit quaternion constraint It can accurately describe the spatial posture of the operator's torso.
[0080] S16. Construct a nonlinear graph optimization model with constraints of global position, global velocity, torso posture quaternion and visual joint coordinates; solve the nonlinear graph optimization model and output a time-series dynamic motion state sequence.
[0081] Specifically, the core of the nonlinear graph optimization model is to transform the problem of estimating dynamic motion states into a graph optimization problem. The graph consists of nodes and edges, where nodes represent motion state variables at each time step, and edges represent the constraints between these variables. Specifically, a node is defined as the motion state vector at each time step. In the formula Let be the global position at time t. Let be the global velocity at time t. Let be the quaternion of the torso posture at time t. Let t represent the global coordinates of n key joints of the human body (each joint coordinate is represented by a 3D vector). The dimension of each node is determined according to the number of joints to ensure a comprehensive representation of the worker's motion state and posture information. The construction of edges is based on three types of constraints: kinematic constraint edges, visual joint constraint edges, and posture-joint consistency constraint edges. Each type of constraint edge is quantified by a residual function to describe the degree of satisfaction of the constraint relationship, providing an optimization objective for graph optimization.
[0082] Kinematic constraint edges are used to connect nodes at adjacent time steps. Constructed based on kinematic principles, they ensure the temporal continuity of motion states. For nodes at adjacent time steps t-1 and t... and The residual function of the kinematically constrained edge is defined as In the formula For time step interval, The acceleration component at time t-1 is calculated from inertial data and used as a known input. This residual function reflects the deviation between the current time step state and the kinematic prediction based on the previous time step state. The smaller the deviation, the better the continuity of the motion state. The residual vector has a dimension of 6, corresponding to the three spatial components of position and velocity.
[0083] Visual joint constraint edges are used to associate the joint coordinates in the associated node with the visual joint coordinates output by S14, ensuring consistency in the observed joint positions. For a node at time t... Corresponding visual joint coordinate observations The residual function of the visual joint constraint edge is defined as In the formula, the residual for each joint is the three-dimensional vector difference between the joint coordinates in the node and the visually observed joint coordinates, with the residual vector having a dimension of 3n (n being the number of joints). This residual function quantifies the deviation between the joint coordinates estimated by the model and the joint coordinates observed by the visual sensor. By minimizing this deviation, the accuracy of joint position estimation is ensured, while the constraint effect of visual observation is introduced to compensate for the lack of information at the joint level in pure inertial and positioning data.
[0084] Posture-joint consistency constraint edges are used to associate the torso posture quaternions and joint coordinates in the associated nodes. Based on the human skeletal kinematics model, this ensures the physical plausibility of posture and joint positions. The human skeletal kinematics model predefines the relative positions of each joint with respect to the torso coordinate system and the bone lengths, such as the relative coordinates of the shoulder joint with respect to the torso center, the upper arm bone length, and the forearm bone length. These parameters are determined through statistical human body size data and support personalized calibration. For nodes at time t... First, through the quaternion of torso posture Transform the relative coordinates of each joint to the global coordinate system of the construction site using the following formula: In the formula The local relative coordinates of the joints with respect to the torso coordinate system (three-dimensional vectors, homogenized into quaternions) ), for The conjugate quaternion, This is quaternion multiplication. Let be the global position at time t. Based on the transformed relative joint global coordinates and bone length constraints, the residual function of the pose-joint consistency constraint edge is defined as... In the formula and For two joints that are skeletally connected (such as shoulder and elbow, elbow and wrist, etc.), the global coordinates are given. Given a preset corresponding bone length, the residual vector has a dimension of m (where m is the number of bone connections). This residual function ensures that the distances between the joint coordinates estimated by the model conform to the physical constraints of human bone length, avoiding unreasonable states where joint positions and postures contradict each other, and improving the physical realism of the motion state sequence.
[0085] Based on the three types of constraint edges mentioned above, the objective function of the nonlinear graph optimization model is defined as minimizing the weighted sum of squared residuals of all constraint edges, expressed as:
[0086]
[0087] In the formula, T is the total number of time steps. , , These are weight matrices for the three types of constraint edges, all of which are diagonal matrices. The values of the diagonal elements are determined based on the measurement accuracy of the corresponding sensor and the importance of the constraint. For example, the weight matrix for visual joint constraints... In the weight matrix of kinematic constraints, the diagonal elements are inversely proportional to the measurement error variance of the visual joint coordinates; joints with higher measurement accuracy (such as the head and shoulders) have larger weights. In this context, the weights of the position and velocity components are set based on the acceleration and angular velocity measurement accuracy of the inertial measurement unit; the weight matrix of the attitude-joint consistency constraint... In this process, the weights of the bone length constraints are set according to the statistical error range of the bone length, ensuring that various constraints can play a role with reasonable priority during the optimization process.
[0088] The Levenberg-Marquardt algorithm is used to solve the nonlinear graph optimization model. This algorithm combines the fast convergence of the Gauss-Newton method with the stability of the gradient descent method, and is suitable for solving nonlinear least squares problems. The specific steps of the solution process are as follows: First, initialize the node states for all time steps. The initial values are derived from the global position, global velocity, and torso pose quaternions output by S15 and the visual joint coordinates output by S14, ensuring that the initial state satisfies the basic constraints. Subsequently, the adjacency matrix and residual Jacobian matrix of the graph are constructed. The Jacobian matrix is calculated using numerical or automatic differentiation methods, reflecting the partial derivatives of each residual with respect to the node state variables. An approximation of the Hessian matrix is then constructed based on the Jacobian matrix and the weight matrix. (in The global residual Jacobian matrix is... (This is the combined weight matrix), and the gradient vector is calculated. (where r is the global residual vector); then, the damping factor is introduced. Solve the incremental equation (where I is the identity matrix,) (This is the increment vector of the node state). The increment vector is then superimposed onto the current node state to obtain a new state estimate. Calculate the objective function value under the new state; adjust the damping factor based on the change in the objective function value. If the objective function value decreases, then accept the new state and decrease it. If the objective function value increases, then the new state is rejected and the value is increased. Repeat the steps of calculating the Jacobian matrix, constructing the Hessian matrix, solving the incremental equation, updating the state, and adjusting the damping factor until the objective function value is less than the preset convergence threshold or the maximum number of iterations is reached. The resulting node state sequence is then obtained. This is the optimized time-series dynamic motion state sequence.
[0089] This time-series dynamic motion state sequence is indexed by time steps. The state at each time step includes global position, global velocity, torso posture quaternions, and global coordinates of all key joints. The temporal resolution of the sequence is consistent with the time step interval of data acquisition, enabling continuous and accurate characterization of the dynamic motion state and posture changes of workers during operations. The sequence also embeds metadata such as timestamps and personnel identifiers, facilitating data association with subsequent feature calculations and risk assessment steps, providing high-quality core data support for comprehensive and accurate safety risk identification.
[0090] In one optional embodiment, the kinematic characteristics and posture stability characteristics of the operator are calculated based on a time-series dynamic motion state sequence, including the following steps:
[0091] S21. Based on the linear acceleration vectors of the current and previous moments in the time-series dynamic motion state sequence, calculate the motion mutation index as a kinematic feature; the formula for calculating the motion mutation index is:
[0092]
[0093] in, Indicates the current moment. express Motion mutation index at time of moment and These represent the extraction from the temporally sequenced dynamic motion state sequence. and The linear acceleration vector at time t. The sampling time interval is the time interval for the time-series dynamic motion state sequence.
[0094] Specifically, the calculation of the motion mutation index focuses on capturing instantaneous and drastic changes in motion state. This index is determined by the ratio of the difference in linear acceleration vectors at consecutive time points to the sampling time interval. The core calculation formula is as follows: ,in For the current moment, and These are the linear acceleration vectors for the current and previous moments, respectively. The sampling time interval is defined. Before calculation, outlier checks are performed on the acceleration vector to remove data that exceeds the physical limits of motion, ensuring the accuracy of the exponent. This results in a continuous sequence of motion mutation exponents to characterize the dynamic changes in the motion state.
[0095] S22. Based on the joint coordinates in the time-series dynamic motion state sequence and the standard human link quality parameters, calculate the three-dimensional coordinates of the human body's overall centroid; project the three-dimensional coordinates vertically onto the local ground to obtain the centroid projection point; and determine the supporting polygon based on the joint coordinates of the two ankle joints in the joint coordinates.
[0096] Specifically, the three-dimensional coordinates of the overall human centroid are calculated based on joint coordinates and standard human link mass parameters. A weighted summation method is used to integrate the centroid coordinates of each limb segment, with the weight representing the mass percentage of each segment. The centroid projection point is obtained by vertically projecting the overall centroid's three-dimensional coordinates onto a local ground plane, which is obtained by segmenting and fitting the fused environmental point cloud. The construction of the support polygons is based on the coordinates of both ankle joints, combined with the characteristics of the foot support area. A convex hull algorithm is used to extract convex polygons, comprehensively covering the current effective support range and providing a spatial benchmark for subsequent stability assessment.
[0097] S23. Calculate the shortest distance from the centroid projection point to the boundary of the supporting polygon, and use it as the first attitude stability feature.
[0098] Specifically, the first attitude stability feature focuses on the spatial relationship between the centroid projection point and the boundary of the supporting polygon, with the core being the calculation of the shortest distance between them. The boundary of the supporting polygon consists of several continuous line segments, each of which is mathematically represented by the two-dimensional coordinates of its adjacent support points. For the centroid projection point... The shortest distance to any boundary line segment is calculated using the formula for the distance from a point to a line. The core formula is: ,in The parameters are the linear equation parameters for the boundary line segments. The minimum calculated distance is taken as the first attitude stability feature after traversing all boundary line segments. This feature directly reflects the safety redundancy of the center of mass within the support range; a larger distance indicates stronger attitude stability. When the distance approaches zero or becomes negative, it indicates that the operator has approached or exceeded the support limit, and the risk of imbalance increases significantly.
[0099] S24. Based on the torso posture quaternions in the temporally sequenced dynamic motion state sequence, calculate the torso tilt angle as a second posture stability feature; the formula for calculating the torso tilt angle is:
[0100]
[0101] in, express The angle of torso tilt at any moment express The trunk posture quaternion in the time-sequential dynamic motion state sequence; where The real part encodes half the cosine value of the rotation angle; , , The virtual components encode the components of the rotation axis in the X, Y, and Z axes of the global coordinate system of the construction site.
[0102] Specifically, the torso tilt angle, as a second posture stability feature, is used to quantify the degree to which the torso deviates from a vertical state. The core calculation formula is as follows: ,in Let the current torso posture quaternion be , For the real part, This is the imaginary component corresponding to the axis of the global coordinate system. The larger this angle value, the more pronounced the torso tilt, and the stronger the destructive effect on posture stability. It is a key indicator for assessing the risk of imbalance.
[0103] S25. Within the set time window, calculate the sample entropy of the torso tilt angle sequence as the third posture stability feature.
[0104] Specifically, the third posture stability feature quantifies the regularity of trunk tilt angle changes through sample entropy. The sample entropy is calculated based on a set time window, and its core function is to measure the degree of repetition of similar patterns in the sequence. Within the specified time window, continuous trunk tilt angle data are processed. The smaller the calculated sample entropy value, the more regular and stable the tilt angle changes; the larger the sample entropy value, the more irregular the tilt angle fluctuations and the worse the posture stability. This feature can effectively capture sudden tilts or persistent unstable states of the trunk.
[0105] The above features, from the three dimensions of motion change, spatial balance, and posture trend, complement each other and comprehensively cover the core influencing factors of worker posture stability, providing accurate and efficient feature input for subsequent risk assessment models, and ensuring the pertinence and reliability of risk identification.
[0106] In one optional embodiment, environmental interaction features are calculated based on fused environmental point cloud and wind speed and direction data, including the following steps:
[0107] S31. Extract the local ground point cloud below each worker's feet from the fused environmental point cloud; calculate the normal direction variance and height variance of all points in the local ground point cloud.
[0108] Specifically, the core of this step is to extract the local ground point cloud beneath the worker's feet and calculate two types of variance parameters to characterize the stability of the ground's physical state. The extraction of the local ground point cloud is based on determining the spatial range using foot joint coordinates (left and right ankle joints and toe joints) in a time-series dynamic motion state sequence: using the minimum bounding box formed by the foot joint coordinates as a basis, a preset buffer distance is extended into three-dimensional space (ensuring coverage of the actual contact area between the feet and the surrounding ground area). Point cloud data within this spatial range is then cropped from the fused environmental point cloud, which is the corresponding local ground point cloud for the worker. The calculation of the normal direction variance requires first solving for the normal vector of each point in the local ground point cloud using Principal Component Analysis (PCA). Specifically, this involves constructing a covariance matrix for the neighborhood point set of each point, solving for the eigenvalues and eigenvectors, and finding the eigenvector corresponding to the smallest eigenvalue, which is the normal direction of that point. Subsequently, the direction cosine variance of the normal vectors of all points is calculated. The larger this variance value, the more drastic the change in ground slope and the worse the flatness. The height variance is calculated by statistically analyzing the global coordinates (height values) of all points in the local ground point cloud. It is obtained by calculating the variance of the coordinate sequence. The larger the value, the more significant the unevenness of the ground. The two types of variance together quantify the physical stability of the ground.
[0109] S32. Based on the variance of the normal direction, the variance of the height, and the ambient humidity data, the ground friction coefficient level is obtained by querying a predefined mapping table.
[0110] Specifically, this step obtains the ground friction coefficient level through multi-parameter mapping, with the core being the establishment of a correlation between physical parameters and friction characteristics. A predefined mapping table is constructed based on extensive experimental data. During the experiments, different ground smoothness (corresponding to normal direction variance), ground undulation (corresponding to height variance), and environmental humidity conditions are simulated. The actual ground friction coefficient is measured through friction tests, and then the continuous friction coefficient is divided into several levels (e.g., high, medium, and low, each corresponding to different safety thresholds), ultimately forming a four-dimensional mapping relationship of "normal direction variance - height variance - environmental humidity - friction coefficient level." Environmental humidity data is synchronously collected by environmental sensors deployed in the work area and, along with the normal direction variance and height variance, serves as an index. Querying the mapping table yields the current ground friction coefficient level, which directly reflects the adhesion level between the worker's feet and the ground; a lower level indicates a higher risk of slipping.
[0111] S33. Calculate the estimated value of wind load impact based on wind speed and direction data, as well as the windward area and posture of personnel estimated from the time-series dynamic motion state sequence.
[0112] Specifically, the calculation of the wind load impact estimate in this step requires integrating multi-source information to quantify the interference of the wind environment. Wind speed and direction data provide basic wind field parameters, where wind speed is a vector value including magnitude and direction; the windward area of personnel is estimated through a time-series dynamic motion state sequence, and the spatial posture of the torso and limbs is determined based on the human posture quaternion. Combined with a standard human body size model (such as height, shoulder width, torso thickness, etc.), the projected area of the human body in the wind vector direction is calculated. The windward area changes dynamically depending on the posture (e.g., the windward area is larger when standing than when bending over). The calculation of the wind load impact estimate is based on aerodynamic principles. The core logic is that the wind load is positively correlated with the square of the wind speed, the windward area, and the drag coefficient. The expression can be simplified to a product of the above parameters. The drag coefficient is determined by fitting experimental data based on the current posture of the worker (e.g., the drag coefficient is greater when standing than when sideways). The final wind load impact estimate quantifies the degree of interference of the wind field on human balance. The larger the value, the more significant the threat of the wind environment to work safety.
[0113] S34. Target detection is performed on the fused environmental point cloud and the synchronously acquired RGB image. By identifying the state of the object held by the operator and analyzing the micro-motion spectrum of the corresponding point cloud, the characteristics of the operation disturbance are obtained.
[0114] Specifically, this study extracts operational disturbance features through multimodal data fusion, with the core being the identification of the state of the handheld object and the analysis of its micro-motion characteristics. First, joint target detection is performed on the fused environmental point cloud and simultaneously acquired RGB images. A deep learning-based detection model (such as the YOLO series) is used, leveraging the complementary geometric features of the point cloud and the texture features of the RGB images to accurately identify whether the operator is holding tools, materials, or other handheld objects, while simultaneously determining the spatial location and approximate type of the object. Subsequently, micro-motion spectrum analysis is performed on the point cloud region corresponding to the handheld object: the coordinate changes of the point cloud in this region are extracted over continuous time steps, and the time-domain micro-motion signal is converted into a frequency-domain spectrum using Fast Fourier Transform (FFT). From the spectrum, characteristic parameters such as peak frequency, spectral width, and energy distribution are extracted. These parameters collectively constitute the operational disturbance features. This feature can quantify the unstable interference caused by the operation of the handheld object (such as inertial disturbance when holding a heavy object, or vibration disturbance when using a tool). The more significant the spectral peak and the more dispersed the energy distribution, the greater the operational disturbance and the more significant its impact on the stability of the work posture.
[0115] S35. Ground friction coefficient level, wind load impact estimate, and operational disturbance characteristics are used as environmental interaction characteristics.
[0116] Specifically, this step integrates the ground friction coefficient level, the estimated wind load impact, and operational disturbance characteristics into environmental interaction features. These three types of features approach the issue from three dimensions: ground adhesion, external wind interference, and operational disturbance, respectively. They complement each other and comprehensively cover the core interaction between the environment and workers. Among them, the ground friction coefficient level reflects the safety of foot-to-ground contact, the estimated wind load impact characterizes the degree of interference from external environmental forces, and the operational disturbance characteristics reflect the dynamic disturbances during the operation. Together, these three features provide a complete environmental input for comprehensive risk assessment, ensuring that risk identification can fully consider the comprehensive impact of complex environmental factors.
[0117] In one optional embodiment, kinematic features, posture stability features, and environmental interaction features are input into a pre-trained risk assessment model to output a comprehensive risk index, including the following steps:
[0118] S41. Combine the kinematic features and attitude stability features into a first feature vector; input the first feature vector into the first risk assessment sub-model for processing to obtain the instability risk probability. .
[0119] Specifically, the kinematic features are centered on the motion mutation index sequence, including its statistical values (such as peak value, mean, and standard deviation) within a set time window; the posture stability features integrate the first, second, and third posture stability features (shortest distance from the centroid projection point to the boundary of the supporting polygon, torso tilt angle, and torso tilt angle sample entropy) to form a multi-dimensional posture representation. The two types of features are concatenated in a preset order to form a first feature vector with fixed dimensions. During the concatenation process, all features must be standardized (using Z-score standardization to eliminate dimensional differences) to ensure that each feature has equal weight in model training and inference. The first risk assessment sub-model uses an optimized machine learning model (such as a gradient boosting decision tree or a lightweight neural network). The model training data comes from instability simulation experiments and actual work data under different operating scenarios, covering various states such as normal motion, slight instability, and severe instability. The training objective is to minimize the cross-entropy loss between the predicted instability probability and the actual state label. During model inference, the first feature vector is input and then forward-propagated to calculate the instability risk probability in the range [0,1]. The higher the value, the higher the likelihood that the worker's current movement and posture will lead to instability.
[0120] S42. Obtain the footstep characteristics of the workers and the proximity of the workers to obstacles. Combine the characteristics, proximity, and the ground friction coefficient level in the environmental interaction characteristics into a second feature vector. Input the second feature vector into the second risk assessment sub-model to obtain the slip risk probability. .
[0121] Specifically, footstep features are extracted from a time-series dynamic motion state sequence. Based on the coordinate changes of foot joints (ankle and toe joints), quantitative parameters such as stride length, stride frequency, foot landing angle, and the angle between the foot and the ground are calculated. These parameters reflect the stability of the worker's walking posture (e.g., sudden changes in stride length or abnormal foot landing angles may increase the risk of slipping). The proximity of the worker to obstacles is determined by fusing environmental point cloud data: first, obstacle point cloud clusters are segmented from the fused environmental point cloud (based on Euclidean clustering and preset obstacle size thresholds). Then, the shortest Euclidean distance between the worker's human point cloud cluster and each obstacle point cloud cluster is calculated, and the minimum value is taken as the final quantitative value of proximity. The smaller the distance, the higher the risk of collision or tripping. The extracted footstep features (statistical form), obstacle proximity values, and ground friction coefficient levels obtained from S32 (converted to numerical codes, such as high=3, medium=2, low=1) are concatenated in sequence to form a second feature vector, which is also standardized and then input into the second risk assessment sub-model. This sub-model is trained based on the causal characteristics of slip risk, focusing on learning the synergistic effect of abnormal footsteps, obstacle proximity, and low friction coefficient, and outputting the probability of slip risk. The value range is [0,1], which directly represents the probability of a slip accident.
[0122] S43. Based on the global position in the time-series dynamic motion state sequence and the boundary of the preset high-risk area in the construction site BIM model, calculate the distance from the worker to the nearest edge or opening as the edge distance; determine whether the worker is located within the preset high-risk area, and generate a high-risk area marker when the worker is within the preset high-risk area; combine the edge distance, the high-risk area marker, and the wind load impact estimate in the environmental interaction features into a third feature vector, and input the third feature vector into the third risk assessment sub-model for processing to obtain the environmental hazard probability. .
[0123] Specifically, the calculation of the edge distance is based on the site BIM model, which pre-stores the boundary coordinates and geometric parameters (such as the polyline coordinates of the edge and the rectangular boundary of the opening) of all preset high-risk areas (such as edges, openings, and deep foundation pit edges). The real-time global position coordinates of the workers are extracted from the time-series dynamic motion state sequence. The edge distance is obtained by calculating the shortest Euclidean distance from these coordinates to the boundary of the high-risk area. A distance of 0 indicates contact with the boundary, while a negative distance indicates entry into the high-risk area. The high-risk area marker is a binary variable; when the global position coordinates of the workers fall within the preset high-risk area boundary range of the BIM model, the marker value is set to 1; otherwise, it is set to 0, used to visually represent whether the workers are in a directly hazardous environment. The edge distance (normalized to eliminate size differences between different high-risk areas), the high-risk area marker (numerical form), and the wind load impact estimate obtained from S33 are concatenated to form a third feature vector, which is then input into the third risk assessment sub-model. This sub-model specifically learns the combined impact of edge environment and wind load disturbance on operational safety (e.g., when the edge distance is close and the wind load is large, the probability of environmental hazards increases significantly). The training data covers simulated operational scenarios under different edge distances and wind field conditions, and outputs the probability of environmental hazards. The value range is [0,1], which quantifies the degree of risk of safety accidents caused by the external environment.
[0124] S44. Based on the worker's movement patterns in the time-series dynamic motion state sequence and predefined scene classification rules, determine the context category of the current work scene; based on the context category, determine the probability of instability risk. Slipping risk probability Probability of environmental hazards Weights are assigned, and the basic probability values of each evidence source obtained based on the weight mapping are fused using DS evidence theory to obtain a comprehensive risk index; among which, the context categories include smooth walking, stationary high-altitude work, movement near the edge, or movement on a narrow beam.
[0125] Specifically, the determination of the context category is based on the worker's movement pattern and predefined rules. Key motion parameters (such as global mean velocity, velocity variance, trajectory curvature, and posture stability index) are extracted from the time-series dynamic motion state sequence and matched with predefined scene classification rules. For example, when the global mean velocity is within a preset stable range, the velocity variance is small, and the trajectory curvature is gentle, it is judged as "stable walking"; when the global velocity is close to 0 and the torso tilt angle is stable within the working posture range, it is judged as "stationary high-altitude work"; when the distance to the edge is less than a preset threshold and the global velocity is not 0, it is judged as "moving near the edge"; when the trajectory width (the span of the trajectory perpendicular to the direction of movement) is less than the narrow space threshold and the average step length is less than a preset value, it is judged as "moving on a narrow beam".
[0126] Based on a defined context category, dynamic weights are assigned to the three risk probabilities. The weighting rules are preset based on the scenario's risk characteristics and calibrated experimentally. For example, in the "smooth walking" scenario, the slip risk probability has the highest weight (e.g., ...). The risk of instability is the second most important factor. ), with the lowest environmental risk weight ( In the scenario of "stationary high-altitude operation", the risk of instability has the highest weight. Environmental risk is the second most significant factor. ), with the lowest risk weight for slipping ( In the "moving near the edge" scenario, environmental hazards have the highest weight. The risk of instability is the second most important factor. ), with the lowest risk weight for slipping ( In the scenario of "movement on a narrow beam," the risk of instability and environmental hazards are equally important (e.g., , The risk weight of slippage is relatively low ( ).
[0127] After weighting, the evidence is fused based on the DS evidence theory. First, each risk probability is mapped to a basic probability assignment (BPA) of the evidence source according to its corresponding weight. The mapping rule is as follows: , ,in For each risk event, the corresponding proposition (such as...) To prevent the risk of instability, To prevent the risk of slipping, (For environmental hazards to occur) To identify the frame (the set of all possible propositions), the DS combination rule is then used to fuse the BPAs of the three evidence sources. The combination formula is as follows:
[0128]
[0129] in, The conflict coefficient is used to measure the degree of conflict between pieces of evidence, ensuring the rationality of the fusion process. Ultimately, the BPA value corresponding to the "risk occurrence" proposition after fusion is the comprehensive risk index, with a value range of [0,1]. The larger the value, the higher the comprehensive safety risk currently faced by the operators, providing a direct quantitative basis for subsequent risk decisions.
[0130] In one optional embodiment, the basic probability assignments of each evidence source obtained based on weight mapping are fused using DS evidence theory to obtain a comprehensive risk index, including the following steps:
[0131] S51, Based on the probability assigned to the instability risk Slipping risk probability Probability of environmental hazards The weights are mapped to the basic probability assignment functions of the corresponding evidence sources. , and And set the recognition framework as ,in This represents a high-risk proposition. This represents a low-risk proposition.
[0132] Specifically, the first step is to establish a recognition framework. In this framework, proposition H is explicitly defined as "high risk," indicating that the current state of the worker presents significant safety hazards, with a high probability of accidents such as falls or slips. Proposition L is explicitly defined as "low risk," indicating that the current state of the worker is safe and stable, and all risk factors are within a controllable range. This binary identification framework is concise and highly targeted, directly addressing the core judgment needs of risk assessment and avoiding the increased complexity of integration caused by multiple proposition frameworks.
[0133] Based on the above identification framework, the probability of instability risk will be... Slipping risk probability Environmental risk probability Each of these, combined with its corresponding weight, is mapped to a basic probability assignment function for three independent sources of evidence. (Evidence of corresponding instability risk) (Evidence of corresponding slip risk) (Corresponding to environmental hazard evidence). Each basic probability assignment function assigns probability values to two propositions in the identification framework, satisfying the probability normalization constraint (i.e., for the same source of evidence, the sum of the basic probability assignments for propositions H and L is 1). The specific mapping rule is: for the source of evidence... It assigns a basic probability value to proposition H. Assigning basic probabilities to proposition L Similarly, sources of evidence The assignment value is , Source of evidence The assignment value is , .in, , , The dynamic weights assigned to S44 based on the context of the work scenario are mapped in a way that reflects the importance of each risk factor in different scenarios. This ensures that the values assigned to the evidence sources reflect both the probability of the risk itself and the risk characteristics of the scenario.
[0134] S52, Assigning basic probability functions , and By performing pairwise combinations and iterative fusion according to Dempster's combination rule, a joint basic probability assignment function is obtained. Among them, for the two sources of evidence and , on the proposition fusion results The calculation method is as follows:
[0135]
[0136] in, Indicate the source of evidence and After fusion, the proposition Basic probability assignment, This represents the Dempster combination rule operator. , and All are identification frames a subset of and Representing the sources of evidence respectively On the proposition Basic probability assignment and sources of evidence On the proposition Basic probability assignment; conflict coefficient Let represent the sum of the basic probability assignments of all propositional combinations that satisfy the condition that the intersection of X and Y is empty, and Iterative fusion refers to the process of computation... To calculate The iterative process.
[0137] Specifically, the iterative fusion process consists of two steps: the first step is to analyze the sources of evidence for instability risk. Evidence sources related to slip risk By fusing the evidence, an intermediate joint source of evidence can be obtained. The second step is to combine intermediate evidence sources. Sources of evidence related to environmental hazards By fusing the evidence, a joint basic probability assignment function covering all evidentiary information is finally obtained. This iterative process follows the principle of "pairwise combination and gradual convergence," which can effectively handle the conflict coordination problem in multi-evidence source fusion. Compared with direct multi-evidence fusion, the calculation logic is clearer and the control of the conflict coefficient is more precise.
[0138] The core formula of Dempster's combination rule is: for any two sources of evidence... and Its fusion result on the identification frame subset C The calculation formula is: In the formula C represents the operator of Dempster's combination rule, symbolizing the evidence fusion operation; C, X, and Y are all identification frames. A subset of, in conjunction with the binary identification framework of this method, includes , , (Complete Collection) and (empty set), but in actual fusion, we only need to focus on the effective combination of non-empty subsets; and Sources of evidence For its subset X, source of evidence The basic probability of its subset Y is assigned a value, which takes the range [0,1]. The sum of the products of all assignments satisfying "the intersection of subset X and subset Y equals subset C" is used to aggregate information from two sources of evidence supporting the same proposition C; K is the conflict coefficient, calculated using the following formula: This represents the sum of the assigned products of two sources of evidence that support contradictory propositions (with an empty intersection). Its value ranges from [0,1) and must satisfy the following conditions: (like This indicates that the evidence is completely conflicting and cannot be reconciled, requiring a return to the risk assessment anomaly signal; the denominator is... This is a normalization factor used to ensure that the sum of the values assigned after fusion is still 1, which conforms to the probability axiom.
[0139] Integration in the first step For example, the specific calculation process is as follows: First, calculate the conflict coefficient. This formula corresponds to the combination of all subsets whose intersection is an empty set (i.e., and , and ); then calculate the fusion result for the proposition separately. and Assignment: (Only when) and At that time, the intersection is ), (Only when) and At that time, the intersection is Similarly, the second step is fusion. First calculate and Conflict coefficient Then, through the same logical calculations described above... and Finally, the joint basic probability assignment function is obtained. .
[0140] S53, Apply the joint basic probability assignment function to the proposition The reliability of the index is used as a comprehensive risk index.
[0141] Specifically, the joint basic probability assignment function Essentially, it is the comprehensive confidence level of the proposition "workers are in a high-risk state" after integrating three types of risk evidence: instability, slippage, and environmental hazards. Its value ranges from [0,1], and its magnitude is positively correlated with the comprehensive risk level. This extraction method directly corresponds to the original design intent of the identification framework, transforming the core results of multi-evidence fusion into an intuitive risk quantification indicator. It maintains the rigor of the DS evidence theory while ensuring the practicality of the comprehensive risk index, providing a clear basis for subsequent risk threshold comparisons and decision signal generation.
[0142] The aforementioned method for identifying safety risks at construction sites involves deploying a multimodal sensor network to synchronously collect data on personnel movement, location, and the environment. Through spatial registration and fusion, a dynamic digital twin is constructed that accurately reflects the personnel's full-body posture and the surrounding environment. From this twin, multidimensional risk features, such as kinematics, posture stability, and environmental interaction, representing signs of instability, are extracted in real time and input into an AI risk assessment model for predictive calculation. Ultimately, when a critical risk is identified, the model proactively predicts the fall trajectory and generates precise intervention commands. This represents a fundamental shift in the assessment of fall risks from post-event alarms to pre-event prediction and early warning, significantly improving the accuracy and timeliness of risk identification. It provides crucial decision-making support for activating active protection systems and effectively enhances the safety protection capabilities for high-altitude operations.
[0143] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0144] Based on the same inventive concept, this application also provides an apparatus for implementing the construction site safety risk identification method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more embodiments of the construction site safety risk identification apparatus provided below can be found in the limitations of the construction site safety risk identification method described above, and will not be repeated here.
[0145] In one exemplary embodiment, such as Figure 3 As shown, a construction site safety risk identification device 30 is provided to implement the methods in the above-described method embodiments. The device includes:
[0146] The multimodal perception data acquisition module 31 is used to simultaneously acquire multimodal perception data through fixed sensing nodes deployed in the construction site operation area and mobile sensing nodes worn by workers. The multimodal perception data includes environmental point cloud data acquired by fixed stereo vision cameras and lidar, inertial data acquired by inertial measurement units in mobile sensing nodes, positioning data acquired by ultra-wideband positioning tags, and wind speed and direction data acquired by environmental sensors.
[0147] The data registration module 32 is used to spatially register multimodal sensing data, generate a temporal dynamic motion state sequence for each worker and a fused environmental point cloud of the work area.
[0148] The dynamic feature analysis module 33 is used to calculate the kinematic characteristics and posture stability characteristics of the operator based on the time-series dynamic motion state sequence; and to calculate the environmental interaction characteristics based on the fusion of environmental point cloud and wind speed and direction data.
[0149] The risk assessment processing module 34 is used to input kinematic features, posture stability features and environmental interaction features into a pre-trained risk assessment model and output a comprehensive risk index.
[0150] The risk decision generation module 35 is used to compare the comprehensive risk index with the preset risk threshold to obtain the comparison result; and to generate a risk identification decision signal based on the comparison result.
[0151] The proactive intervention decision module 36 is used to predict the fall trajectory of a person based on a time-series dynamic motion state sequence when the risk identification decision signal indicates a critical risk, and to generate a proactive intervention instruction including the target person's identifier, suggested intervention type, and intervention parameters.
[0152] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.
[0153] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.
[0154] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0155] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A construction site safety risk identification method characterized by, The method comprises: S1, synchronously collecting multi-modal perception data by fixed sensing nodes deployed in a construction site operation area and mobile sensing nodes worn by operation personnel; wherein the multi-modal perception data comprises environment point cloud data collected by a fixed stereo vision camera pair and a laser radar, inertial data collected by an inertial measurement unit in the mobile sensing nodes, positioning data collected by an ultra-wideband positioning tag, and wind speed and direction data collected by an environment sensor; S2, spatially registering the multi-modal perception data to generate a time-sequenced dynamic motion state sequence of each operation personnel and a fused environment point cloud of the operation area; S3, calculating kinematic features and posture stability features of the operation personnel based on the time-sequenced dynamic motion state sequence, and calculating environment interaction features based on the fused environment point cloud and the wind speed and direction data; S4, inputting the kinematic features, the posture stability features and the environment interaction features into a pre-trained risk assessment model to output a comprehensive risk index; S5, comparing the comprehensive risk index with a preset risk threshold to obtain a comparison result, and generating a risk identification decision signal according to the comparison result; S6, when the risk identification decision signal indicates a critical risk, predicting a falling trajectory of the personnel based on the time-sequenced dynamic motion state sequence, and generating an active intervention instruction including a target personnel identifier, a recommended intervention type and intervention parameters.
2. The method of claim 1, wherein, The spatial registration of the multi-modal perception data to generate a time-sequenced dynamic motion state sequence of each operation personnel and a fused environment point cloud of the operation area comprises: S11, obtaining an initial transformation matrix of each fixed sensor from its own coordinate system to a global coordinate system of the construction site according to pre-measured fixed sensor installation parameters; S12, converting stereo vision point cloud and laser radar point cloud in the environment point cloud data to the global coordinate system of the construction site based on the initial transformation matrix; performing an iterative closest point registration algorithm on the stereo vision point cloud and the laser radar point cloud converted to the global coordinate system of the construction site in the spatial overlapping area to obtain an accurate transformation matrix of each fixed sensor; S13, performing fusion processing on the environment point cloud data collected by all fixed sensors based on the accurate transformation matrix to generate the fused environment point cloud; S14, performing Euclidean clustering on the fused environment point cloud to obtain independent human body point cloud clusters corresponding to different operation personnel, and assigning a personnel identifier to each independent human body point cloud cluster; inputting each independent human body point cloud cluster into a pre-trained human body posture estimation neural network to output visual joint coordinates of all joints of the corresponding personnel identifier in the global coordinate system of the construction site; S15, inputting the positioning data corresponding to each personnel identifier and the position data calculated based on the center of the independent human body point cloud cluster into an extended Kalman filter for fusion filtering processing to obtain a fused global position and global speed; inputting the inertial data corresponding to each personnel identifier into a complementary filter for posture solving processing to obtain a torso attitude quaternion; S16, construct a nonlinear graph optimization model with the global position, the global speed, the torso posture quaternion, and the visual joint coordinates as constraints; solve the nonlinear graph optimization model to output the time-sequenced dynamic motion state sequence.
3. The method of claim 2, wherein, The kinematic feature and the posture stability feature of the work personnel are calculated based on the time-sequenced dynamic motion state sequence, including: S21, calculate a motion mutation index as the kinematic feature according to the linear acceleration vectors of the current time and the previous time in the time-sequenced dynamic motion state sequence; the calculation formula of the motion mutation index is: wherein, denotes the current time instant, denotes the motion jerk index at the time instant, and denotes the linear acceleration vector at the time instant, and denotes the linear acceleration vector at the time instant, is the sampling time interval of the time-sequential dynamic motion state sequence; S22, calculate a three-dimensional coordinate of the overall center of mass of the human body based on the joint coordinates in the time-sequenced dynamic motion state sequence and the standard human body link quality parameters; obtain a center of mass projection point by vertically projecting the three-dimensional coordinate to a local ground; and determine a support polygon based on the double-ankle joint coordinates in the joint coordinates; S23, calculate the shortest distance from the center of mass projection point to the boundary of the support polygon as a first posture stability feature; S24, calculate a torso tilt angle as a second posture stability feature based on the torso posture quaternion in the time-sequenced dynamic motion state sequence; the calculation formula of the torso tilt angle is: wherein, represents the trunk inclination angle at the time instant, represents the trunk orientation quaternion in the time-sequential dynamic motion state sequence at the time instant; wherein is the real part, encoding the half cosine value of the rotation angle; , , are the imaginary parts, respectively encoding the components of the rotation axis in the global coordinate system X, Y, Z axes directions. S25, calculate a sample entropy of the torso tilt angle sequence as a third posture stability feature within a set time window.
4. The method according to any one of claims 1 to 3, characterized in that, The environmental interaction feature is calculated based on the fused environmental point cloud and the wind speed and direction data, including: S31, extract local ground point clouds under the feet of each work personnel from the fused environmental point cloud; calculate the normal direction variance and height variance of all points in the local ground point cloud; S32, obtain a ground friction coefficient level by querying a pre-defined mapping table according to the normal direction variance, the height variance, and environmental humidity data; S33, calculate a wind load influence estimate value according to the wind speed and direction data and the personnel windward area and posture estimated from the time-sequenced dynamic motion state sequence; S34, perform target detection on the fused environmental point cloud and the synchronously collected RGB image, obtain an operation disturbance feature by identifying the state of the handheld object of the work personnel and analyzing the micro-motion spectrum of the corresponding point cloud; S35, take the ground friction coefficient level, the wind load influence estimate value, and the operation disturbance feature as the environmental interaction feature.
5. The method of claim 4, wherein, The kinematic feature, the posture stability feature, and the environmental interaction feature are input into a pre-trained risk assessment model to output a comprehensive risk index, including: S41, combine the kinematic characteristics and the posture stability characteristics into a first feature vector; input the first feature vector into a first risk assessment sub-model for processing to obtain an instability risk probability ; S42, obtain the footstep feature of the worker and the proximity of the worker to the obstacle, combine the feature, the proximity, and the ground friction coefficient level in the environment interaction feature into a second feature vector; input the second feature vector into a second risk assessment sub-model to obtain a slip and trip risk probability ; S43、based on the global position in the time-sequenced dynamic motion state sequence and the boundary of the preset high-risk area in the construction site BIM model, the distance of the worker to the nearest edge or hole is calculated as the edge distance; and it is judged whether the worker is located in the preset high-risk area, and a high-risk area mark is generated when the worker is located in the preset high-risk area; the edge distance, the high-risk area mark and the wind load influence estimation value in the environmental interaction feature are combined into a third feature vector, and the third feature vector is input into a third risk assessment sub-model for processing to obtain an environmental risk probability ; S44, determining a context category of a current work scene based on the work personnel motion pattern in the time-sequenced dynamic motion state sequence and a pre-defined scene classification rule; assigning weights to the instability risk probability , the slip and fall risk probability and the environment-induced risk probability , and fusing basic probability assignments of each evidence source mapped based on the weights by using a D-S evidence theory to obtain the comprehensive risk index; wherein the context category includes stable walking, stationary high-altitude work, edge moving or moving on a narrow beam.
6. The method of claim 5, wherein, The basic probability assignments of each evidence source based on the weight mapping are fused by using the D-S evidence theory to obtain the comprehensive risk index, including: S51, mapping the weights of the destabilization risk probability , the slip and fall risk probability and the environmental risk probability to the basic probability assignment functions of the corresponding evidence sources respectively , and , and setting the recognition framework as , where represents the high-risk proposition, represents the low-risk proposition; S52, the basic probability assignment function , and are combined in pairs, and iteratively fused according to the Dempster combination rule to obtain a joint basic probability assignment function ; wherein, for two evidence sources and , the calculation method of the fusion result of the proposition is: wherein, represents a piece of evidence , represents the basic probability assignment of a proposition after fusion, represents a Dempster combination rule operator, , and are subsets of a discernment frame , and represent the basic probability assignment of a proposition by evidence and the basic probability assignment of a proposition by evidence , respectively; a conflict coefficient , represents the sum of the products of the basic probability assignments of all combinations of propositions satisfying the intersection of X and Y being empty, and ; the iterative fusion refers to an iterative process from calculating to calculating . S53, the joint basic probability assignment function to the plausibility of the proposition as the combined risk index.
7. A construction site safety risk identification apparatus for implementing the method of any one of claims 1 to 6, characterized by, The device includes: A multi-modal perception data collection module is configured to synchronously collect multi-modal perception data through fixed sensor nodes deployed in a construction site operation area and mobile sensor nodes worn by workers; wherein the multi-modal perception data includes environment point cloud data collected by a fixed stereo vision camera pair and a laser radar, inertial data collected by an inertial measurement unit in the mobile sensor nodes, positioning data collected by an ultra-wideband positioning tag, and wind speed and direction data collected by an environment sensor; A data registration module is configured to perform spatial registration on the multi-modal perception data to generate a time-sequenced dynamic motion state sequence of each worker and a fused environment point cloud of the operation area; A dynamic feature analysis module is configured to calculate kinematic features and posture stability features of the workers based on the time-sequenced dynamic motion state sequence, and to calculate environment interaction features based on the fused environment point cloud and the wind speed and direction data; A risk assessment processing module is configured to input the kinematic features, the posture stability features, and the environment interaction features into a pre-trained risk assessment model to output a comprehensive risk index; A risk decision generation module is configured to compare the comprehensive risk index with a preset risk threshold to obtain a comparison result, and to generate a risk identification decision signal according to the comparison result; An active intervention decision module is configured to predict a falling trajectory of a worker based on the time-sequenced dynamic motion state sequence when the risk identification decision signal indicates a critical risk, and to generate an active intervention instruction including a target worker identification, a recommended intervention type, and intervention parameters.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-6.