A construction site safety early warning system based on UWB and AI cameras

The construction site safety early warning system, which combines UWB with AI cameras, achieves high-precision three-dimensional positioning, cross-modal identity binding, and dynamic hazardous area management at tunnel construction sites. It solves the problems of large positioning errors, difficulty in identity association, and lagging boundary updates in existing technologies, and provides accurate multi-level early warning results.

CN122135477APending Publication Date: 2026-06-02GUANGDONG TONGCHUANG INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG TONGCHUANG INTELLIGENT TECH CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, UWB positioning at tunnel construction sites suffers from large non-line-of-sight errors, difficulty in associating visual behavior recognition with worker identity, lag in updating hazardous area boundaries, and insufficient accuracy in multimodal risk assessment, resulting in inadequate timeliness of site safety management.

Method used

The construction site safety early warning system, which combines UWB and AI cameras, acquires multi-source sensor data through a data acquisition module, identifies non-line-of-sight states using time-frequency domain feature extraction and deep residual networks, and performs 3D positioning by combining trajectory estimation. It achieves cross-modal identity binding through a behavior binding module, dynamically updates the boundaries of dangerous areas using instance segmentation and point cloud reconstruction, and distributes early warnings using a multi-level risk assessment model.

Benefits of technology

It improved the accuracy of 3D positioning of workers, solved the problem of difficulty in associating worker identities, realized the dynamic updating of dangerous area boundaries, and provided accurate and timely multi-level early warning results, providing an effective decision-making basis for construction site safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135477A_ABST
    Figure CN122135477A_ABST
Patent Text Reader

Abstract

This invention relates to the field of construction safety monitoring technology, and discloses a construction site safety early warning system based on UWB and AI cameras. Through multi-source sensor calibration and synchronization, it collects UWB ranging and video frames, and outputs aligned multi-source data. Based on time-frequency domain features and residual networks, it identifies non-line-of-sight locations and compensates for deviations, combining adaptive unscented Kalman filtering to output worker 3D coordinates. It detects targets and behaviors in video frames, achieving cross-modal binding through trajectory constraints and depth metric learning. It analyzes the BIM model to obtain hazardous zone boundaries, using instance segmentation and point cloud reconstruction for correction, and sets dynamic warning zones based on location covariance. It outputs tiered warnings through three-level risk assessment and attention fusion. Edge nodes trigger alarms locally, and edge-cloud collaboratively synchronizes data to the cloud, outputting records and reports. This invention can achieve multi-modal fusion perception and tiered early warning of worker location, behavior, and hazardous zones in complex environments such as tunnel construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of construction safety monitoring technology, and more specifically, to a construction site safety early warning system based on UWB and AI cameras. Background Technology

[0002] Tunnel construction sites are high-risk working environments, with processes such as tunnel boring machine advancement, segment hoisting, and high-pressure hydraulic line operation occurring simultaneously. The boundaries of hazardous areas are constantly and dynamically changing with the progress of construction. Workers working intensively in narrow tunnels are highly susceptible to accidents due to accidentally entering hazardous areas, fatigue, or sudden falls. Traditional construction site safety management relies mainly on manual patrols and fixed fences, which cannot monitor worker locations and behaviors in real time, nor can they dynamically update hazardous area boundaries according to construction progress, resulting in a significant lack of timeliness in safety management.

[0003] In existing technologies, some construction sites use a single UWB positioning system to track worker locations. However, the dense metal equipment and complex multipath propagation within tunnels, coupled with non-line-of-sight propagation leading to significant ranging deviations, make it difficult to meet the positioning accuracy requirements for safety early warning. Other solutions employ video surveillance to identify worker behavior, but video systems cannot directly acquire workers' three-dimensional coordinates, and identity association is difficult in scenarios with densely packed and overlapping workers, making it hard to reliably link behavior recognition results to specific worker identities. None of the aforementioned single-modal solutions can simultaneously solve the problem of coordinating accurate positioning and behavior recognition.

[0004] Furthermore, existing hazardous area management schemes typically rely on manual input of static boundaries from BIM models, leading to update lags and input errors, and failing to reflect the dynamic changes in the actual positions of construction machinery. At the risk assessment level, existing methods often employ simple rule-based judgments or linear weighting, which cannot effectively capture the non-linear interactions between multi-dimensional features such as location, behavior, and historical violations, resulting in insufficient early warning accuracy. Therefore, there is an urgent need for a construction site safety early warning system that can integrate UWB positioning with AI visual perception, dynamically update hazardous area boundaries, and achieve precise hierarchical early warnings through multimodal deep fusion. Summary of the Invention

[0005] This invention provides a construction site safety early warning system based on UWB and AI cameras, which solves the technical problems in related technologies such as large non-line-of-sight errors in UWB positioning, difficulty in associating visual behavior recognition with worker identity, lag in updating dangerous area boundaries, and insufficient accuracy of multimodal risk assessment.

[0006] This invention provides a construction site safety early warning system based on UWB and AI cameras, comprising: The data acquisition module collects UWB tag ranging data from workers and camera video frame data to obtain a multi-source sensor dataset. The positioning and calculation module uses time-frequency domain feature extraction and depth residual network to identify non-line-of-sight states from UWB tag ranging data, and combines trajectory estimation to obtain the worker's three-dimensional coordinate sequence. The behavior binding module performs target detection and behavior recognition on video frames. It combines the worker's 3D coordinates with trajectory consistency constraints and metric learning to achieve cross-modal identity binding and outputs a behavior label sequence. The fence update module obtains the initial boundary of the danger zone, uses video frames for instance segmentation and point cloud reconstruction for verification and correction, combines worker position covariance statistics to set a buffer warning zone, and outputs a dynamic set of danger zone boundaries. The fusion assessment module inputs the worker's three-dimensional coordinates, behavioral labels, and dynamic hazardous area boundaries into the fusion assessment model, and outputs multi-level early warning results through a three-level risk assessment. The early warning distribution module receives early warning results from edge nodes, executes local responses to trigger alarms on the worker side, synchronizes data to the cloud, and outputs alarm records and situation reports.

[0007] In a preferred embodiment, the data acquisition module includes: A unified three-dimensional coordinate system for the construction site is established based on the site measurement benchmark. A group of UWB base stations is deployed at preset intervals along the longitudinal direction in the tunnel. Each group of base stations has no less than 4 base stations, and the spatial distribution of a single group of base stations on the cross section is guaranteed to be non-coplanar. The three-dimensional installation coordinates of each UWB base station are obtained by actual measurement with a total station and entered into the base station coordinate database. The internal parameters of the multiple AI cameras deployed along the tunnel were calibrated using the calibration board method, and the external parameters of each camera were calibrated to the unified three-dimensional coordinate system of the construction site through a joint checkerboard calibration field. Using the system clock of the edge computing node as the master clock reference, a precise time protocol synchronization signal is broadcast to all UWB base stations and AI cameras. The edge computing node aligns the UWB ranging data set and video frame set at the same moment with a preset time window to form a multi-source sensing dataset. Each worker's UWB tag ID is bound to the worker's identity information and entered into the worker information database. Each UWB base station collects bilateral and bidirectional ranging data in a cyclical manner at a preset ranging polling frequency. Each ranging sampling outputs the estimated flight time from the corresponding tag to the base station and the original CIR sampling data.

[0008] In a preferred embodiment, the positioning calculation module includes: Based on the original CIR sampling data of each UWB ranging link in the multi-source sensing dataset, joint time-frequency domain feature extraction is performed. After normalizing the CIR sampling sequence in the time domain, the peak amplitude of the first path, root mean square delay spread, peak-to-average ratio, and number of effective multipaths are extracted. In the frequency domain, a short-time Fourier transform is performed on the CIR sequence to extract the frequency domain energy concentration, main bandwidth, and spectral entropy. The time-domain feature vector and the frequency-domain feature vector are concatenated to form a joint time-frequency domain feature vector, which is then input into the NLOS state classifier constructed by the deep residual network. The deep residual network consists of an input layer, three residual blocks, a global average pooling layer, and an output layer. The output layer uses the Softmax activation function to output three probability distributions: LOS state, mild NLOS state, and severe NLOS state. After NLOS identification is completed, no ranging correction is performed on links in the LOS state. For links in the mild NLOS state, the flight time estimate is corrected using a mean deviation compensation method based on historical statistics. For links in the severe NLOS state, a nonlinear deviation regression model based on a multilayer perceptron is used for compensation. Each state corresponds to a preset weight coefficient, and the corrected flight time estimate and corresponding weight coefficient of each ranging link are output.

[0009] In a preferred embodiment, the positioning calculation module further includes: Based on the corrected flight time estimate and corresponding weight coefficients, a set of three-dimensional positioning equations based on arrival time is established. The weighted least squares linearized iterative solution method is adopted, with the weight coefficients as the weights of each equation, to solve the initial estimate of the three-dimensional coordinates of the worker tag. The extrapolated coordinates of the worker's trajectory at the previous moment are used as the initial point of iteration. When the residual exceeds the preset solution quality threshold, the positioning quality flag is marked as low confidence. Based on the initial 3D coordinate estimation sequence, an adaptive unscented Kalman filter is used for trajectory smoothing. The state vector contains the worker's 3D position coordinates and 3D velocity components. At each time step, the covariance of the innovation sequence is calculated and compared with the theoretical innovation covariance. If the difference exceeds the preset adaptive adjustment threshold, the process noise covariance matrix is ​​adaptively adjusted. Combined with the positioning quality flag, the measurement noise covariance is dynamically adjusted. The smoothed 3D coordinate sequence and 3D velocity vector estimation sequence of each worker are output as the worker's 3D coordinate sequence.

[0010] In a preferred embodiment, the behavior binding module includes: Based on video frames from a multi-source sensor dataset, a target detection model is used to perform worker target detection. The target detection model introduces a multi-scale feature fusion module of a bidirectional feature pyramid network into the neck network and adds a safety helmet color classification sub-task branch to the head detection part, outputting a set of worker bounding boxes and safety helmet color categories. Simultaneously, compliance checks are performed on the wearing status of safety helmets and safety vests. When the confidence level of the safety helmet test is lower than the preset compliance threshold, a compliance flag is marked.

[0011] In a preferred embodiment, the behavior binding module further includes: Based on the worker bounding box set, worker 3D coordinate sequence and camera intrinsic and extrinsic parameters, cross-modal worker identity binding is performed. The worker's 3D coordinates are projected onto the image plane using the camera projection matrix. The two-dimensional Euclidean distance between the worker's 3D coordinates and the center point coordinates of the visually detected bounding box is calculated to construct a preliminary geometric cost matrix. For each candidate matching pair, extract the UWB projection trajectory sequence and visual trajectory sequence within the past N consecutive frames, calculate the dynamic temporal warping distance as the trajectory consistency cost, and weight and fuse it with the preliminary geometric cost to obtain the comprehensive cost matrix. For each visually detected worker bounding box, extract the appearance image region, input it into a Siamese-based deep metric learning network to obtain the appearance feature vector, calculate the cosine similarity with the standard appearance feature vector in the worker information database, convert it into the appearance cost and the comprehensive cost matrix, and then fuse them to obtain the final multi-constraint fusion cost matrix. The Hungarian algorithm is used to perform optimal binary matching on the final multi-constraint fusion cost matrix. Matching pairs with costs lower than the preset binding cost threshold are taken as valid binding results, and a cross-modal identity association mapping table is output.

[0012] In a preferred embodiment, the behavior binding module further includes: Based on the cross-modal identity association mapping table, the corresponding bounding box region image sequence is extracted from the video frame for each bound worker. The sequence is then input into the pose estimation network to extract the coordinate sequence of key points of the worker's skeleton. A spatiotemporal graph structure is constructed. The spatial dimension defines the graph topology structure based on the skeleton connection relationship, and the temporal dimension constructs temporal edges based on the temporal connection of the same joint node in adjacent frames. The spatiotemporal graph is input into the ST-GCN network. The joint co-motion features in the spatial dimension are aggregated through graph convolution operations, and the temporal evolution of the action sequence is captured through temporal convolution operations. The probability distribution of the behavior category is output, and the category with the highest probability is taken as the behavior recognition result. The behavior label sequence contains the behavior category label and behavior confidence score of each worker.

[0013] In a preferred embodiment, the fence update module includes: Based on the construction progress update data pushed by the construction site management system, the set of currently active construction nodes in the BIM model is parsed, the corresponding set of dangerous area geometry is extracted and the coordinates are transformed to the unified three-dimensional coordinate system of the construction site. Dangerous areas are divided into three categories: prohibited dangerous areas, controlled operation areas and warning buffer zones. Mask R-CNN instance segmentation network is used to perform pixel-level segmentation of key construction machinery in video frames to obtain contour masks. Key point detection network is used to extract image coordinates of key mechanical components. Multi-view triangulation is used to reconstruct the three-dimensional spatial coordinate point cloud of key mechanical components. The reconstruction results are compared with the mechanical position of BIM analysis. When the deviation exceeds the preset verification trigger threshold, boundary correction is triggered. Based on the estimated position covariance matrix corresponding to the worker's three-dimensional coordinate sequence, a dynamic buffer warning zone is added to the outside of each danger zone. The width of the buffer warning zone is not less than a preset multiple of the current positioning standard deviation estimate, and the dynamic danger zone boundary set is output.

[0014] In a preferred embodiment, the fusion evaluation module includes: The first level is based on the worker's three-dimensional coordinate sequence and the dynamic dangerous area boundary set. It uses a three-dimensional polyhedron point inclusion test algorithm to determine the type of area where the worker is located. For workers who are located in the buffer warning zone and are approaching the dangerous area, the estimated time to reach the dangerous area is calculated. When the estimated time to reach the dangerous area is less than the preset TTB warning threshold, the location risk component is increased and the location risk component value is output. The second level is based on the behavior label sequence. Combined with the preset behavior risk level mapping table, the behavior category is mapped to the behavior risk component. The behavior confidence score is used as a correction coefficient to weight the behavior risk score and output the behavior risk component value. The third level inputs the location risk component, behavioral risk component, worker's historical violation weight, dwell time, job type, and fatigue level into the feature embedding subnetwork and maps them to a unified dimension embedding space. The interaction relationship between multimodal features is captured through a multi-head self-attention mechanism. Then, a gated recurrent unit is introduced to perform time-series modeling of the historical risk state sequence and output a normalized comprehensive risk score. This score is compared with the preset monitoring-level threshold, early warning-level threshold, and alarm-level threshold to output multi-level early warning results.

[0015] In a preferred embodiment, the early warning distribution module includes: Edge computing nodes maintain priority response queues. For alarm-level early warning events, vibration commands are sent to worker UWB tags via the UWB network's reverse downlink control channel, and voice broadcasts are triggered by local loudspeakers. For warning-level early warning events, mild vibration alerts are sent to worker UWB tags, and pop-up warning information is pushed to the handheld terminals of on-site management personnel. For monitoring-level early warning events, they are recorded in the local event database and highlighted on the local monitoring screen. Edge computing nodes upload event logs and worker location snapshots in batches to the cloud safety management platform via an encrypted transmission channel at a preset synchronization cycle. Alarm-level events trigger an immediate push mechanism. The cloud platform updates worker locations and warning level indicators in the 3D construction site model in real time and sends alarm notifications to the safety manager via instant messaging applications for alarm-level events. The cloud platform extracts historical data from the event database at a preset report generation cycle, summarizes and statistically analyzes the data by worker, hazardous area, and time dimensions, generates and archives structured safety situation reports, and outputs alarm records and situation reports.

[0016] The beneficial effects of this invention are as follows: By combining time-frequency domain joint feature extraction and deep residual networks, three-class classification recognition of non-line-of-sight states of UWB ranging links is performed. An adaptive bias compensation strategy is adopted for different degrees of non-line-of-sight, and trajectory smoothing is performed by combining adaptive unscented Kalman filtering, which effectively improves the 3D positioning accuracy of workers in dense metal tunnel environments. In cross-modal identity binding, the problem of reliable association between visual detection results and UWB worker identities in dense worker occlusion scenarios is solved by fusing geometric projection, temporal trajectory consistency constraints and deep metric learning constraints, so that behavior recognition results can be accurately attributed to specific workers. By combining BIM-driven construction progress with visual-assisted verification of instance segmentation and point cloud reconstruction, dynamic adaptive updates of hazardous area boundaries were achieved, solving the problem of lagging traditional static boundaries. At the risk assessment level, a multimodal deep fusion model based on attention mechanism was adopted. Through multi-head self-attention mechanism, the nonlinear interaction relationship of multi-dimensional features such as location, behavior, and historical violations was captured. Gated cyclic units were introduced to model the temporal evolution of risk states, outputting interpretable hierarchical early warning results, providing accurate and timely decision-making basis for construction site safety management. Attached Figure Description

[0017] Figure 1 This is a block diagram of a construction site safety early warning system based on UWB and AI cameras according to the present invention; Figure 2 This is a detailed flowchart of a construction site safety early warning system based on UWB and AI cameras according to the present invention. Detailed Implementation

[0018] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0019] At least one embodiment of the present invention discloses a construction site safety early warning system based on UWB and AI cameras, such as... Figures 1 to 2 As shown, it includes the following steps: The data acquisition module collects UWB tag ranging data from workers and camera video frame data to obtain a multi-source sensor dataset. S11. Establishment of a unified 3D coordinate system and sensor spatial calibration at the construction site: Based on the site's as-built measurement benchmark, a unified 3D coordinate system is established. The origin of the coordinate system is set at a fixed measurement benchmark at the tunnel entrance. The X-axis extends longitudinally along the tunnel, the Y-axis extends laterally along the tunnel's width, and the Z-axis points vertically upwards, forming a right-handed coordinate system. A group of UWB base stations (anchor nodes) is deployed at regular intervals along the longitudinal direction within the tunnel. Each group has no fewer than four base stations, and the spatial distribution of each group of base stations on the cross-section is ensured to be non-coplanar, meeting the geometrical dilution of precision (GDOP) requirements for 3D coordinate calculation. The 3D installation coordinates of each UWB base station are obtained through on-site total station measurements and entered into the base station coordinate database. For multiple AI cameras deployed along the tunnel's extension direction (one camera deployed at regular intervals, with camera installation positions avoiding areas obstructed by other equipment), an intrinsic parameter calibration method was used for each camera (obtaining focal length, principal point coordinates, and distortion coefficients). Then, the extrinsic parameters (rotation matrix and translation vector) of each camera were calibrated to the unified 3D coordinate system of the construction site using a joint checkerboard calibration field, uniquely determining the transformation relationship between the camera coordinate system and the construction site coordinate system. During the calibration process, active light sources were added to the calibration board to supplement the low-light environment of the tunnel, ensuring the accuracy of corner point extraction. Existing geometric data in the BIM model were also aligned to the unified 3D coordinate system of the construction site through coordinate transformation, forming a complete set of spatial calibration parameters, including the 3D coordinates of all base stations, the intrinsic and extrinsic parameters of all cameras, and the BIM coordinate transformation matrix.

[0020] S12, Multi-source sensor data time synchronization protocol configuration; using the system clock of the edge computing node as the master clock reference, a Precision Time Protocol (PTP, IEEE 1588) synchronization signal is broadcast to all UWB base stations and AI camera data acquisition devices within the construction site's local area network. This controls the deviation between each device's clock and the master clock to the sub-microsecond level, meeting the stringent clock accuracy requirements of UWB ranging time difference positioning. After completing each ranging sampling, the UWB base station uploads the ranging message containing the local timestamp to the edge computing node via the construction site Ethernet switch; after completing each frame image acquisition, the AI ​​camera uploads the video frame (or hardware-encoded H.265 compressed video stream) containing the acquisition timestamp to the edge computing node. When receiving multi-source data, the edge computing node records the arrival timestamp and device reporting timestamp for each UWB ranging data and each video frame. The data is then time-aligned in the data buffer queue according to a unified time axis. The UWB ranging data set and video frame set at the same moment are aligned with a set time window (the specific window size is determined according to the target sampling frequency and motion speed characteristics) to form a timestamp-aligned multi-source sensor dataset for use in subsequent steps.

[0021] S13, Worker UWB Tag Information Registration and Data Acquisition Initialization: Each worker entering the site must complete the registration of their UWB tag (blind node) before entering, binding the tag ID with the worker's identity information (name, job type, and work group) and entering it into the worker information database. The UWB tag is worn in the worker's safety helmet's internal compartment (without affecting normal wearing) or the vest's chest pocket, ensuring the tag's spatial stability in the worker's head and chest area and reducing the impact of human body obstruction on the antenna radiation pattern. After the system is powered on, each UWB base station establishes a ranging link with the worker tags within its effective range in sequence, and cyclically collects two-way ranging (TWR, Two-Way Ranging) data at a set ranging polling frequency (the ranging frequency is determined based on the worker's walking speed characteristics and positioning accuracy requirements). Each ranging sample outputs the estimated flight time from the tag to the base station and the original CIR sampling data. The above output, together with the timestamp information configured in step S12, constitutes the timestamp-aligned multi-source sensor dataset of the final output of step S1, which includes: the original UWB ranging dataset (including tag ID, base station ID, timestamp, time of flight estimation, CIR data) and the timestamp-aligned video frame set of each AI camera.

[0022] In some embodiments, due to the large number of workers and high density of UWB tags during peak tunnel construction periods, the polling-based TWR protocol experiences a decrease in single-tag ranging update rate and an increase in positioning latency after the number of tags exceeds a certain scale. A ranging protocol combining Time Division Multiple Access (TDMA) and priority scheduling can be adopted to assign higher polling priority to worker tags closer to the current high-risk work area, so as to ensure high-frequency ranging updates for workers in key areas under resource-constrained conditions. The aim is to maintain the real-time positioning of workers in high-risk areas under large-scale tag concurrency scenarios. Specifically, the edge computing node dynamically delineates the priority tag set (i.e., tags currently in the danger zone or warning buffer zone) based on the danger zone boundary data updated in real time in step S4, and issues the current polling scheduling strategy to each base station through the UWB network controller, so that the priority tags can obtain effective ranging link updates from at least 4 base stations in each ranging cycle, ensuring the continuity of 3D positioning calculation.

[0023] The positioning and calculation module uses time-frequency domain feature extraction and depth residual network to identify non-line-of-sight states from UWB tag ranging data, and combines trajectory estimation to obtain the worker's three-dimensional coordinate sequence. S21, NLOS recognition and adaptive ranging deviation compensation based on joint time-frequency domain feature extraction and deep residual network; In view of the characteristics of dense metal equipment, complex multipath propagation and dynamic changes of channel characteristics with the construction progress in tunnel construction scenarios, this step adopts an NLOS (Non-Line-of-Sight) recognition method based on joint time-frequency domain feature extraction and deep residual network. Based on the raw CIR (Channel Impulse Response) sampling data of each UWB ranging link output by S1, joint time-frequency domain feature extraction is first performed: In the time domain, after normalizing the CIR sampling sequence, traditional statistical features such as first-path peak amplitude, root mean square delay spread (RMS Delay Spread), peak-to-average ratio (PAR), and effective multipath number are extracted; in the frequency domain, a short-time Fourier transform (STFT) is performed on the CIR sequence to obtain the power spectral density distribution of the channel frequency response, and frequency domain features such as frequency domain energy concentration, main bandwidth, and spectral entropy are extracted. Frequency domain features can effectively reflect the frequency-selective fading characteristics caused by metallic reflectors within the tunnel and have a stronger ability to distinguish NLOS states. The time-domain feature vector and the frequency-domain feature vector are concatenated to form a joint time-frequency domain feature vector, which is used as the input for subsequent deep networks.

[0024] To address the diverse NLOS propagation patterns in tunnel construction environments (including reflections from the tunnel boring machine's metal shell, occlusion by the rebar cage, and diffraction by concrete segments), a deep residual network (ResNet) is employed to construct an NLOS state classifier. The ResNet network structure comprises an input layer, three residual blocks (each containing two fully connected layers and skip connections), a global average pooling layer, and an output layer. The skip connections within the residual blocks effectively mitigate the vanishing gradient problem during deep network training, enabling the network to learn deeper levels of NLOS propagation pattern features. The network output layer uses a Softmax activation function, outputting three state probability distributions: LOS state, mild NLOS state (single reflection), and severe NLOS state (multiple reflections or complete occlusion). Compared to traditional binary classification methods, the three-class output provides a more refined characterization of the NLOS severity, offering a more accurate basis for subsequent ranging bias compensation.

[0025] The ResNet classifier is trained using supervised learning. Training data is obtained as follows: In the initial stage of system deployment, an optical total station is used to measure the reference ground truth coordinates of worker labels at multiple typical locations within the tunnel. Simultaneously, the raw CIR data of each ranging link at the corresponding time is recorded. The deviation between the UWB ranging results and the total station reference ground truth is used as the basis for generating NLOS state labels (deviation less than a set LOS threshold is labeled as LOS, deviation between the LOS threshold and the severe NLOS threshold is labeled as mild NLOS, and deviation exceeding the severe NLOS threshold is labeled as severe NLOS), forming the initial training dataset. The training process employs the cross-entropy loss function and the Adam optimizer, and data augmentation (adding Gaussian noise and time jitter to the CIR sequences) enhances the model's generalization ability.

[0026] After NLOS identification, adaptive ranging bias compensation is performed for different degrees of NLOS: For links with a LOS output from ResNet, no ranging correction is performed, and the weight coefficient is set to 1; for links with mild NLOS, a mean bias compensation method based on historical statistics is used, which uses the average ranging bias accumulated during the training phase in mild NLOS scenarios to correct the estimated flight time, and the weight coefficient is set to 0.7; for links with severe NLOS, due to the large uncertainty of ranging bias, a nonlinear bias regression model based on a deep neural network is used for compensation. This regression model takes the joint feature vector of the time and frequency domains as input and outputs the estimated ranging bias of the link. The regression model adopts a multi-layer perceptron (MLP) structure, and the training label is the deviation between the total station reference value and the original UWB ranging result, and the weight coefficient is set to 0.4. After adaptive compensation processing, the corrected flight time estimate and corresponding weight coefficient for each ranging link are output.

[0027] In some embodiments, due to the dynamic changes in the shape of the metal tail shell during the tunnel boring machine's advancement and the gradual increase in newly poured concrete segments, the statistical distribution of CIR channel features continuously drifts with the construction progress. This may cause a decrease in the recognition accuracy of the pre-trained ResNet classifier after long-term operation. An online incremental learning mechanism based on cross-modal visual positioning verification can be adopted to continuously collect new samples with high confidence during system operation and perform small-batch online updates to the ResNet classifier. The aim is to ensure that the NLOS classifier maintains stable recognition capabilities during the dynamic evolution of the construction environment. Specifically, the worker's position obtained through visual positioning in step S3 is used as a parameter. For verification, when the confidence level of the visual positioning result exceeds a set high confidence threshold and the deviation between the result and the UWB positioning result exceeds a set verification threshold, the CIR time-frequency domain joint feature vector at the corresponding time and the NLOS state label inferred from the deviation value are used as new training samples and stored in the online sample buffer pool. When the number of samples in the buffer pool reaches the set batch size, the incremental update process of the ResNet classifier is triggered, and the network parameters are fine-tuned with a small learning rate to avoid catastrophic forgetting. After the update is completed, the buffer pool is cleared and new samples continue to be accumulated, forming a continuously adaptive closed-loop learning mechanism that can maintain the classifier's adaptability to changes in the construction environment without manual intervention.

[0028] S22, Weighted 3D Coordinate Calculation: Based on the corrected time-of-flight estimates and corresponding weight coefficients of each worker tag to each available base station output from S21, a system of 3D positioning equations based on Time of Arrival (ToA) is established. Specifically, the coordinates of the i-th base station are taken as known quantities, and the 3D coordinates of the worker tag are taken as unknown quantities. The corrected distance estimate of the link is obtained by multiplying the speed of light by the corrected time-of-flight estimate, and a distance equation is constructed with each base station as the center and the corrected distance as the radius. Since the system of equations is a nonlinear overdetermined system, a weighted least squares (WLS) linearization iterative solution method is adopted. The weight coefficients output from S21 are used as the weights of each equation. Through Taylor expansion linearization and iterative convergence, the initial estimated values ​​of the 3D coordinates of the worker tag are solved. To ensure the rationality of the initial iteration point, the extrapolated coordinates of the worker's trajectory at the previous moment are used as the initial iteration point. After the solution is completed, the weighted root mean square value of the residual is calculated. When the residual exceeds the set solution quality threshold, the quality flag of this positioning result is marked as low confidence, so that the subsequent UKF steps can reduce the weight of the measurement value at this moment during fusion.

[0029] S23, Adaptive Unscented Kalman Filter (UKF) Trajectory Smoothing Estimation: Based on the initial estimation sequence of the worker's 3D coordinates output by S22, an adaptive UKF is used to perform temporal estimation and trajectory smoothing of the worker's motion state. The UKF state vector contains the worker's 3D position coordinates and 3D velocity components, totaling 6 dimensions. The process noise covariance matrix is ​​initialized with a priori values ​​set according to the worker's walking acceleration characteristics, and the measurement noise covariance matrix is ​​initialized with a priori values ​​set according to the statistical results of the solution accuracy in S22. The specific implementation of the adaptive mechanism is as follows: At each time step, the innovation sequence (i.e., the difference between the measured value and the state prediction value) is calculated. The covariance of the innovation sequence is statistically estimated using a sliding window. The statistical estimation result is compared with the theoretical innovation covariance of the UKF. If the difference exceeds the adaptive adjustment threshold, the process noise covariance matrix is ​​adaptively amplified or reduced according to the direction of the difference. This allows the filter to automatically switch gain between two states: rapid change of direction movement (higher process noise) and stable uniform movement (lower process noise), maintaining the dynamic tracking accuracy of the filter. The UKF prediction step uses a uniformly accelerated kinematics model to update the state over time. The update step uses the worker's 3D coordinates output from step S22 as the measurement input, and dynamically adjusts the measurement noise covariance based on the positioning quality flags from step S22 (amplifying the measurement noise covariance at low confidence moments to reduce the impact of the measured value on the filtering result at that moment). After UKF processing, a smoothed 3D coordinate sequence and a 3D velocity vector estimation sequence for each worker are output as the final output of step S2.

[0030] The behavior binding module performs target detection and behavior recognition on video frames. It combines the worker's 3D coordinates with trajectory consistency constraints and metric learning to achieve cross-modal identity binding and outputs a behavior label sequence. S31 improves YOLOv8 worker target detection and safety helmet color classification; based on the timestamps output from S1, video frames are aligned, and improved YOLOv8 target detection inference is performed on the current frame image of each AI camera. The specific structure of the improved YOLOv8 model is as follows: On the basis of the original YOLOv8 backbone network, a multi-scale feature fusion module based on Bidirectional Feature Pyramid Network (BiFPN) is introduced into the Neck network to enhance the feature representation ability of small targets (far-away workers) and occluded targets (workers partially obscured by equipment); a safety helmet color classification sub-task branch is added to the detection head part, which takes the pixel features of the safety helmet area within the worker's bounding box as input and outputs the safety helmet color category (preset color categories according to the construction site team code, such as red, yellow, blue, white, etc. to represent different types of work teams). The color classification branch and the target detection branch share the backbone feature extraction layer to realize multi-task joint inference. For low-light tunnel scenarios, during model training, the training data is augmented with low-light enhancement data (including brightness reduction, noise addition, and Gamma transformation) to specifically improve the model's robustness to low-light conditions. After detection and inference are completed, the set of bounding boxes for all detected workers in the current video frame is output (including bounding box coordinates, detection confidence, helmet color category, and its confidence). Simultaneously, compliance checks are performed on helmet wearing status (workers detected but no helmet areas detected, or helmet detection confidence below a set compliance threshold) and safety vest wearing status, and compliance flags are appended to the corresponding worker's detection result.

[0031] S32, Cross-modal worker identity binding based on temporal trajectory consistency constraints and deep metric learning; addressing the challenges of dense worker populations, frequent cross-occlusion, and spatiotemporal discrepancies between UWB localization and visual detection in tunnel construction scenarios, this step employs a cross-modal worker identity binding method based on temporal trajectory consistency constraints and deep metric learning. Based on the worker detection bounding box set output from S31, the worker 3D coordinate sequence output from S2, and the camera intrinsic and extrinsic parameters calibrated in S1, cross-modal worker identity binding is performed.

[0032] Perform preliminary geometric projection matching: For all UWB-tracked workers within the current camera's field of view (determined by their 3D coordinates to determine if they are within the camera's effective field of view cone), use the calibrated camera projection matrix (obtained by multiplying the intrinsic and extrinsic parameter matrices) to project the 3D coordinates of each worker from the construction site coordinate system onto the camera image plane, obtaining the coordinates of the image projection center point of each UWB-tracked worker. Calculate the 2D Euclidean distance between the center point coordinates of the visual detection bounding box output by S31 and the coordinates of the UWB projection center point to construct the preliminary geometric cost matrix.

[0033] To enhance binding robustness, temporal trajectory consistency constraints are introduced: Addressing the issue of mismatches in single-frame geometric matching within densely populated worker scenarios, temporal consistency verification is performed using worker motion trajectory information from multiple consecutive frames. Specifically, for each candidate matching pair (UWB worker ID and visual detection box), the UWB 3D trajectory sequence and the image plane trajectory sequence of the visual detection box center point are extracted from the worker's trajectory over the past N consecutive frames (N is determined based on the worker's movement speed and frame rate, typically 5 to 10 frames). The UWB 3D trajectory sequence is then projected onto the image plane using a camera projection matrix to obtain the UWB projected trajectory sequence. The Dynamic Time Warping (DTW) distance between the UWB projected trajectory sequence and the visual trajectory sequence is calculated. The DTW distance effectively measures the similarity of the two trajectories in their temporal evolution patterns and is robust to local temporal scaling of the trajectories. The DTW distance is normalized and used as the trajectory consistency cost. It is then weighted and fused with the preliminary geometric cost to obtain the comprehensive cost matrix. The weight coefficients are adaptively adjusted according to the worker density of the current scene (the higher the worker density, the greater the weight of the trajectory consistency cost, so as to enhance the ability to suppress mismatches in dense scenes).

[0034] To address visual appearance features such as helmet color and worker posture, an appearance similarity constraint based on deep metric learning is introduced. Specifically, for each visually detected worker bounding box, the worker's appearance image region within the bounding box is extracted and input into a pre-trained deep metric learning network (using a Siamese Network structure). This network, pre-trained on a construction site worker appearance dataset, maps worker appearance images to a high-dimensional feature embedding space, ensuring that the appearance features of the same worker in different frames and poses are close in the embedding space, while the appearance features of different workers are far apart. For each UWB worker ID, the standard appearance feature vector collected when the worker enters the site (extracted via the Siamese Network) is stored in the worker information database. The cosine similarity between the appearance feature vector of the current visual detection box and the appearance feature vectors corresponding to each UWB worker ID in the database is calculated. This similarity is converted into an appearance cost (higher similarity means lower cost), and then weighted and fused with the comprehensive cost matrix to obtain the final multi-constraint fusion cost matrix.

[0035] The Hungarian Algorithm is used to perform optimal binary matching on the final multi-constraint fusion cost matrix. Matches with the lowest cost and below a set binding cost threshold are considered valid binding results, and the visually detected bounding boxes are associated with their corresponding UWB worker IDs. For workers whose binding cost exceeds the set threshold (due to camera field of view obstruction or UWB positioning deviation), their binding status is marked as low confidence, and they are given reduced weight in subsequent risk assessments of behavior recognition results. After binding is complete, a cross-modal identity association mapping table is output, with the UWB worker ID as the key. Each record contains the UWB worker ID, the coordinates of the detected bounding box in the current video frame, the helmet color, a compliance flag, the binding confidence, and an appearance feature vector.

[0036] In some embodiments, due to the presence of deep blind spots in certain areas of the tunnel that cannot be covered by the field of view of multiple cameras (such as narrow passages on both sides of the tunnel boring machine), workers located in the blind spots cannot be effectively detected by the cameras, causing S32 binding to fail. A behavior state-assisted inference mechanism based on UWB trajectory motion features and generative adversarial networks can be adopted. That is, when a worker is continuously located in the visual blind spot within a continuous time window and only UWB positioning data is available, the three-dimensional velocity vector sequence and position coordinate sequence of the worker output by S2 are first used to extract temporal motion features such as velocity amplitude statistics, acceleration change, motion direction deflection angle, trajectory curvature, and dwell point density through feature engineering. These features are then input into a pre-trained bidirectional long short-term memory network (Bi-LSTM) to encode the worker's motion pattern. Bi-LSTM can simultaneously capture the forward and backward temporal dependencies of the motion sequence and output a motion pattern embedding vector. For the missing visual behavior information in the blind spot, a conditional generative adversarial network (CGAN) is used. The CGAN network takes motion pattern embedding vectors as conditional input and generates a probability distribution of possible behavioral categories of the worker within blind spots. The generator is trained on historical data in visually covered areas to learn the cross-modal mapping between UWB motion features and visual behavioral categories. A discriminator distinguishes the generated behavioral distribution from the real behavioral distribution. Adversarial training enables the generator to output behavioral inferences that conform to the real distribution. The goal is to maintain the ability to perceive dangerous worker behaviors even within visual blind spots. Specifically, the CGAN-generated behavioral category probability distribution includes categories for falls (sudden change in speed to near zero and a sudden drop in the Z-axis position) and fatigue-induced stillness (prolonged period of inactivity). The probability estimates of typical dangerous behaviors (such as near-zero amplitude) are compared with the set behavior judgment thresholds. Results with confidence levels exceeding the thresholds proceed to the subsequent risk assessment stage, while results with confidence levels below the thresholds are marked as blind zone uncertain states. During risk assessment, a conservative early warning strategy is triggered (i.e., maintaining a continuous level of attention to workers within the blind zone). Simultaneously, to improve the inference accuracy of CGAN in blind zone scenarios, when a worker moves from the blind zone to an area with visual coverage, the system automatically associates the worker's UWB motion features within the blind zone with the visual behavior labels actually observed after the worker moves out of the blind zone. This association serves as a new training sample for online incremental updates to CGAN, forming a continuously self-improving closed-loop learning mechanism.

[0037] S33, Multi-Frame Temporal ST-GCN Worker Behavior Recognition: Based on the cross-modal identity association mapping table output by S32, for each successfully bounded worker, the corresponding bounding box region image sequence in the video frame is extracted. This sequence is then input into a pose estimation network (HRNet, High-Resolution Net) to extract the worker's skeletal keypoint coordinate sequence (keypoints include 17 keypoints: top of head, neck, shoulders, elbows, wrists, hips, knees, and ankles), forming a two-dimensional skeletal sequence data for each worker. The skeletal keypoint coordinate sequence is constructed into a spatiotemporal graph structure: spatially, the graph topology between joint nodes is defined by skeletal connections; temporally, temporal edges are constructed by temporal connections between the same joint nodes in adjacent frames. The constructed spatiotemporal graph is input into the ST-GCN network for behavior recognition inference. ST-GCN aggregates the co-motion features between joints in the spatial dimension through graph convolution operations and captures the temporal evolution of actions in the temporal dimension through temporal convolution operations. Finally, it outputs the probability distribution of behavior categories within the worker's current time window, and takes the category with the highest probability as the current behavior recognition result. The behavior category set includes: normal standing, normal walking, normal head-down operation, fatigue and stillness (long-term inactivity), falling, rapid climbing, running, and waving arms (distress signal), etc. The behavior recognition results are merged with the cross-modal identity association mapping table output by S32, and the behavior category label, behavior confidence score, and skeletal key point sequence of each worker are output, forming the final output result of step S3.

[0038] The fence update module obtains the initial boundary of the danger zone, uses video frames for instance segmentation and point cloud reconstruction for verification and correction, combines worker position covariance statistics to set a buffer warning zone, and outputs a dynamic set of danger zone boundaries. S41, Automatic Resolution of Hazardous Area Boundaries Driven by BIM Construction Progress; Based on construction progress updates pushed by the site management system at a set update frequency (the update frequency is determined according to the construction progress speed, generally in hours or shifts), the system dynamically resolves the current construction status of the site's BIM model. Construction progress information in the BIM model is stored in the form of a construction activity node tree. Each activity node corresponds to a specific construction procedure (e.g., advancing the tunnel boring machine to XXX meters, completing the Nth layer of formwork support, etc.), and is associated with the geometry of the area in a hazardous state in 3D space during that procedure (represented by polyhedra or polygonal prisms). The system resolves the set of currently active construction nodes, extracts the corresponding set of hazardous area geometry, transforms its coordinates to the unified 3D coordinate system established in S1, and outputs the set of polygonal boundaries of the hazardous area at the current moment. Hazardous areas are divided into three categories: the first category is a prohibited hazardous area (machine rotation coverage area, high-pressure pipeline operation area), where workers are prohibited from entering; the second category is a controlled operation area (only specific types of workers are allowed to enter, requiring continuous monitoring); and the third category is a warning buffer zone (a buffer zone outside the hazardous area, used to trigger early warnings). Various hazardous areas have preset attribute labels in the BIM model, which are extracted together during parsing.

[0039] S42, Visual Perception-Assisted Verification and Correction of Hazardous Area Boundaries Based on Instance Segmentation and Point Cloud Reconstruction: Addressing the issues of time lag in BIM model construction progress updates and discrepancies between hazardous area boundaries and the actual site conditions due to manual input errors, this step employs a visual perception-assisted verification method based on instance segmentation and point cloud reconstruction. This method, building upon the BIM-analyzed boundaries output in S41, utilizes AI camera images to accurately perceive and verify the actual spatial positions of key construction machinery.

[0040] Instance segmentation and key point detection of critical construction machinery: Utilizing the multi-class extension capability of the improved YOLOv8 target detection model in S31, target detection is performed on critical construction machinery (such as the outline of a tunnel boring machine, the end of a tower crane boom, the bucket of an excavator, and the boom of a concrete pump truck) in video frames to obtain the bounding box coordinates of these machines in the image. Based on the detection, a Mask R-CNN instance segmentation network is further used to perform pixel-level precise segmentation of the detected machinery targets, obtaining accurate contour masks of the machinery targets. Simultaneously, for machinery components with clearly defined hazardous operating areas (such as the center of the tunnel boring machine cutterhead, the position of the tower crane hook, and the tip of the excavator bucket teeth), a key point detection network is used to extract the image coordinates of these key components. The key point detection network uses a heatmap regression method to output the sub-pixel-level precise coordinates of each key point in the image plane.

[0041] 3D Point Cloud Reconstruction with Multi-View Geometric Constraints: Addressing the issue of overlapping views of the same mechanical target from multiple cameras within a tunnel, multi-view geometric constraints are used to reconstruct its 3D spatial location. Specifically, for mechanical targets simultaneously observed by multiple cameras (cross-camera target correspondence is achieved through appearance feature matching of instance segmentation masks and temporal trajectory association), the set of segmentation mask boundary points and key point coordinates of the mechanical target from each viewpoint are extracted. Using the intrinsic and extrinsic parameters of each camera calibrated by S1, the multi-view image planar coordinates are back-projected to the site's 3D coordinate system using triangulation to reconstruct the 3D spatial coordinate point cloud of the key mechanical components. For mechanical targets observed only by a single camera, combined with prior information on the known geometric dimensions of the machine (obtained from the BIM model or equipment database), the 3D position of the key mechanical components in the site coordinate system is estimated using homography transformation combined with scale constraints. After the 3D point cloud reconstruction is completed, outlier filtering and smoothing are performed on the point cloud to obtain a robust 3D position estimate of the key mechanical components.

[0042] Deviation detection and adaptive correction between BIM boundaries and visual perception results: The spatial distance between the 3D positions of key mechanical components reconstructed from visual perception and the corresponding mechanical positions from BIM analysis is compared, and a position deviation vector is calculated. When the magnitude of the deviation vector exceeds a set verification trigger threshold, it is determined that there is lag or error in the BIM boundary, triggering the boundary correction process. The correction strategy adopts an adaptive boundary adjustment method based on deep reinforcement learning: the current BIM boundary, visually perceived mechanical positions, and historical correction records are used as state inputs. A pre-trained Deep Q-Network (DQN) outputs boundary adjustment actions (including discrete action spaces such as boundary translation direction, translation distance, expansion or contraction ratio, etc.). The DQN is trained offline on historical verification data, and the reward function is designed as the matching accuracy between the corrected boundary and the actual hazardous area. After executing the boundary adjustment action, the corrected hazardous area boundary is output, and an abnormal event is reported to the site management system, prompting the BIM model to be manually verified and updated. For blind spots (camera blind spots) that cannot be covered by visual perception, the corresponding hazardous area boundary maintains the BIM analysis result and is not visually corrected.

[0043] In some embodiments, due to the impact of construction pollutants such as smoke and water mist in the tunnel on the image quality of the camera, the reliability of mechanical target detection may decrease during periods of severe pollution, leading to the failure of visual-assisted verification. An anomaly clustering detection method based on worker tag trajectories can be used as a supplementary verification method. That is, when the trajectories of a large number of worker tags show a group movement pattern of obvious outward avoidance near a certain area, it is inferred that there is an actual hazard source (mechanical operation or hazardous gas leak) in that area, triggering a conservative strategy of temporarily expanding the hazard boundary of that area. The aim is to still be able to make a reasonable conservative estimate of the boundary of the hazard area under the harsh environmental conditions where visual-assisted verification fails. Specifically, using the velocity direction vector of all worker 3D coordinates output by S2 within a continuous time window, the mean shift clustering algorithm is used to detect the deflection direction of the group movement trend. When the spatial statistical results of the deflection direction converge to a specific area, that area is marked as a temporary hazard expansion area, and the expansion range is added to the current hazard area boundary set.

[0044] S43, Dynamic Setting of Buffer Warning Zone Based on Positioning Error Statistics; Based on the worker position covariance matrix estimate output by the adaptive UKF in S2, statistical analysis is performed on the current positioning uncertainty in each direction of the unified coordinate system of the construction site. Each time the boundary of a hazardous area is updated, a dynamic buffer warning zone is added outside each hazardous area, based on the positioning uncertainty of each hazardous area boundary in different directions (measured by the estimated standard deviation of the position covariance matrix in the corresponding direction). The width of the buffer warning zone is set according to the following: When a worker's trajectory enters the buffer warning zone, considering the positioning uncertainty, there is a certain probability that the worker's actual position is already inside the hazardous area. Therefore, the width of the buffer warning zone should not be less than a certain multiple of the current positioning standard deviation estimate (the specific multiple is set according to the site management safety level requirements) to achieve a balance between the probability of missed detection and the probability of false detection. The dynamic buffer zone width is updated in real time with the positioning uncertainty estimated by the UKF. During periods of severe NLOS, as the positioning uncertainty increases, the buffer zone automatically widens accordingly, conservatively expanding the warning range to control the risk of missed detection. The output of S43 is the set of three-dimensional boundary polygons corresponding to each dangerous area and the width parameter of the outer buffer warning zone, which together constitute the final output of step S4.

[0045] The fusion assessment module inputs the worker's three-dimensional coordinates, behavioral labels, and dynamic hazardous area boundaries into the fusion assessment model, and outputs multi-level early warning results through a three-level risk assessment. S51, Level 1: Location risk assessment based on dynamic hazard zone boundaries; based on the real-time 3D coordinates of workers output by S2 and the current dynamic hazard zone 3D boundary set output by S4, a 3D spatial positional relationship judgment is performed on each worker. Specifically, for each type of hazard zone, a point inclusion test algorithm based on 3D polyhedra is used to determine whether the worker's 3D coordinates are located inside the hazard zone, inside the buffer warning zone, or within the safe zone. The scoring rules for the location risk component are as follows: when the worker's coordinates are inside the prohibited hazard zone, the location risk component takes the highest location risk score; when located inside the controlled operating area and the worker's job type permissions do not match (job type permissions are verified through the worker information database), the second highest location risk score is taken; when located inside the buffer warning zone, a medium location risk score is taken, and combined with the worker's movement speed direction relative to the boundary (using the velocity vector output by S2 to determine whether they are approaching the hazard zone boundary), if they are approaching, the risk score is appropriately increased; when located within the safe zone, the location risk component takes the safe score. For workers approaching the boundary of the danger zone, the estimated time to reach the danger zone (TTB) based on the current velocity vector is further calculated. When the TTB is less than the set TTB warning threshold, the location risk component is upgraded to the warning level to achieve early warning of impending boundary crossing. The first level outputs the location risk component value and location hazard level label for each worker.

[0046] S52, Second Level: Behavioral Risk Assessment Based on a Behavioral Category Mapping Table; Based on the behavioral category label sequence and corresponding behavioral confidence score of each worker output from S3, combined with a pre-set behavioral hazard level mapping table, the behavioral identification results are mapped to behavioral risk components. The construction principles of the behavioral hazard level mapping table are as follows: Direct personal safety hazards such as falling and waving arms (distress signals) are mapped to the highest behavioral risk score; running (which may cause collisions or accidental falls) and rapid climbing (risk of falling from heights) are mapped to relatively high behavioral risk scores; fatigue and stillness (prolonged inactivity, which may lead to loss of consciousness or operational errors due to mental fatigue) are mapped to medium behavioral risk scores; normal standing, walking, and operation are mapped to low behavioral risk scores. During the mapping process, the behavioral confidence score is used as a correction coefficient to weight the mapped behavioral risk scores. The higher the confidence score, the greater the weight of the mapped risk score; the lower the confidence score, the more conservatively the risk score approaches the medium behavioral risk score, avoiding the direct triggering of high-level warnings by erroneous identification results with low confidence. For workers marked as visually invisible in the blind zone in S32, the motion characteristic behavior classification results of the S32 alternative scheme are adopted, and the corresponding behavioral risk scores are weighted down according to the blind zone uncertainty correction coefficient. The second level outputs the behavioral risk component value and behavioral hazard level label for each worker.

[0047] S53, Multimodal Deep Fusion Risk Scoring and Multi-Level Early Warning Output Based on Attention Mechanism: Addressing the problem that traditional linear weighted fusion methods cannot effectively capture the complex nonlinear interactions between multidimensional features such as location, behavior, and history, this step employs an attention-based fusion evaluation model. Based on the location risk component output from S51 and the behavior risk component output from S52, combined with multidimensional contextual features obtained from the worker information database, including the worker's historical violation statistics, cumulative violation risk weights, the worker's continuous stay duration near the current high-risk area (obtained from the time integral of the worker's coordinate sequence), the worker's current job category, and the current shift's fatigue estimate (calculated based on on-duty hours), a multimodal feature input vector is constructed.

[0048] Deep embedding representation of multimodal features is performed: Heterogeneous features such as location risk components, behavioral risk components, historical violation weights, dwell time, job type, and fatigue are input into corresponding feature embedding sub-networks. Each sub-network adopts a two-layer fully connected network structure to map the original features to an embedding space of uniform dimension, making features from different modalities comparable. During feature embedding, normalized linear embedding is used for numerical features (such as location risk components and dwell time), and table-based embedding is used for categorical features (such as job type) using a learnable embedding matrix. The embedding matrix automatically learns the association pattern between each job type and risk assessment during training.

[0049] A multi-head self-attention mechanism is introduced to capture the interaction relationships between multimodal features: the embedding vectors of each modality feature are concatenated into a feature sequence, which is then input into the multi-head self-attention layer. The self-attention mechanism automatically learns the contribution of different modal features to the final risk score and the synergistic enhancement or inhibition relationships between features by calculating the attention weights between elements in the feature sequence. For example, when both the worker's location risk component and behavioral risk component are high (i.e., a dangerous behavior scenario occurs in a high-risk area), the self-attention mechanism can automatically assign higher attention weights to these two features and capture their coupling gain effect through feature interaction. Conversely, when a worker has few historical violations and low current fatigue, the self-attention mechanism can appropriately reduce the weight of the location risk component, reflecting the worker's higher reliability. The multi-head attention mechanism uses four attention heads, each focusing on different aspects of feature interaction. The multi-head outputs are concatenated and then fused through a feedforward network.

[0050] Execution of temporal context-aware risk score prediction: Addressing the temporal evolution of worker risk states (e.g., the process of a worker gradually approaching a danger zone from a safe area, or the accumulation of fatigue with on-duty hours), a Gated Recurrent Unit (GRU) is introduced to model the worker's historical risk state sequence within a continuous time window, based on the fused features output from the self-attention layer. The GRU inputs the fused feature vector at the current time step and the hidden state vector at the previous time step, and outputs an updated hidden state vector at the current time step, which encodes the temporal evolution trend information of the worker's risk state. The hidden state vector output by the GRU is then input into the final risk score prediction layer (two fully connected layers + a sigmoid activation function), outputting a normalized comprehensive risk score value ranging from 0 to 1.

[0051] The process involves multi-level warning threshold judgment and event log generation: The comprehensive risk score is compared with preset three-level risk thresholds (monitoring threshold, warning threshold, and alarm threshold), outputting the corresponding warning level: Monitoring level (risk score below the monitoring threshold) corresponds to a minor potential risk, requiring recording and attention; Warning level (risk score between the monitoring and warning thresholds) corresponds to a significant risk, requiring a notification; Alarm level (risk score exceeding the alarm threshold) corresponds to an urgent danger, requiring immediate response. Each time a warning result is generated, an event log containing complete fields such as worker ID, warning level, trigger time, worker's current 3D coordinates, behavior category, name of the hazardous area, comprehensive risk score, modal characteristic values, and attention weight distribution is recorded and written to the event database, forming the final output of step S5. The recording of attention weight distribution provides interpretability support for post-event analysis, helping safety managers understand the main reasons for warning triggers.

[0052] In some embodiments, due to the large number of workers operating simultaneously on construction sites, the risk patterns of different shifts, different types of work, and different construction stages vary significantly. A single assessment model may not be able to generalize effectively when faced with changes in the scenario. A rapid scenario adaptation mechanism based on meta-learning can be adopted, which involves pre-training a meta-learning model on historical data from multiple different construction sites and different construction stages, enabling the model to quickly adapt to new scenarios. The goal is to enable the assessment model to quickly adapt to a new construction stage or a new construction project with a small number of new scenario samples. Specifically, when the system detects a change in construction stage (triggered by changes in BIM construction progress nodes) or a decrease in the early warning accuracy of the assessment model within a continuous time window, the rapid meta-learning adaptation process is automatically triggered. The model parameters are fine-tuned with a small number of newly accumulated labeled samples (using manually verified early warning event results as labels) in the current scenario. The fine-tuning process uses the optimal initialization parameters and adaptive learning rate strategy learned during the meta-learning training stage, enabling the model to quickly recover its early warning accuracy with the support of 5 to 10 new scenario samples, avoiding the problem of traditional retraining methods requiring a large amount of new scenario data.

[0053] The early warning distribution module receives the early warning result at the edge node, executes a local response to trigger alarms on the worker end, synchronizes data to the cloud, and outputs alarm records and situation reports. S61 enables millisecond-level local early warning response and worker-side alarm triggering at edge computing nodes. Based on the real-time multi-level early warning results for each worker output by S5, the edge computing node immediately enters the local response processing flow upon receiving an early warning event. The core objective of local response is to transmit alarm signals to on-site workers and management personnel with the shortest possible latency without relying on cloud network connectivity, ensuring local safety fallback capabilities in the event of tunnel network interruption or cloud communication delays.

[0054] The specific processing procedure for local response is as follows: The edge computing node maintains a priority response queue. Alarm-level warning events output by S5 enter the queue with the highest priority, and warning-level and monitoring-level events are arranged in order. For alarm-level warning events, the edge computing node sends a vibration command to the corresponding worker's UWB tag through the reverse downlink control channel of the UWB network. The UWB tag has a built-in miniature vibration motor, which immediately triggers continuous vibration upon receiving the vibration command (the vibration pattern is three short, continuous vibrations followed by a pause and loop), allowing the worker to receive an emergency alarm through tactile perception, unaffected by tunnel noise interference. At the same time, the edge computing node triggers a voice broadcast through the local broadcast speaker network deployed in the tunnel (connected to the edge node via the tunnel construction site Ethernet), which includes the worker's number and corresponding hazard warning information, so that surrounding workers and on-site management personnel are simultaneously informed of the alarm. For early warning events, the edge computing node sends a mild vibration alert (single, short vibration) to the worker's UWB tag and pushes a pop-up warning message to the nearest on-site management personnel's handheld terminal (a tablet or handheld PDA device connected to the edge node via LAN). The message includes the worker's name, current location, behavior category, and risk description. For monitoring-level early warning events, the edge computing node only records the event to the local event database without triggering vibration or voice alarms. However, the corresponding worker's location icon is highlighted on the on-site management personnel's local monitoring screen, enabling continuous visual monitoring of workers at potential risk.

[0055] The edge computing node records the complete response process for this event, including the warning level, corresponding worker ID, alarm triggering method (vibration or voice), response processing time, and total response delay. This record is written to a local log file as the output of S61, and is used for cloud synchronization in subsequent S62 steps.

[0056] In some embodiments, due to the risk of downlink frame collision in the reverse downlink control channel of UWB tags in tunnels when the tag density is extremely high, vibration alarm commands cannot be correctly received by worker tags, resulting in an increased alarm miss rate. An alarm retransmission protocol based on an acknowledgment mechanism can be adopted. After the edge computing node issues a vibration command, it waits for the ACK (Acknowledgement) flag attached to the worker's UWB tag in the next ranging polling cycle. If no ACK is received within the set retransmission timeout period, the vibration command is automatically retransmitted with the highest priority, and the number of retransmissions does not exceed the set maximum number of retransmissions. The purpose is to ensure the reliable delivery of alarm-level vibration commands in high-density concurrent tag scenarios. Specifically, when the number of retransmissions reaches the upper limit and no ACK is received, the edge computing node automatically upgrades the worker's alarm method to on-site broadcast voice alarm, records the alarm anomaly event in the local log, and reports the tag communication failure information to the cloud, triggering the operation and maintenance troubleshooting process.

[0057] S62 enables edge-to-cloud data synchronization and real-time cloud-based security posture management. Based on the worker's local vibration alarm trigger records output by S61 and the complete event logs output by S5, edge computing nodes upload event logs, real-time worker location snapshots, and the current status of hazardous area boundaries to the cloud-based security management platform in batches via an encrypted transmission channel (TLS, Transport Layer Security) at a set synchronization period (the synchronization period is determined comprehensively based on cloud bandwidth conditions and real-time requirements, generally between seconds and minutes). For alarm-level events, there is no synchronization period limitation; an immediate push mechanism is triggered, and the data is uploaded to the cloud separately with minimal network transmission delay after the event occurs, ensuring that cloud management personnel can perceive on-site emergencies immediately.

[0058] After receiving data uploaded from edge nodes, the cloud-based safety management platform performs the following processing: It updates the received worker location, behavior, and warning level data to the cloud-based real-time situation database; it also refreshes the location icon, behavior status label, and warning level color indicator (yellow for monitoring level, orange for warning level, and flashing red for alarm level) of each worker in the 3D construction site model (based on cloud-based BIM model rendering) on ​​the cloud-based web monitoring interface in real time, enabling remote construction site safety managers to obtain a real-time panoramic view of the safety situation consistent with the actual site conditions through the cloud interface; the cloud platform triggers a secondary cloud-based push mechanism for received alarm-level events, sending alarm notifications to the construction site safety manager and project management manager via SMS and enterprise instant messaging applications (such as WeChat Work, DingTalk, etc.). The notification content includes event details, a screenshot of the current worker's location, and operational suggestions; thirdly, the cloud platform performs online statistical analysis on historical event data from multiple time periods, calculating the frequency of various dangerous behaviors, the number of worker intrusions in each dangerous area, and the time distribution, forming a hazard heat map overlaid on the cloud-based 3D construction site model. This assists construction site management in identifying high-risk periods and areas where safety problems repeatedly occur, providing data support for management decisions.

[0059] The output of S62 is the update record of the real-time security situation database in the cloud (including worker location snapshots, early warning event upload times, and cloud push records) and the refresh completion status of the real-time situation visualization interface, which are used for the generation of the security situation report in the subsequent S63 step.

[0060] In some embodiments, due to the unstable network environment at the tunnel construction site, the network link between the edge node and the cloud may be intermittently interrupted, resulting in data gaps in the cloud situational data and affecting the continuous perception of the site situation by cloud management personnel. An offline fault-tolerant mechanism based on local priority caching and breakpoint resumption can be adopted. During network interruptions, the edge computing node continuously writes event logs and location snapshot data to local high-reliability storage media (using RAID, Redundant Array of Independent Disks, or a dual solid-state drive backup solution) and maintains complete situational display capabilities on the local monitoring screen. When the network link is restored, the edge node automatically detects the start timestamp of the data that was not uploaded locally and re-uploads it to the cloud in batches according to time sequence. After receiving the re-uploaded data, the cloud platform automatically fills in the data gaps of the corresponding time period to ensure the integrity of historical event logs. The purpose is to ensure the continuous integrity of historical data in the cloud in the unstable underground tunnel construction environment. Specifically, during the re-upload process, alarm-level historical events are prioritized for transmission, ensuring that the integrity of safety responsibility records takes precedence over general monitoring records.

[0061] S63, automatic generation and archiving of periodic security posture reports in the cloud; based on historical event data accumulated in the real-time cloud security posture database output by S62, the cloud security management platform automatically triggers the security posture report generation process according to the set report generation cycle (such as after the end of each daily shift or weekly summary). The specific processing steps of the report generation process are as follows: The cloud platform extracts complete behavioral trajectory data and early warning event records of all workers within the current reporting period from the event database. It summarizes and statistically analyzes each worker's total on-duty time, number of times they entered dangerous areas, frequency of various behaviors, number of vibration alarm triggers, and number of early warning events at all levels within the current period, according to worker ID. It also summarizes and statistically analyzes the total number of times and total cumulative time that each dangerous area was entered by workers within the current period, identifying high-frequency intrusion areas. Finally, it analyzes the distribution of early warning event density for each hour of the day, identifying high-risk time periods, according to time.

[0062] Based on the above statistical results, the cloud platform calls the automatic report template engine to populate the statistical figures, charts (including bar charts of intrusion frequency in hazardous areas, line charts of warning density for different time periods, and worker individual risk ranking tables) and text descriptions in the report template, generating a structured safety situation report document. The report format is a standard PDF document, including four parts: a cover (construction site name, reporting period, generation time), an execution summary (number of major warning events in this period, list of high-risk workers), detailed statistical analysis, a screenshot of the hazard heat map, and management recommendations. The management recommendations are automatically generated by the rule engine based on the statistical results. The rule base is preset by the construction site safety regulations. For example, when a worker intrusion event occurs in an area more than the threshold number specified in the safety regulations within this period, the system automatically outputs a recommendation to physically reinforce the fence of the area and conduct targeted safety education.

[0063] The generated safety situation report PDF document is stored in the cloud document archive, and the report link is sent to the email and instant messaging accounts of the site safety manager and project management manager via a cloud push mechanism. The output of S63 is the safety situation report document for this cycle and its cloud archive record, which together with the local vibration alarm trigger record output of S61 and the cloud early warning push record output of S62 constitute the final complete output of step S6.

[0064] In some embodiments, due to significant differences in management procedures and safety priorities among different construction site projects, a unified report template and rule engine may not be sufficiently targeted to some sites, and management recommendations may lack scenario-specificity. A knowledge-enhanced report generation method based on a historical safety accident case database can be adopted. This involves pre-entering historical safety accident cases (including early warning event patterns, worker behavior characteristics, and hazardous area distribution patterns) of the current construction site and similar tunnel construction sites into a knowledge base. During report generation, the statistical characteristics of the current period are matched with the characteristics of historical cases in the knowledge base. If a highly similar historical accident precursor pattern is matched, specific case references and targeted prevention and control measures are added to the management recommendations. The aim is to make the management recommendations in the automatically generated safety situation report more persuasive and targeted with case support, thereby increasing management's attention to the report and their willingness to implement it. Specifically, the similarity matching uses a cosine similarity calculation method based on event pattern feature vectors. Historical cases with similarity exceeding a set case matching threshold are included in the report reference, and the top three cases with the highest similarity are appended to the end of the report in the form of a brief summary.

[0065] In one embodiment of the invention, it is applied to the high-risk construction project of underground tunnel construction in a section of an urban subway. The construction section is a shield tunnel with a diameter of approximately six meters, and the tunnel face is about 800 meters from the tunnel entrance. The work is carried out in three shifts, with about 30 to 50 workers working in the tunnel at the same time during each shift. The main types of workers include shield operators, steelworkers, segment assemblers, electricians, and safety inspectors. There are four dynamic danger zones at the construction site: the shield machine propulsion zone, the high-pressure hydraulic pipeline operation zone, the segment storage and hoisting zone, and the temporary electrical equipment zone. The boundaries of each zone shift by tens to hundreds of meters each shift as the construction progresses.

[0066] The system deployment plan is as follows: Four UWB base stations are deployed approximately every 30 meters along the tunnel's longitudinal direction, for a total of approximately 27 sets of base stations. One wide-angle AI camera is deployed approximately every 40 meters along the tunnel's longitudinal direction, for a total of approximately 20 cameras. All cameras are equipped with active lighting. An edge computing server node is set up at the tunnel entrance, interconnected with all base stations and cameras via the construction site's 10 Gigabit Ethernet ring network. The edge node maintains synchronization with the cloud-based safety management platform via a 4G backup link. All workers complete UWB tag registration before entering the site, and the tags are worn in the internal compartment of their safety helmets.

[0067] Taking the actual operation data of the early shift on a certain workday as an example, Table 1 shows some typical worker status data recorded by the system during this shift, and Table 2 shows some typical sample records of early warning events triggered during this shift.

[0068] Table 1, Example of Real-Time Worker Status Data

[0069] Table 2. Examples of typical early warning event data

[0070] In the above application example, the system achieved real-time perception and graded response for five workers representing risk states. Worker W-0067 engaged in rapid climbing behavior within the prohibited danger zone, triggering the highest-level alarm response; worker W-0041, although not yet in the danger zone, had his risk score exceed the warning threshold due to the superposition of fatigued stationary behavior and historical violation weights, prompting an early warning; worker W-0082 was within the safe zone but his movement trend was approaching the danger boundary, and the system achieved advanced perception through the TTB warning mechanism. The graded response results in these different scenarios demonstrate the comprehensive perception and accurate early warning capabilities of the multimodal fusion three-level evaluation model of this invention in complex tunnel construction scenarios.

[0071] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A construction site safety early warning system based on UWB and AI cameras, characterized in that, include: The data acquisition module collects UWB tag ranging data from workers and camera video frame data to obtain a multi-source sensor dataset. The positioning and calculation module uses time-frequency domain feature extraction and depth residual network to identify non-line-of-sight states from UWB tag ranging data, and combines trajectory estimation to obtain the worker's three-dimensional coordinate sequence. The behavior binding module performs target detection and behavior recognition on video frames. It combines the worker's 3D coordinates with trajectory consistency constraints and metric learning to achieve cross-modal identity binding and outputs a behavior label sequence. The fence update module obtains the initial boundary of the danger zone, uses video frames for instance segmentation and point cloud reconstruction for verification and correction, combines worker position covariance statistics to set a buffer warning zone, and outputs a dynamic set of danger zone boundaries. The fusion assessment module inputs the worker's three-dimensional coordinates, behavioral labels, and dynamic hazardous area boundaries into the fusion assessment model, and outputs multi-level early warning results through a three-level risk assessment. The early warning distribution module receives early warning results from edge nodes, executes local responses to trigger alarms on the worker side, synchronizes data to the cloud, and outputs alarm records and situation reports.

2. The construction site safety early warning system based on UWB and AI cameras according to claim 1, characterized in that, The data acquisition module includes: A unified three-dimensional coordinate system for the construction site is established based on the site measurement benchmark. A group of UWB base stations is deployed at preset intervals along the longitudinal direction in the tunnel. Each group of base stations has no less than 4 base stations, and the spatial distribution of a single group of base stations on the cross section is guaranteed to be non-coplanar. The three-dimensional installation coordinates of each UWB base station are obtained by actual measurement with a total station and entered into the base station coordinate database. The internal parameters of the multiple AI cameras deployed along the tunnel were calibrated using the calibration board method, and the external parameters of each camera were calibrated to the unified three-dimensional coordinate system of the construction site through a joint checkerboard calibration field. Using the system clock of the edge computing node as the master clock reference, a precise time protocol synchronization signal is broadcast to all UWB base stations and AI cameras. The edge computing node aligns the UWB ranging data set and video frame set at the same moment with a preset time window to form a multi-source sensing dataset. Each worker's UWB tag ID is bound to the worker's identity information and entered into the worker information database. Each UWB base station collects bilateral and bidirectional ranging data in a cyclical manner at a preset ranging polling frequency. Each ranging sampling outputs the estimated flight time from the corresponding tag to the base station and the original CIR sampling data.

3. A construction site safety early warning system based on UWB and AI cameras according to claim 1, characterized in that, The positioning calculation module includes: Based on the original CIR sampling data of each UWB ranging link in the multi-source sensing dataset, joint time-frequency domain feature extraction is performed. After normalizing the CIR sampling sequence in the time domain, the peak amplitude of the first path, root mean square delay spread, peak-to-average ratio, and number of effective multipaths are extracted. In the frequency domain, a short-time Fourier transform is performed on the CIR sequence to extract the frequency domain energy concentration, main bandwidth, and spectral entropy. The time-domain feature vector and the frequency-domain feature vector are concatenated to form a joint time-frequency domain feature vector, which is then input into the NLOS state classifier constructed by the deep residual network. The deep residual network consists of an input layer, three residual blocks, a global average pooling layer, and an output layer. The output layer uses the Softmax activation function to output three probability distributions: LOS state, mild NLOS state, and severe NLOS state. After NLOS identification is completed, no ranging correction is performed on links in the LOS state. For links in the mild NLOS state, the flight time estimate is corrected using a mean deviation compensation method based on historical statistics. For links in the severe NLOS state, a nonlinear deviation regression model based on a multilayer perceptron is used for compensation. Each state corresponds to a preset weight coefficient, and the corrected flight time estimate and corresponding weight coefficient of each ranging link are output.

4. A construction site safety early warning system based on UWB and AI cameras according to claim 3, characterized in that, The positioning calculation module also includes: Based on the corrected flight time estimate and corresponding weight coefficients, a set of three-dimensional positioning equations based on arrival time is established. The weighted least squares linearized iterative solution method is adopted, with the weight coefficients as the weights of each equation, to solve the initial estimate of the three-dimensional coordinates of the worker tag. The extrapolated coordinates of the worker's trajectory at the previous moment are used as the initial point of iteration. When the residual exceeds the preset solution quality threshold, the positioning quality flag is marked as low confidence. Based on the initial 3D coordinate estimation sequence, an adaptive unscented Kalman filter is used for trajectory smoothing. The state vector contains the worker's 3D position coordinates and 3D velocity components. At each time step, the covariance of the innovation sequence is calculated and compared with the theoretical innovation covariance. If the difference exceeds the preset adaptive adjustment threshold, the process noise covariance matrix is ​​adaptively adjusted. Combined with the positioning quality flag, the measurement noise covariance is dynamically adjusted. The smoothed 3D coordinate sequence and 3D velocity vector estimation sequence of each worker are output as the worker's 3D coordinate sequence.

5. A construction site safety early warning system based on UWB and AI cameras according to claim 1, characterized in that, The behavior binding module includes: Based on video frames from a multi-source sensor dataset, a target detection model is used to perform worker target detection. The target detection model introduces a multi-scale feature fusion module of a bidirectional feature pyramid network into the neck network and adds a safety helmet color classification sub-task branch to the head detection part, outputting a set of worker bounding boxes and safety helmet color categories. Simultaneously, compliance checks are performed on the wearing status of safety helmets and safety vests. When the confidence level of the safety helmet test is lower than the preset compliance threshold, a compliance flag is marked.

6. A construction site safety early warning system based on UWB and AI cameras according to claim 5, characterized in that, The behavior binding module also includes: Based on the worker bounding box set, worker 3D coordinate sequence and camera intrinsic and extrinsic parameters, cross-modal worker identity binding is performed. The worker's 3D coordinates are projected onto the image plane using the camera projection matrix. The two-dimensional Euclidean distance between the worker's 3D coordinates and the center point coordinates of the visually detected bounding box is calculated to construct a preliminary geometric cost matrix. For each candidate matching pair, extract the UWB projection trajectory sequence and visual trajectory sequence within the past N consecutive frames, calculate the dynamic temporal warping distance as the trajectory consistency cost, and weight and fuse it with the preliminary geometric cost to obtain the comprehensive cost matrix. For each visually detected worker bounding box, extract the appearance image region, input it into a Siamese-based deep metric learning network to obtain the appearance feature vector, calculate the cosine similarity with the standard appearance feature vector in the worker information database, convert it into the appearance cost and the comprehensive cost matrix, and then fuse them to obtain the final multi-constraint fusion cost matrix. The Hungarian algorithm is used to perform optimal binary matching on the final multi-constraint fusion cost matrix. Matching pairs with costs lower than the preset binding cost threshold are taken as valid binding results, and a cross-modal identity association mapping table is output.

7. A construction site safety early warning system based on UWB and AI cameras according to claim 6, characterized in that, The behavior binding module also includes: Based on the cross-modal identity association mapping table, the corresponding bounding box region image sequence is extracted from the video frame for each bound worker. The sequence is then input into the pose estimation network to extract the coordinate sequence of key points of the worker's skeleton. A spatiotemporal graph structure is constructed. The spatial dimension defines the graph topology structure based on the skeleton connection relationship, and the temporal dimension constructs temporal edges based on the temporal connection of the same joint node in adjacent frames. The spatiotemporal graph is input into the ST-GCN network. The joint co-motion features in the spatial dimension are aggregated through graph convolution operations, and the temporal evolution of the action sequence is captured through temporal convolution operations. The probability distribution of the behavior category is output, and the category with the highest probability is taken as the behavior recognition result. The behavior label sequence contains the behavior category label and behavior confidence score of each worker.

8. A construction site safety early warning system based on UWB and AI cameras according to claim 1, characterized in that, The fence update module includes: Based on the construction progress update data pushed by the construction site management system, the set of currently active construction nodes in the BIM model is parsed, the corresponding set of dangerous area geometry is extracted and the coordinates are transformed to the unified three-dimensional coordinate system of the construction site. Dangerous areas are divided into three categories: prohibited dangerous areas, controlled operation areas and warning buffer zones. Mask R-CNN instance segmentation network is used to perform pixel-level segmentation of key construction machinery in video frames to obtain contour masks. Key point detection network is used to extract image coordinates of key mechanical components. Multi-view triangulation is used to reconstruct the three-dimensional spatial coordinate point cloud of key mechanical components. The reconstruction results are compared with the mechanical position of BIM analysis. When the deviation exceeds the preset verification trigger threshold, boundary correction is triggered. Based on the estimated position covariance matrix corresponding to the worker's three-dimensional coordinate sequence, a dynamic buffer warning zone is added to the outside of each danger zone. The width of the buffer warning zone is not less than a preset multiple of the current positioning standard deviation estimate, and the dynamic danger zone boundary set is output.

9. A construction site safety early warning system based on UWB and AI cameras according to claim 1, characterized in that, The fusion evaluation module includes: The first level is based on the worker's three-dimensional coordinate sequence and the dynamic dangerous area boundary set. It uses a three-dimensional polyhedron point inclusion test algorithm to determine the type of area where the worker is located. For workers who are located in the buffer warning zone and are approaching the dangerous area, the estimated time to reach the dangerous area is calculated. When the estimated time to reach the dangerous area is less than the preset TTB warning threshold, the location risk component is increased and the location risk component value is output. The second level is based on the behavior label sequence. Combined with the preset behavior risk level mapping table, the behavior category is mapped to the behavior risk component. The behavior confidence score is used as a correction coefficient to weight the behavior risk score and output the behavior risk component value. The third level inputs the location risk component, behavioral risk component, worker's historical violation weight, dwell time, job type, and fatigue level into the feature embedding subnetwork and maps them to a unified dimension embedding space. The interaction relationship between multimodal features is captured through a multi-head self-attention mechanism. Then, a gated recurrent unit is introduced to perform time-series modeling of the historical risk state sequence and output a normalized comprehensive risk score. This score is compared with the preset monitoring-level threshold, early warning-level threshold, and alarm-level threshold to output multi-level early warning results.

10. A construction site safety early warning system based on UWB and AI cameras according to claim 1, characterized in that, The early warning distribution module includes: Edge computing nodes maintain priority response queues. For alarm-level early warning events, vibration commands are sent to worker UWB tags via the UWB network's reverse downlink control channel, and voice broadcasts are triggered by local loudspeakers. For warning-level early warning events, mild vibration alerts are sent to worker UWB tags, and pop-up warning information is pushed to the handheld terminals of on-site management personnel. For monitoring-level early warning events, they are recorded in the local event database and highlighted on the local monitoring screen. Edge computing nodes upload event logs and worker location snapshots in batches to the cloud safety management platform via an encrypted transmission channel at a preset synchronization cycle. Alarm-level events trigger an immediate push mechanism. The cloud platform updates worker locations and warning level indicators in the 3D construction site model in real time and sends alarm notifications to the safety manager via instant messaging applications for alarm-level events. The cloud platform extracts historical data from the event database at a preset report generation cycle, summarizes and statistically analyzes the data by worker, hazardous area, and time dimensions, generates and archives structured safety situation reports, and outputs alarm records and situation reports.