Building construction safety monitoring method and system based on multi-source sensing fusion
By performing spatiotemporal benchmark alignment, parallel feature extraction, and BIM semantic evaluation on multi-source sensor data from construction sites, the problems of asynchronous data from heterogeneous sensors and signal interference were solved, enabling real-time, accurate, and reliable early warning and intervention for construction safety monitoring.
Patent Information
- Application Number
- CN202511099877.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-14
AI Technical Summary
Existing construction safety monitoring methods rely on manual inspections, which suffer from delays, limited coverage, strong subjectivity, and high labor costs. They are difficult to achieve real-time, comprehensive, and accurate early warning and intervention. Furthermore, the asynchronicity of multi-source sensor data and signal interference make it difficult to guarantee positioning accuracy and continuity.
By aligning the original video stream and UWB positioning data packets with a spatiotemporal reference, performing parallel feature extraction and preliminary entity association, using BIM semantic data for dynamic confidence assessment, conducting adaptive fusion decision-making and state optimization, and outputting high-confidence object states.
It significantly improves the accuracy, real-time performance, and reliability of construction safety monitoring, enabling stable and precise safety monitoring and early warning at construction sites.
Smart Images

Figure CN120953809A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent monitoring, and more specifically, to a method and system for monitoring building construction safety based on multi-source sensor fusion. Background Technology
[0002] As a pillar of the national economy, the construction industry is characterized by complex production environments, diverse operations, and inherently high risks. Traditionally, construction safety has relied primarily on manual inspections, on-site supervision, and post-event traceability. However, these methods typically suffer from drawbacks such as delays, limited coverage, strong subjectivity, and high labor costs, making it difficult to achieve real-time, comprehensive, and accurate early warning and intervention for hazardous situations at construction sites. Therefore, a more efficient and intelligent safety monitoring solution is urgently needed.
[0003] Sensors based on positioning technologies (such as UWB and GPS) can provide accurate location information for objects or people. UWB excels in indoor or localized positioning due to its high accuracy and penetration, while GPS is suitable for large-scale outdoor positioning. However, in reinforced concrete walls and construction sites with dense large metal equipment, UWB and GPS signals are highly susceptible to interference from shielding, reflection, and multipath effects, leading to signal interruptions, data jumps, or instantaneous drifts of several meters, making it difficult to guarantee the accuracy and continuity of positioning results. Furthermore, these wearable positioning devices typically only output a single three-dimensional coordinate point, lacking the semantic granularity to describe a person's posture, orientation, or the actual position of their upper and lower body. Even more challenging is the fact that different types of sensors often report data at different frequencies through their own independent communication networks. For example, visual data and UWB data may have transmission delays ranging from milliseconds to seconds and timestamp mismatches, resulting in inherent asynchronicity in the spatiotemporal reference of the raw data. This poses a significant challenge to the accurate fusion of multi-source data. Existing single-source sensors or simple multi-source data overlay methods often cannot effectively solve these inherent environmental constraints, sensor limitations, and data asynchronicity problems, making it difficult to provide stable, reliable, and high-confidence construction safety situation awareness capabilities, thus limiting the large-scale promotion and application of intelligent safety monitoring solutions in actual engineering projects.
[0004] Therefore, an optimized method for monitoring building construction safety based on multi-source sensor fusion is needed. Summary of the Invention
[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a construction safety monitoring method and system based on multi-source sensor fusion. First, it precisely aligns the original video stream and UWB positioning data packets using spatiotemporal references to solve the inherent asynchronicity and conflict problems of heterogeneous sensor data, ensuring the accuracy of subsequent data fusion. Subsequently, it performs parallel feature extraction and preliminary entity association on the aligned data, and innovatively utilizes BIM semantic data to conduct dynamic confidence assessment of the preliminary fusion objects based on scene constraints. This fully leverages prior knowledge of the building environment to dynamically optimize the reliability of sensor data, thereby compensating for the shortcomings of single sensors being susceptible to environmental influences, occlusion, or signal interference leading to information loss. Finally, based on the obtained confidence scores, the system performs adaptive fusion decisions and state optimization, outputting high-confidence object states. This approach significantly improves the accuracy, real-time performance, and reliability of construction safety monitoring.
[0006] According to one aspect of this application, a method for monitoring building construction safety based on multi-source sensor fusion is provided, comprising: The acquired raw video stream and raw UWB positioning data packet are spatiotemporally aligned to obtain aligned image data and aligned UWB data; Parallel feature extraction and preliminary entity association are performed on aligned image data and aligned UWB data to obtain a preliminary list of fused objects; Based on BIM semantic data, a dynamic confidence assessment based on scene constraints is performed on each preliminary fusion object in the preliminary fusion object list to obtain a confidence score. Based on the confidence score, adaptive fusion decision and state optimization are performed on the preliminary fusion object list to obtain the optimized object state. The optimized object state is input into the rule engine to obtain alarm events.
[0007] According to another aspect of this application, a construction safety monitoring system based on multi-source sensor fusion is provided, comprising: The spatiotemporal reference alignment module is used to perform spatiotemporal reference alignment on the acquired raw video stream and raw UWB positioning data packets to obtain aligned image data and aligned UWB data. The preliminary data fusion module is used to perform parallel feature extraction and preliminary entity association on aligned image data and aligned UWB data to obtain a preliminary fusion object list. The dynamic confidence assessment module is used to perform dynamic confidence assessment on each preliminary fusion object in the preliminary fusion object list based on scene constraints, based on BIM semantic data, to obtain a confidence score. An adaptive optimization module is used to perform adaptive fusion decision-making and state optimization on the preliminary fusion object list based on the confidence score to obtain the optimized object state. The alarm module is used to input the optimized object status into the rule engine to obtain alarm events.
[0008] Compared with existing technologies, this application provides a construction safety monitoring method and system based on multi-source sensor fusion. First, it precisely aligns the original video stream and UWB positioning data packets using spatiotemporal references to address the inherent asynchronicity and conflict issues of heterogeneous sensor data, ensuring the accuracy of subsequent data fusion. Then, it performs parallel feature extraction and preliminary entity association on the aligned data. Innovatively, it utilizes BIM semantic data to conduct dynamic confidence assessment of the preliminary fusion objects based on scene constraints, fully leveraging prior knowledge of the building environment to dynamically optimize the reliability of sensor data. This compensates for the shortcomings of single sensors, which are susceptible to environmental influences, occlusion, or signal interference leading to information loss. Finally, based on the obtained confidence scores, the system performs adaptive fusion decisions and state optimization, outputting high-confidence object states. This approach significantly improves the accuracy, real-time performance, and reliability of construction safety monitoring. Attached Figure Description
[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 This is a flowchart of a construction safety monitoring method based on multi-source sensor fusion according to an embodiment of this application; Figure 2 This is a schematic diagram of the data flow in the construction safety monitoring method based on multi-source sensor fusion according to an embodiment of this application; Figure 3 This is a block diagram of a building construction safety monitoring system based on multi-source sensor fusion according to an embodiment of this application. Detailed Implementation
[0011] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0012] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0013] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.
[0014] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0015] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0016] The technical solution of this application proposes a method for monitoring building construction safety based on multi-source sensor fusion. Figure 1 This is a flowchart of a construction safety monitoring method based on multi-source sensor fusion according to an embodiment of this application. Figure 2 This is a system architecture diagram of a construction safety monitoring method based on multi-source sensor fusion according to an embodiment of this application. Figure 1 and Figure 2 As shown, the construction safety monitoring method based on multi-source sensor fusion according to an embodiment of this application includes the following steps: S1, performing spatiotemporal reference alignment on the acquired original video stream and original UWB positioning data packet to obtain aligned image data and aligned UWB data; S2, performing parallel feature extraction and preliminary entity association on the aligned image data and aligned UWB data to obtain a preliminary fusion object list; S3, based on BIM semantic data, performing dynamic confidence assessment on each preliminary fusion object in the preliminary fusion object list based on scene constraints to obtain a confidence score; S4, based on the confidence score, performing adaptive fusion decision and state optimization on the preliminary fusion object list to obtain an optimized object state; S5, inputting the optimized object state into a rule engine to obtain an alarm event.
[0017] Specifically, in step S1, the acquired raw video stream and raw UWB positioning data packets are spatiotemporally aligned to obtain aligned image data and aligned UWB data. It should be understood that in actual construction environments, due to differences in the working principles of different types of sensors, communication networks, and data transmission paths, the acquired raw data often suffers from inherent asynchronicity. For example, different sensors typically report data at different frequencies through different networks (e.g., 4G / 5G, Wi-Fi, LoRa), which can lead to latency differences ranging from milliseconds to seconds when the data reaches the cloud processing center. For instance, when processing data from a specific time slice, the received camera data frame might correspond to t=10.4s, while the UWB positioning data might correspond to t=10.6s. If this inherent timing mismatch is not properly handled, it will amplify data conflicts, severely impacting the accuracy of subsequent data fusion, the reliability of the system, and the effectiveness of decision-making. Therefore, to ensure the logical consistency and physical accuracy of subsequent data processing and fusion, the technical solution of this application aligns the acquired raw video stream and raw UWB positioning data packet with a spatiotemporal reference. This achieves calibration and synchronization of the raw video stream data and the raw UWB positioning data packet in the time and space dimensions, resulting in aligned image data and aligned UWB data. The aligned image data includes a timestamp, frame ID, and image, while the aligned UWB data includes a timestamp, target ID, and three-dimensional coordinates. This alignment process ensures that all data processed subsequently has a unified temporal and spatial reference, laying a solid foundation for subsequent parallel feature extraction and entity association.
[0018] Spatiotemporal reference alignment refers to synchronizing data from different sensors with timestamps and transforming spatial coordinate systems to create a unified reference benchmark in time and space, thereby enabling effective data fusion.
[0019] In practice, the raw video stream may be captured by an industrial camera equipped with a precise clock synchronization module and timestamped in real time with microsecond-level precision. Meanwhile, the UWB positioning data packets are acquired via a UWB tag worn by personnel or equipment, using a UWB base station network to obtain its 3D location information with the same precise timestamp. During the alignment process, the backend processing system loads pre-completion joint calibration parameters for the visual sensor and the UWB positioning system. Using these parameters, it precisely projects the physical 3D coordinates of the UWB onto the image plane corresponding to each video frame. Simultaneously, it performs synchronous adjustments based on the timestamps of all data to ensure that at any given point in time, the physical world described by the visual data and the positioning data are in a consistent spatiotemporal state.
[0020] Specifically, in step S2, parallel feature extraction and preliminary entity association are performed on the aligned image data and aligned UWB data to obtain a preliminary fusion object list. It should be understood that although the previous step completed the spatiotemporal alignment of the data, the two are still independent data streams. Data from a single sensor often has limitations; for example, visual sensors are susceptible to occlusion and lighting conditions, lacking accurate depth information. While UWB positioning data can provide precise three-dimensional coordinates, it lacks semantic information and visual features of the target and may also exhibit jumps due to environmental interference. To overcome these inherent defects of single-source data and provide more comprehensive and reliable entity information for subsequent intelligent security decisions, it is necessary to perform initial fusion of heterogeneous data. Therefore, in the technical solution of this application, by extracting meaningful features from their respective modalities in parallel and attempting to associate these features, key entities (such as personnel and equipment) in the construction site can be initially identified and constructed, laying the foundation for subsequent refined evaluation based on BIM data. This parallel processing and preliminary association method aims to fully utilize the advantages of each sensor and compensate for its shortcomings, thereby forming a richer and more robust entity description.
[0021] In practice, firstly, the aligned image data is input into the visual processing branch to obtain a set of tracked video targets. This visual processing branch includes a trained object detection model and a multi-object tracker. Specifically, the object detection model is responsible for identifying various target entities in the image frames, such as construction workers and machinery, and providing their position (usually a bounding box) and category information in the image. For each frame, the model outputs the detected targets and their confidence scores. Subsequently, the multi-object tracker, based on these real-time detection results and combining the target's motion trajectory and feature information, associates the same target across different frames, thereby assigning a unique ID to each continuously detected and identified target and maintaining its continuity in the video sequence. In this way, a set of tracked video targets is formed, containing multiple tracked targets and their corresponding information in the video (such as ID, category, bounding box, timestamp, etc.).
[0022] Simultaneously, the aligned UWB data is input into the UWB processing branch to obtain a structured UWB location point data queue. In this parallel processing path, the UWB processing branch receives spatiotemporally aligned UWB location data, which typically includes timestamps, target IDs, and 3D coordinates. During this process, the UWB processing branch parses, organizes, and optimizes these raw UWB location data packets, potentially including preprocessing such as noise filtering and outlier removal, transforming them into a unified, easily processed structured UWB location point data queue. Each structured UWB location point data typically includes a unique identifier for the target (if the UWB tag has an ID), its precise location coordinates in 3D space, and the corresponding timestamp.
[0023] Next, the structured UWB location point data queue is projected across modal coordinate space based on camera parameter calibration to obtain a projected UWB point set. Since visual data is a two-dimensional image, while UWB data is three-dimensional coordinates, to effectively correlate the two, the three-dimensional UWB location points need to be projected onto a two-dimensional image plane. This projection process relies on pre-defined camera calibration parameters, including the camera's intrinsic parameters (focal length, principal point, distortion coefficients, etc.) and extrinsic parameters (camera pose and position in the three-dimensional world coordinate system). Using these parameters, each UWB location point (three-dimensional coordinates) can be accurately converted into its pixel coordinates on the corresponding video frame, forming a projected UWB point set, where each point corresponds to the estimated location of an entity represented by a UWB label on the image.
[0024] Then, a candidate association pair list is generated based on spatiotemporal proximity driven by the set of tracked video targets and the set of projected UWB points. That is, visual targets and UWB points that may correspond to the same physical entity are identified based on temporal and spatial proximity. Specifically, firstly, it is determined whether the difference between the timestamp of each projected UWB point in the set of projected UWB points and the timestamp of the corresponding tracked video target in the set of tracked video targets is within a preset range to obtain a first judgment result; this ensures that the visual and UWB data to be associated are synchronous or approximately synchronous in time. Secondly, it is determined whether the pixel coordinates of each projected UWB point in the set of projected UWB points are within the detection box of the corresponding tracked video target in the set of tracked video targets to obtain a second judgment result; this considers spatial proximity, i.e., whether the projection of the UWB point on the image falls within the area defined by the visual target. Furthermore, if both the first and second judgment results are true, the corresponding projected UWB point and the tracked video target are determined as a candidate association pair.
[0025] Further, affinity assessment and association clarification are performed on each candidate association pair in the candidate association pair list to obtain the preliminary fusion object list. That is, after identifying all possible candidate association pairs, each candidate association pair in the candidate association pair list is further refined and confirmed. Here, affinity assessment may involve more complex similarity measures, such as considering factors such as UWB signal quality, visual target features (e.g., color, size), and kinematic consistency, to calculate an affinity score for each candidate association pair. Based on these scores, the system will perform association clarification, that is, adopt a certain matching strategy (e.g., Hungarian algorithm, greedy matching, or threshold-based decision) to determine the optimal pairing relationship. Each successfully associated visual target and UWB point will be merged into a preliminary fusion object to obtain a preliminary fusion object list, where each object contains preliminary information from both visual and UWB modalities, such as its visual ID, UWB ID, bounding box in the image, three-dimensional spatial coordinates, and timestamp.
[0026] Taking the solution in this application as an example, in practical applications, for instance, the visual processing branch can use pre-trained deep learning models such as YOLOv7 or MaskR-CNN to detect targets of people and equipment, and combine them with multi-target tracking algorithms such as ByteTrack to maintain the target ID. The UWB processing branch may use a Kalman filter to smooth and filter the raw 3D localization data of UWB to reduce the impact of instantaneous drift. In cross-modal projection, functions provided by computer vision libraries such as OpenCV can be used to perform 3D to 2D perspective projection, utilizing calibrated camera intrinsic and extrinsic parameter matrices. For candidate associations driven by spatiotemporal proximity, a time difference threshold (e.g., 0.1 seconds) and a pixel distance threshold can be set. For example, if the Euclidean distance between the projected pixel coordinates of a UWB point and the center point of the video target bounding box is less than a certain preset value, or if the point falls inside its bounding box, then the association is considered. In the affinity assessment and association clarification stage, the Received Strength Indication (RSSI) of the UWB signal can be further considered as a reliability indicator, as well as visual features such as the target's size change and color histogram similarity in the image. The comprehensive affinity can be calculated by weighted summation and other methods, and finally, strategies such as nearest neighbor matching can be used to complete the initial fusion.
[0027] Specifically, in step S3, based on BIM semantic data, a dynamic confidence assessment based on scene constraints is performed on each preliminary fusion object in the preliminary fusion object list to obtain a confidence score. It should be understood that in a building construction safety monitoring scenario, the preliminary fusion object list may contain target entities formed by associating visual and UWB data. However, due to sensor noise, environmental occlusion, or dynamic interference, some association results may have the risk of mismatch or insufficient confidence. Therefore, to further enhance the reliability and robustness of the fusion results, in the technical solution of this application, a dynamic confidence assessment based on scene constraints is performed on each preliminary fusion object in the preliminary fusion object list based on BIM semantic data. This utilizes rich scene context information to conduct in-depth logical verification and authenticity assessment of these preliminary fusion objects. Here, performing dynamic confidence assessment based on BIM semantic data (including building structure, entity geometric attributes, and spatial constraint information) can quantify the physical rationality and scene consistency of the fusion objects, thereby selecting highly reliable targets and providing data quality assurance for subsequent adaptive fusion decisions, avoiding false alarms.
[0028] In specific implementation, firstly, the detection score, size factor, image quality factor, and occlusion factor of each preliminary fusion object in the preliminary fusion object list are calculated. The detection score refers to the confidence score given by the target detection model, representing the certainty of target recognition; the size factor reflects the size of the target in the image; for example, targets that are too small or too large may have lower reliability; the image quality factor is used to evaluate the image's sharpness, brightness, contrast, etc., measuring the input quality of the image data; the occlusion factor is used to evaluate the degree of target occlusion; targets that are partially or completely occluded generally have lower reliability. In a specific example, the detection score... The raw confidence score (range 0~1) output by the object detection model; size factor The ratio between the area of the target detection box and the total image area can be calculated. If the area ratio is less than a threshold, the target is considered too small, and the factor value is reduced proportionally. Additionally, there is the image quality factor. The occlusion factor can be calculated based on image sharpness and lighting conditions, using ambiguity metrics (such as Laplacian variance) and brightness histogram distribution. If the ambiguity exceeds a threshold or the brightness deviates from the standard range, the factor value decreases linearly. Subsequently, the occlusion factor can be calculated based on the proportion of occluded pixels within the target detection bounding box. This can be expressed as a formula: , in, Indicates the occlusion factor. Indicates the number of pixels that are obscured. This indicates the total number of pixels in the detection box.
[0029] Next, based on the detection factor, size factor, image quality factor, and occlusion factor, the visual confidence score of each preliminary fusion object is calculated. That is, by weighting or logically combining the above multiple visually relevant factors, a unified visual dimension confidence score is obtained to measure the reliability of the preliminary fusion object provided by the visual data. In a specific example, the visual confidence score of each preliminary fusion object can be calculated using the following formula: , in, This represents the visual confidence score.
[0030] Furthermore, based on BIM semantic data, signal quality assessment, volume collision detection, gravity and surface constraint detection, and kinematic consistency detection are performed on each preliminary fusion object in the preliminary fusion object list to obtain a UWB confidence score. Specifically, firstly, an initial penalty factor is set to 1; that is, at the start of the evaluation, the UWB confidence is at its highest level, and subsequent scene inconsistencies will be reflected by reducing the penalty factor. Next, based on the signal quality indicators of the projected UWB points in each preliminary fusion object, a basic confidence score is calculated. This part focuses on the intrinsic quality of the UWB signal itself, such as Received Signal Strength (RSSI), Signal-to-Noise Ratio (SNR), and positioning error accuracy, which can reflect the original reliability of the UWB positioning data. In a specific example, the signal quality indicators (such as SNR and multipath error) of the projected UWB points in the preliminary fusion object can be extracted and converted into a basic confidence score according to a preset mapping table. (Example: SNR > 30dB) =0.9; SNR < 10dB =0.3). Furthermore, in the BIM semantic model, it is queried whether the projected UWB point is located within any geometry marked as an entity. If so, the penalty factor is determined to be 0.1 times the initial penalty factor. Here, the BIM model precisely defines the geometric volume of entities such as walls, beams, columns, and equipment in the building structure. If a UWB location point (representing a person or equipment) is detected as being located inside these impenetrable entities, it indicates a serious deviation or error in the UWB location data. In this case, the penalty factor will significantly decrease to 0.1, indicating a substantial reduction in its UWB confidence. In a specific example, this process can be expressed by the formula: , in, This indicates the initialization penalty factor. Punishment factor; Secondly, in the BIM semantic model, the system queries for walkable surfaces located directly above the projected UWB point. Gravity and surface constraint checks are performed on these walkable surfaces based on their Z-coordinates. If found, the penalty factor is determined to be 0.2 times the initial penalty factor. In a specific example of this application, the penalty factor can be determined as 0.2 times the initial penalty factor if the difference between the Z-coordinate of the walkable surface and the ground Z-coordinate is calculated. Alternatively, the penalty factor can be determined as 0.2 times the initial penalty factor if the difference between the Z-coordinate of the walkable surface and the ground Z-coordinate is calculated. If the difference is less than zero, the penalty factor is also determined to be 0.2 times the initial penalty factor. This means that if the UWB positioning point is located in the air (above a walkable surface directly below it exceeding the maximum human height, e.g., a person may not be able to stand 3 meters above the ground) or below the ground (i.e., the Z-coordinate is less than zero, e.g., underground, but BIM does not provide underground space information), these situations are also considered unreasonable, and the UWB confidence level will be penalized accordingly.
[0031] In the technical solution of this application, firstly, the difference between the Z-coordinate of the surface marked as passable and the Z-coordinate of the ground is calculated. This process is specifically expressed by the formula: , in, This represents the Z-coordinate of the passable surface. Represents the Z-coordinate of the ground; Next, if any of the following conditions are met, the penalty factor P is updated, and this process is specifically expressed by the formula: like Maximum human height: ; like : ; Then, based on the projected UWB point and the projected UWB point of the previous timestamp, the instantaneous velocity is calculated. If the instantaneous velocity is greater than the sprint limit, the penalty factor is set to 0.3 times the initial penalty factor; that is, the target's movement speed is calculated based on continuous UWB positioning points. If the calculated instantaneous velocity exceeds the reasonable movement limit achievable by humans or specific equipment (e.g., faster than the human sprint limit), the positioning data is considered to have a jump or anomaly, thus reducing its confidence level.
[0032] Subsequently, the penalty factor is multiplied by the baseline confidence level to obtain the UWB confidence score. In other words, the baseline confidence level is corrected using the penalty factor generated from the preceding series of comparisons with the BIM model and physical common sense, thus obtaining the final UWB confidence score, which reflects the reliability of the UWB positioning data within the scene context.
[0033] Ultimately, the visual confidence score and the UWB confidence score constitute the overall confidence score. In a specific example of this application, this can be achieved through a weighted average or fusion function, combining the confidence scores from the visual modality and the UWB modality to form a comprehensive and unified confidence score, which represents the system's final credibility assessment of the initial fused object state.
[0034] Specifically, in step S4, based on the confidence score, adaptive fusion decision-making and state optimization are performed on the preliminary fusion object list to obtain optimized object states. It should be understood that while the aforementioned preliminary fusion and confidence assessment provide multimodal information and reliability judgments for the objects, the results may still be probabilistic or contain residual uncertainties. For example, a preliminary fusion object may have high confidence in the visual modality but low confidence in the UWB modality (and vice versa), or in some extreme cases, its total confidence score may be in a fuzzy intermediate range. To extract a unique, deterministic object state that best reflects the true situation from this information with varying degrees of reliability, and to further eliminate potential noise and jitter, thereby providing stable and reliable input for the accurate judgment of the subsequent rule engine, the technical solution of this application performs adaptive fusion decision-making and state optimization on the preliminary fusion object list based on the confidence score to transform multi-source, multi-level information into a unified and high-confidence entity representation, ensuring that the system can make intelligent and robust judgments in complex and ever-changing environments.
[0035] In practice, the initial fusion object list undergoes adaptive fusion decision-making. Specifically, for objects with extremely high confidence scores, the system directly adopts their current state as the optimized object state, or performs only slight smoothing. For objects with moderate confidence scores, the system may employ more complex fusion algorithms, such as weighted average, Kalman filtering, or particle filtering, where weights or filtering parameters are dynamically adjusted based on the confidence score. This means that if the visual confidence score is high while the UWB confidence score is low, the visual data will be given greater weight during fusion, and vice versa. Furthermore, for objects with confidence scores below a certain preset threshold, or those deemed extremely unreasonable under BIM scenario constraints (with extremely low penalty factors), the system may make a decision to ignore them to prevent erroneous or unreliable information from being passed on. State optimization further ensures the accuracy and stability of the final object state. This process includes refined estimation of attributes such as the object's position (e.g., eliminating jumps in UWB positioning and jitter in visual tracking through filtering algorithms), posture, and velocity. For example, during long-term tracking, motion models can be used to predict and correct the trajectory of an object to compensate for the impact of short-term data quality degradation and ensure a smooth transition of its state. The result of this process is that each initially fused object is refined into an optimized object state, containing the most reliable and accurate attribute information for that object at the current moment, such as its unique ID, its 3D coordinates in the real world, its possible category (person, equipment, etc.), and other key state parameters derived through fusion and optimization.
[0036] Taking the solution in this application as an example, in practical applications, adaptive fusion decision-making can be achieved by establishing a decision tree or neural network model. The input is the visual confidence score, UWB confidence score, and other relevant indicators, and the output is the specific fusion strategy for the object (e.g., using only visual information, using only UWB information, fusion by weight, or discarding). During the state optimization process, appropriate Kalman filter variants (such as extended Kalman filter or unscented Kalman filter) can be selected for different types of objects (such as personnel, specific equipment) and their motion characteristics. The process noise and measurement noise covariance can be dynamically adjusted according to the confidence score, thereby achieving smoothing and prediction of the object's position, velocity, and even attitude. For example, when the UWB confidence is high, the measurement noise covariance of the filter can be set to be smaller, trusting the UWB positioning data more; while when the video confidence is high, more reliance can be placed on visual tracking data, and the bounding box information it provides can be used to assist in positioning accuracy. Through iterative optimization, it is ensured that each entity can be stably and accurately tracked and identified in complex construction environments.
[0037] Specifically, in step S5, the optimized object state is input into the rule engine to obtain alarm events. It should be understood that the core of construction safety lies in the timely detection of potential hazards and the issuance of early warnings, thereby taking intervention measures to prevent accidents. Simply obtaining the optimized object state is insufficient to realize the ultimate value of safety monitoring. As an intelligent decision-making component, the rule engine can intelligently judge the real-time updated optimized object state based on a series of rules such as preset safety specifications, operating procedures, hazardous area definitions, and abnormal behavior patterns. Therefore, in the technical solution of this application, the optimized object state is input into the rule engine to obtain alarm events. In this process, the rule engine combines objective on-site status information with subjective safety management requirements, thereby automatically identifying any violations of safety regulations, personnel entering hazardous areas, abnormal operating conditions of equipment, etc., and triggering alarms in a timely manner. This achieves a leap from passive monitoring to proactive early warning, maximizing the protection of the lives and property of construction personnel and improving the inherent safety level of the construction site.
[0038] In practice, the optimized state of each object (including its unique identifier, precise 3D position, type attributes, motion state, and other most reliable information) is input into a pre-configured rule engine. Rules within the rule engine are typically represented in IF-THEN format. Each "IF" condition corresponds to a judgment logic for a potential hazard or unsafe situation, while the "THEN" part defines the alarm action or message to be triggered when the condition is met. For example, a rule might include: Area intrusion rule: IF "Person A" enters "Danger Area B" THEN trigger "Area Intrusion Alarm"; Fall from Height Risk Rule: IF “Construction Worker” is on “High-Altitude Platform C” and “Hard Helmet Wearing” or “Safety Belt Fastening” is not detected THEN “High-Altitude Operation Violation Alarm” is triggered; Abnormal equipment operation rules: IF "Spatial conflict" occurs between the "boom range" of "tower crane D" and the "position" of "ground worker E" THEN trigger "equipment collision warning"; Fatigue driving rule: IF “Driver F”’s “behavioral characteristics” are in an “inefficient” or “inattentive” state for a continuous long period of time (e.g., more than 4 hours) THEN trigger “Fatigue driving warning”.
[0039] This means the rule engine evaluates these rules in real-time and cyclically. Whenever a new optimized object state is received, it checks whether these states satisfy the "IF" condition of any rule. If the state of one or more objects meets the triggering condition of a preset rule, the rule engine generates a corresponding alarm event according to the "THEN" action defined by that rule. This alarm event can include, but is not limited to: sending audible and visual alarms, sending SMS or app notifications to management personnel, recording event logs, triggering video recording, and linking on-site emergency devices. In this way, the system achieves automated and intelligent identification and response to various hazardous situations at the construction site, transforming traditional manual inspections and post-event handling into real-time early warning and intervention.
[0040] Taking the solution in this application as an example, the rule engine can be implemented based on an open-source rule engine framework (such as Drools or OpenL). Rule definitions can be configured through a graphical interface or a domain-specific language (DSL), allowing safety managers to easily adjust and extend them according to actual needs. Alarm events can be linked with on-site audible and visual alarms, LED displays, and the backend operation and maintenance platform. For example, when the rule engine determines that a construction worker wearing a UWB tag has entered a "restricted area" defined in the BIM model, it immediately triggers an alarm: a voice alarm is played on speakers near the restricted area, a camera immediately zooms in and records the area, and a real-time notification is pushed to the safety manager's mobile app. The notification includes the violator's ID, the time of the violation, the location of the violation, and on-site photos / video clips, ensuring that information is quickly delivered to the relevant personnel for timely intervention to prevent accidents.
[0041] In summary, the construction safety monitoring method based on multi-source sensor fusion according to the embodiments of this application is explained. First, it precisely aligns the original video stream and UWB positioning data packets using spatiotemporal references to address the inherent asynchronicity and conflict issues of heterogeneous sensor data, ensuring the accuracy of subsequent data fusion. Then, it performs parallel feature extraction and preliminary entity association on the aligned data, and innovatively utilizes BIM semantic data to conduct dynamic confidence assessment of the preliminary fusion objects based on scene constraints. This fully leverages prior knowledge of the building environment to dynamically optimize the reliability of sensor data, thereby compensating for the shortcomings of single sensors being susceptible to environmental influences, occlusion, or signal interference leading to information loss. Finally, based on the obtained confidence scores, the system performs adaptive fusion decisions and state optimization, outputting high-confidence object states. In this way, the accuracy, real-time performance, and reliability of construction safety monitoring are significantly improved.
[0042] Furthermore, a construction safety monitoring system based on multi-source sensor fusion is also provided.
[0043] Figure 3 This is a block diagram of a building construction safety monitoring system based on multi-source sensor fusion according to an embodiment of this application. Figure 3 As shown, the construction safety monitoring system 300 based on multi-source sensor fusion according to an embodiment of this application includes: a spatiotemporal reference alignment module 310, used to perform spatiotemporal reference alignment on the acquired original video stream and original UWB positioning data packet to obtain aligned image data and aligned UWB data; a preliminary data fusion module 320, used to perform parallel feature extraction and preliminary entity association on the aligned image data and aligned UWB data to obtain a preliminary fusion object list; a dynamic confidence evaluation module 330, used to perform dynamic confidence evaluation on each preliminary fusion object in the preliminary fusion object list based on scene constraints based on BIM semantic data to obtain a confidence score; an adaptive optimization module 340, used to perform adaptive fusion decision and state optimization on the preliminary fusion object list based on the confidence score to obtain an optimized object state; and an alarm module 350, used to input the optimized object state into a rule engine to obtain an alarm event.
[0044] As described above, the construction safety monitoring system 300 based on multi-source sensor fusion according to the embodiments of this application can be implemented in various wireless terminals, such as servers with construction safety monitoring algorithms based on multi-source sensor fusion. In one possible implementation, the construction safety monitoring system 300 based on multi-source sensor fusion according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the construction safety monitoring system 300 based on multi-source sensor fusion can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the construction safety monitoring system 300 based on multi-source sensor fusion can also be one of many hardware modules of the wireless terminal.
[0045] Alternatively, in another example, the construction safety monitoring system 300 based on multi-source sensor fusion and the wireless terminal can also be separate devices, and the construction safety monitoring system 300 based on multi-source sensor fusion can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.
[0046] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for monitoring building construction safety based on multi-source sensor fusion, characterized in that, include: The acquired raw video stream and raw UWB positioning data packet are spatiotemporally aligned to obtain aligned image data and aligned UWB data; Parallel feature extraction and preliminary entity association are performed on aligned image data and aligned UWB data to obtain a preliminary list of fused objects; Based on BIM semantic data, a dynamic confidence assessment based on scene constraints is performed on each preliminary fusion object in the preliminary fusion object list to obtain a confidence score. Based on the confidence score, adaptive fusion decision and state optimization are performed on the preliminary fusion object list to obtain the optimized object state. The optimized object state is input into the rule engine to obtain alarm events.
2. The construction safety monitoring method based on multi-source sensor fusion according to claim 1, characterized in that, The aligned image data includes a timestamp, frame ID, and image; the aligned UWB data includes a timestamp, target ID, and three-dimensional coordinates.
3. The construction safety monitoring method based on multi-source sensor fusion according to claim 1, characterized in that, Parallel feature extraction and preliminary entity association are performed on the aligned image data and aligned UWB data to obtain a preliminary list of fused objects, including: The aligned image data is input into the visual processing branch to obtain the set of tracked video targets; The aligned UWB data is input into the UWB processing branch to obtain a structured UWB location point data queue. A cross-modal coordinate space projection based on camera parameter calibration is performed on the structured UWB positioning point data queue to obtain the projected UWB point set; Spatiotemporal proximity-driven candidate association pair generation is performed on the set of tracked video targets and the set of projected UWB points to obtain a candidate association pair list; Affinity assessment and association clarification are performed on each candidate association pair in the candidate association pair list to obtain the preliminary fusion object list.
4. The construction safety monitoring method based on multi-source sensor fusion according to claim 3, characterized in that, The vision processing branch includes a trained object detection model and a multi-object tracker.
5. The construction safety monitoring method based on multi-source sensor fusion according to claim 3, characterized in that, Spatiotemporal proximity-driven candidate association pair generation is performed on the set of tracked video targets and the set of projected UWB points to obtain a candidate association pair list, including: Determine whether the difference between the timestamp of each projected UWB point in the projected UWB point set and the timestamp of the corresponding tracked video target in the tracked video target set is within a preset range to obtain a first determination result; Determine whether the pixel coordinates of each projected UWB point in the projected UWB point set are within the detection box of the corresponding tracked video target in the tracked video target set to obtain a second determination result; If both the first and second judgment results are true, the corresponding projected UWB point and the tracked video target are identified as candidate association pairs.
6. The construction safety monitoring method based on multi-source sensor fusion according to claim 1, characterized in that, Based on BIM semantic data, a dynamic confidence assessment based on scene constraints is performed on each preliminary fusion object in the preliminary fusion object list to obtain a confidence score, including: Calculate the detection factor, size factor, image quality factor, and occlusion factor for each preliminary fusion object in the preliminary fusion object list; Based on the detection factor, size factor, image quality factor, and occlusion factor, the visual confidence score of each preliminary fusion object is calculated; Based on BIM semantic data, signal quality assessment, volume collision detection, gravity and surface constraint detection, and kinematic consistency detection are performed on each preliminary fusion object in the preliminary fusion object list to obtain a UWB confidence score; and The confidence score is composed of the visual confidence score and the UWB confidence score.
7. The construction safety monitoring method based on multi-source sensor fusion according to claim 1, characterized in that, Based on BIM semantic data, signal quality assessment, volume collision detection, gravity and surface constraint detection, and kinematic consistency detection are performed on each preliminary fusion object in the preliminary fusion object list to obtain a UWB confidence score, including: The initial penalty factor is set to 1; Based on the signal quality index of the projected UWB points in each of the preliminary fusion objects, the basic confidence level is calculated. In the BIM semantic model, it is queried whether the projected UWB point is located within any geometry marked as an entity. If so, the penalty factor is determined to be 0.1 times the initial penalty factor. In the BIM semantic model, the surface marked as accessible located directly above the projected UWB point is queried, and gravity and surface constraint detection is performed on the surface marked as accessible based on its Z coordinate. If so, the penalty factor is determined to be 0.2 times the initial penalty factor. Based on the projected UWB point and the projected UWB point of the previous timestamp, the instantaneous velocity is calculated. If the instantaneous velocity is greater than the sprint limit speed, the penalty factor is determined to be 0.3 times the initial penalty factor. The penalty factor is multiplied by the base confidence level to obtain the UWB confidence score.
8. The construction safety monitoring method based on multi-source sensor fusion according to claim 7, characterized in that, Gravity and surface constraint detection is performed based on the Z-coordinate of the marked accessible surface, including: Calculate the difference between the Z-coordinate of the marked accessible surface and the Z-coordinate of the ground. If this difference is greater than the maximum human height, determine the penalty factor to be 0.2 times the initial penalty factor; or Calculate the difference between the Z-coordinate of the surface marked as passable and the Z-coordinate of the ground. If the difference is less than zero, determine the penalty factor to be 0.2 times the initial penalty factor.
9. A construction safety monitoring system based on multi-source sensor fusion, characterized in that, include: The spatiotemporal reference alignment module is used to perform spatiotemporal reference alignment on the acquired raw video stream and raw UWB positioning data packets to obtain aligned image data and aligned UWB data. The preliminary data fusion module is used to perform parallel feature extraction and preliminary entity association on aligned image data and aligned UWB data to obtain a preliminary fusion object list. The dynamic confidence assessment module is used to perform dynamic confidence assessment on each preliminary fusion object in the preliminary fusion object list based on scene constraints, based on BIM semantic data, to obtain a confidence score. An adaptive optimization module is used to perform adaptive fusion decision-making and state optimization on the preliminary fusion object list based on the confidence score to obtain the optimized object state. The alarm module is used to input the optimized object status into the rule engine to obtain alarm events.
Citation Information
Cited By
Building construction site comprehensive intelligent management method and system and storage medium
CN121998243A