Multi-modal construction site safety early warning method and system based on dynamic weight distribution

Through the multi-modal construction site safety warning method with dynamic weight allocation, combined with sensors, visual and text data, the weight is dynamically adjusted, and the problem of insufficient accuracy and adaptability of existing construction site safety warning technologies is solved, and efficient and accurate identification and early warning of safety hazards is achieved.

CN120599780AActive Publication Date: 2025-09-05浙江蓝宸数联科技有限公司

Patent Information

Application Number
CN202510776251.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-05
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing construction site safety warning technology has insufficient early warning accuracy, real-timeness and adaptability, and cannot effectively correlate personnel behavior and management documents, and the fixed weight cannot dynamically adapt to the risk modal priority of different construction stages.

Method used

The multimodal construction site safety warning method is adopted with dynamic weight allocation. By deploying the multimodal data acquisition module to collect sensor, visual and text data in real time, build a multimodal correlation model, dynamically adjust the weight, perform hierarchical warnings based on abnormal detection results, and optimize the correlation relationship through the closed-loop correction model.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of safety hazard identification, can adapt to changes in the construction site environment, ensure that the warning results are highly consistent with the actual risks, reduce the missed and false alarm rates, and improve the reliability and response speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599780A_ABST
    Figure CN120599780A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of construction site safety early warning, in particular to a multi-modal construction site safety early warning method and system for dynamic weight allocation, and the method comprises the following steps: deploying a multi-modal data collection module, and collecting and transmitting construction site data to a central processing unit in real time; constructing a correlation model according to the historical accident database; carrying out preprocessing and anomaly detection on real-time data, and dynamically adjusting a multi-modal weight according to a result; inputting the adjusted weight into the model, calculating a comprehensive early warning score and performing graded early warning; and correcting the model according to the early warning and accident matching degree. According to the method, the multi-modal data acquisition module is deployed to acquire sensor, visual and text data, and a dynamic weight distribution mechanism is combined, so that the weight of each modal can be adjusted in real time according to an anomaly detection result, the accuracy and comprehensiveness of potential safety hazard identification are remarkably improved, and the problems of one-sidedness and missing report of single-modal monitoring are effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of construction site safety early warning, and in particular to a multi-modal construction site safety early warning method and system with dynamic weight distribution. Background Art

[0002] With the development of construction site informatization, information technology has penetrated into every aspect of construction site production. However, existing construction site safety early warning technologies have many problems that need to be solved, resulting in insufficient accuracy, real-time performance, and adaptability.

[0003] First, traditional methods often rely on a single data source. For example, sensor modalities only collect environmental or equipment data and are unable to associate key risk factors such as human behavior and management documents. They are prone to missing accidents caused by "unsafe human behavior" and "management defects." Visual modalities rely solely on camera monitoring and lack linkage analysis with equipment operating parameters and environmental data, resulting in a high false alarm rate. At the same time, text data such as safety inspection records and construction plans are not structured and cannot be mapped to dynamic construction scenarios in real time, forming "data islands."

[0004] Secondly, the existing multimodal fusion technology uses fixed weights and does not take into account the real-time abnormal dynamic impact and differences in construction stages. For example, when the sensor value suddenly changes or high-risk behavior is detected by vision, the corresponding modal weight is not increased in a targeted manner, resulting in the weakening of key risk signals. The risk modal priorities in different construction stages are different, and the static weights cannot be dynamically adapted.

[0005] In view of this, a multimodal construction site safety early warning method and system with dynamic weight distribution is proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a multi-modal construction site safety early warning method and system with dynamic weight distribution to solve the problems of insufficient early warning accuracy, real-time performance and adaptability in existing early warning technologies.

[0007] To solve the above technical problems, the present invention provides a multi-modal construction site safety early warning method with dynamic weight distribution, comprising the following steps: S1, by deploying a multimodal data acquisition module, collects multimodal data of the construction site in real time and transmits it to the central processing unit; S2. Based on the historical accident database, a correlation model between multimodal data and accident types is constructed; S3. Preprocess and detect anomalies in the multimodal data collected in real time, and dynamically adjust the multimodal weights based on the anomaly detection results; S4. Input the adjusted multimodal weights into the association model, calculate the comprehensive warning score and trigger a graded warning; S5. Modify the association model based on the matching degree between the warning results and the actual accident.

[0008] As a further improvement of the present technical solution, in step S1, the multimodal data includes sensor modal data, visual modal data and text modal data, wherein: Sensor modal data includes dust, noise, and gas concentration data collected by environmental sensors, mechanical vibration and load data collected by equipment sensors, and personnel and equipment location data collected by positioning sensors. Visual modality data includes real-time camera video and human behavior characteristics, safety equipment wearing status, and dangerous area intrusion information extracted through target detection and posture estimation; Text modal data includes construction plans, safety inspection records, hidden danger rectification reports, and structured safety indicators extracted through natural language processing.

[0009] As a further improvement of the present technical solution, in step S2, the association model is to establish a mapping relationship table between multimodal features and accident types by statistically analyzing the sensor features, visual features, and text features corresponding to each accident type; calculate the causal relationship strength between sensor feature nodes, visual feature nodes, and text feature nodes, and construct a correlation network including multimodal nodes; The initial weight is calculated based on the triggering frequency of each modal data in historical accidents and the data quality index. The data quality index is quantified by sensor signal strength, video image clarity, and text field missing rate.

[0010] As a further improvement of the present technical solution, in step S3, the preprocessing includes: De-noise the sensor modal data and detect the short-term variation of the denoised value. If it exceeds a preset threshold value that is dynamically adjusted according to the construction stage, it is marked as an abnormal sensor feature node. The preset threshold value is obtained by matching the current construction stage with a standard library of equipment operating parameters. The standard library of equipment operating parameters generates dynamic threshold intervals based on the construction stage type, equipment model, and environmental parameters. Perform lighting processing and occlusion detection on visual modality data. When failure to wear safety equipment, intrusion into a dangerous area, or illegal operation is detected, it is marked as an abnormal visual feature node. Failure to wear safety equipment includes not wearing a hard hat, protective gloves, or a high-altitude work safety belt. Keyword extraction and timeliness evaluation are performed on text modal data. When a security risk keyword match is detected and the timeliness is lower than the preset standard, it is marked as an abnormal text feature node.

[0011] As a further improvement of the present technical solution, in step S3, dynamically adjusting the multimodal weights includes: For the marked abnormal sensor feature nodes, abnormal visual feature nodes, and abnormal text feature nodes, the weight increase is positively correlated with the value change, target danger level, and keyword matching degree, and the weight increment is transmitted to the associated sensor feature nodes, visual feature nodes, or text feature nodes according to the causal relationship strength through the correlation network; For unassociated feature nodes, the weight is reduced according to a preset decay rate, which is set according to the modality type.

[0012] As a further improvement of the present technical solution, in step S3, a real-time priority queue for weight adjustment tasks is established, and the priority is calculated based on the type of anomaly, the proportion of numerical change, and the associated accident level; the weight adjustment and conduction operations are performed in order of priority, and tasks with a priority higher than a preset threshold trigger real-time priority scheduling of computing resources.

[0013] As a further improvement of this technical solution, the priority calculation method includes: For sensor modal anomaly tasks, the priority is the weighted sum of the ratio of the value change amplitude to the sensor range and the severity level of the associated accident; For visual modality anomaly tasks, the priority is determined by the proximity of the target to the hazard area and the data acquisition time attenuation coefficient; For text modality anomaly tasks, the priority is determined by the matching degree between keywords and accident types and the timeliness of the text.

[0014] As a further improvement of the present technical solution, in step S4, the correlation network is split into logical subgraphs according to the sensor modality, visual modality, and text modality, and the weight transmission within each subgraph is processed in parallel through distributed computing nodes; for cross-modal association nodes, the weight is updated using an asynchronous message passing mechanism; Sensor modal data, visual modal data, and text modal data correspond to the output warning probabilities of sensor feature nodes, visual feature nodes, and text feature nodes respectively. A comprehensive warning score is generated by weighted average. The corresponding level warning is triggered by comparing the comprehensive warning score with the preset multi-level warning threshold.

[0015] As a further improvement of the present technical solution, in step S5, matching data between the warning results and actual accidents are collected, false alarms, missed alarms and causal modes are marked, and a corrected sample set is formed. Based on the corrected sample set, the causal relationship strength of the correctly associated nodes is enhanced, and the strength of the incorrectly associated nodes is weakened.

[0016] A multi-modal construction site safety warning system with dynamic weight distribution, which is used to implement the multi-modal construction site safety warning method with dynamic weight distribution, includes: Multimodal data acquisition module, including sensor array, camera group, and text data interface; The central processing unit is connected to the multimodal data acquisition module and includes a data preprocessing module, a correlation network construction module, a dynamic weight allocation module, a correlation warning module and a closed-loop correction module; The early warning output module is connected to the central processing unit and includes an audible and visual alarm, a mobile terminal push interface and a large visual screen.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. In this multimodal construction site safety early warning method and system with dynamic weight allocation, a correlation network is used to construct a mapping relationship between multimodal features and accident types, and each modal subgraph is processed in parallel through distributed computing nodes. At the same time, a real-time priority queue for weight adjustment tasks is established, which greatly shortens the calculation time and breaks through the bottleneck of slow response and low efficiency of traditional fixed weight early warning methods.

[0018] 2. In the multimodal construction site safety early warning method and system with dynamic weight distribution, by deploying a multimodal data acquisition module to collect three types of data: sensor, vision, and text, combined with a dynamic weight distribution mechanism, the weight of each modality can be adjusted in real time according to the abnormal detection results, significantly improving the accuracy and comprehensiveness of safety hazard identification, and effectively avoiding the one-sidedness and underreporting problems of single-modality monitoring.

[0019] 3. In the multimodal construction site safety early warning method and system with dynamic weight distribution, a correction sample set is formed by collecting matching data between early warning results and actual accidents, and the association model is corrected in a closed loop. The causal relationship strength of the correct association nodes is continuously enhanced, and the incorrect association is weakened, ensuring that the early warning system can continuously adapt to changes in the construction site environment, so that the early warning results are highly consistent with the actual safety risks, and the long-term reliability of the system is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 Schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all secondary embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0022] At present, with the development of construction site informatization, information technology has penetrated into every aspect of construction site production. However, there are many problems that need to be solved in the existing construction site safety early warning technology, resulting in insufficient accuracy, real-time performance and adaptability of early warning. For this reason, see Figure 1 As shown, one of the purposes of the present invention is to provide a multi-modal construction site safety early warning method with dynamic weight distribution, and the multi-modal construction site safety early warning method with dynamic weight distribution includes the following steps: S1, by deploying a multimodal data acquisition module, collects multimodal data of the construction site in real time and transmits it to the central processing unit; S2. Based on the historical accident database, a correlation model between multimodal data and accident types is constructed; S3. Preprocess and detect anomalies in the multimodal data collected in real time, and dynamically adjust the multimodal weights based on the anomaly detection results; S4. Input the adjusted multimodal weights into the association model, calculate the comprehensive warning score and trigger a graded warning; S5. Modify the association model based on the matching degree between the warning results and the actual accident.

[0023] In this multimodal construction site safety early warning method with dynamic weight allocation, a correlation network is used to construct a mapping relationship between multimodal features and accident types, and each modal subgraph is processed in parallel through distributed computing nodes. At the same time, a real-time priority queue for weight adjustment tasks is established, which greatly shortens the calculation time and breaks through the bottleneck of slow response and low efficiency of traditional fixed weight early warning methods. In addition, by deploying a multimodal data acquisition module to collect three types of data: sensor, vision, and text, combined with a dynamic weight allocation mechanism, the weights of each modality can be adjusted in real time according to the anomaly detection results, significantly improving the accuracy and comprehensiveness of safety hazard identification, and effectively avoiding the one-sidedness and underreporting problems of single modality monitoring. At the same time, by collecting matching data between warning results and actual accidents to form a revised sample set, the association model is closed-loop corrected, continuously enhancing the causal relationship strength of correctly associated nodes and weakening incorrect associations, ensuring that the early warning system can continuously adapt to changes in the construction site environment, making the warning results highly consistent with actual safety risks, and improving the long-term reliability of the system.

[0024] Considering that traditional construction site safety early warning technology only relies on single-modal data, such as only monitoring environmental parameters through sensors or only monitoring personnel behavior through cameras, there are problems such as one-sided monitoring dimensions, isolated information, and incomplete hidden danger identification. Specifically, sensors cannot capture behavioral risks such as "personnel not wearing safety equipment", and cameras find it difficult to correlate mechanical operating states such as "equipment load exceeding the limit" in real time. The safety hazards in text records cannot be dynamically matched with real-time scenes, resulting in a high underreporting rate and incomplete risk assessment.

[0025] Therefore, in step S1, multimodal data fusion technology is used to incorporate sensor modality, visual modality, and text modality into the collection system, and comprehensive risk perception is achieved through the following technologies: Regarding sensor modalities, environmental sensors (dust, noise, and gas concentration sensors) are deployed to collect construction environment parameters in real time. Equipment sensors (mechanical vibration and load sensors) monitor the operating status of machinery. Positioning sensors (UWB or GPS positioning) track the location of personnel and equipment. This enables real-time quantitative monitoring of "environmental risks (such as excessive dust and harmful gas leakage)", "equipment risks (such as abnormal vibration and excessive load)", and "location risks (such as personnel entering high-risk areas and equipment collision risks)". This provides basic data on the "unsafe state of objects" for the early warning system, solving the problem that traditional single sensors cannot cover multi-dimensional equipment and environmental risks. For visual modalities, the system uses real-time video streams collected by cameras and combines them with computer vision technologies such as target detection (YOLO algorithm) and pose estimation (OpenPose) to extract behavioral features such as "personnel illegal operations (such as throwing objects from high places)", "missing safety equipment (not wearing a hard hat / safety belt)", and "intrusion into dangerous areas (entering a foundation pit without permission)". This technology fills the gap in "unsafe human behavior" that traditional sensors cannot monitor. For example, when the sensor detects that the tower crane load is approaching the threshold, the vision module simultaneously verifies whether the operator is operating in accordance with regulations, avoiding accidents caused by the combined risks of "equipment abnormality + personnel violation". The system also uses pose estimation to identify the safety belt wearing status of workers working at heights, directly linking it to the risk of "falling from height", upgrading the early warning system from "object monitoring" to "human-object coordinated monitoring". For text modalities, natural language processing (NLP) technology is used to extract structured indicators such as "hazard type," "rectification time limit," and "responsible party" from unstructured texts such as construction plans, safety inspection records, and hazard rectification reports (for example, locating the hazard location through named entity recognition and extracting the rectification deadline through time expressions). This converts static management data into dynamic risk indicators. For example, if the text modality detects that "the hazard of loose scaffolding bolts in a certain area has not been rectified within 48 hours," combined with the "frequent entry and exit of personnel in the area" in real-time visual data, a higher-level warning can be triggered. By evaluating the timeliness of text data (such as hazard rectification beyond the deadline), a "historical hazard-current status" correlation analysis is formed with real-time sensor and visual data to solve the problem of "disconnection between recording and execution" in traditional management data. Through the technical integration of three types of modal data, data complementarity, full-scene coverage, and precise early warning are achieved, breaking through the bottleneck of "one-sided data and fragmented scenes" of the traditional early warning system. Through the integration of multiple technologies, "full-factor perception and full risk correlation" are achieved, providing a solid data foundation for subsequent dynamic weight allocation and intelligent early warning.

[0026] Considering that traditional construction site safety early warning models are mostly based on independent analysis of single-modal data, lack correlation modeling between multimodal features, and fail to consider differences in data quality and accident trigger frequency, fixed weight processing is still used when sensor signals are noisy, video images are blurry, or text records are missing. This leads to inaccurate mapping between features and accident types, and the model's early warning accuracy fluctuates significantly with data quality. Therefore, in step S2, multimodal correlation modeling technology is used to construct a highly interpretable and robust correlation model through statistical feature mapping, causal relationship analysis, and data quality weighting. The specific technical path and results are as follows: Traditional models rely solely on the direct "feature-accident" correlation of a single modality, ignoring the combined effects of multimodal features. For example, "abnormal equipment vibration, personnel failure to evacuate the danger zone, and overdue hazard rectification" all indicate a high-risk accident. Therefore, based on a historical accident database, we compile statistics for sensor features (such as vibration amplitude and gas concentration), visual features (such as failure to wear a safety belt and intrusion into a danger zone), and text features (such as the keyword "scaffolding hazard" and records of overdue rectification) corresponding to each accident type (such as falls from height, object strikes, and mechanical injuries). This creates a "feature-accident" mapping matrix. This matrix enables cross-indexing of multimodal features. For example, when "tower crane load sensor value exceeds limit" (a sensor feature) and "operator entering under the crane boom without wearing a safety helmet" (a visual feature) are simultaneously triggered, the mapping table quickly identifies the accident type as "mechanical injury," avoiding misjudgments based on a single feature and giving the model initial multi-dimensional risk association capabilities. Traditional models assume that each modal feature is independent and ignore the potential causal relationship between features. For example, "missing maintenance plan" in management text may lead to abnormal equipment sensor data. This defect leads to the lack of risk transmission path analysis and fails to capture the chain risk of "management defect → equipment failure → casualties". Therefore, Granger causality test or structural causal model (SCM) is used to calculate the causal relationship strength between the three types of feature nodes: sensor, vision, and text. (value range [0,1]), for example, the causal influence strength of the "dust concentration exceeds the standard" node on the "worker not wearing a dust mask" node, or the transmission probability of the "hazard rectification report missing" node to the "equipment failure" node, and construct a directed acyclic graph (DAG) containing cross-modal associations to clarify the causal transmission path of multimodal features. For example, when the system recognizes that "edge protection is not mentioned in the safety inspection record" (text feature), the correlation network will automatically enhance the association weight of "personnel fall risk area" (visual feature), enabling the model to simulate the causal chain of "management loopholes → behavioral risks → accident occurrence", improving the foresight and logical interpretability of early warnings; In addition, traditional models do not distinguish between differences in data quality and accident triggering frequency, and treat all features equally. Therefore, in step S2, the initial weights are calculated based on the triggering frequency of each modal data in historical accidents and the data quality index. Among them, the data quality index is quantified by sensor signal strength, video image clarity, and text field missing rate. For features that appear frequently in historical accidents, higher initial weights are assigned to ensure that core risk signals are not ignored; features with low data quality are automatically downgraded, such as sensor data with signal strength below the threshold, visual features corresponding to blurred video images, and text features with missing fields, to avoid model misjudgments due to low-quality data and effectively improve the reliability and robustness of the model's initial input.

[0027] Considering the potential for false causal associations between multimodal feature nodes, such as changes in ambient temperature that simultaneously affect sensor signals and human behavior, leading to the misassociation of nodes with indirect causal relationships, and the fact that traditional correlation networks are prone to computational delays due to the explosion of nodes in complex scenarios, resulting in incorrect warnings or delayed responses, we employ causal denoising technology and a hierarchical response optimization architecture when constructing correlation networks to improve network performance in terms of both association accuracy and computational efficiency. In terms of causal relationship denoising technology, first, the initially calculated causal relationship strength (such as the causal strength from node A to node B) is verified through counterfactual reasoning: assuming that node A does not experience an anomaly, does the probability of node B experiencing an anomaly decrease significantly? If the decrease is less than 10%, it is considered a false association and removed. At the same time, domain expert knowledge is introduced to construct a "causal prior rule library." For example, it is clarified that "text hidden danger rectification expires" must occur before "equipment failure," and there is no reverse causal relationship. This is used to logically verify the causal relationships automatically calculated by the algorithm and filter out associations with inconsistent temporal order. In this way, false causal edges, such as the accidental co-occurrence of low sensor battery power and illegal personnel operation, are successfully eliminated, making the causal relationships in the network more consistent with actual physical logic. For example, the one-way causal chain from "not wearing a hard hat" to "object impact injury" is strengthened, effectively avoiding weight misconduction caused by erroneous associations. For example, a construction site mistakenly associated "canteen lunch hours" with "equipment downtime." After denoising, such misleading warnings are reduced. Second, the minimum survival algorithm in graph theory is used to The Multi-Stage Tree (MST) algorithm thins out overly dense local subgraphs (e.g., fully connected nodes with more than 10 nodes), retaining core causal paths such as "excessive load → abnormal vibration → loose bolts" while removing irrelevant edges such as "excessive load → operator's mood." Furthermore, a dynamic threshold for causal relationship strength is set (initial threshold 0.4, gradually adjusted during model training), retaining only edges with strength greater than the threshold. For cross-modal association edges (e.g., sensor → vision node), the threshold is increased by 0.1, as cross-modal physical associations are more complex. This series of operations significantly reduces the number of network edges and computational complexity, while preserving true causal relationships and avoiding confusion in warning logic caused by "flooding of low-probability associations." For example, a construction site once experienced simultaneous alarms from over 30 feature nodes due to excessive associations. After pruning, this problem was completely resolved. In terms of the hierarchical response optimization architecture, the correlation network is first divided into intra-modal sub-networks and inter-modal bridge nodes. The intra-modal sub-network covers sensor, visual, and textual node associations. The sensor sub-network is grouped by equipment type (crane / elevator), the visual sub-network is partitioned by risk area (high altitude / foundation pit), and the textual sub-network is classified by hazard type (equipment / management). Each sub-network independently performs weight transfer on distributed computing nodes to achieve parallel processing. Inter-modal transfer is implemented through asynchronous message queues. Cross-modal weight transfer is only activated when an intra-modal anomaly triggers a bridge node (e.g., an anomaly in a crane load sensor triggers a risk in the "crane operation area" visual node). This avoids meaningless cross-modal calculations and effectively improves computational efficiency. Second, causal edges are labeled with accident level weights, such as a +30% weight for edges pointing to "mass casualties" and a -10% weight for edges pointing to "minor injuries," thereby constructing prioritized transfer paths. Dijkstra is used. The algorithm variant automatically plans a "high-accident-level priority transmission path" for early warning tasks. For example, when "combustible gas concentration exceeds the standard" (sensor node) is detected, it is preferentially transmitted to strongly associated high-risk nodes such as "personnel in the hot work area are not wearing protective masks" (visual node) and "firefighting equipment inspection is overdue" (text node), rather than first transmitting to low-priority nodes such as "irrelevant equipment vibration"; when simulating a gas leak scenario at a chemical construction site, the system prioritizes activating the "gas concentration → personnel protection → fire management" causal chain, quickly completing multi-modal weight aggregation and triggering a first-level early warning, responding earlier than the traditional non-priority network, gaining critical and valuable time for emergency response, and significantly improving the weight transmission efficiency of high-risk scenarios.

[0028] Considering that traditional models fail to distinguish between high-frequency key features and low-frequency interference features when processing multimodal data and lack effective filtering mechanisms for noisy data (such as sensor signal drift and garbled text formatting), they are often misled by low-quality data. For example, noise in equipment vibration signals can be mistakenly identified as a fault, leading to distorted warning results. Therefore, initial weights are assigned by statistically analyzing the frequency of occurrence of each modal feature in historical accidents. For example, if "not wearing a seatbelt" occurs in 80% of height-fall accidents, this feature is given a higher initial weight to highlight its importance. Furthermore, a data quality indicator system is constructed based on the characteristics of different modal data. For sensor modalities, signal strength is evaluated. For example, if the GPS positioning signal is weak, the weight of location data is reduced accordingly. For visual modalities, image clarity is evaluated. If video recognition accuracy decreases in low-light environments, the system dynamically reduces its weight. For text modalities, the reliability of text data is reduced based on the field missing rate. If key fields are missing in hazard reports, the credibility of the text data is reduced. This approach effectively improves the reliability of the model's initial input. For example, high-frequency visual features like "not wearing a safety belt while working at height" are assigned a high weight to ensure that core risk signals are not overlooked. Sensor data with signal strength below a threshold is automatically downgraded to avoid false alarms caused by equipment failure or signal noise. This significantly improves the model's robustness in complex scenarios with unstable data quality, reduces false alarms caused by data quality issues, and improves the accuracy and stability of the early warning system.

[0029] Taking into account the noise interference, environmental changes and data timeliness differences in the multimodal data of the construction site, for example, the sensor signal is susceptible to abnormal fluctuations caused by electromagnetic interference, the camera misses violations in low-light or occluded scenes, and the hidden danger information in the text record may lose its reference value due to overdue processing. At the same time, the traditional fixed threshold preprocessing method cannot adapt to the dynamic changes in the construction stage (such as the significant differences in equipment operating parameters between pile foundation construction and high-altitude operations), resulting in a high false alarm rate of anomaly detection and a high risk of missed detection. Therefore, in step S3, the sub-modal adaptive preprocessing technology is adopted to achieve accurate anomaly marking of multimodal data through dynamic threshold matching, environmental robustness processing and timeliness analysis. The specific technical path and effects are as follows: For sensor modal data, wavelet transform or Kalman filtering is used to reduce the noise of the sensor original signal, filter out high-frequency noise (such as electromagnetic interference) and low-frequency drift (such as sensor temperature drift); calculate the short-term change amplitude of the value based on the sliding window ( =|current value - average value of the previous 5 minutes|), suppressing occasional fluctuation interference; at the same time, building a three-dimensional standard library of "construction stage - equipment model - environmental parameters". For example: During the pile foundation construction phase, the vibration threshold of the rotary drilling rig is dynamically adjusted according to the soil hardness (obtained through environmental sensors); during high-altitude operations, the basket load threshold is combined with the wind speed (environmental parameter) and the basket model (equipment parameter) to generate a dynamic range. By matching the current construction stage (such as process information obtained by the BIM system) with equipment operating parameters in real time, the corresponding threshold interval is called; thereby improving the accuracy of abnormal marking and adapting to complex working conditions; For visual modality data, we use environmental adaptive preprocessing technology. To address environmental interference, we use histogram equalization and CycleGAN generative adversarial networks to enhance the contrast of low-light images, ensuring high differentiation between the helmet (red / yellow) and the background. To address occlusion issues, we use deformable convolution based on YOLOv8 + Deformable DETR to capture the outline of safety equipment under partial occlusion (for example, only 1 / 3 of the helmet edge is exposed). To address behavioral complexity issues, we use OpenPose to extract joint coordinates and combine it with an LSTM neural network to analyze posture sequences of more than 5 consecutive frames (for example, determining whether "height workers have not been wearing their safety belts") for posture estimation. Semantic segmentation is used to delineate ROIs (regions of interest) such as foundation pits and tower crane operating areas. Optical flow is used to track the movement of people, distinguishing between "brief passes" and "illegal stops," and detect intrusions into dangerous areas. For text modal data, the BERT-wwm pre-training model is used to identify domain-specific hidden danger words such as "fall risk" and "edge loss", and the keyword weight is calculated in combination with TF-IDF (for example, the higher the frequency of "unaccepted" in the text, the greater the risk weight), to achieve keyword extraction; through the knowledge graph, the "tower crane model QTZ80" is linked to the corresponding equipment of the sensor modality to achieve cross-modal data association of "text hidden danger-real-time equipment status"; at the same time, spaCy is used to identify time entities in the text such as "rectification deadline 2025-04-20" and "inspection date 2025-04-15" to define the timeliness coefficient ,in, is the current time, The time when hidden dangers should be rectified. is the attenuation factor), when S<0.5, it is marked as “overdue hidden danger”; Through the modal adaptive preprocessing technology, the "false elimination and true retention, dynamic adaptation, and multi-source alignment" of multimodal data are achieved. Compared with traditional fixed preprocessing methods, the accuracy of anomaly detection is comprehensively improved, laying a solid data foundation for subsequent dynamic weight allocation and comprehensive early warning, and ensuring the reliability and real-time performance of the entire early warning system from the source.

[0030] Considering that traditional multimodal early warning models use fixed weight allocation, they are unable to respond in real time to the dynamic risk differences of different modal anomalies (for example, the urgency of a sudden change in sensor values ​​is different from the risk level of a textual hidden danger exceeding its expiration date), and that unconnected redundant feature nodes introduce noise interference and consume computing resources, resulting in delayed warnings or confusion about priorities. Therefore, step S3 uses dynamic weight adaptive adjustment technology. Through the three-in-one mechanism of "anomaly-driven weight enhancement - irrelevant node attenuation - causal network transmission", it achieves precise focusing and coordinated response of multimodal risk signals. The specific technical logic and effects are as follows: The weight increase of the marked abnormal sensor feature nodes, abnormal visual feature nodes, and abnormal text feature nodes is positively correlated with the value change amplitude, target danger level, and keyword matching degree respectively; For the sensor modality, define the weight boost function ,in is the value change, is the sensor range, Sensitivity coefficient for the construction phase (e.g., load sensor for equipment during high altitude operation) Set to 1.5, and set to 1 during pile foundation construction. For example, if the value of a tower crane load sensor suddenly increases by 80% (close to the fracture threshold), the weight increase will reach 80% × 1.5 = 120%, far exceeding the weight increase of conventional abnormalities, ensuring that high-risk signals are handled first. For visual modalities, a "target hazard level matrix" is established to divide violations into different risk levels (level 5 for falling from height, level 3 for not wearing a helmet, and level 1 for not wearing gloves), and the weight increase = Level value × Time decay coefficient. For example, if the violation of "not wearing a helmet" occurs frequently recently, the Time decay coefficient will increase, further increasing the weight of this behavior; For text modalities, the keyword matching degree (value ranges from 0 to 1) is calculated through the BERT model, and the weight increase is calculated by combining the timeliness score (an additional 30% weight is added for overdue hidden dangers). = Matching degree × (1 + aging correction coefficient). For example, when the keyword "scaffolding bolts loose" is detected and the rectification period has expired, the weight increase will increase significantly. This strategy allows for precise amplification of risk signals. For example, if the vibration sensor readings of a construction site elevator suddenly spike (high risk), the weight is increased, triggering the visual module to detect whether a person is inside the elevator. Compared to a fixed-weight model, this improves the response speed to high-risk anomalies and avoids the drowning out of critical signals caused by average weight distribution. The traditional model only independently weights the single-modal abnormal nodes, ignoring the "cross-modal causal relationship" (such as "equipment load exceeding the limit" may cause "people to misjudge the safe distance and enter the danger zone"), and cannot form a "ripple effect" of risk transmission, resulting in the omission of multi-modal superposition risks; Therefore, through the correlation network (causal relationship strength ∈[0,1]), the weight increment is calculated according to the causal strength The proportion is transmitted to the related nodes; for example, the causal strength of "not wearing a high-altitude safety belt" (visual node, weight increased by 50%) with the "height fall accident" node is 0.8, and the causal strength with "safety net not set up" (text node) is 0.6, then the latter will receive a weight increase of 50% × 0.6 = 30%; Considering that the traditional model does not perform weight decay on irrelevant feature nodes, resulting in computational redundancy and random fluctuations of irrelevant features (such as irrelevant sensor data far away from the work area) that may mislead the warning results, the weights of unrelated feature nodes are reduced according to the preset decay rate, and the decay rate is set according to the modality type; for sensor modality, the unrelated nodes are decayed at 5% / minute (because the device status data is highly real-time, the value of outdated signals decreases rapidly); for visual modality, the unrelated routine behavior nodes (such as "normal walking") decay at a rate of 3% / minute, and the dangerous behavior nodes decay at a rate of 1% / minute (to retain short-term risk memory); for text modality, historical hidden danger nodes are decayed at a rate of attenuation( is the number of days until the rectification deadline). The weight of nodes that exceed the deadline by more than 30 days approaches 0. At the same time, nodes with a weight less than 0.1 after attenuation are automatically blocked and do not participate in subsequent comprehensive early warning calculations, thereby improving computing efficiency and reducing noise interference.

[0031] Considering the significant differences in risk levels and urgency of multimodal abnormal events on construction sites (e.g., a sudden sensor value spike could trigger an immediate accident, while an overdue textual hazard represents a gradual risk), traditional indiscriminate sequential processing can easily lead to "high-risk tasks being blocked by low-priority tasks," resulting in delayed or even missed critical warnings. Furthermore, computing resources cannot be allocated on demand, resulting in inefficiency and waste. Therefore, step S3 employs real-time priority scheduling technology. By building a dynamic priority queue and resource triggering mechanism, this achieves "risk-level-driven task scheduling and on-demand computing resource allocation." The specific technical approach and results are as follows: Establish a real-time priority queue for weight adjustment tasks. The priority is calculated based on the anomaly type, the percentage of value change, and the associated accident level. Weight adjustment and transmission operations are performed in priority order. Tasks with a priority higher than the preset threshold trigger real-time priority scheduling of computing resources. For sensor modes, the priority is the weighted sum of the ratio of the value change amplitude to the sensor range and the severity level of the associated accident. The calculation formula is priority , where β+γ=1, is the value change, is the range, and Sr is the severity level of the associated accident, which is 1-5. For example, if the load of the tower crane suddenly increases by 80% (close to the fracture threshold, =0.8) and is associated with a "mass casualties" accident level of 5, =0.6×0.8+0.4×5=2.48 (highest priority); For the visual modality, its priority is determined by the proximity between the target and the dangerous area and the data acquisition time attenuation coefficient. The calculation formula is priority ,in, is the distance between the target and the danger zone, Safety threshold distance is the data collection time interval, For example, a person is only 1 meter away from the edge of the foundation pit (d / D=0.2), and the latest frame detection of the real-time video stream ( =0.1), =0.8×0.9=0.72 (medium-high priority); For text mode, its priority is determined by the matching degree between keywords and accident types and the timeliness of text. The calculation formula is priority =Keyword matching degree×(1+time correction coefficient), the correction coefficient is +0.1 for every day the hidden danger exceeds the expiration date, with a maximum of +0.5; Through this priority quantification model, the risk urgency of abnormal tasks can be accurately quantified, ensuring that high-risk tasks are handled first and avoiding delays in key warnings.

[0032] Based on the above priority quantification model, a real-time priority queue is established: the priority queue is implemented using efficient data structures such as binary heaps or Fibonacci heaps, ensuring that the insertion / extraction time complexity of the highest priority task is O(logn), supporting the concurrent processing of tens of thousands of tasks; the system performs weight adjustment and transmission operations in order of priority to ensure that high-risk abnormal tasks are processed first, effectively avoiding the situation where "high-risk tasks are blocked by low-priority tasks".

[0033] There are two major problems in the traditional fixed resource allocation mode. First, high-priority tasks need to wait for the completion of the current computing tasks, resulting in delays in key warnings. For example, an urgent task where the sensor value exceeds the range by 50% may not trigger a warning in time due to waiting for text data preprocessing; second, distributed computing nodes have uneven loads, with some nodes idle and some nodes overloaded. To solve these problems, this technology uses threshold-triggered resource scheduling technology to set priority thresholds. =1.5 (presettable), when the task When a task is executed, it triggers: local resource preemption, suspending low-priority processes and releasing CPU / GPU resources to prioritize urgent tasks; distributed expansion, dynamically applying for additional computing nodes through the Kubernetes cluster to achieve cross-node parallel weight transmission. For example, the sensor subgraph and the visual subgraph are calculated synchronously on different nodes; at the same time, resource usage isolation technology is used to allocate dedicated cache space for high-priority tasks, such as partitioning the GPU video memory, to avoid bandwidth competition caused by sharing resources with regular tasks, thereby ensuring the computing resource requirements of high-priority tasks; Considering that traditional construction site safety warnings often lead to false alarms and missed alarms due to the independent judgment of a single modality (e.g., relying solely on sensor value excursions without considering human behavioral risks), and lacking refined differentiation of risk levels (e.g., uniformly triggering warnings of the same level, leading to misallocation of emergency resources), they cannot meet the actual needs of "precise warnings and graded responses." Therefore, step S4 utilizes multimodal probability fusion and graded warning technology. Through weighted aggregation of warning probabilities and intelligent matching of multi-level thresholds, this technology achieves a transition from "single indicator judgment" to "comprehensive risk assessment." The specific technical logic and effects are as follows: Traditional early warning systems directly use raw data (such as sensor values ​​and visual inspection results) for judgment, which has significant flaws: First, data is not comparable. Continuous sensor values, discrete labels from visual inspection, and matching scores from text analysis cannot be directly combined for calculations. Second, single-modality limitations exist. A single modality is prone to misjudgments due to data noise or environmental interference, such as temporary sensor failures leading to abnormal values ​​or missed violations due to camera obstructions. Third, confidence levels are missing. The reliability of the detection results is not quantitatively assessed. For example, the reliability of helmet recognition results from visual inspection in low-light environments is low, but they are still weighted equally in the judgment. Therefore, the data from each modality must first be probabilistically converted: For sensor modalities, the probability Ps of the current sensor value deviating from the normal range is calculated based on the Gaussian mixture model (GMM). For example, by analyzing the distribution of historical data, the possibility of the current tower crane load value exceeding the safe range is determined; For visual modalities, the confidence score Pv output by the target detection algorithm (such as YOLOv8) is directly used as the probability of the corresponding behavior (such as not wearing a seat belt) occurring, with a value range of 0 to 1; For text modalities, the BERT classifier is used to calculate the matching probability Pt between the hidden danger keywords in the text and the accident type, and the final probability is adjusted based on the timeliness score. Through this step, heterogeneous multimodal data are uniformly converted into probability values ​​in the range of 0-1, laying the foundation for subsequent fusion calculations; Considering that the traditional simple average weight (such as the equal weight addition of each modality) ignores the real-time risk difference, the risk contribution of each modality is different in different construction stages and scenarios. For example, the visual modality (safety belt wearing status) is more critical than the sensor modality during high-altitude operations. At the same time, the risk level of a slight abnormality of a single modality is different from that of a simultaneous abnormality of multiple modalities. However, fixed weights cannot reflect this difference. Therefore, the real-time adjustment weight output in step S3 is directly used. , , It includes abnormal node enhancement, causal conduction increment, and unrelated node attenuation), ensuring that the weight reflects the actual importance of each mode under the current working conditions; the comprehensive warning score formula is: , the formula uses normalization to ensure that changes in the sum of weights do not affect the comparability of scores, and achieves accurate aggregation of risk signals through dynamic weights. For example, when the sensor detects abnormal equipment vibration ( =0.7), and visually detected that the personnel had not evacuated the danger zone ( =0.8), the text shows that the relevant hidden dangers have not been rectified ( =0.6), the dynamic weight will strengthen the contribution of these high-risk modes, significantly increase the comprehensive score S, and accurately reflect the superimposed risks.

[0034] Traditional early warning systems use a single threshold to trigger alarms, which has obvious drawbacks. First, there is a lack of risk grading. Minor anomalies (such as dust concentration exceeding the standard by 10%) and severe anomalies (gas concentration exceeding the standard by 100%) trigger the same level of warning, which makes safety officers overwhelmed. Second, the response measures are single and cannot link different treatment plans based on risk level. For example, high-risk scenarios require immediate shutdown, while low-risk scenarios only require recording and reminder. Therefore, this embodiment has established a three-level threshold system: Level 1 warning (S≥0.8): Red warning, triggering the sound and light alarm, automatic equipment shutdown, and SMS notification to the project manager (response time <1 second); Level 2 warning (0.6≤S<0.8): Yellow warning, push APP pop-up window, voice broadcast, and arrange special inspection by security supervisor (response time <5 seconds); Level 3 warning (0.4≤S<0.6): Blue warning, the system records and archives it, and generates an inspection work order (response time <30 seconds).

[0035] At the same time, the system automatically optimizes threshold boundaries based on historical accident data and updates them quarterly to adapt to long-term changes in site personnel proficiency, equipment aging, etc.

[0036] Considering that traditional construction site safety early warning models rely on initial training data and lack a feedback mechanism for actual accidents, they are prone to problems such as "accumulation of false positives and missed negatives" and "failure to adapt to new scenarios" during long-term use (for example, risk characteristics introduced by new processes are not recognized by the model, and old related nodes still dominate the early warning logic), making it impossible to achieve closed-loop optimization of "data-model-effect". Therefore, in step S5, model self-evolution technology based on a modified sample set is adopted. Through the closed-loop process of "abnormal data annotation-causal relationship recalibration-model parameter iteration", an intelligent early warning system that can self-learn and continuously optimize is constructed. The specific technical path and results are as follows: Traditional models have significant flaws when processing actual accident data. First, false alarms are improperly handled. False alarms caused by sensor noise, for example, are not effectively labeled, causing the model to consistently misassociate between "noise signature" and "accident type." For example, sensor electromagnetic interference noise may be mistakenly associated with equipment failure. Second, the causes of missed alarms are unclear. For missed alarms, the causal modality is not traced. For example, when multimodal superposition risks are not identified, factors such as "missing management text" are not incorporated into the association model, leading to repeated missed detections of similar risks. To address these issues, a revised sample set is constructed and structured annotation is performed. Structured annotation includes: Warning result types are clearly labeled as correct warnings, false alarms (no accident triggering), and missed alarms (accidents but no warnings), clearly distinguishing the warning effects of the model; Causing modality labeling accurately determines whether it is a single modality misjudgment (such as missing a seat belt due to visual occlusion) or a multi-modal linkage failure (such as a sensor anomaly not being transmitted to the associated text node), thereby locating the root cause of the problem; The contribution of feature nodes is annotated. The SHAP value is used to analyze the impact of each feature node on the warning results, accurately locating nodes that are "over-correlated" (contribution > reasonable threshold) or "under-correlated" (contribution < reasonable threshold). For example, if the SHAP value of a feature node is too high, it means that it is over-correlated in the warning, which may lead to false positives; conversely, if the SHAP value is too low, there may be under-correlation, resulting in false negatives. Through the above annotations, the corrected sample set can comprehensively and accurately reflect the false positives and false negatives of the model and their causes, providing strong data support for subsequent model optimization.

[0037] In traditional models, causal relationships are fixed and cannot adapt to new scenarios and actual accident feedback. Therefore, the strength of causal relationships is dynamically adjusted based on the revised sample set. For the "correctly associated node" (such as "not wearing a safety belt - falling from a height"), follow the formula: , enhancing the strength of causal relationships (A is the sample contribution, obtained through SHAP value analysis, and η is the learning rate). For example, if the association between "not wearing a seatbelt" and "falling from a height" is correct and has a high contribution in the corrected sample, this formula can further strengthen the causal relationship between the two, allowing the model to pay more attention to such associations in subsequent warnings; For "error association nodes" (such as "irrelevant device vibration-object impact"), according to the formula , weakening the strength of the causal relationship (B is the number of false positives) to a minimum of 0.1. For example, if "irrelevant equipment vibration" and "object impact" are repeatedly falsely associated in the correction sample, this formula can gradually weaken the relationship between the two, preventing the model from being misled by false associations. To more effectively process the revised sample set, the corrected samples are fed into a graph neural network (GNN). Specifically, graph convolution is used to update the edge weights of the entire correlation network. Additional weight is assigned to missing edges corresponding to missed incidents, allowing them to be captured by the model in subsequent warnings. Redundant edges with a high frequency of false alarms are suppressed to reduce their impact on warning results. Furthermore, the application of GNNs enables the model to better handle the complex relationships between multimodal data. By optimizing the correlation network, the model can more quickly adapt to changes in risk characteristics brought about by new scenarios and processes, reducing the occurrence of false alarms and missed alarms.

[0038] See also Figure 2 As shown, the second object of the present invention is to provide a multi-modal construction site safety warning system with dynamic weight distribution, which is used to implement the multi-modal construction site safety warning method with dynamic weight distribution, including: Multimodal data acquisition module, including sensor array, camera group, and text data interface; The central processing unit is connected to the multimodal data acquisition module and includes a data preprocessing module, a correlation network construction module, a dynamic weight allocation module, a correlation warning module and a closed-loop correction module; The early warning output module is connected to the central processing unit and includes an audible and visual alarm, a mobile terminal push interface and a large visual screen.

[0039] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-modal construction site safety early warning method with dynamic weight distribution, characterized in that: The following steps are involved: S1, by deploying a multimodal data acquisition module, collects multimodal data of the construction site in real time and transmits it to the central processing unit; S2. Based on the historical accident database, a correlation model between multimodal data and accident types is constructed; S3. Preprocess and detect anomalies in the multimodal data collected in real time, and dynamically adjust the multimodal weights based on the anomaly detection results; S4. Input the adjusted multimodal weights into the association model, calculate the comprehensive warning score and trigger a graded warning; S5. Modify the association model based on the matching degree between the warning results and the actual accident.

2. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 1 is characterized in that: In step S1, the multimodal data includes sensor modal data, visual modal data, and text modal data, wherein: Sensor modal data includes dust, noise, and gas concentration data collected by environmental sensors, mechanical vibration and load data collected by equipment sensors, and personnel and equipment location data collected by positioning sensors. Visual modality data includes real-time camera video and human behavior characteristics, safety equipment wearing status, and dangerous area intrusion information extracted through target detection and posture estimation; Text modal data includes construction plans, safety inspection records, hidden danger rectification reports, and structured safety indicators extracted through natural language processing.

3. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 1 is characterized in that: In step S2, the association model is to establish a mapping relationship table between multimodal features and accident types by counting the sensor features, visual features and text features corresponding to each accident type; calculate the causal relationship strength between sensor feature nodes, visual feature nodes and text feature nodes, and construct a correlation network including multimodal nodes; The initial weight is calculated based on the triggering frequency of each modal data in historical accidents and the data quality index. The data quality index is quantified by sensor signal strength, video image clarity, and text field missing rate.

4. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 3 is characterized in that: In step S3, the pre-processing includes: De-noise the sensor modal data and detect the short-term variation of the denoised value. If it exceeds a preset threshold value that is dynamically adjusted according to the construction stage, it is marked as an abnormal sensor feature node. The preset threshold value is obtained by matching the current construction stage with a standard library of equipment operating parameters. The standard library of equipment operating parameters generates dynamic threshold intervals based on the construction stage type, equipment model, and environmental parameters. Perform lighting processing and occlusion detection on visual modality data. When failure to wear safety equipment, intrusion into a dangerous area, or illegal operation is detected, it is marked as an abnormal visual feature node. Failure to wear safety equipment includes not wearing a hard hat, protective gloves, or a high-altitude work safety belt. Keyword extraction and timeliness evaluation are performed on text modal data. When a security risk keyword match is detected and the timeliness is lower than the preset standard, it is marked as an abnormal text feature node.

5. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 4 is characterized in that: In step S3, dynamically adjusting the multimodal weights includes: For the marked abnormal sensor feature nodes, abnormal visual feature nodes, and abnormal text feature nodes, the weight increase is positively correlated with the value change, target danger level, and keyword matching degree, and the weight increment is transmitted to the associated sensor feature nodes, visual feature nodes, or text feature nodes according to the causal relationship strength through the correlation network; For unassociated feature nodes, the weight is reduced according to a preset decay rate, which is set according to the modality type.

6. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 1 is characterized by: In step S3, a real-time priority queue for weight adjustment tasks is established, and the priority is calculated based on the type of anomaly, the proportion of the numerical change, and the associated accident level; the weight adjustment and conduction operations are performed in order of priority, and tasks with a priority higher than a preset threshold trigger real-time priority scheduling of computing resources.

7. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 6 is characterized by: The priority calculation method includes: For sensor modal anomaly tasks, the priority is the weighted sum of the ratio of the value change amplitude to the sensor range and the severity level of the associated accident; For visual modality anomaly tasks, the priority is determined by the proximity of the target to the hazard area and the data acquisition time attenuation coefficient; For text modality anomaly tasks, the priority is determined by the matching degree between keywords and accident types and the timeliness of the text.

8. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 1 is characterized by: In step S4, the correlation network is split into logical subgraphs according to sensor modality, visual modality, and text modality, and weight propagation within each subgraph is processed in parallel through distributed computing nodes; for cross-modal association nodes, the weights are updated using an asynchronous message passing mechanism; Sensor modal data, visual modal data, and text modal data correspond to the output warning probabilities of sensor feature nodes, visual feature nodes, and text feature nodes respectively. A comprehensive warning score is generated by weighted average. The corresponding level warning is triggered by comparing the comprehensive warning score with the preset multi-level warning threshold.

9. The multi-modal construction site safety early warning method with dynamic weight distribution according to claim 1, characterized in that: In step S5, matching data between warning results and actual accidents are collected, false positives, missed negatives and causal modes are marked, and a revised sample set is formed. Based on the revised sample set, the causal relationship strength of the correct associated nodes is enhanced, and the strength of the incorrect associated nodes is weakened.

10. A multi-modal construction site safety warning system with dynamic weight distribution, wherein the multi-modal construction site safety warning system with dynamic weight distribution is used to implement the multi-modal construction site safety warning method with dynamic weight distribution according to any one of claims 1 to 9, characterized in that: include: Multimodal data acquisition module, including sensor array, camera group, and text data interface; The central processing unit is connected to the multimodal data acquisition module and includes a data preprocessing module, a correlation network construction module, a dynamic weight allocation module, a correlation warning module and a closed-loop correction module; The early warning output module is connected to the central processing unit and includes an audible and visual alarm, a mobile terminal push interface and a large visual screen.

Citation Information

Patent Citations

  • Mountainous area high pier cast-in-place safety detection method based on multi-source data fusion

    CN119202599A

  • Intelligent monitoring and early warning method and system for cofferdam construction process

    CN119250547A

  • Sensing data fusion and abnormal dynamic weight adjustment method for transformer fire

    CN119649535A

  • Open caisson construction soil gushing dynamic early warning system based on multi-parameter fusion

    CN119992809A

  • Deformation dynamic prediction and early warning method for tunnel under construction based on multi-source data fusion

    WO2025112473A1

Cited By

  • Online early warning method for necking defect of drawing forming of spherical tank connecting shell

    CN120976224A

  • A method for online early warning of necking defects of a spherical tank connecting shell drawn forming

    CN120976224B

  • Intelligent cruise ship passenger safety early warning system and method based on fusion of large model and visual identification

    CN121600595A

  • Building construction safety monitoring system with automatic detection and danger early warning functions

    CN121921946A

  • Coal mine power supply line geological disaster susceptibility evaluation factor weight dynamic analysis method fusing attention mechanism

    CN122133931A