Multi-modal inspection data intelligent analysis and closed-loop management system for construction sites

By constructing a multimodal data intelligent analysis and closed-loop management system, the problems of difficulty in integrating multimodal inspection data and lack of management closed loop at construction sites have been solved. It has realized the unified representation of heterogeneous data and cross-modal collaborative reasoning, thereby improving the management efficiency and risk warning capabilities of construction sites.

CN121303862BActive Publication Date: 2026-04-10SICHUAN INSITITUTE OF BUILDING RES
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN INSITITUTE OF BUILDING RES
Filing Date
2025-12-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The existing system is unable to achieve deep correlation and collaborative reasoning of cross-modal features due to difficulties in integrating multimodal inspection data at construction sites, insufficient analysis capabilities, and the lack of management closed loop. As a result, inspection problems rely on manual processing, with delayed response and low rectification efficiency.

Method used

A multimodal data intelligent analysis and closed-loop management system is constructed. Through multimodal data acquisition module, data fusion and feature extraction module, intelligent analysis and risk assessment module, decision generation and task distribution module, and closed-loop tracking and verification module, it realizes unified representation of heterogeneous data, cross-modal collaborative reasoning and full-process automated management.

Benefits of technology

It enables intelligent analysis and cross-modal collaborative reasoning of multimodal inspection data at construction sites, improving management efficiency and risk warning capabilities, ensuring the real-time nature of management decisions and the effectiveness of rectification, and eliminating false rectification and inadequate rectification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303862B_ABST
    Figure CN121303862B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal inspection data intelligent analysis and closed-loop management system for construction site, belong to building construction management technical field.The system includes multi-modal data acquisition, fusion and feature extraction, intelligent analysis, decision distribution and closed-loop verification module.System to the visual, sensing and positioning heterogeneous data collected in space-time alignment, unified representation is generated using cross-modal attention fusion;Through joint inference, modal contradiction is resolved and risk is evaluated, and disposal tasks are automatically generated;After task completion, data review process is automatically started, and improvement degree is calculated by comparing data before and after rectification to verify the effect, and unqualified is automatically rolled back.The application effectively solves the problem of heterogeneous data integration, realizes the whole-process closed loop from hazard identification to automatic verification, and significantly improves the risk response speed and management efficiency of construction site.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of building construction management, and particularly relates to a multi-modal inspection data intelligent analysis and closed-loop management system for a construction site. BACKGROUND

[0002] In the field of building and engineering management, inspection and safety management of the construction site is the core link to ensure the quality of the project and the safety of the construction. The traditional inspection method mainly relies on manual recording and experience judgment, which is difficult to meet the real-time collection and intelligent analysis needs of massive and heterogeneous data in complex construction sites, and it is urgent to introduce intelligent technology to improve management efficiency and risk warning capability.

[0003] Among them, the multi-modal inspection data intelligent analysis and closed-loop management system for the construction site aims to integrate visual, sensor, positioning and other data sources to build an intelligent management platform integrating data collection, analysis, decision-making and feedback. The system is committed to realizing the automation of the inspection process, the precision of risk identification and the intelligence of management decision-making, so as to effectively improve the overall management level of the construction site.

[0004] In the prior art, the collection and fusion of multi-modal data face significant challenges: the format difference and semantic gap of heterogeneous data sources (such as video images, sensor readings, location information) lead to difficulty in information integration, and traditional analysis methods are difficult to realize deep correlation and collaborative reasoning of cross-modal features. In addition, most of the existing systems focus on data collection and display, lack real-time intelligent analysis and risk warning capability of inspection data, and cannot effectively support the closed-loop execution of management decision-making. The problems found in the inspection process often rely on manual reporting and processing, resulting in delayed response and low processing efficiency, and it is difficult to track and evaluate the whole process of rectification. Therefore, to build an efficient system that can realize intelligent analysis of multi-modal inspection data and support management closed loop has become a technical problem to be solved in the intelligent management of construction site. SUMMARY

[0005] The technical problem to be solved by the present application is to overcome the defects of difficult integration of multi-modal inspection data, insufficient analysis capability and lack of management closed loop in the prior art, and to provide a multi-modal inspection data intelligent analysis and closed-loop management system for a construction site. The system realizes intelligent analysis and cross-modal collaborative reasoning of inspection data by building a unified multi-modal data representation framework and deep correlation analysis mechanism, and establishes a whole-process automated management closed loop from problem discovery to rectification verification.

[0006] The technical scheme of the present application is a multi-modal inspection data intelligent analysis and closed-loop management system for construction sites, comprising a multi-modal data acquisition module, a data fusion and feature extraction module, an intelligent analysis and risk assessment module, a decision generation and task distribution module, and a closed-loop tracking and verification module, which work cooperatively to realize the whole-process closed-loop management from hazard identification to rectification verification:

[0007] The multi-modal data acquisition module is used to capture heterogeneous data of the construction site in real time, including video stream data obtained by the visual acquisition unit, environmental monitoring data obtained by the distributed sensor network, and personnel and equipment trajectory data obtained by the fusion positioning technology.

[0008] The data fusion and feature extraction module is used to obtain the heterogeneous data, and map the data of each acquisition unit to the same high-precision time axis through hardware clock synchronization and software timestamp correction, complete three-dimensional spatial position mapping based on a unified coordinate system, and then use a multi-path parallel neural network architecture to extract object contour and spatial relationship features from visual data, time-frequency domain features from sensor data, and mobile mode features from positioning data. Then, the cross-modal attention fusion layer is used to calculate the correlation weight of different modal features in the current context environment, and based on the weight, each modal feature is adaptively weighted and fused to generate a unified feature representation with consistent semantics.

[0009] The intelligent analysis and risk assessment module is used to input the unified feature representation into the multi-modal joint reasoning unit, complement and resolve contradictions of cross-modal information through a multi-layer perception machine model, output a comprehensive state description vector, and output an evaluation result containing risk type and level in combination with a preset rule base and a machine learning model.

[0010] The decision generation and task distribution module is used to automatically generate disposal tasks and assign them to the responsible person according to the evaluation result.

[0011] The closed-loop tracking and verification module is used to automatically start the multi-modal data review process after the responsible person feeds back the completion of the task. The process instructs the multi-modal data acquisition module to reacquire the latest data of the same area, compares the multi-modal data before and after the task is generated, quantitatively calculates the improvement degree of each key indicator, and only when the improvement degree reaches the preset threshold, the rectification is determined to be qualified, otherwise the task is automatically rolled back to the decision generation and task distribution module for secondary processing.

[0012] Further, the specific hardware configuration and parameter settings of each acquisition unit in the multi-modal data acquisition module are as follows:

[0013] The visual acquisition unit is composed of a high-definition network camera array deployed in key areas of the construction site (it should be noted that in addition to this positioning arrangement of camera array, mobile video acquisition equipment can be added to obtain video stream data), each camera is configured to acquire video stream data with a resolution of 1080P or above at a constant rate of 30 frames per second or higher, the key areas include aerial work platforms, material storage areas, deep foundation pit edges and construction passages; the camera is internally integrated with an automatic exposure control algorithm and a wide dynamic range processing chip, configured to dynamically adjust the shutter speed and gain according to the scene average brightness and the brightness of the specific attention area, and to expand the dynamic range by synthesizing multiple frames of images with different exposure to adapt to the lighting changes from indoor shadows to outdoor strong light in the construction site, the specific attention area includes the region of interest containing workers, mechanical equipment or construction materials;

[0014] The sensor network is composed of distributed Internet of Things nodes deployed in a mesh topology, the node spacing is set between 10 meters and 50 meters, each node is integrated with a three-axis vibration sensor, an A-weighted noise sensor and a laser scattering dust concentration sensor; the sampling frequency of all sensors is uniformly configured to 100 Hz, the three-axis vibration sensor has a range of ±2g and a resolution of 0.001g, the A-weighted noise sensor has a measurement range of 30 decibels to 130 decibels, and the laser scattering dust concentration sensor has a measurement accuracy of ±10% reading;

[0015] The positioning data is obtained by a hybrid positioning scheme that combines the Beidou satellite navigation system and the ultra-wideband positioning technology; in outdoor areas where satellite signals can be received, centimeter-level positioning data is obtained using the carrier phase difference technology of the Beidou satellite navigation system; in indoor or underground areas where satellite signals are blocked, the transmission time of extremely narrow pulse signals between the ultra-wideband base station network deployed on site and the tags worn by personnel is used to calculate the three-dimensional coordinates by combining the triangulation method and the Kalman filtering algorithm, and the update frequency of the positioning data is configured to 1 to 10 times per second.

[0016] Further, the specific process of mapping each acquisition unit data to the same high-precision time axis and completing the three-dimensional spatial position mapping by the data fusion and feature extraction module includes:

[0017] In terms of time alignment, the system configures the network time protocol to keep all cameras, Internet of Things nodes and positioning tags synchronized with the system master clock, and controls the local clock error within milliseconds; at the same time, each data packet carries a hardware timestamp of the acquisition time during data transmission; the strategy for software timestamp correction is to compensate the received data timestamp by calculating the fixed delay and jitter in the data transmission link, and map the data of different acquisition units to the globally unified high-precision time axis;

[0018] In terms of spatial alignment, the system pre-establishes a unified coordinate system with a fixed reference point in the construction site as the origin; for visual data, the image pixel coordinates are inversely projected into the unified coordinate system using camera calibration parameters and known poses; for sensor data, the spatial attributes are associated based on the known physical installation positions of the Internet of Things nodes; for positioning data, the three-dimensional coordinates of the target in the unified coordinate system are directly obtained, thereby ensuring that all modal data are correlated in the same space-time framework.

[0019] Further, the specific construction of the multi-path parallel neural network architecture includes:

[0020] The visual feature extraction branch adopts a pre-trained deep convolutional neural network as the backbone network, and the backbone network selects ResNet50 or EfficientNet architecture; after the input image is normalized, the features from low-level edges, corner points to middle-level textures, components, and then to high-level object contours are extracted through convolution layers, and the dimensions are reduced through pooling layers, and finally a fixed-length visual feature vector is output by the fully connected layer or the global average pooling layer, which encodes the semantic information of whether there is a person without a safety helmet, the edge protection state, and the material stacking condition in the image;

[0021] The sensor feature extraction branch adopts a time series convolution network, which is configured with causal convolution and dilated convolution layers for processing the time series segment after space-time alignment, and the length of the segment is a 5-second sequence containing 500 data points; the time series convolution network extracts the time domain fluctuations and frequency domain periodicity features of the signal through multi-layer convolution operations, identifies abnormal vibration, continuous high noise or dust surge patterns, and finally outputs a fixed-dimension sensor feature vector through a global maximum pooling layer;

[0022] The positioning feature extraction branch adopts a graph neural network, which discretizes the construction site into a graph structure, where the nodes represent key position areas and the edges represent movement paths; the positioning feature extraction branch aggregates the information of the node itself and its neighbor nodes through the message passing mechanism, learns the deep patterns contained in the movement trajectory sequence of personnel and equipment, including gathering behavior, abnormal movement and dangerous area residence time, and outputs a positioning feature vector that encodes behavior patterns and spatial use characteristics.

[0023] Further, the cross-modal attention fusion layer receives the visual feature vector, the sensor feature vector and the positioning feature vector as input, and performs the following attention weight calculation process:

[0024] For the i-th modal feature vector , its attention score is calculated, and the calculation formula is:

[0025]

[0026] wherein, and are learnable weight matrices and bias vectors, respectively, is a learnable weight vector, is a hyperbolic tangent activation function, i.e. the calculated attention weight of the i-th modality; the weights of all modalities are normalized by the Softmax function, ensuring that their sum is 1; finally, the unified feature representation after weighted fusion is obtained by weighted summation: .

[0027] Further, the specific structure and functional logic of the multi-modal joint inference unit are as follows: the multi-modal joint inference unit is composed of a plurality of stacked fully connected layers and a nonlinear activation function; the first fully connected layer is configured with 512 neurons, which is used to map the input unified feature representation to a high-dimensional hidden space to learn the interaction relationship between modalities; the second fully connected layer is configured with 256 neurons, which is used to further refine and abstract feature information; the third fully connected layer is configured with 128 neurons, which is used to output the comprehensive state description vector;

[0028] The multi-modal joint inference unit is configured to perform cross-modal contradiction resolution: when the visual feature shows that the area is safe but the positioning feature shows that there is an abnormal gathering of people, or the visual feature represents water accumulation and the vibration feature shows that the equipment is abnormal, the inference unit integrates the weights and feature values of each modality, infers the potential risks that have not been captured by a single modality based on the trained multi-layer perceptron model, and outputs a comprehensive state description vector that accurately reflects the actual situation on the scene.

[0029] Further, the intelligent analysis and risk assessment module includes a risk quantification unit, which contains a risk assessment rule base and a machine learning model;

[0030] The risk assessment rule base contains a number of (e.g. more than 200) logical rules developed by experts, covering safety protection, construction quality and environmental monitoring dimensions. The rule engine traverses the rule base, matches the comprehensive state description vector with the rule conditions, and calculates the rule-based risk base score. The machine learning model uses the gradient boosting decision tree algorithm, takes the comprehensive state description vector as input, and is trained using a large number of data sets containing historical accident records and expert annotations. By integrating the prediction results of multiple weak decision trees, it outputs a data-driven risk probability prediction value;

[0031] The risk quantification unit can weight and fuse the risk base score and the risk probability prediction value, output a final risk level score, and the score range is 0-100 points; the system preset threshold interval divides the score into three levels: 0-30 points for low risk, 31-70 points for medium risk, and 71-100 points for high risk; and finally outputs an evaluation result containing risk type description, quantitative score and belonging level interval.

[0032] Further, the decision generation and task distribution module includes a decision reasoning engine and a task dispatcher.

[0033] The decision reasoning engine adopts a hybrid model of decision tree and case reasoning; for a risk scenario that can be matched to a predefined processing scheme template through the decision tree, the risk type and the on-site context information are judged from the root node according to the tree structure until the leaf node matches the predefined processing scheme template; for the case that fails to match the predefined processing scheme template through the decision tree, the feature similarity between the current scenario and the cases in the historical case library is calculated, the most similar case is retrieved and its processing scheme is adapted;

[0034] The task dispatcher adopts the Hungarian algorithm for optimal allocation of tasks and personnel; the Hungarian algorithm is based on a responsibility area mapping table and a personnel skill matrix, and comprehensively considers the personnel skill qualification matching degree, the distance between the current location and the task site, and the current task load, to construct a cost matrix and calculate the allocation scheme that makes the total task completion time shortest or the total load most balanced;

[0035] The generated task instructions are sent in a structured message form, and the content explicitly includes: operation steps that need to be executed immediately, a specific list of tools and materials required, a mandatory completion deadline set based on the risk level, and explicit acceptance criteria.

[0036] Further, the task tracking method in the closed-loop tracking and verification module includes:

[0037] A dynamic task status board is established, and the "received", "in progress", "completed" or "obstacle encountered" status feedback by the responsible person is received through a mobile terminal;

[0038] The real-time position data obtained by the positioning acquisition unit is combined with electronic fence technology to automatically determine whether the responsible person has actually arrived at the task site;

[0039] A countdown is set according to the completion deadline in the task instruction, and a reminder message is automatically sent when the remaining time is below a preset threshold; if the task is overdue, a hierarchical early warning mechanism is triggered, and an alarm is automatically sent to the direct supervisor and higher-level management personnel of the responsible person in stages until the task status is updated.

[0040] Further, the rectification effect verification and system evolution mechanism in the closed-loop tracking and verification module comprises:

[0041] The multi-modal data review process is to call an index calculation algorithm: for visual data, image difference analysis or target detection algorithm is used to check whether the hidden danger target disappears; for sensor data, the value is compared to see whether it falls below the preset standard threshold value, which is the upper limit of the value pre-configured in the system according to the construction safety or environmental monitoring requirements; for positioning data, the analyst analyzes whether the time of staying in the dangerous area is reduced;

[0042] The improvement degree is equal to the absolute value of the difference between the index value before rectification and the index value after rectification, divided by the absolute value of the difference between the index value before rectification and the ideal target value; the improvement degree threshold for rectification qualification is set to 85%, and if the improvement degree of all key indicators reaches the threshold, it is determined to be qualified;

[0043] The system is also configured with a knowledge updating and optimization function, that is, decision cases and rectification verification results generated during system operation are continuously collected, an incremental learning algorithm is used to update the gradient boosting decision tree model in the risk quantification unit and the decision template library, and the model optimization effect is verified through an A / B test framework.

[0044] Compared with the prior art, the present application has the following advantages:

[0045] (1) Through the strategy of combining hardware clock synchronization and software time stamp correction, and the construction of a unified coordinate system, the visual, sensing and positioning three heterogeneous data are precisely aligned in millisecond time and centimeter space, effectively overcoming the "cause and effect inversion" or "correlation error" problem caused by the asynchronous data of traditional systems (for example, the vibration error at A is incorrectly associated with the image at B). Combined with Beidou / UWB (Ultra Wide Band) hybrid positioning and high-frequency sensor (100Hz), the robustness and accuracy of data sensing in complex environments on the construction site (such as shielding, light changes) are ensured.

[0046] (2) Breakthrough cross-modal deep semantic understanding and contradiction resolution capability. The present application introduces a deep analysis architecture containing a cross-modal attention fusion layer and a multi-modal joint reasoning unit. The attention mechanism can automatically adjust the weight of each modality according to the current scene (for example, automatically increase the sensor weight when the line of sight is blocked), significantly reducing the false alarm or false alarm caused by single modality interference (such as dust blocking the lens), combined with visual CNN, time series TCN and spatial GNN, it can identify complex hidden dangers that cannot be detected by a single modality (such as: abnormal aggregation of personnel trajectory + weak vibration = potential collapse precursor), intelligently process data conflicts between different sensors, and output a comprehensive state judgment that is more consistent with the objective facts on the scene.

[0047] (3) The application adopts a double risk assessment mechanism of "rule + machine learning", and utilizes the Hungarian algorithm for task allocation. The rigid bottom line of the expert rule is retained, and the implicit rules in the data are mined by using the gradient boosting decision tree, so that the fine risk quantification of 0-100 points is realized. The Hungarian algorithm comprehensively considers the skills, distance and load, realizes the global optimal matching of personnel and tasks, and greatly shortens the response time from finding problems to personnel on site.

[0048] (4) The application constructs a unique closed-loop system of "automatic trigger review + quantitative improvement degree calculation + rollback mechanism". The traditional mode of relying on manual photo uploading of rectification results is completely changed, the system automatically calls the on-site equipment to reacquire data and calculates the improvement degree (which needs to be > 85%), and the phenomenon of "false rectification" or "rectification not in place" is eliminated. Combined with the electronic fence and the hierarchical early warning mechanism, it is ensured that the responsible person must be truly on duty and complete the task on time, and the rigid execution of the construction site safety management is realized.

[0049] (5) The application integrates the incremental learning and A / B test framework. The system can learn from historical disposal cases and review data as the construction period advances, automatically update the risk model and decision template, so that it becomes more and more suitable for the environmental characteristics and work habits of a specific construction site, and solves the decline problem of the traditional static system "getting worse and worse".

[0050] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the following embodiments of the present application are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some embodiments of the present application, and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0052] Figure 1 is the overall technical scheme architecture schematic diagram of the multi-modal inspection data intelligent analysis and closed-loop management system for construction site proposed by the present application;

[0053] Figure 2 is the core principle framework schematic diagram of the multi-modal data fusion and feature extraction module in the present application;

[0054] Figure 3 is the multi-level interaction relationship and data flow schematic diagram of the intelligent analysis and risk assessment module in the present application;

[0055] Figure 4is the full-process automatic management closed-loop logic flow framework from problem discovery to rectification verification in the present application. DETAILED DESCRIPTION

[0056] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments of the present application.

[0057] Please refer to the accompanying Figure 1 The embodiments of the present application specifically describe a specific technical implementation of a multi-modal inspection data intelligent analysis and closed-loop management system for construction sites. The system aims to realize the collection, fusion, analysis, decision-making and closed-loop tracking of multi-source heterogeneous data in construction sites through highly integrated modular design. Its core architecture includes a multi-modal data collection module, a data fusion and feature extraction module, an intelligent analysis and risk assessment module, a decision-making and task distribution module and a closed-loop tracking and verification module. The modules work together to form a complete automatic management closed loop from data perception to action verification.

[0058] The multi-modal data collection module is the basis for the system to perceive the physical world. The module specifically includes a visual collection unit, a sensor network unit and a positioning collection unit. The visual collection unit is composed of high-definition network camera arrays deployed in key areas of the construction site, such as aerial work platforms, material storage areas, deep foundation pit edges and construction passages (it should be noted that in addition to this kind of positioning arrangement of camera arrays, mobile video collection devices can also be added to obtain video stream data, such as inspection personnel wearing mobile video collection devices to collect the required video). These cameras collect 1080P or higher resolution video stream data at a constant rate of 30 frames per second. To cope with the complex and variable lighting conditions in the construction site from dawn to dusk, from indoor shadows to outdoor strong light, each camera is equipped with advanced automatic exposure control algorithms and wide dynamic range processing chips. Automatic exposure control can dynamically adjust shutter speed and gain according to scene average brightness and brightness of specific areas of interest, ensuring that images are not overexposed or underexposed, and specific areas of interest include workers, machinery or construction materials of interest. Wide dynamic range processing effectively expands the dynamic range of images by capturing multiple frames of images with different exposure times in the same scene and synthesizing them, so that both dark and bright details can be clearly presented in environments with strong contrast between light and dark. The cameras are connected to the system core through gigabit Ethernet or 5G wireless networks, and the transmission uses H.265 encoding format to save bandwidth, and is accompanied by millisecond-level timestamp information.

[0059] The sensor network unit is composed of a large number of distributed Internet of Things sensor nodes. These nodes are usually deployed in a mesh topology throughout the construction area, with a node spacing of 10 to 50 meters according to the monitoring accuracy requirements. Each Internet of Things node is a ruggedized enclosure device that integrates multiple sensors, including a three-axis vibration sensor, an A-weighted noise sensor, and a laser scattering dust concentration sensor. The sampling frequency of all sensors is uniformly configured at 100 Hz to ensure that the collected data has a consistent time basis. The three-axis vibration sensor is used to monitor structural vibrations and equipment operation vibrations, with a range of ±2g and a resolution of 0.001g. The A-weighted noise sensor is used to monitor environmental noise levels, with a measurement range of 30 to 130 decibels, accurately reflecting the impact of construction activities on the acoustic environment. The laser scattering dust concentration sensor is used to monitor the concentration of PM2.5 and PM10 in the air, with a measurement accuracy of ±10% reading. Each sensor node has a built-in microprocessor for preliminary data filtering and analog-to-digital conversion, and sends data packets to the aggregation gateway through the LoRaWAN or NB-IoT low-power wide-area network protocol. The data packet contains the node's unique identifier, sensor type, measurement value, battery voltage, and precise timestamp.

[0060] The positioning acquisition unit adopts a hybrid positioning scheme that combines the Beidou satellite navigation system and ultra-wideband technology to achieve centimeter-level accurate positioning of personnel and important mobile equipment such as tower cranes, pump trucks, excavators, etc. For outdoor areas that can receive satellite signals, the Beidou satellite navigation system is mainly relied on. By receiving signals from multiple satellites, the positioning accuracy is improved from meter-level to centimeter-level using carrier phase differential technology. For indoor and underground areas where satellite signals are blocked, an ultra-wideband positioning base station network is deployed in the construction site. Personnel wear ultra-wideband tags, and equipment is installed with ultra-wideband modules. These tags and modules measure accurate distances by transmitting and receiving nanosecond-level extremely narrow pulse signals with the base stations. The positioning engine server fuses the absolute coordinates of Beidou and the relative distance information of ultra-wideband, and through the triangulation algorithm and Kalman filtering algorithm, it calculates the three-dimensional coordinates of the target in real time. The positioning data is updated at a frequency of 1 to 10 times per second and transmitted in real time to the system server through the wireless local area network in the construction site. The data fusion and feature extraction module is responsible for standardizing and deeply processing heterogeneous data from the above three acquisition units. This module further includes a time-space alignment submodule and a feature extraction submodule. The primary task of the time-space alignment submodule is to solve the inconsistency of multi-source data in time and space. In terms of time alignment, the system adopts a combination of hardware clock synchronization and software timestamp correction. All acquisition devices, including cameras, Internet of Things nodes, and positioning tags, are synchronized with the system master clock through the Network Time Protocol before deployment, controlling the local clock error within milliseconds. During data stream transmission, each data packet carries a hardware timestamp of its acquisition time. Software timestamp correction compensates for the network transmission delay by calculating the fixed delay and jitter in the data transmission link, and finally maps all data to a global, high-precision time axis. In terms of spatial alignment, the system establishes a unified construction site coordinate system. The origin of this coordinate system is usually set at a fixed reference point in the construction site, such as a corner point on the general plan. The physical positions of all acquisition devices are measured by precise measuring instruments such as total stations during deployment and recorded in the unified coordinate system. The pixel positions in visual data can be back-projected into the world coordinate system through camera calibration parameters combined with known camera poses. The positions of sensor nodes are known, and their monitoring data naturally have spatial attributes. Positioning data directly provides the three-dimensional coordinates of targets in the unified coordinate system. Through this mechanism, any modality of data collected at any time can be correlated and queried in a unified time-space framework.

[0061] The feature extraction submodule is one of the core technical components of the system. Its core principle framework is described in the patent application entitled "A Method for Real-Time Construction Site Monitoring and Control Based on Multi-Modal Data Fusion" filed on the same day. Figure 2The sub-module adopts a multi-path parallel deep neural network architecture, which learns features from visual data, sensor data and positioning data respectively, and finally generates a unified high-level feature representation through cross-modal fusion mechanism. The visual feature extraction branch is designed for processing high-definition video stream. It adopts a pre-trained deep convolutional neural network as the backbone network, such as ResNet50 or EfficientNet. The single frame image input into the network is first normalized to the range of 0 to 1, and then passes through a series of convolutional layers, pooling layers and activation function layers. The convolutional layer extracts hierarchical features from low-level edges, corner points, to mid-level textures, components, to high-level object contours, spatial relationships by sliding its convolutional kernel on the image. The pooling layer is used to reduce the spatial dimension of the feature map, increase the translation invariance of the feature and reduce the computational amount. Finally, the fully connected layer or global average pooling layer of the network compresses the extracted features into a fixed length feature vector, which concisely expresses the visual semantic information in the image, such as whether there is a person without safety helmet, whether the edge protection is set, whether the material stacking is neat, etc.

[0062] The sensor feature extraction branch is designed for processing time series sensor data. Since vibration, noise and dust data are all continuous signals that change over time, this branch adopts a time series convolutional network as the core model. The time series convolutional network can effectively capture the long-term dependencies in time series data while maintaining computational efficiency by using causal convolution and dilated convolution techniques. The input to this branch is the time series segment uploaded by each sensor node after spatio-temporal alignment, such as a 5-second sequence containing 500 data points. The sequence is first standardized by subtracting the mean and dividing by the standard deviation. Then, the multi-layer convolution operation of the time series convolutional network gradually extracts the time and frequency domain features of the signal. Low-level convolution may capture local fluctuations and periodicity of the signal, and high-level convolution can identify more complex patterns, such as abnormal vibration events, persistent high noise periods or sudden increase in dust concentration. Finally, through a global max pooling layer or attention pooling layer, the features of the entire time series are summarized into a fixed dimension feature vector, which represents the comprehensive state of the environment monitored by the sensor in that time period.

[0063] The positioning feature extraction branch focuses on analyzing the movement trajectories and spatial distribution patterns of personnel and equipment. This branch employs a graph neural network for processing. First, the construction site is discretized into a graph structure, where nodes represent key location points or areas, and edges represent connectivity or movement paths between locations. The continuous positioning data of each personnel or equipment is transformed into a sequence of movement trajectories on this graph. The graph neural network learns the deep patterns embedded in the trajectories, such as personnel gathering behavior, equipment abnormal movement, and personnel residence time in dangerous areas, through a message passing mechanism that aggregates information from the node itself and its neighbor nodes. Finally, for each tracked target, the graph neural network outputs a feature vector encoding its behavior patterns and spatial usage characteristics.

[0064] After extracting the feature vectors of each modality in the three branches, the cross-modal attention fusion layer of the feature extraction submodule begins to work. This layer receives the visual feature vector, sensor feature vector, and positioning feature vector as input. The core idea is to use the attention mechanism to dynamically calculate the importance weights of different modal features in a specific context. Specifically, for each modality's feature, the fusion layer calculates an attention score that reflects the relevance of that modality's feature to the overall scene state that needs to be understood. For example, when evaluating the safety risk of an area, if the visual feature shows that there is personnel activity in the area, and the positioning feature shows that the personnel is approaching a dangerous source, then the attention scores of the visual and positioning features will be higher; conversely, if the sensor feature shows that the vibration in the area is normal, then its attention score may be relatively lower. The calculation of attention scores is usually done through a small neural network that takes the concatenation vector of all modal features as input and outputs the weight of each modality. The calculation formula is as follows:

[0065]

[0066] where, represents the feature vector of the i-th modality, and are learnable weight matrices and bias vectors, is a learnable weight vector, is the hyperbolic tangent activation function, is the attention weight of the i-th modality calculated. All modal weights are normalized by the Softmax function to ensure that their sum is 1. Finally, the unified feature representation is obtained by weighted summation: This unified feature representation integrates multi-modal information, has stronger semantic consistency and representation ability, and lays a solid foundation for subsequent intelligent analysis.

[0067] The intelligent analysis and risk assessment module receives the unified feature representation from the feature extraction sub-module and performs deep analysis to identify risks and assess levels. The module includes a multi-modal joint inference unit and a risk quantification unit. The core task of the multi-modal joint inference unit is to perform deeper joint understanding and inference on the fused features, achieving complementarity and contradiction resolution of cross-modal information. This unit is usually implemented by a multi-layer perceptron model. The multi-layer perceptron is stacked by multiple fully connected layers and nonlinear activation functions such as ReLU functions. Unified feature representation As input, it first enters the first fully connected layer, which has 512 neurons, which maps the input features to a higher-dimensional hidden space where the complex interaction between different modal features is learned. Then, a nonlinear transformation is introduced through the ReLU activation function. Then, the features pass through the second fully connected layer, which has 256 neurons, to further refine and abstract information. There may be a third fully connected layer with 128 neurons for more detailed inference. Through this multi-layer nonlinear transformation, the multi-layer perceptron can learn very complex mapping functions, for example, when visual features represent the possibility of water accumulation and vibration sensor features show that the area has abnormal device vibration, the joint inference unit can integrate these two pieces of information to infer that there is a risk of foundation instability or leakage, which may not be accurately determined by a single modality. Contradiction resolution is also achieved in this process, for example, if a safe area is visually identified, but positioning data shows that people frequently gather abnormally, the joint inference unit will tend to believe that there may be potential risk points that have not been captured by vision, thus outputting a comprehensive state description vector that needs further attention.

[0068] The risk quantification unit performs accurate risk level determination based on the comprehensive state description vector output by the multi-modal joint inference unit. This unit adopts a dual determination mechanism combining a pre-set risk assessment rule base with a machine learning model. The risk assessment rule base is a set of logical rules developed by domain experts, including safety engineers, quality supervisors, etc. These rules cover multiple dimensions such as safety protection (e.g. high-altitude work safety belt wearing, edge hole protection), construction quality (e.g. concrete pouring quality, steel bar binding spacing), and environmental monitoring (e.g. dust control, noise pollution). Rules usually exist in the form of "if conditions then conclusions", for example, "if visual recognition shows that personnel are not wearing safety helmets and positioning shows that the personnel are in a high-altitude work area, then the safety risk level increases by 30 points". The rule engine will traverse all relevant rules, match the information in the comprehensive state description vector with the rule conditions, and accumulate a rule-based risk base score.

[0069] Meanwhile, the machine learning model works in parallel. The system adopts gradient boosting decision tree algorithm as the core prediction model. The model is trained on a large amount of historical data, including records of various accidents that have occurred in the past, hazard reports, and expert risk level annotations for a large number of scenarios. Gradient boosting decision trees iteratively train a series of weak decision trees, each tree attempting to correct the prediction error of the previous tree, and finally combining the prediction results of these trees to form a powerful ensemble model. The model takes the comprehensive state description vector as input and directly outputs a data-driven risk probability prediction value. The risk quantification unit weights and fuses the risk base calculated by the rule engine and the risk probability value predicted by the gradient boosting decision tree model. The weight can be adjusted according to the actual situation, for example, in a data-rich scenario, more reliance can be placed on model prediction; in a special scenario lacking historical data, more reliance can be placed on expert rules. The final result after fusion is quantified as a risk level score between 0 and 100 points. The system has preset clear threshold intervals, for example, 0-30 points for low risk, 31-70 points for medium risk, and 71-100 points for high risk. The risk quantification unit finally outputs the risk type description of each identified risk, such as "insufficient protection for high-altitude work", as well as the quantified risk level score and the corresponding level interval. Please refer to the attached Figure 3 The figure clearly shows the multi-level interaction relationship and data flow from multi-modal data input, through feature extraction and joint reasoning to risk quantification.

[0070] The decision generation and task distribution module automatically generates response strategies and assigns execution based on the risk information output by the intelligent analysis and risk assessment module. The module includes a decision reasoning engine and a task dispatcher. The decision reasoning engine has a hybrid model combining decision trees and case reasoning. The decision tree model is suitable for handling decision paths with clear structure and explicit logic (i.e., able to match to a pre-defined handling scheme template through decision trees). It judges from the root node according to the input risk type, such as safety, quality, risk level, low, medium, high, and on-site context information, such as weather conditions, work stages, and involved trades, and follows different branch paths downward until the leaf node, which corresponds to a pre-defined handling scheme template. The case reasoning model is suitable for handling more complex and novel situations that lack clear rules (i.e., unable to match to a pre-defined handling scheme template through decision trees). It retrieves and matches the current risk scenario with similar cases in the historical case library. The case library stores a large number of past successfully handled hazard cases, each containing problem description, handling measures, and final effect. The decision reasoning engine calculates the feature similarity between the current scenario and the historical cases, finds the most similar cases, and then adapts and modifies the handling schemes of these cases to generate customized schemes suitable for the current scenario.

[0071] Regardless of whether it is through decision trees or case-based reasoning, a detailed task instruction is ultimately generated. This instruction usually contains the following key parts: specific operation steps, which clearly tell the responsible person what to do, such as "stop work immediately and install missing protective barriers"; required resources, which list the tools, materials, or personnel cooperation needed to perform the task, such as "5 sets of protective barriers, 2 wrenches, and 2 workers"; completion time limit, which sets a clear completion time according to the risk level and task complexity, such as "high-risk tasks must be completed within 2 hours"; and acceptance criteria, which describe the effects that need to be achieved after the task is completed, such as "protective barriers are installed firmly and the height meets the 1.2-meter standard".

[0072] The task dispatcher is responsible for distributing the generated task instructions to the most suitable responsible person. It maintains two core mapping tables: the responsibility area mapping table and the personnel skills matrix. The responsibility area mapping table defines the main responsible units or individuals for different geographical or functional areas of the construction site. The personnel skills matrix records the skills and qualifications, current task load, and location of each site personnel, including managers and workers. After receiving the task instruction, the task dispatcher first filters out the possible responsible units or personnel according to the location of the risk occurrence, combined with the responsibility area mapping table. Then, combined with the skills required for the task, such as electrician, welder, and scaffolder, it is matched in the personnel skills matrix. In order to select the best from among the multiple qualified candidates, the task dispatcher uses the Hungarian algorithm, a classic combinatorial optimization algorithm that can solve the optimal allocation problem between tasks and personnel in polynomial time. The optimization goal is usually the shortest total task completion time, the highest total task execution efficiency, or the most balanced total task load. The algorithm considers factors such as skill matching degree, distance between current location and task location, and current task load to calculate an optimal allocation scheme. Once the allocation is determined, the task dispatcher immediately sends the task instruction in the form of a structured message to the corresponding responsible person's mobile intelligent terminal, such as a smartphone or tablet computer, through an integrated message push interface. The application program will remind the responsible person in the form of pop-up windows, sounds, and vibrations, and allow the responsible person to confirm receipt, report progress, or feedback problems.

[0073] The closed-loop tracking and verification module ensures that the problems discovered can be thoroughly solved. The module includes a task status monitoring submodule and a rectification verification submodule. The task status monitoring submodule establishes a dynamic task status board. The person in charge reports the task status through the mobile terminal, including "received", "in progress", "completed" or "encountered obstacles". At the same time, the system obtains the location information of the person in charge in real time through the positioning acquisition unit, and combined with the electronic fence technology, it can judge whether the person in charge is at the task site. The task status monitoring submodule sets a countdown reminder according to the completion time limit set in the task instruction. When the remaining time of the task is less than the preset threshold, for example, 30 minutes remaining, the system will automatically send a gentle reminder message to the person in charge. If the task is overdue, the system will automatically trigger different levels of early warning reminders, first reminding the person in charge again, and then reporting to his direct supervisor or even higher-level management personnel, to ensure that the problem is not forgotten or delayed.

[0074] The rectification verification submodule is the final link of the management closed loop, and its logic flow is described in the attached Figure 4 When the person in charge marks the task status as "completed", the submodule does not immediately consider that the problem has been solved, but automatically starts a multi-modal data review process. The core idea of this process is: by comparing the multi-modal data of the same area or the same target before and after the task is executed, the actual effect of the rectification measures is quantitatively evaluated. The system will retrieve the relevant data of the risk point from the historical database at the time when it was identified, that is, before the task was generated for a period of time, including visual data such as the on-site photos or video clips at that time, sensor data such as the vibration, noise, dust readings at that time, and positioning data such as the personnel and equipment distribution at that time. At the same time, the system will instruct the corresponding multi-modal data acquisition module to collect the current latest data immediately after the task is completed.

[0075] The rectification verification submodule is built-in with index calculation algorithms for quantitative comparison. For visual data, image difference analysis algorithms can be used to calculate the difference between the images before and after rectification, or target detection algorithms can be run to check if the previously identified hidden hazards, such as personnel without safety helmets or disorganized material piles, have disappeared or met the specifications. For sensor data, the values before and after rectification are directly compared, such as whether the noise decibel value has decreased to below the preset standard threshold (i.e., the upper limit value pre-configured in the system according to construction safety or environmental monitoring requirements) or whether the dust concentration has significantly decreased. For positioning data, the system can analyze whether the personnel's stay time in the dangerous area has decreased or the equipment has returned to the safe operation path. The system pre-sets an improvement threshold for each index that needs to be verified, usually set at 85%. The calculation method of improvement is: improvement equals the absolute value of the difference between the index value before rectification and the index value after rectification, divided by the absolute value of the difference between the index value before rectification and the ideal target value. If the improvement of all key indicators reaches or exceeds the threshold of 85%, the rectification verification submodule determines that the task rectification is qualified, the closed-loop process ends, and the system records the case and archives it. If the improvement of any key indicator does not reach the threshold, such as visual analysis finding that the protective fence is installed but the height is insufficient, or the sensor shows that the noise is still over standard, then the rectification verification submodule will determine that the rectification is unqualified. At this time, the system will automatically return the task together with the review results to the decision generation and task distribution module, triggering the secondary processing process. The decision generation and task distribution module will generate more specific supplementary task instructions or upgrade the task level according to the new situation.

[0076] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-modal inspection data intelligent analysis and closed-loop management system for construction sites, characterized in that, The system comprises a multi-modal data acquisition module, a data fusion and feature extraction module, an intelligent analysis and risk assessment module, a decision generation and task distribution module, and a closed-loop tracking and verification module, which work together to realize the whole-process closed-loop management from hazard identification to rectification verification. The multi-modal data acquisition module is used to capture heterogeneous data of the construction site in real time, including video stream data obtained by a visual acquisition unit, environmental monitoring data obtained by a distributed sensor network, and personnel and equipment trajectory data obtained by a fusion positioning technology. The data fusion and feature extraction module is used to obtain the heterogeneous data, map the data of each acquisition unit to the same high-precision time axis through hardware clock synchronization and software timestamp correction, complete three-dimensional space position mapping based on a unified coordinate system, and then use a multi-path parallel neural network architecture to extract object contour and spatial relationship features from visual data, time-frequency domain features from sensor data, and mobile mode features from positioning data. The intelligent analysis and risk assessment module is used to input the unified feature representation into a multi-modal joint reasoning unit, complement and resolve contradictions of cross-modal information through a multi-layer perception machine model, output a comprehensive state description vector, and output an evaluation result containing risk type and level in combination with a preset rule base and a machine learning model. The decision generation and task distribution module is used to automatically generate disposal tasks and assign them to responsible persons according to the evaluation result. The closed-loop tracking and verification module is used to automatically start a multi-modal data review process after the responsible person feeds back the completion of the task. The decision generation and task distribution module comprises a decision reasoning engine and a task dispatcher. The decision reasoning engine adopts a hybrid model of decision tree and case reasoning. For a risk scene with clear structure, the root node is determined according to the risk type and the on-site context information in a tree structure until the leaf node matches the predefined processing scheme template. For complex and novel situations, the feature similarity between the current scene and the cases in the historical case library is calculated, the most similar case is retrieved, and its processing scheme is adapted. The task dispatcher uses the Hungarian algorithm to optimally allocate tasks and personnel. The Hungarian algorithm constructs a cost matrix based on a responsibility area mapping table and a personnel skill matrix, considering the matching degree of personnel skill qualification, the distance between the current location and the task location, and the current task load, to calculate the allocation scheme that minimizes the total task completion time or balances the total load. The generated task instructions are sent in the form of structured messages, and the content explicitly includes: operation steps that need to be performed immediately, specific tool and material list required, mandatory completion time limit set based on risk level, and explicit acceptance criteria; The rectification effect verification and system evolution mechanism in the closed-loop tracking and verification module includes: The multi-modal data review process calls an index calculation algorithm: for visual data, image difference analysis or object detection algorithm is used to check whether the hidden danger target disappears; for sensor data, the numerical value is compared to see whether it falls below the legal standard; for positioning data, the analyst's residence time in the danger area is analyzed to see whether it is reduced; The improvement degree is equal to the absolute value of the difference between the index value before rectification and the index value after rectification, divided by the absolute value of the difference between the index value before rectification and the ideal target value; the improvement degree threshold for rectification qualification is set to 85%, and if the improvement degree of all key indicators reaches the threshold, it is determined to be qualified; The system is also configured with knowledge updating and optimization function, that is, continuously collecting decision cases and rectification verification results generated during system operation, updating the gradient boosting decision tree model in the risk quantification unit and the decision template library using incremental learning algorithm, and verifying the model optimization effect through A / B test framework.

2. The system of claim 1, wherein, The specific hardware configuration and parameter setting of each acquisition unit in the multi-modal data acquisition module are as follows: The visual acquisition unit is composed of a high-definition network camera array deployed in key areas of the construction site; the camera is integrated with an automatic exposure control algorithm and a wide dynamic range processing chip, configured to dynamically adjust the shutter speed and gain according to the scene average brightness and the brightness of the specific attention area, and expand the dynamic range by synthesizing multiple images with different exposure times to adapt to the lighting changes from indoor shadows to outdoor strong light in the construction site; The sensor network is composed of distributed Internet of Things nodes deployed in a mesh topology, with a node spacing of 10-50 meters; each node is integrated with a three-axis vibration sensor, an A-weighted noise sensor, and a laser scattering dust concentration sensor; the sampling frequency of all sensors is uniformly configured to 100 Hz, among which the three-axis vibration sensor has a range of ±2g and a resolution of 0.001g, the A-weighted noise sensor measures a range of 30-130 decibels, and the laser scattering dust concentration sensor has a measurement accuracy of ±10% reading; The positioning data is obtained by a hybrid positioning scheme that combines the Beidou satellite navigation system and ultra-wideband positioning technology; in outdoor open areas, centimeter-level positioning data is obtained using the carrier phase difference technology of the Beidou satellite navigation system; in indoor or satellite signal blocked areas, the three-dimensional coordinates are calculated using the extremely narrow pulse signal transmission time between the ultra-wideband base station network deployed on site and the tags worn by personnel, combined with the triangulation method and Kalman filtering algorithm; the update frequency of the positioning data is configured to 1-10 times per second, with an accuracy better than 20 cm.

3. The system of claim 1, wherein, The spatio-temporal alignment strategy of the data fusion and feature extraction module during data preprocessing includes: In terms of time alignment, the system configures the network time protocol to synchronize all cameras, Internet of Things nodes and positioning tags with the system master clock, control the local clock error within milliseconds, and carry the hardware timestamp of the acquisition time in each data packet during data transmission. The software timestamp correction strategy is to fine-tune the received data timestamp by estimating the fixed delay and jitter in the data transmission link, and map the data of different acquisition units to a globally unified high-precision time axis. In terms of spatial alignment, the system pre-establishes a unified coordinate system with the fixed reference point in the construction site as the origin. For visual data, the image pixel coordinates are inversely projected into the unified coordinate system using camera calibration parameters and known poses. For sensor data, the spatial attributes are associated based on the known physical installation position of the Internet of Things nodes. For positioning data, the three-dimensional coordinates of the target in the unified coordinate system are directly obtained, thereby ensuring that all modal data are correlated in the same space-time framework.

4. The system of claim 1, wherein, The specific construction of the multi-path parallel neural network architecture includes: The visual feature extraction branch uses a pre-trained deep convolutional neural network as the backbone network, and the backbone network selects ResNet50 or EfficientNet architecture. After normalization processing, the input image is processed through convolution layers to extract features from low-level edges and corner points to medium-level textures and components, and then to high-level object contours. Through the pooling layer, the dimension is reduced, and finally a fixed-length visual feature vector is output by the fully connected layer or the global average pooling layer. This vector encodes the semantic information of whether there is a person without a safety helmet, the edge protection state, and the material stacking condition in the image. The sensor feature extraction branch uses a time series convolution network configured with causal convolution and dilated convolution layers to process the time series segment after spatio-temporal alignment. The length of the segment is a 5-second sequence containing 500 data points. The time series convolution network extracts the time domain fluctuations and frequency domain periodicity features of the signal through multiple convolution operations, identifies abnormal vibration, continuous high noise or dust surge patterns, and finally outputs a fixed-dimension sensor feature vector through a global maximum pooling layer. The positioning feature extraction branch uses a graph neural network to discretize the construction site into a graph structure, where the nodes represent key location areas and the edges represent movement paths. Through the message passing mechanism, the positioning feature extraction branch aggregates the information of the node itself and its neighbor nodes, learns the deep patterns contained in the movement trajectory sequence of personnel and equipment, including gathering behavior, abnormal movement and dangerous area residence time, and outputs a positioning feature vector that encodes behavior patterns and spatial usage characteristics.

5. The system of claim 4, wherein, The cross-modal attention fusion layer receives the visual feature vector, sensor feature vector and positioning feature vector as input, and performs the following attention weight calculation process: For the i-th modality, the feature vector , its attention score is calculated, and the calculation formula is: where, and are learnable weight matrices and bias vectors, respectively, is a learnable weight vector, is a hyperbolic tangent activation function, is the calculated attention weight of the i-th modality; the weights of all modalities are normalized by the Softmax function to ensure that their sum is 1; finally, the unified feature representation after weighted fusion is obtained by weighted summation: .

6. The system of claim 1, wherein, The specific structure and function logic of the multi-modal joint inference unit are as follows: the multi-modal joint inference unit is composed of a plurality of stacked fully connected layers and a nonlinear activation function; the first fully connected layer is configured with 512 neurons, which is used to map the input unified feature representation to a high-dimensional hidden space to learn the interaction relationship between modalities; the second fully connected layer is configured with 256 neurons, which is used to further refine and abstract feature information; the third fully connected layer is configured with 128 neurons, which is used to output the comprehensive state description vector; The multi-modal joint inference unit is configured to perform cross-modal contradiction resolution: when the visual feature shows that the area is safe but the positioning feature shows that the personnel are abnormally gathered, or the visual feature suggests that there is water and the vibration feature shows that the equipment is abnormal, the inference unit integrates the modal weight and feature value to infer the potential risk that is not captured by a single modality, and outputs a comprehensive state description vector that accurately reflects the actual situation on the scene.

7. The system of claim 1, wherein, The intelligent analysis and risk assessment module includes a risk quantification unit, which includes a risk assessment rule base and a machine learning model; The risk assessment rule base includes a number of logical rules developed by experts, covering safety protection, construction quality and environmental monitoring dimensions. The rule engine traverses the rule base, matches the comprehensive state description vector with the rule conditions, and calculates the rule-based risk base score. The machine learning model uses the gradient boosting decision tree algorithm, takes the comprehensive state description vector as input, and is trained using a large number of data sets containing historical accident records and expert annotations. By integrating the prediction results of multiple weak decision trees, the data-driven risk probability prediction value is outputted; The risk quantification unit can weight and fuse the risk base score and the risk probability prediction value to output the final risk level score, with a score range of 0 to 100 points; The system predefines a threshold interval to divide the score into three levels: 0 to 30 points for low risk, 31 to 70 points for medium risk, and 71 to 100 points for high risk. The final output includes the risk type description, quantitative score and corresponding level interval.

8. The system of claim 1, wherein, The task tracking method in the closed-loop tracking and verification module includes: Establish a dynamic task status board to receive feedback from the responsible person through a mobile terminal, including "received", "in progress", "completed" or "encountered obstacles"; Use real-time location data obtained by the positioning acquisition unit in combination with electronic fence technology to automatically determine whether the responsible person has actually arrived at the task site; Set a countdown according to the completion time limit in the task instruction, and automatically send a reminder message when the remaining time is below the preset threshold. If the task is overdue, trigger the hierarchical early warning mechanism to automatically send an alarm to the direct supervisor and higher-level management personnel of the responsible person until the task status is updated.

Citation Information

Patent Citations

  • Building construction quality safety risk management system

    CN120235455A