Construction project quality risk prediction method and system based on big data and artificial intelligence

By constructing a construction project quality risk prediction system based on big data and artificial intelligence, the problems of insufficient data integration and analysis model generalization ability have been solved. It has achieved efficient integration and intelligent analysis of multi-source heterogeneous data, improved the accuracy and interpretability of risk prediction, and promoted proactive prevention of engineering quality risk management.

CN121998188APending Publication Date: 2026-05-08GUANGZHOU ZHONGTIAN ENG TESTING SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU ZHONGTIAN ENG TESTING SERVICE CO LTD
Filing Date
2026-01-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for predicting quality risks in construction projects suffer from weak data integration capabilities, limited generalization ability of analytical models, delayed risk warnings, and a lack of interpretability. They are unable to achieve systematic and forward-looking risk warnings and interventions, and cannot meet the needs of modern engineering refined management.

Method used

A construction project quality risk prediction system based on big data and artificial intelligence is constructed, including a data acquisition and integration module, a data middle platform and feature engineering module, an intelligent risk prediction model module, a risk visualization and decision support module, and a mobile terminal early warning and collaborative handling module, to achieve efficient integration, intelligent analysis and real-time early warning of multi-source heterogeneous data.

Benefits of technology

It has achieved efficient integration and governance of multi-source heterogeneous data, improved the generalization ability and prediction accuracy of risk prediction models, promoted the transformation of engineering quality risk management from passive response to proactive prevention, and provided interpretable risk warning and disposal suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998188A_ABST
    Figure CN121998188A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, discloses a construction project quality risk prediction method and system based on big data and artificial intelligence, and aims at solving the problems that a traditional method depends on manpower, information fragmentation is caused, and risk early warning is lagged. The method comprises the following steps: collecting and cleaning multi-source heterogeneous engineering data; extracting a multi-dimensional feature vector from the standardized data based on the knowledge graph; calculating the features by using a hybrid intelligent model, and outputting risk prediction probability and grade; visualizing the result and generating a structured early warning report; and a report is pushed through the mobile terminal and feedback is tracked and disposed to form closed-loop management. The system comprises a data acquisition and integration module, a data center and feature engineering module, an intelligent risk prediction model module, a risk visualization and decision support module and a mobile terminal early warning and co-processing module. Prospective and intelligent prediction and closed-loop management and control of engineering quality risks can be realized, and risk identification precision and response efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically relating to a method and system for predicting construction project quality risks based on big data and artificial intelligence. Background Technology

[0002] With the accelerated digital transformation of the construction industry, construction quality risk management has become a core element in ensuring project safety and sustainable development. Modern engineering construction involves a large amount of multi-source heterogeneous data, including construction process records, environmental monitoring information, material supply chain data, BIM model parameters, and real-time status indicators collected by various IoT sensors. This data collectively constitutes the basic resources for project quality assessment and risk identification. However, the industry currently generally lacks the ability to effectively integrate and deeply mine this data, resulting in quality risk identification still heavily relying on manual experience and periodic on-site inspections, making it difficult to achieve systematic and forward-looking risk warnings and interventions.

[0003] Among them, the construction project quality risk prediction technology based on big data and artificial intelligence aims to break down information silos throughout the entire project lifecycle by constructing a unified data platform architecture, integrating structured and unstructured data, and using machine learning and time-series modeling methods to dynamically model and predict trends of key quality indicators. This technology focuses on automatically extracting risk features from complex, high-dimensional, and time-varying engineering data, identifying potential quality hazard patterns, and thus supporting intelligent risk assessment and decision-making response mechanisms.

[0004] Existing technologies have significant shortcomings in predicting construction project quality risks: First, data integration capabilities are weak, making it difficult to efficiently aggregate heterogeneous data from multiple channels such as BIM platforms, IoT devices, and project management systems, resulting in severe information fragmentation. Second, the generalization ability of analytical models is limited; most systems still rely on rule engines or static threshold judgments, failing to adapt to the dynamic changes in different engineering scenarios. Third, risk warnings are delayed and lack interpretability, failing to establish a closed-loop logical link from raw data to risk levels and then to disposal recommendations. Finally, the lack of real-time push and collaborative disposal mechanisms for mobile devices leads to low risk response efficiency, making it difficult to meet the needs of modern refined engineering management. These problems severely restrict the ability to shift from "post-event correction" to "pre-event prevention" in engineering quality, urgently requiring an integrated risk prediction system that combines big data governance and intelligent algorithm-driven approaches to solve these problems. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and to provide a method and system for predicting construction project quality risks based on big data and artificial intelligence, which can effectively solve the problems in the background art.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A construction project quality risk prediction system based on big data and artificial intelligence, comprising the following components: The data acquisition and integration module is used to acquire raw engineering data from multiple heterogeneous data sources in real time or near real time, and to clean, convert, and standardize the raw engineering data to generate data to be analyzed in a unified format. The heterogeneous data sources include at least an Internet of Things sensor network, a building information modeling platform, a project management system, and a materials supply chain database. The data acquisition and integration module is internally deployed with a data adapter cluster, each adapter corresponding to a specific data source protocol to achieve non-intrusive data access. The data platform and feature engineering module are used to receive and store the data to be analyzed, and to build a thematic data warehouse for engineering quality risk analysis. Based on a preset engineering quality risk knowledge graph, the data platform and feature engineering module extracts multi-dimensional feature vectors from the thematic data warehouse. The feature vectors include static features, time-series features, and correlation features. The static features represent the inherent attributes of engineering entities, the time-series features represent the dynamic change process of key monitoring indicators, and the correlation features represent the spatial and logical relationships between different entities or indicators. The intelligent risk prediction model module is used to receive the multi-dimensional feature vector, construct and train a risk prediction model through ensemble learning and deep learning algorithms, and output the predicted probability and risk level for a specific quality risk event. The intelligent risk prediction model module contains multiple sub-models, which are specifically modeled for different types of quality risks. The sub-models use the gradient boosting tree algorithm to classify and predict static and related features, and use the long short-term memory network algorithm to perform trend inference and anomaly detection on time-series features. The risk visualization and decision support module is used to map the predicted probability and risk level into a visualized risk heat map, trend curve and early warning signal, and automatically generate a structured early warning report containing the risk location, cause analysis and treatment suggestions when the risk level exceeds a preset threshold; the risk visualization and decision support module provides a web-based interactive dashboard to support users to perform multi-dimensional drill-down analysis on the risk prediction results; The mobile terminal early warning and collaborative handling module is used to distribute the structured early warning report to the mobile terminals of relevant responsible personnel in real time through message push service, and to receive handling feedback information from the mobile terminals, forming a closed-loop management process of risk warning, task assignment, on-site handling, and result feedback; the mobile terminal early warning and collaborative handling module integrates geofencing and personnel positioning functions to ensure that early warning information can accurately reach the on-site responsible persons.

[0007] Preferably, the IoT sensor network in the data acquisition and integration module is deployed at key locations on the construction site, and the sensor types include strain gauges, inclinometers, temperature and humidity sensors, pressure sensors, and image acquisition devices; the data adapter cluster supports multiple data exchange protocols such as MQTT, OPC UA, RESTful API, and direct database connection.

[0008] Preferably, the thematic data warehouse constructed by the data platform and feature engineering module is organized using a star schema. The fact table records monitoring indicator values ​​with timestamps and spatial locations as dimensions, and the dimension table includes engineering stage dimension, responsible unit dimension, component type dimension, and risk type dimension. The feature engineering process specifically includes: filling missing values ​​and smoothing and denoising the original monitoring data; extracting mean, variance, slope, and periodic features from time-series data based on sliding window technology; and mining the correlation strength between entities based on the engineering quality risk knowledge graph through graph traversal algorithm and quantifying it into correlation feature values.

[0009] Preferably, the gradient boosting tree sub-model in the intelligent risk prediction model module has an objective function defined as minimizing the weighted sum of prediction error and model complexity, specifically as follows:

[0010] in, Represents the loss function; Representative sample The true risk label; Representative model for samples The predicted probability; It acts on the first Decision Tree The regularization term; This represents the total number of decision trees during the gradient boosting process.

[0011] Preferably, the long short-term memory network sub-model in the intelligent risk prediction model module is used to process continuous time-series data monitored by sensors. Its cell state update mechanism can effectively capture long-term dependencies, and the calculation formula for its gating unit is as follows:

[0012]

[0013]

[0014]

[0015]

[0016]

[0017] in, , , These represent the forget gate, input gate, and output gate, respectively. In cellular state, In hidden state, for Input characteristics at time step and For model parameters, It is the sigmoid activation function.

[0018] Preferably, the structured early warning report generated by the risk visualization and decision support module includes at least: risk event identifier, three-dimensional coordinates or two-dimensional planar icon annotation of the risk location, risk level, confidence level, key indicators that trigger the early warning and their historical change curves, the most likely cause inferred based on the knowledge graph, and a list of recommended disposal measures matched from the historical disposal case library.

[0019] This solution also provides a construction project quality risk prediction method based on big data and artificial intelligence. The specific steps of this method are as follows: S110, through a cluster of data adapters deployed in the data acquisition and integration module, collects multi-source heterogeneous raw engineering data in parallel from IoT sensor networks, building information modeling platforms, project management systems, and material supply chain databases; S120, In the data acquisition and integration module, the original engineering data is cleaned, outliers and duplicate records are removed, and data of different formats are uniformly encoded and converted in units to generate a standardized data stream to be analyzed. S130, the data stream to be analyzed is input into the data platform and feature engineering module, and stored in the corresponding fact table and dimension table of the subject data warehouse according to the preset data model. Feature engineering is performed based on the engineering quality risk knowledge graph to extract multi-dimensional feature vectors containing static features, time-series features and correlation features. S140, the multi-dimensional feature vector is input into the pre-trained intelligent risk prediction model module. The gradient boosting tree sub-model and the long short-term memory network sub-model in the model module respectively perform fusion calculation on the features, output the predicted probability for various quality risks, and classify the risk level according to the probability value. S150, In the risk visualization and decision support module, the risk level and prediction results are visualized and rendered to generate a risk heat map and trend analysis chart. When the risk level exceeds the preset threshold, an early warning is automatically triggered and a structured report containing cause analysis and treatment suggestions is generated. S160, through the mobile terminal early warning and collaborative handling module, pushes the structured early warning report to the mobile terminal of relevant responsible personnel, and tracks and records handling feedback information to complete the closed-loop process from risk identification to handling verification.

[0020] Preferably, in step S120, the cleaning process employs outlier removal based on the statistical three sigma principle and duplicate record deduplication based on timestamps and data source identifiers; the unified encoding and unit conversion are performed based on a preset engineering data standardization dictionary.

[0021] Preferably, in step S130, the extraction of time series features from time series data based on the sliding window technique specifically includes: sliding a sliding window of preset length and step size on the time series data, and calculating the mean, variance, and slope obtained by linear fitting of the data in each window.

[0022] Preferably, in step S160, the push process integrates geofencing and personnel positioning functions. When the warning involves a specific construction area, the warning information is pushed to the mobile terminal of the person in charge who is currently located within the geofence.

[0023] Compared with the prior art, the present invention has the following beneficial effects: By building a unified data platform and modular data acquisition adapters, we have achieved efficient integration and governance of multi-source heterogeneous data from IoT, BIM, project management, and other sources, fundamentally solving the problem of information fragmentation and providing a complete and consistent data foundation for high-quality risk analysis.

[0024] By adopting a hybrid modeling strategy that combines gradient boosting trees and long short-term memory networks, we can make full use of the discriminative power of static and related features, accurately capture the dynamic evolution of time series data, and significantly improve the generalization ability and prediction accuracy of risk prediction models in different engineering scenarios.

[0025] A complete technical chain has been established, from raw data to risk feature extraction, model prediction, visual early warning and mobile collaborative handling. This has enabled the interpretable presentation of risk prediction results and the intelligent generation of handling suggestions, promoting the transformation of engineering quality risk management from passive response to proactive prevention. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the overall technical solution architecture of the construction project quality risk prediction method and system based on big data and artificial intelligence proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the hybrid intelligent risk prediction model that combines gradient boosting tree and long short-term memory network in this invention; Figure 3This is a logical flowchart of data acquisition, platformization, and multi-dimensional feature engineering in this invention; Figure 4 This is a schematic diagram of the multi-level interaction relationship and data flow of risk visualization, decision support and mobile terminal collaborative processing in this invention; Figure 5 This is a flowchart illustrating the complete methodology of this invention, from raw data to closed-loop management of risk warning. Detailed Implementation

[0027] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the specific embodiments according to the present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments.

[0028] In construction projects of large urban complexes or super high-rise buildings, the construction site environment is complex, with numerous overlapping processes and many participating units. Traditional quality management methods, which rely on manual inspections and post-event records, are insufficient for real-time perception and early warning of potential risks such as concrete pouring quality, deep foundation pit support stability, and the safety of high-formwork systems. The construction engineering quality risk prediction method and system based on big data and artificial intelligence described in this invention aims to construct a complete technical system from data perception and intelligent analysis to decision execution, achieving proactive prediction and closed-loop management of engineering quality risks.

[0029] See Figure 1 The overall architecture of this system consists of a data acquisition and integration module, a data middle platform and feature engineering module, an intelligent risk prediction model module, a risk visualization and decision support module, and a mobile terminal early warning and collaborative handling module, forming a data-driven integrated platform for risk identification, assessment, early warning and handling.

[0030] First, the data acquisition and integration module, serving as the system's data entry point, is responsible for comprehensive and uninterrupted data collection from multiple heterogeneous data sources at the construction site and the management backend. At the construction site, an IoT sensor network is strategically deployed in key risk areas. For example, in the concrete pouring area, strain gauges are embedded in critical reinforcing bars to collect internal stress changes during concrete solidification at millisecond-level frequencies; inclinometers are installed on the core tube climbing formwork to monitor the inclination angle of the formwork under construction loads in real time; temperature and humidity sensor arrays are distributed inside and on the surface of large-volume concrete, continuously recording temperature gradients and humidity distribution caused by hydration heat; pressure sensors are installed on the waist beams and supports of deep foundation pit support piles to monitor the dynamic balance between soil pressure and support axial force; and high-definition image acquisition equipment captures visual information such as the tightness of the formwork support, the spacing of the reinforcing bars, and the thickness of the protective layer. These sensors transmit the raw engineering data to the data acquisition and integration module in real time via an industrial wireless network using the MQTT protocol. Simultaneously, this module connects to other management information systems in a non-intrusive manner through a deployed cluster of data adapters. Specifically, the data adapter for the Building Information Modeling (BIM) platform periodically retrieves precise geometric dimensions, material strength grades, design loads, and the latest construction schedule milestones for engineering components by parsing IFC standard documents or calling the platform's RESTful API. The data adapter for the project management system extracts weather conditions, work team information, quality inspection records (including pass / fail items), and the qualification certificate status of special operations personnel from construction logs via direct database connection or Web Service interface. The adapter for the materials supply chain database synchronizes factory certificates of conformity, third-party testing reports, transportation trajectories, and on-site re-inspection results for major building materials such as steel bars, cement, and admixtures. All raw data accessed through different protocols undergoes real-time cleaning before entering the module's internal processing pipeline. This cleaning process includes outlier removal based on the statistical three-sigma principle and duplicate record deduplication based on timestamps and data source identifiers. Subsequently, for issues such as inconsistent units (e.g., MPa and kPa being used interchangeably for pressure units) and inconsistent coding (e.g., different component numbering rules) that may exist in different data sources, mandatory format conversion and unified coding are performed based on the preset engineering data standardization dictionary. Finally, a standardized data stream to be analyzed with a unified time base and spatial reference is generated and pushed to the downstream module.

[0031] The data stream to be analyzed is continuously input into the data platform and feature engineering module for in-depth processing and organization. See also Figure 3The core of this module is to build a thematic data warehouse for engineering quality risk analysis. This data warehouse is organized using a star schema. Its fact table uses millisecond-level timestamps and two-dimensional or three-dimensional spatial locations based on the construction coordinate system as core dimensions, recording the instantaneous values ​​of each monitoring indicator at a specific time and location, such as "2023-10-27 14:30:25.123, coordinates (X1, Y1, Z1), internal temperature of concrete, 42.5℃". Around the fact table, multiple dimension tables are constructed, including engineering stage dimensions (such as earthwork excavation, foundation slab pouring, main structure construction, etc.), responsible unit dimensions (general contractor, subcontractor, supervision unit), component type dimensions (column, beam, slab, wall), and risk type dimensions (strength risk, stability risk, process risk, etc.). All standardized data to be analyzed is sorted according to its data tags and stored in corresponding data table partitions, forming ordered and interconnected data assets. Based on this, characteristic engineering processes are activated. This process strictly follows a pre-defined engineering quality risk knowledge graph. This graph uses engineering entities (e.g., "shear wall in area A of the 5th floor"), technological activities (e.g., "C50 concrete pouring"), materials (e.g., "PO 42.5 cement"), and environmental factors (e.g., "average daily temperature") as nodes, and relationships such as "located in," "used," "affected by," and "may lead to" as edges, constructing a semantic risk association network. Feature extraction first preprocesses the original monitoring data sequence. For data gaps caused by transient sensor failures or network packet loss, predicted values ​​based on a time-series autoregressive model are used for filling. For high-frequency random noise in the signal, a Savitzky-Golay filter is used for smoothing, effectively reducing noise while preserving trend characteristics. For temporal features, extraction is based on the sliding window technique: a sliding window with a length of 1 hour and a step size of 5 minutes is used to slide across the concrete temperature time series data. The mean (reflecting the average temperature level), variance (reflecting the severity of temperature fluctuations), and slope obtained through linear fitting (reflecting the trend of rising or falling temperatures) of the data within each window are calculated. Fourier transform is then used to analyze whether there are hidden periodic features related to the curing watering cycle in the data. For static features, they are directly extracted from the dimension table, such as the design concrete strength grade of the component, the reinforcement ratio, and the credit rating of the responsible unit. The most crucial aspect is the mining of correlation features. The system searches the engineering quality risk knowledge graph using a graph traversal algorithm. For example, it calculates the path length and relationship strength between the node "the floor slab currently being poured" and the node "the rainfall record of the area in the previous three days," and then combines this with the actual rainfall data to quantify the "humid environment influence coefficient" as a correlation feature; or it analyzes the average relationship strength between the node "a certain batch of cement" and the node group "the measured strength of components that have used this batch of cement in 28 days" to obtain the "material batch reliability index."Ultimately, for each risk analysis target (such as "the risk of insufficient concrete strength in the beams and slabs of the 8th floor, area B"), the system will assemble a feature vector containing dozens or even hundreds of dimensions, covering static attributes, dynamic temporal patterns and complex correlations, providing high-information-density input for subsequent intelligent prediction.

[0032] See Figure 2 The intelligent risk prediction model module receives multi-dimensional feature vectors from upstream sources and calculates risk probabilities using its internally integrated hybrid model. This module contains multiple specialized sub-models, each modeling different types of quality risks, such as insufficient concrete strength, instability of deep foundation pit support structures, deformation of high-formwork systems, and welding defects in steel structures. Each sub-model employs an architecture combining gradient boosting trees and long short-term memory networks to accommodate the characteristics of different feature types. For static features (such as design strength and reinforcement ratio) and correlated features (such as material correlation index and environmental correlation coefficient) in the feature vectors, the gradient boosting tree sub-model processes them. The model iteratively constructs a series of decision trees, with each new tree dedicated to correcting the residuals predicted by the previous tree. Its training process is guided by minimizing a specific objective function, defined as the weighted sum of prediction error and model complexity, specifically expressed as:

[0033] in, The loss function is represented by the logarithmic loss function, which is used in this embodiment to be suitable for probabilistic prediction. Representative sample The true risk label (0 indicates no risk, 1 indicates risk); Representative model for samples The predicted probability; It acts on the first Decision Tree The regularization term controls model complexity and prevents overfitting by limiting the number of leaf nodes and the weight values ​​of leaf nodes in the tree. This represents the total number of decision trees used in the gradient boosting process. The model automatically assesses the importance of each static and correlated feature and provides non-linear, high-precision classification predictions. Simultaneously, for temporal features in the feature vector (such as the mean concrete temperature sequence and the slope sequence of the support structure displacement), a Long Short-Term Memory (LSTM) network sub-model is activated. This network is specifically designed to handle sequence data with long-term dependencies. Its core lies in the cell state update mechanism, which finely regulates the flow of information through three gating units (forget gate, input gate, and output gate). At each time step... The calculation process is as follows: Forget Gate The decision is based on the cell state at the previous moment. What information is discarded in the input gate? Decide which new candidate information to include Store the cell state; then update the cell state. Output gate Based on the updated cell state, the output hidden state is determined. The calculation formula for the gate control unit is:

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] in, for The feature vector input at each time step (such as the temporal features extracted by the sliding window); This is the hidden state from the previous moment; , , , and the corresponding bias , , , These are the parameters that the model needs to train; The sigmoid activation function maps the output to a range of 0-1, controlling the degree to which the door opens or closes. The activation function is hyperbolic tangent. Through this mechanism, the network can capture subtle but crucial trend anomalies from sensor monitoring data spanning days or even weeks, such as "an early plateau in the concrete temperature rise curve may indicate weak later strength growth" or "a sustained, slow increase in support displacement rate may point to potential instability." Finally, the outputs of the two sub-models (the risk probability based on static and correlated features from the gradient boosting tree, and the risk probability based on temporal trends from the long short-term memory network) are integrated through a weighted fusion layer to generate a comprehensive predicted probability for the specific risk event. The final risk level is then determined based on a preset probability threshold range (e.g., 0-0.3 for low risk, 0.3-0.7 for medium risk, and 0.7-1.0 for high risk).

[0040] The risk visualization and decision support module is responsible for transforming the above intelligent analysis results into intuitive and actionable decision-making information. (See also...) Figure 4This module provides a web-based interactive dashboard. Once the predictive model outputs the risk level, the system automatically performs visualization rendering. For spatially distributed risks, such as concrete strength risks on different floors, the system maps the risk level to the corresponding 3D scene of the building information model, generating a risk heatmap represented by color gradients (e.g., green, yellow, red). Users can rotate and zoom 360 degrees to view areas of risk concentration. For risks that evolve over time, such as foundation pit support displacement risks, the system plots historical trend curves of displacement at key monitoring points and extends the model's predicted future trend lines, overlaying the time nodes of risk level changes onto the curves. Once the system determines that the risk level of a certain risk event exceeds a preset threshold (e.g., a probability of insufficient concrete strength risk greater than 0.7 is classified as high risk), the early warning mechanism is immediately triggered. The system does not simply issue an alarm but automatically calls upon the knowledge graph and historical case library to generate a structured early warning report. The report is comprehensive, including: a unique risk event identifier; the precise coordinates of the risk location in the 3D model and a prominent label on the 2D construction plan; the determined risk level and the model prediction confidence level; a list of key indicators that triggered the warning and screenshots of their historical change curves over the past 24 hours; the most likely causal chain inferred from a knowledge graph (e.g., "Cement batch A used in the current component, historical data shows that under similar curing conditions, the 28-day strength compliance rate of this batch of cement is only 85%; at the same time, the average humidity of the area in the past 48 hours is lower than the curing requirements, which may affect early hydration"); and a list of recommended treatment measures matched from a historical case database for similar causes and risk levels (e.g., "Immediately increase the frequency of mulching and watering curing in this area; plan to conduct early strength rebound testing on components cast with this batch of cement after 7 days; notify the materials department to pay attention to the subsequent use of this batch of cement"). Users can customize warning thresholds for different risk types on the dashboard and can perform horizontal comparisons of various risks under the same project to generate a comprehensive risk score for the entire project, providing support for project managers' macro-level decision-making.

[0041] The mobile early warning and collaborative response module ensures that risk information can penetrate management levels and reach frontline workers, forming a closed-loop management system. Once a structured early warning report is generated in the risk visualization and decision support module, this module immediately distributes the core content of the report (such as risk location, level, brief causes, and primary response recommendations) to the mobile terminals of relevant responsible personnel in real time via integrated push notifications, SMS messages, or internal enterprise communication tools. The distribution logic is intelligent: the system integrates geofencing and personnel positioning functions. When an early warning involves a specific construction area, the message will be prioritized and pushed to the terminals of construction workers, quality inspectors, or team leaders currently within that geofence, ensuring accurate information delivery to on-site personnel. Responsible personnel can view the complete early warning report on their mobile terminals and must provide feedback on the response based on on-site verification. Feedback information includes: confirmation of the actual on-site situation (whether it matches the early warning), the response measures taken, the response completion time, the preliminary results after the response (e.g., "Cultivation has been strengthened, surface humidity has reached the standard"), and supporting photos taken on-site. These feedback messages are transmitted back to the system in real time and linked to the original warning event. The system backend automatically tracks the entire process of "dispatch-receive-handling-feedback" for each warning event, escalating alerts for events that fail to respond or handle within the time limit. Finally, the warning event can be closed in loop once the feedback confirms that the risk has been eliminated or effectively controlled. All interactive data throughout the process, including warning content, recipients, response time, handling measures, and feedback results, is structured and stored to form a valuable case study library. This library is used to optimize future warning rules and handling suggestion matching algorithms, thereby enabling continuous self-evolution of the system's capabilities.

[0042] Based on the above system, the complete process of the construction project quality risk prediction method based on big data and artificial intelligence corresponding to this invention can be found in [link to relevant documentation]. Figure 5 The specific steps are as follows: Step S110 involves using a cluster of data adapters deployed in the data acquisition and integration module to collect multi-source, heterogeneous raw engineering data in parallel from the IoT sensor network, building information modeling platform, project management system, and materials supply chain database. In this step, the IoT sensor network collects physical world data at millisecond to second-level frequencies; the building information modeling platform provides design and management data at hourly or daily frequencies; and the project management system and materials database provide management process data in real-time or on a scheduled basis based on business occurrences. All adapters operate in parallel to ensure the timeliness and completeness of data acquisition.

[0043] Step S120 involves cleaning the raw engineering data in the data acquisition and integration module, removing outliers and duplicate records, and uniformly encoding and converting units for data of different formats to generate a standardized data stream for analysis. The cleaning process combines a rule engine and a statistical model; outlier removal is based not only on fixed thresholds but also on adaptive thresholds that consider dynamic changes in working conditions. Uniform encoding strictly follows the project's unified component coding system (WBS) and material coding system, and unit conversion is performed according to a conversion table between the International System of Units (SI) and commonly used engineering units to ensure unambiguous downstream processing.

[0044] Step S130: The data stream to be analyzed is input into the data platform and feature engineering module. Based on a preset data model, it is stored in the corresponding fact table and dimension table of the thematic data warehouse. Feature engineering is then performed based on the engineering quality risk knowledge graph to extract multi-dimensional feature vectors containing static features, time-series features, and correlation features. The storage process employs columnar storage and partitioning techniques to optimize query efficiency. Feature engineering is a dynamic and continuous process; whenever new data to be analyzed flows in, it triggers an update of the feature vector of the corresponding entity or indicator, ensuring that the features input to the prediction model always reflect the latest state.

[0045] Step S140: The multi-dimensional feature vector is input into the pre-trained intelligent risk prediction model module. The gradient boosting tree sub-model and the long short-term memory network sub-model in the model module respectively perform fusion calculations on the features, outputting the predicted probability for various quality risks, and classifying the risk level according to the probability value. The model calculation is executed in an event-driven or timed manner. After the start of key processes such as concrete pouring, the relevant model will enter a high-frequency calculation mode (e.g., once every 15 minutes); for routine monitoring, a comprehensive evaluation is performed at a lower frequency (e.g., once a day). The calculation result is a probability value for each preset risk type. The system converts it into low, medium, and high discrete risk levels according to the configured risk level-probability mapping table.

[0046] Step S150: In the risk visualization and decision support module, the risk level and prediction results are visualized and rendered to generate a risk heatmap and trend analysis charts. When the risk level exceeds a preset threshold, an early warning is automatically triggered, and a structured report containing causal analysis and remedial recommendations is generated. The visualization rendering uses a front-end graphics library for real-time drawing and supports large-screen display and viewing on personal computers. Early warning triggering is the core decision point, and threshold management supports differentiated settings based on risk type, project stage, and responsible unit. The generation of the structured report is a knowledge-driven process that requires real-time querying of knowledge graphs and case libraries, and assembly is completed within hundreds of milliseconds.

[0047] Step S160: Through the mobile terminal early warning and collaborative handling module, the structured early warning report is pushed to the mobile terminals of relevant responsible personnel, and handling feedback information is tracked and recorded, completing the closed-loop process from risk identification to handling verification. The push service must ensure message reachability and real-time performance under high concurrency, and a message queue is used for peak shaving and valley filling. Feedback information is collected through a simple mobile terminal form, reducing the operational threshold for frontline personnel. Closed-loop verification not only relies on manual feedback, but also, where possible, the system continuously monitors relevant sensor indicators and judges whether the handling measures are truly effective through data trends, thereby achieving an intelligent closed loop of "system early warning - manual handling - data verification".

[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A construction project quality risk prediction system based on big data and artificial intelligence, characterized in that, The system includes the following components: The data acquisition and integration module is used to acquire raw engineering data from multiple heterogeneous data sources in real time or near real time, and to clean, convert, and standardize the raw engineering data to generate data to be analyzed in a unified format. The heterogeneous data sources include at least an Internet of Things sensor network, a building information modeling platform, a project management system, and a materials supply chain database. The data acquisition and integration module is internally deployed with a data adapter cluster, each adapter corresponding to a specific data source protocol to achieve non-intrusive data access. The data platform and feature engineering module are used to receive and store the data to be analyzed, and to build a thematic data warehouse for engineering quality risk analysis. Based on a preset engineering quality risk knowledge graph, the data platform and feature engineering module extracts multi-dimensional feature vectors from the thematic data warehouse. The feature vectors include static features, time-series features, and correlation features. The static features represent the inherent attributes of engineering entities, the time-series features represent the dynamic change process of key monitoring indicators, and the correlation features represent the spatial and logical relationships between different entities or indicators. The intelligent risk prediction model module is used to receive the multi-dimensional feature vector, construct and train a risk prediction model through ensemble learning and deep learning algorithms, and output the predicted probability and risk level for a specific quality risk event. The intelligent risk prediction model module contains multiple sub-models, which are specifically modeled for different types of quality risks. The sub-models use the gradient boosting tree algorithm to classify and predict static and related features, and use the long short-term memory network algorithm to perform trend inference and anomaly detection on time-series features. The risk visualization and decision support module is used to map the predicted probability and risk level into a visualized risk heat map, trend curve and early warning signal, and automatically generate a structured early warning report containing the risk location, cause analysis and treatment suggestions when the risk level exceeds a preset threshold; the risk visualization and decision support module provides a web-based interactive dashboard to support users to perform multi-dimensional drill-down analysis on the risk prediction results; The mobile terminal early warning and collaborative handling module is used to distribute the structured early warning report to the mobile terminals of relevant responsible personnel in real time through message push service, and to receive handling feedback information from the mobile terminals, forming a closed-loop management process of risk warning, task assignment, on-site handling, and result feedback; the mobile terminal early warning and collaborative handling module integrates geofencing and personnel positioning functions to ensure that early warning information can accurately reach the on-site responsible persons.

2. The construction project quality risk prediction system based on big data and artificial intelligence according to claim 1, characterized in that, The IoT sensor network in the data acquisition and integration module is deployed at key locations on the construction site. The sensor types include strain gauges, inclinometers, temperature and humidity sensors, pressure sensors, and image acquisition devices. The data adapter cluster supports multiple data exchange protocols, including MQTT, OPC UA, RESTful API, and direct database connection.

3. The construction project quality risk prediction system based on big data and artificial intelligence according to claim 1, characterized in that, The data warehouse constructed by the data platform and feature engineering module is organized using a star schema. The fact table records the monitoring indicator values ​​with timestamps and spatial locations as dimensions. The dimension table includes engineering stage dimension, responsible unit dimension, component type dimension, and risk type dimension. The feature engineering process specifically includes: filling missing values ​​and smoothing and denoising the original monitoring data; extracting mean, variance, slope and periodic features from time series data based on sliding window technology; and mining the correlation strength between entities based on the engineering quality risk knowledge graph through graph traversal algorithm and quantifying it into correlation feature values.

4. The construction project quality risk prediction system based on big data and artificial intelligence according to claim 1, characterized in that, The gradient boosting tree sub-model in the intelligent risk prediction model module has an objective function defined as minimizing the weighted sum of prediction error and model complexity, as shown in the following formula: , in, Represents the loss function; Representative sample The true risk label; Representative model for samples The predicted probability; It acts on the first Decision Tree The regularization term; This represents the total number of decision trees during the gradient boosting process.

5. The construction project quality risk prediction system based on big data and artificial intelligence according to claim 1, characterized in that, The long short-term memory network sub-model in the intelligent risk prediction model module is used to process continuous time-series data monitored by sensors. Its cell state update mechanism can effectively capture long-term dependencies, and the calculation formula for its gating unit is as follows: , , , , , , in, , , These represent the forget gate, input gate, and output gate, respectively. In cellular state, In hidden state, for Input characteristics at time step and For model parameters, It is the sigmoid activation function.

6. The construction project quality risk prediction system based on big data and artificial intelligence according to claim 1, characterized in that, The structured early warning report generated by the risk visualization and decision support module includes at least the following: risk event identifier, three-dimensional coordinates or two-dimensional planar icon of the risk location, risk level, confidence level, key indicators that trigger the early warning and their historical change curves, the most likely cause inferred from the knowledge graph, and a list of recommended handling measures matched from the historical handling case library.

7. A method for predicting construction project quality risks based on big data and artificial intelligence, characterized in that, The specific steps of this method are as follows: S110, through a cluster of data adapters deployed in the data acquisition and integration module, collects multi-source heterogeneous raw engineering data in parallel from IoT sensor networks, building information modeling platforms, project management systems, and material supply chain databases; S120, In the data acquisition and integration module, the original engineering data is cleaned, outliers and duplicate records are removed, and data of different formats are uniformly encoded and converted in units to generate a standardized data stream to be analyzed. S130, the data stream to be analyzed is input into the data platform and feature engineering module, and stored in the corresponding fact table and dimension table of the subject data warehouse according to the preset data model. Feature engineering is performed based on the engineering quality risk knowledge graph to extract multi-dimensional feature vectors containing static features, time-series features and correlation features. S140, the multi-dimensional feature vector is input into the pre-trained intelligent risk prediction model module. The gradient boosting tree sub-model and the long short-term memory network sub-model in the model module respectively perform fusion calculation on the features, output the predicted probability for various quality risks, and classify the risk level according to the probability value. S150, In the risk visualization and decision support module, the risk level and prediction results are visualized and rendered to generate a risk heat map and trend analysis chart. When the risk level exceeds the preset threshold, an early warning is automatically triggered and a structured report containing cause analysis and treatment suggestions is generated. S160, through the mobile terminal early warning and collaborative handling module, pushes the structured early warning report to the mobile terminal of relevant responsible personnel, and tracks and records handling feedback information to complete the closed-loop process from risk identification to handling verification.

8. The construction project quality risk prediction method based on big data and artificial intelligence according to claim 7, characterized in that, In step S120, the cleaning process employs outlier removal based on the statistical three sigma principle and duplicate record deduplication based on timestamps and data source identifiers; the unified encoding and unit conversion are performed based on a preset engineering data standardization dictionary.

9. The construction project quality risk prediction method based on big data and artificial intelligence according to claim 7, characterized in that, In step S130, the extraction of time series features from time series data based on the sliding window technique specifically includes: sliding a sliding window of preset length and step size on the time series data, and calculating the mean, variance, and slope obtained by linear fitting of the data in each window.

10. The construction project quality risk prediction method based on big data and artificial intelligence according to claim 7, characterized in that, In step S160, the push process integrates geofencing and personnel positioning functions. When the warning involves a specific construction area, the warning information is pushed to the mobile terminal of the person in charge who is currently located within the geofence.