Construction settlement dynamic monitoring method based on data analysis
By constructing a three-level variable causal graph and a global causal graph, the problem of difficulty in updating causal relationships in construction settlement monitoring was solved, enabling real-time dynamic monitoring and decision support of the construction environment, and improving the interpretability and adaptability of the monitoring system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TAIZHOU UNIV
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-28
AI Technical Summary
Existing construction settlement monitoring technologies rely on finite element models and traditional data analysis methods, which are difficult to adapt to real-time changing construction environments, cannot effectively handle multi-scale correlations, and lack interpretability and causal evidence, making it difficult to support engineers' intervention decisions.
A data analysis-based dynamic monitoring method for construction settlement is adopted. By constructing a three-level variable causal graph, introducing dummy variables, forming an extended graph, updating the skeleton graph and performing rapid independence tests of triples, constructing a global causal graph, performing anomaly detection and visualization, and updating causal relationships in combination with prior confidence.
It achieves structured modeling and dynamic correlation identification of multi-level observation variables in the construction settlement area, updates the causal structure in real time, automatically determines conflicting causal relationships, supports engineers' decision-making based on causal evidence, and realizes the self-evolution of the knowledge system.
Smart Images

Figure CN121936731A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of construction engineering monitoring and data-driven modeling technology, and in particular to a method for dynamic monitoring of construction settlement based on data analysis. Background Technology
[0002] Traditional settlement monitoring mainly relies on methods such as manual leveling, total station observation, or distributed fiber optic sensing. In recent years, with the development of sensor networks, the Internet of Things, and big data technologies, construction monitoring has gradually evolved towards automation, networking, and data-driven approaches.
[0003] Current construction settlement monitoring technologies mainly rely on finite element analysis or empirical fitting at the model level, making it difficult to achieve dynamic cognitive updates by "learning physical causality from data". On the one hand, finite element models require a large amount of prior geological and boundary condition information, and their parameter calibration process is highly dependent on human experience, making it difficult to adapt to the real-time changing construction environment. On the other hand, traditional data analysis methods often assume linear independence or fixed lag relationships between variables, which cannot effectively handle multi-scale correlations between different observation levels. At the same time, existing monitoring systems are mostly based on threshold triggering, and the judgment of anomalies lacks interpretability and causal mechanisms, making it difficult to support engineers in making intervention decisions based on causal evidence. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a data analysis-based method for dynamic monitoring of construction settlement to address two issues: First, finite element models require a large amount of prior geological and boundary condition information, and their parameter calibration process is highly dependent on human experience, making it difficult to adapt to real-time changing construction environments. Second, traditional data analysis methods often assume linear independence or fixed lag relationships between variables, failing to effectively handle multi-scale correlations between different observation levels. Furthermore, existing monitoring systems are mostly based on threshold triggering, lacking interpretability and causal mechanisms for anomaly determination, making it difficult to support engineers in making intervention decisions based on causal evidence.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for dynamic monitoring of construction settlement based on data analysis, comprising, The observed variables in the construction settlement area were divided into three levels of variables, an initial cause-effect graph was constructed, and dummy variables were introduced to form an extended graph; Based on the extended graph, a skeleton graph is constructed and divided into overlapping modules. The variables of the overlapping modules are set as triples. A fast conditional independence tester for triples is set to update the skeleton graph. Binary variables are defined for conflicting edges in the updated skeleton graph, and DAG constraints are applied to non-conflicting edges. A multi-objective function is constructed for the conflict edges, including objective functions that maximize global logical consistency and minimize deviations from local opinions. The objective functions are solved to determine the direction of the conflict edges and obtain the global causal graph. Based on the global causal graph, a structural equation model is constructed, causal residuals are calculated and anomaly detection is performed, and E-value detection is performed on the edge relationships of the global causal graph to generate data evidence. Based on data evidence and prior confidence, the edges in the global causal graph are updated, and a visual interface is constructed to display the global causal graph and structural equation model of the construction settlement area in real time, and a natural language report on the monitoring of the construction settlement area is generated.
[0007] As a preferred embodiment of the data analysis-based dynamic monitoring method for construction settlement described in this invention, the three-level variables refer to the core driving layer, the associated observation layer, and the environmental background layer. The core driving layer consists of variables defined by physical mechanisms that directly affect the system behavior, including a set of construction disturbance variables, a set of ground state variables, and a set of settlement index variables. The associated observation layer includes variables that have a strong physical correlation with or are easily observable to the variables of the core driving layer. These variables are connected to the variables of the core driving layer through known deterministic or statistical mapping relationships. The environmental background layer includes external factors that affect the system but are not directly related to the results of construction activities; Based on physical laws, functional dependencies, and expert consensus, an initial causal graph among variables in the three-level variables is constructed, and each edge of the initial causal graph is assigned a value based on the data source to obtain the edge confidence. For each edge in the initial causal graph after assignment, the causal time lag is estimated using the mutual information maximum time lag method based on historical data; Based on historical construction settlement accident cases, a set of dummy variables is defined, where each dummy variable represents a known but unobservable potential risk source. Each dummy variable is added as a hidden node to the initial causal graph to form an extended graph; Each virtual variable node in the extended graph is assigned a subset of observable variables that it may jointly influence by a domain expert, and corresponding directed edges are added accordingly. The virtual variable nodes retain only the topological connection structure and are not assigned values or dynamic models.
[0008] As a preferred embodiment of the data analysis-based dynamic monitoring method for construction settlement described in this invention, the following steps are included: constructing a skeleton diagram and dividing it into overlapping modules, setting the variables of the overlapping modules as triples, setting a fast conditional independence tester for the triples, and updating the skeleton diagram: Perform a permutation test on the maximum mutual information of causal time lag for each edge in the extended graph, calculate the significance of the edges, verify the significance, delete the edges that fail the verification, and construct a skeleton graph. Divide the skeleton graph into several overlapping modules. Select two non-dummy variables from the three-level variables of each overlapping module as variable to be tested 1 and variable to be tested 2. Then select the adjacent nodes related to variable to be tested 1 and variable to be tested 2 from the three-level variables as the set of condition variables. The variable to be tested 1, the variable to be tested 2, and the set of condition variables are combined to form a triple; A fast conditional independence tester for triples is set up. When a new batch of data is received, the recursive update protocol of the fast conditional independence tester is updated in parallel for all triples in the modules to obtain the conditional distribution probability of each test variable 1 and test variable 2 under the conditional variable set. Edges of the skeleton graph are retained or removed based on conditional distribution probability.
[0009] As a preferred embodiment of the data analysis-based construction settlement dynamic monitoring method of the present invention, wherein: the recursive update protocol of the rapid conditional independence tester includes: Construct two regression models, use the condition variable set to predict test variable 1 and test variable 2 respectively, and combine the two regression models into matrix form to calculate the prediction error. For the merged regression model, perform RLS state updates, including updates to the Kalman gain, model parameters, and uncertainty matrix; Based on RLS state update, the conditional partial correlation coefficient and the standard error of the effective sample size are calculated, and Fisher's Z-transform is applied to the conditional partial correlation coefficient. Based on the standard error and Fisher's Z-transform results, the conditional distribution probability is calculated.
[0010] As a preferred embodiment of the data analysis-based dynamic monitoring method for construction settlement described in this invention, the step of maximizing the global logical consistency objective function to obtain a global causal graph includes: The objective function for maximizing global logical consistency refers to enumerating all variables to construct edge triplets of a triangle, and defining a consistency score for each edge triplet. The triangle weights are set as the average confidence of the three variable edges in the edge triplet, and the objective function that maximizes global logical consistency is defined in combination with the consistency score. For multi-objective functions, a multi-objective particle swarm optimization algorithm is used to solve the problem, obtain the direction of the edges, and generate a global causal graph.
[0011] As a preferred embodiment of the data analysis-based dynamic monitoring method for construction settlement described in this invention, the following steps are included: constructing a structural equation model based on a global causal graph, calculating causal residuals and performing anomaly detection, performing E-value detection on the edge relationships of the global causal graph, and generating data evidence, including: Based on the causal relationships identified in the global causal graph, a structural equation model is established for each target variable. For each time point, the structural equation model is used to predict the value of the target variable, and the difference between the actual value and the predicted value is calculated. Two rules are used to monitor for anomalies in the causal residual sequence: persistent shift detection and variance surge detection. If either rule triggers an alarm, it indicates that there is a problem with the model and the attribution and repair process needs to be initiated. The sensitivity of each causal edge in the global causal graph to the effects of unobserved confounding factors is quantified using the E-value. Use abnormal situations and sensitive data as evidence.
[0012] As a preferred embodiment of the data analysis-based dynamic monitoring method for construction settlement described in this invention, the step of updating the edges in the global causal graph based on data evidence and prior confidence includes: Based on data evidence, the anomaly strength is calculated, a data-driven confidence score is defined, and based on expert domain knowledge, a prior confidence score decay function is set to calculate the overall confidence score. Newly discovered relationships need to undergo a trial period during which they must maintain a high level of confidence before they can be formally included in the core knowledge base; otherwise, they will be automatically discarded. The high confidence level refers to the fact that its overall confidence level value must be continuously greater than the preset overall confidence level threshold during the trial period; Existing edge relationships that have been in a "paused" state for multiple cycles without expert intervention will be moved to the "historical dormant zone".
[0013] As a preferred embodiment of the data analysis-based dynamic monitoring method for construction settlement described in this invention, the step of constructing a visual interface to display the global causal graph and structural equation model of the construction settlement area in real time, and generating a natural language report on the monitoring of the construction settlement area, includes: The node size of the global causal graph reflects the recent activity level of the variable, the thickness of the edges represents the strength of the causal effect, and the color of the edges is composed of three superimposed parts. Meanwhile, the structural equation model is overlaid on the construction plan in the form of a heat map to represent the interpretability of the area; Based on the results displayed by the model and causal graph, a natural language report is automatically generated, which includes: the overall health level of the current network, a list of high-risk causal relationships, items that require expert review, and a description of the location of abnormal areas, so that managers can quickly grasp the system status.
[0014] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the data analysis-based dynamic monitoring method for construction settlement as described in the first aspect of the present invention.
[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the data analysis-based dynamic monitoring method for construction settlement as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: By combining a causal graph modeling method with a multi-source data analysis mechanism, this invention achieves structured modeling and dynamic correlation identification of multi-level observation variables in the construction settlement area. By combining a fast conditional independence test algorithm with a modular skeleton graph update mechanism, it achieves real-time updating of the causal structure under multi-source heterogeneous data streams, maintaining the logical consistency of the monitoring model. By combining a multi-objective function optimization algorithm with DAG constraints and directional confidence matrices, it realizes automatic determination of conflicting causal relationships and global optimal causal direction reasoning. By combining a prior knowledge decay function with a data-driven confidence fusion model, it achieves time-series dynamic evaluation of the reliability of causal edges, realizing the self-evolution of the knowledge system. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Fig. 1 This is a flowchart of the construction settlement dynamic monitoring method based on data analysis in Example 1.
[0019] Fig. 2 This is a schematic diagram of the three-level variables and extended graph structure in Example 1.
[0020] Fig. 3 This is the skeleton graph and conflict edge optimization system architecture diagram in Example 1.
[0021] Fig. 4 This is a diagram of the structural equation model and confidence update system architecture in Example 1. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0025] Example 1, referring to Figs. 1 to 4 This is the first embodiment of the present invention, which provides a method for dynamic monitoring of construction settlement based on data analysis, including the following steps: S1. Divide the observed variables of the construction settlement area into three levels of variables, construct an initial cause-effect graph, introduce dummy variables, and form an extended graph; Preferably, the observed variables in the construction settlement area are divided into three levels of variables, including the core driving layer, the related observation layer, and the environmental background layer; The core driving layer consists of variables defined by physical mechanisms that directly affect the system behavior, including a set of construction disturbance variables, a set of ground state variables, and a set of settlement index variables. The set of construction disturbance variables refers to the excavation volume variable, which directly reflects the scale and progress of the excavation project. The set of formation state variables includes pore water pressure, effective stress of soil, and physical property parameters of the formation. The effective stress of soil is determined by both pore water pressure and total stress, reflecting the soil's ability to withstand external loads. The physical property parameters of the formation include elastic modulus, Poisson's ratio, etc. These parameters are used to describe the mechanical behavior of materials and are expressed through the conservation equations of continuum mechanics. The settlement index variable set includes variables of settlement at a certain point on the ground surface and soil compression. The settlement at a certain point on the ground surface is the integral result of vertical strain at different depths, which directly reflects the degree of ground subsidence. The associated observation layer includes variables that have a strong physical correlation with or are easy to observe to the variables of the core driving layer. These variables are connected to the variables of the core driving layer through known deterministic or statistical mapping relationships. For example, secondary variables related to construction disturbance may include excavator GPS trajectory, secondary variables related to stratum state may be dewatering well water level, and secondary variables related to settlement index may be measurement data based on ground monitoring points, etc. The environmental background layer includes external factors that affect the system but are not directly related to the results of construction activities, such as natural conditions and surrounding vibration source logs. The surrounding vibration source log refers to other factors in the surrounding environment that may cause soil vibration; Based on physical laws, functional dependencies, and expert consensus, an initial causal graph among variables in the three-level variables is constructed, and each edge of the initial causal graph is assigned a value based on the data source to obtain the edge confidence. The physical laws refer to consulting literature such as rock and soil mechanics, fluid mechanics, and consolidation theory to confirm whether a certain variable appears in the governing equation of another variable. The functional dependency refers to the existence of deterministic causality if a variable is a direct measurement, transformation, or aggregation of another variable. Each edge of the causal graph is assigned a value based on the data source, such as physical laws, functional dependencies, and expert consensus, with the assigned value ranges being (0.6~1.0), (0.4~0.8), and (0.1~0.5) respectively. Physical laws, functional dependencies, and expert consensus are divided into high, medium, and low ranges. The numerical variations within these ranges are adjusted based on the universality of the equations and the degree of experimental verification, noise levels and sensor accuracy, consistency of expert opinions, and expert qualifications. For each edge in the initial causal graph after assignment, the causal time lag is estimated using the maximum mutual information method based on historical data (the mutual information between the excavation volume at the current time and the surface settlement variable at different lag times, and the lag time corresponding to the maximum mutual information is selected). Based on historical construction settlement accident cases, a set of dummy variables is defined, each dummy variable representing a known but unobservable potential risk source (such as "unknown weak interlayer" or "hidden seepage channel"). Each dummy variable is added as a hidden node to the initial causal graph to form an extended graph; Each virtual variable node in the extended graph is assigned a subset of observable variables that it may jointly influence by a domain expert, and corresponding directed edges are added accordingly. The virtual variable nodes retain only the topological connection structure and are not assigned values or dynamic models, in order to characterize the potential common cause structure.
[0026] S2. Based on the extended graph, construct a skeleton graph and divide it into overlapping modules. Set the variables of the overlapping modules as triples, set a fast conditional independence tester for triples, update the skeleton graph, define binary variables for conflicting edges in the updated skeleton graph, and apply DAG constraints to non-conflicting edges. Preferably, the maximum mutual information of the causal time lag of each edge in the extended graph is subjected to a permutation test, the significance of the edge is calculated and verified, the edge that fails the verification is deleted, and a skeleton graph is constructed. The skeleton graph is an undirected graph; Divide the skeleton graph into several overlapping modules. Select two non-dummy variables from the three-level variables of each overlapping module as variable to be tested 1 and variable to be tested 2. Then select the adjacent nodes related to variable to be tested 1 and variable to be tested 2 from the three-level variables as the set of condition variables. The variable to be tested 1, the variable to be tested 2, and the set of condition variables are combined to form a triple; A fast conditional independence tester (FPC) for triples is set up. When a new batch of data is received, the recursive update protocol of the fast conditional independence tester is updated in parallel for all triples in the modules to obtain the conditional distribution probability of each test variable 1 and test variable 2 under the conditional variable set. Edges in the skeleton graph are retained or removed based on the conditional distribution probability (the calculated conditional distribution probability is compared with a preset threshold, and edges with a probability less than the threshold are retained, otherwise they are deleted).
[0027] Furthermore, the recursive update protocol of the fast conditional independence tester includes: Construct two regression models, using condition variable sets to predict test variable 1 and test variable 2 respectively. Combine the two regression models into a matrix form and calculate the prediction error using the following formula: , in, Let this be the residual vector at the current time step. and These are the variable to be tested 1 and variable to be tested 2 at the current time. The set of condition variables at the current moment. The regression coefficient matrix of the previous time step is the transpose of the matrix. ,in b and b represent the parameters of the regression model for variable 1 and variable 2, respectively; RLS state update includes updates to the Kalman gain, model parameters, and uncertainty matrix. The Kalman gain of the new data is calculated, and the model parameters and uncertainty matrix are updated simultaneously. The formula is: , , , in, The Kalman gain at the current moment. This is the inverse information matrix from the previous time step, reflecting the model's response to the parameters. The uncertainty lies in the fact that the initial values are set based on experience. This is the forgetting factor, a set value such as 0.99. The smaller the value, the more important new data is; the larger the value, the longer the memory lasts. and These are the model parameters and inverse information matrix at the current time, respectively; The residual covariance is updated using the following formula: , in, Let the residual covariance be the value at the current moment. The residual covariance of the previous time step. and Let V be the variances of variable 1 and variable 2 under test at the current time. and These are the covariances of test variable 1 and test variable 2 at the current time and the covariances of test variable 2 and test variable 1, respectively. These values are estimated values predicted using the set of conditional variables. Calculate the conditional partial correlation coefficient and perform a statistical significance test; , in, This represents the conditional partial correlation coefficient at the current moment; Calculate the effective sample size and standard error; , in, This represents the number of valid samples at the current moment, initially set to 0. For the conditional partial correlation coefficient, perform the Fisher Z-transform, formula: , in, for The result after performing the Fisher Z-transform; Based on the standard error and Fisher's Z-transform results, the conditional probability distribution is calculated. , , in, The standard distance at the current moment. For the number of condition variables, The standard normal cumulative distribution function is... This represents the conditional probability distribution at the current moment.
[0028] S3. Construct a multi-objective function for the conflict edges, including objective functions that maximize global logical consistency and minimize deviations from local opinions. Solve the objective functions to determine the direction of the conflict edges and obtain the global causal graph. Preferably, based on the updated skeleton graph, a direction confidence matrix is constructed. If the module does not involve conflicting edges, the direction confidence is set to zero; otherwise, the direction confidence is set to the edge confidence or its difference (1 - edge confidence) based on the edge direction. The conflicting edge refers to the inconsistent direction of the edge between two variables in the overlapping area between different modules; Define a binary variable (one set to 1 and the other set to -1) for each conflicting edge, consider the direction of all non-conflicting edges as determined, and apply DAG constraints to all directed edges; The DAG constraint requires that all directed edges satisfy the topological order of the starting point being less than the topological order of the ending point, as shown in the formula: : , , : , in, It is the index of the directed edge. and These represent the nodes, or variables, at the two ends of the directed edge. For the set of conflicting edges, For binary variables, and These refer to the nodes respectively. and The topological order of the sorted edges determines the direction of non-collision edges, so the sorting is based on their directions. Let the set of edges include the set of conflicting edges and the set of non-conflicting edges. The meaning of "point" is "side". The direction of; Construct a multi-objective function, including an objective function that maximizes global logical consistency and minimizes deviations from local opinions; The objective function for maximizing global logical consistency refers to enumerating all variables to construct edge triplets of a triangle, and defining a consistency score for each edge triplet, as shown in the formula: , in, For consistency score, and The penalty values are set, such as 1 and 5, where the requirements are... The set value is greater than The value is used to highlight the stronger penalty for loops. 2 and 3 refer to transitivity, directed cycle and other cases respectively, 0 is the edge that has not appeared, and k is the third node that is constructed as the edge of the triangle, i.e., the variable; Let the triangle weights be the average confidence of the three variable edges in the edge triplet. Define the objective function to maximize global logical consistency, as follows: , in, This is a binary variable vector containing the binary variable values of all conflicting edges. The triangle weights of the three nodes that form the sides of the triangle. The index of the triangular node; The objective function for minimizing deviations from local opinions is defined by the following formula: , in, For the number of modules, For indicator functions, An index for the number of modules. In the module The weight of edge e is represented here by the average confidence level. In the module The direction of the middle edge e, Indicate the direction of edge e and The directions of the middle edges are inconsistent; For multi-objective functions, a multi-objective particle swarm optimization algorithm is used to solve the problem, obtain the direction of the edges, and generate a global causal graph. The multi-objective function problem takes the form, as shown in the formula: , , , , in, This represents the index of the node being traversed. The number of nodes.
[0029] S4. Based on the global causal graph, construct a structural equation model, calculate causal residuals and perform anomaly detection, perform E-value detection on the edge relationships of the global causal graph, and generate data evidence. Preferably, a structural equation model is established for each target variable based on the causal relationships identified in the global causal graph; The structural equation model refers to the linear structural equation established for each target variable by extracting all its parent nodes from the global causal graph and based on the lag time of each causal edge. The structural equation model takes into account all factors that directly affect the target variable and their time lag effects; For each time point, the structural equation model is used to predict the value of the target variable, and the difference between the actual value and the predicted value (i.e., the causal residual) is calculated. Two rules are used to monitor for anomalies in the causal residual sequence: persistent shift detection and variance surge detection. If either rule triggers an alarm, it indicates that there is a problem with the model and the attribution and repair process needs to be initiated. The continuous offset detection refers to checking whether the average residual over a recent period of time deviates significantly from zero, and whether this deviation has a consistent directionality; The variance surge detection refers to comparing the degree of fluctuation of the residuals in the current time period with the standard fluctuation level of historical data to determine whether there has been an abnormal increase. The E-value is used to quantify the sensitivity (causal effect value) of each causal edge in the global causal graph to the influence of unobserved confounding factors, providing an interpretable robustness certificate for each causal conclusion (e.g., the existence of an unobserved factor with an association strength greater than 2.5 with both exposure and outcome is required to refute the effect). The aforementioned initiation attribution and repair process refers to the system automatically executing the following closed-loop response mechanism when residual abnormalities are triggered, including: Freeze the use of relevant causal edges in the decision-making system to prevent the spread of erroneous reasoning; Based on associated external logs (such as construction records, environmental monitoring, etc.), the maximum time-delay mutual information between candidate variables and residual sequences is calculated to screen potential new confounding factors. Highly correlated candidate variables are added as new nodes to the causal graph, triggering a new round of causal discovery process, thereby enabling the model to self-update online and evolve its knowledge. Use abnormal situations and sensitive data as evidence.
[0030] S5. Based on data evidence and combined with prior confidence, update the edges in the global causal graph, construct a visual interface to display the global causal graph and structural equation model of the construction settlement area in real time, and generate a natural language report on the monitoring of the construction settlement area. Furthermore, based on data evidence, the anomaly strength is calculated using the following formula: , in, In the cycle variables within and The combined anomaly intensity between the edges, The standardized offset is the result value of the continuous offset detection. Here, F is the variance statistic, and the result is the variance surge detection result. Define data-driven confidence level, formula:
[0031] in, In the cycle variables within and Data-driven confidence between edges The E-valuation is based on the lower limit of the confidence interval. To perform lower bound processing, values exceeding the lower bound are set as boundary values. The confidence interval is calculated based on the lower bound of the 95% frequency school confidence interval for causal effect estimation. This is an anomaly sensitivity hyperparameter used to control the rate at which the anomaly intensity decays with the confidence level. It is a hyperparameter set based on experimental methods. Based on expert domain knowledge, a prior confidence decay function is defined, as follows:
[0032] in, In the cycle variables within and Prior confidence of the edges between them For variables and The last time the edge was marked between them. The a priori decay time constant is expressed in periods. The prior confidence level refers to the relationship between edges in the global causal graph, which is reviewed and scored by experts and used as the prior confidence level. Over time, if a relationship is not verified as "normal" by data for a long period of time, its prior trust will automatically and slowly decay to prevent outdated knowledge from dominating for a long time. Meanwhile, the system provides a human feedback interface: experts can mark any relationship as "confirmed valid," "questionable," or "invalid." Once feedback is received, the system immediately adjusts its prior confidence level and freezes automatic updates for a period of time to respect human judgment. If the prior trust in a relationship decays to a low level, the system will mark it as "awaiting expert review" and prompt manual intervention to confirm its validity. The overall confidence level is calculated by fusing prior knowledge and data evidence, using the following formula: , in, In the cycle variables within and The overall confidence level of the edges between them; Newly discovered relationships need to undergo a trial period (e.g., 7 cycles), during which they must maintain a high level of confidence before they can be formally included in the core knowledge base; otherwise, they will be automatically discarded. The high confidence level refers to the fact that its overall confidence level value must be continuously greater than the preset overall confidence level threshold during the trial period; If an existing relationship is in a "paused" state for multiple cycles (e.g., 10 cycles) without expert intervention, it will be moved to the "historical dormant zone" and will no longer participate in real-time inference, but the record will be retained for future reference.
[0033] Furthermore, a visual interface is constructed to display the global cause-effect graph and structural equation model of the construction settlement area in real time, and to generate a natural language report on the monitoring of the construction settlement area. The node size of the global causal graph reflects the recent activity level of the variable, the thickness of the edges represents the strength of the causal effect, and the edge color is composed of three parts: red indicates low overall confidence (warning), green indicates strong prior knowledge, and blue indicates strong data evidence. Meanwhile, the structural equation model is overlaid on the construction plan in the form of a heat map. The darker the red, the worse the model's explanatory power in that area, which may indicate the existence of unmodeled factors. Based on the results displayed by the model and causal graph, a natural language report is automatically generated, which includes: the overall health level of the current network, a list of high-risk causal relationships, items that require expert review, and a description of the location of abnormal areas, so that managers can quickly grasp the system status.
[0034] This embodiment also provides a computer device applicable to the data analysis-based dynamic monitoring method for construction settlement, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the data analysis-based dynamic monitoring method for construction settlement as proposed in the above embodiment.
[0035] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0036] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the data analysis-based dynamic monitoring method for construction settlement as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0037] In summary, this invention achieves structured modeling and dynamic correlation identification of multi-level observation variables in construction settlement areas through: a causal graph modeling method combined with a multi-source data analysis mechanism; real-time updating of the causal structure under multi-source heterogeneous data streams, maintaining the logical consistency of the monitoring model through a fast conditional independence test algorithm combined with a modular skeleton graph update mechanism; automatic determination of conflicting causal relationships and global optimal causal direction reasoning through a multi-objective function optimization algorithm combined with DAG constraints and directional confidence matrices; and time-series dynamic evaluation of the reliability of causal edges through a prior knowledge decay function combined with a data-driven confidence fusion model, realizing the self-evolution of the knowledge system.
[0038] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for dynamic monitoring of construction settlement based on data analysis, characterized in that: include, The observed variables in the construction settlement area were divided into three levels of variables, an initial cause-effect graph was constructed, and dummy variables were introduced to form an extended graph; Based on the extended graph, a skeleton graph is constructed and divided into overlapping modules. The variables of the overlapping modules are set as triples. A fast conditional independence tester for triples is set to update the skeleton graph. Binary variables are defined for conflicting edges in the updated skeleton graph, and DAG constraints are applied to non-conflicting edges. A multi-objective function is constructed for the conflict edges, including objective functions that maximize global logical consistency and minimize deviations from local opinions. The objective functions are solved to determine the direction of the conflict edges and obtain the global causal graph. Based on the global causal graph, a structural equation model is constructed, causal residuals are calculated and anomaly detection is performed, and E-value detection is performed on the edge relationships of the global causal graph to generate data evidence. Based on data evidence and prior confidence, the edges in the global causal graph are updated, and a visual interface is constructed to display the global causal graph and structural equation model of the construction settlement area in real time, and a natural language report on the monitoring of the construction settlement area is generated.
2. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 1, characterized in that: The three-level variables refer to the core driving layer, the related observation layer, and the environmental background layer; The core driving layer consists of variables defined by physical mechanisms that directly affect the system behavior, including a set of construction disturbance variables, a set of ground state variables, and a set of settlement index variables. The associated observation layer includes variables that have a strong physical correlation with or are easily observable to the variables of the core driving layer. These variables are connected to the variables of the core driving layer through known deterministic or statistical mapping relationships. The environmental background layer includes external factors that affect the system but are not directly related to the results of construction activities; Based on physical laws, functional dependencies, and expert consensus, an initial causal graph among variables in the three-level variables is constructed, and each edge of the initial causal graph is assigned a value based on the data source to obtain the edge confidence. For each edge in the initial causal graph after assignment, the causal time lag is estimated using the mutual information maximum time lag method based on historical data; Based on historical construction settlement accident cases, a set of dummy variables is defined, where each dummy variable represents a known but unobservable potential risk source. Each dummy variable is added as a hidden node to the initial causal graph to form an extended graph; Each virtual variable node in the extended graph is assigned a subset of observable variables that it may jointly influence by a domain expert, and corresponding directed edges are added accordingly. The virtual variable nodes retain only the topological connection structure and are not assigned values or dynamic models.
3. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 2, characterized in that: The process of constructing a skeleton graph and dividing it into overlapping modules, setting variables for overlapping modules as triples, setting a fast conditional independence tester for triples, and updating the skeleton graph includes: Perform a permutation test on the maximum mutual information of causal time lag for each edge in the extended graph, calculate the significance of the edges, verify the significance, delete the edges that fail the verification, and construct a skeleton graph. Divide the skeleton graph into several overlapping modules. Select two non-dummy variables from the three-level variables of each overlapping module as variable to be tested 1 and variable to be tested 2. Then select the adjacent nodes related to variable to be tested 1 and variable to be tested 2 from the three-level variables as the set of condition variables. The variable to be tested 1, the variable to be tested 2, and the set of condition variables are combined to form a triple; A fast conditional independence tester for triples is set up. When a new batch of data is received, the recursive update protocol of the fast conditional independence tester is updated in parallel for all triples in the modules to obtain the conditional distribution probability of each test variable 1 and test variable 2 under the conditional variable set. Edges of the skeleton graph are retained or removed based on conditional distribution probability.
4. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 3, characterized in that: The recursive update protocol of the fast conditional independence tester includes: Construct two regression models, use the condition variable set to predict test variable 1 and test variable 2 respectively, and combine the two regression models into matrix form to calculate the prediction error. For the merged regression model, perform RLS state updates, including updates to the Kalman gain, model parameters, and uncertainty matrix; Based on RLS state update, the conditional partial correlation coefficient and the standard error of the effective sample size are calculated, and Fisher's Z-transform is applied to the conditional partial correlation coefficient. Based on the standard error and Fisher's Z-transform results, the conditional distribution probability is calculated.
5. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 4, characterized in that: The objective function for maximizing global logical consistency yields a global causal graph, including: The objective function for maximizing global logical consistency refers to enumerating all variables to construct edge triplets of a triangle, and defining a consistency score for each edge triplet. The triangle weights are set as the average confidence of the three variable edges in the edge triplet, and the objective function that maximizes global logical consistency is defined in combination with the consistency score. For multi-objective functions, a multi-objective particle swarm optimization algorithm is used to solve the problem, obtain the direction of the edges, and generate a global causal graph.
6. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 5, characterized in that: The process involves constructing a structural equation model based on a global causal graph, calculating causal residuals and performing anomaly detection, performing E-value detection on the edge relationships of the global causal graph, and generating data evidence, including: Based on the causal relationships identified in the global causal graph, a structural equation model is established for each target variable. For each time point, the structural equation model is used to predict the value of the target variable, and the difference between the actual value and the predicted value is calculated. Two rules are used to monitor for anomalies in the causal residual sequence: persistent shift detection and variance surge detection. If either rule triggers an alarm, it indicates that there is a problem with the model and the attribution and repair process needs to be initiated. The sensitivity of each causal edge in the global causal graph to the effects of unobserved confounding factors is quantified using the E-value. Use abnormal situations and sensitive data as evidence.
7. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 6, characterized in that: The update of edges in the global causal graph based on data evidence and prior confidence includes: Based on data evidence, the anomaly strength is calculated, a data-driven confidence score is defined, and based on expert domain knowledge, a prior confidence decay function is set to calculate the overall confidence score. Newly discovered relationships need to undergo a trial period during which they must maintain a high level of confidence before they can be formally included in the core knowledge base; otherwise, they will be automatically discarded. The high confidence level refers to the fact that its overall confidence level value must be continuously greater than the preset overall confidence level threshold during the trial period; If existing edge relationships are in a "paused" state for multiple cycles without expert intervention, they will be moved to the "historical dormant zone".
8. The method for dynamic monitoring of construction settlement based on data analysis as described in claim 7, characterized in that: The constructed visualization interface displays the global cause-effect graph and structural equation model of the construction settlement area in real time, and generates a natural language report on the monitoring of the construction settlement area, including: The node size of the global causal graph reflects the recent activity level of the variable, the thickness of the edges represents the strength of the causal effect, and the color of the edges is composed of three superimposed parts. Meanwhile, the structural equation model is overlaid on the construction plan in the form of a heat map to represent the interpretability of the area; Based on the results displayed by the model and causal graph, a natural language report is automatically generated, which includes: the overall health level of the current network, a list of high-risk causal relationships, items that require expert review, and a description of the location of abnormal areas, so that managers can quickly grasp the system status.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the data analysis-based dynamic monitoring method for construction settlement as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the data analysis-based dynamic monitoring method for construction settlement as described in any one of claims 1 to 7.