Data analysis methods for digital management of engineering consulting

By employing techniques such as spatiotemporal grid standardization, dynamic weight allocation, and multi-scenario simulation, the problems of low data integration efficiency and insufficient accuracy in engineering consulting data analysis have been solved. This has enabled unified data mapping and risk prediction, thereby improving the intelligence and scientific level of engineering consulting management.

CN120653941BActive Publication Date: 2025-11-14ZHONGSHAN LUCHENG ENG MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511159045.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-14
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing engineering consulting data analysis methods are insufficient in terms of the comprehensiveness, accuracy, and dynamic adaptability of data processing, making it difficult to meet the needs of digital and intelligent management in engineering consulting. In particular, they have obvious defects in low data integration efficiency, insufficient analysis accuracy, incomplete detection of inconsistent data formats, single time dimension processing, and lack of multi-scenario simulation and dynamic adjustment capabilities.

Method used

The spatiotemporal grid standardization method is used to map heterogeneous data to a unified framework. Format conflicts are handled through dynamic weight allocation rules. Missing values ​​and abnormal fluctuations are handled by combining time series autoregressive models and spatial interpolation algorithms. A digital mirror is constructed and multi-scenario simulations are performed to generate dynamic adjustment schemes, thereby achieving real-time data synchronization and risk prediction.

Benefits of technology

It improves the efficiency and accuracy of data analysis integration, ensures data format consistency and logical consistency, realizes the continuity of data in the spatiotemporal dimensions, provides comprehensive and accurate decision support for engineering consulting project management, and enhances the dynamic adaptability and intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653941B_ABST
    Figure CN120653941B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of data analysis and discloses a data analysis method for intelligent management of engineering consulting. The method includes: mapping heterogeneous data from multiple projects to a unified three-dimensional spatiotemporal grid through spatiotemporal grid partitioning to form a grid-mapped dataset; detecting data format and attribute conflicts based on neighborhood correlation analysis and calibrating the data through dynamic weight allocation; handling missing values ​​and abnormal fluctuations using a time-series autoregressive model and smoothing techniques, and constructing a continuous and complete dataset by combining linear interpolation and multi-dimensional data fusion; extracting key indicators to construct a digital mirror containing spatial location, time series, and attribute correlations, synchronizing real-time data, and correcting outlier grid points; and constructing multi-scenario operating conditions based on the mirror data to achieve dynamic adjustment of the plan. This method integrates heterogeneous data through spatiotemporal grids and combines digital twin and multi-scenario simulation technologies to achieve full-cycle management of engineering data and intelligent risk decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis, and in particular to a data analysis method for intelligent management of engineering consulting. Background Technology

[0002] In today's era of rapid digital development, the engineering consulting industry is accelerating its transformation towards intelligent and digital management. Engineering consulting projects often involve multiple stages, including planning, design, construction, and operation, each generating massive amounts of heterogeneous data, such as geospatial data, project progress data, and equipment monitoring data. This data contains critical information about project operation, and efficient analysis can assist in project decision-making, optimize resource allocation, and predict potential risks. Therefore, data analysis has become a core component of intelligent and digital management in engineering consulting.

[0003] Currently, traditional engineering consulting data analysis methods often combine manual processing with simple statistical tools, which struggles to handle complex heterogeneous data, resulting in low data integration efficiency and insufficient analytical accuracy. While some methods have attempted to use techniques like grid partitioning and time series analysis for data processing with technological advancements, numerous shortcomings remain. For example, in the data standardization stage, existing technologies struggle to accurately map heterogeneous data from different project sources and with varying format standards into a unified framework, leading to difficulties in data fusion. In the data quality inspection stage, the detection of inconsistent data formats or attribute conflicts is not comprehensive or accurate enough, failing to effectively identify potential data contradictions. Regarding data integrity processing, the handling of missing values ​​and abnormal fluctuations in the time dimension is relatively simplistic and insufficient to meet the needs of complex engineering scenarios. Furthermore, existing technologies lack the ability to perform multi-scenario simulations and dynamic adjustments for project risk prediction and configuration optimization, failing to provide comprehensive and accurate decision support for engineering consulting.

[0004] In summary, existing engineering consulting data analysis methods have significant shortcomings in terms of the comprehensiveness, accuracy, and dynamic adaptability of data processing, making it difficult to meet the growing demand for intelligent management in engineering consulting. There is an urgent need for a more efficient, accurate data analysis method with dynamic adjustment capabilities to improve the intelligence level and scientific decision-making of engineering consulting project management. Summary of the Invention

[0005] The main objective of this invention is to provide a data analysis method for the digital management of engineering consulting, aiming to solve the technical problem that existing engineering consulting data analysis methods have significant shortcomings in terms of the comprehensiveness, accuracy, and dynamic adaptability of data processing, making it difficult to meet the growing demand for digital management of engineering consulting.

[0006] To achieve the above-mentioned objectives, the first aspect of this invention proposes a data analysis method for intelligent management of engineering consulting, the method comprising:

[0007] The raw business data streams are obtained from multiple project stages. Using a pre-established spatiotemporal grid partitioning method, the heterogeneous data from different projects are mapped to a unified three-dimensional spatiotemporal grid, resulting in a preliminary standardized grid mapping dataset.

[0008] For the grid mapping dataset, the data format inconsistency or attribute conflict is detected by calculating the format compatibility and business attribute logic consistency of grid points, including data type, length, and precision. If there is a data format inconsistency or attribute conflict, the data priority is adjusted according to the reliability of the data source and the update time through a preset dynamic weight allocation rule to obtain a grid calibration dataset with a consistent format.

[0009] For the grid calibration dataset, periodic features are identified based on a time series autoregressive model, missing values ​​in the time dimension are filled in by a spatial interpolation algorithm, and abnormal fluctuation points are removed by a preset spatiotemporal constraint rule to obtain a time-stable dataset.

[0010] The grid point timestamp interval of the time-series stable dataset is analyzed. If it exceeds the preset time threshold, linear interpolation is used to generate time points. The intermediate data points are optimized by combining multidimensional data fusion technology to construct a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions.

[0011] Key project indicators for each monitoring period are extracted from the complete grid dataset. A digital image of the project containing spatial location, time series, and attribute associations is constructed using virtual entity modeling technology. Data from multiple periods are uniformly mapped to a benchmark time axis based on the project duration to obtain an aligned image mapping dataset.

[0012] Real-time data synchronization is performed on the mirror mapping dataset. The latest monitoring data is obtained through the sensor interface. If key indicators are missing, a weighted average is performed based on the spatial distance of neighboring grid points and the correlation of historical data. The weight coefficients are dynamically adjusted in combination with the data fluctuation amplitude to supplement the missing values ​​and obtain the mirror updated dataset.

[0013] For the mirror update dataset, a density-based spatial clustering algorithm is used to identify outlier grid points. If their spatial distribution deviates from the core cluster and the time series fluctuation exceeds a preset range, outliers are corrected using time series smoothing techniques including moving average and exponential smoothing to obtain a verified mirror stable dataset.

[0014] Based on the aforementioned mirror-stable dataset, multi-scenario operating conditions are constructed to meet the project risk prediction needs. By using methods including Monte Carlo simulation and finite element analysis, project behavior patterns under different combinations of key indicators are simulated to identify potential risk points that exceed the risk threshold and generate a structured simulation analysis dataset.

[0015] For the simulation analysis dataset, the differences in operating indicators under different scenarios are quantified by techniques including principal component analysis and cluster comparison. Combined with the regression model of project configuration parameters and indicator performance, a dynamic adjustment plan including equipment scheduling, resource allocation and process adjustment is generated.

[0016] Beneficial effects:

[0017] The data analysis method for intelligent management of engineering consulting in this invention solves the problem of low integration efficiency by constructing a standardized spatiotemporal grid system to map heterogeneous data from multiple stages to a unified framework; it eliminates data format conflicts and attribute contradictions by using format compatibility verification and dynamic weight allocation, thereby improving analysis accuracy; it achieves continuous consistency of data in the spatiotemporal dimensions by relying on time series optimization and digital mirror modeling, laying the foundation for real-time monitoring; and it accurately identifies potential risks and generates resource scheduling plans through multi-scenario simulation and dynamic adjustment algorithms, enabling engineering consulting management to form a complete closed loop from data collection to decision support, effectively enhancing the system's dynamic adaptability to complex working conditions, promoting the transformation of project management towards intelligence and science, and realizing full-cycle management of engineering data and intelligent risk decision-making. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a data analysis method for intelligent management of engineering consulting, according to an embodiment of the invention. The realization of the purpose, functional characteristics, and advantages of the invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0020] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any modules and all combinations of one or more associated listed items.

[0021] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0022] Reference Figure 1 This invention provides a data analysis method for intelligent management of engineering consulting, comprising:

[0023] S1: Map heterogeneous data from different projects onto a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset;

[0024] S2: Detect data format and attribute conflicts in the grid mapping dataset, and adjust the data priority in the grid mapping dataset according to the preset dynamic weight allocation rules to obtain a grid calibration dataset with consistent format;

[0025] S3: For the grid calibration dataset, use a time series autoregressive model and smoothing techniques to process missing values ​​and abnormal fluctuations to obtain a time-stable dataset;

[0026] S4: Analyze the grid point timestamp intervals of the time-series stable dataset, and construct a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions by combining linear interpolation and multidimensional data fusion.

[0027] S5: Extract key indicators from the complete grid dataset, construct a digital mirror of the project including spatial location, time series, and attribute association, and map multi-period data to a benchmark time axis based on the project duration to obtain an aligned mirror mapping dataset;

[0028] S6: Obtain the latest monitoring data through the sensor interface. If key indicators are missing, perform a weighted average based on the spatial distance of neighboring grid points and the correlation of historical data, and dynamically adjust the weight coefficients in combination with the data fluctuation amplitude to supplement the missing values. Update the mirror mapping dataset to obtain the mirror update dataset.

[0029] S7: Identify outlier grid points in the mirror update dataset, correct the outlier grid points, and obtain the verified mirror stable dataset;

[0030] S8: Based on the aforementioned mirror-stabilized dataset, construct multi-scenario operating conditions to meet the project risk prediction needs and generate a structured simulation analysis dataset;

[0031] S9: Based on the simulation analysis dataset, generate a dynamic adjustment scheme that includes equipment scheduling, resource allocation, and process adjustment.

[0032] This embodiment relates to a data analysis method for intelligent management of engineering consulting. Its core lies in solving the problems of low efficiency in integrating heterogeneous data, insufficient analytical accuracy, and lack of dynamic adaptability in traditional methods through multi-dimensional data processing and dynamic modeling. Specifically:

[0033] Step S1 primarily involves spatiotemporal grid standardization. This involves acquiring raw business data streams from multiple project stages and using a pre-established spatiotemporal grid partitioning method to map heterogeneous data from different projects onto a unified three-dimensional spatiotemporal grid, resulting in a preliminary standardized grid mapping dataset. "Raw business data streams" refer to heterogeneous data collected from the planning, design, construction, and operation stages of engineering consulting projects. This data includes spatial coordinates (such as latitude, longitude, and altitude), timestamps (such as UTC time), and business attributes (such as progress percentage and equipment operating parameters). The formats may include CSV, JSON, and GIS vector data. The "spatiotemporal grid partitioning method" divides the three-dimensional space (X, Y, and Z axes) into grids of fixed or dynamic granularity, with time as the fourth dimension, forming a unified spatiotemporal coordinate system. Mapping functions (such as projection transformation and time unit normalization) are used to transform data from different data sources to this coordinate system. For example, the spatial coordinates of a building BIM model and the time series data from meteorological monitoring are mapped to the same grid, achieving a "preliminary standardized grid mapping dataset." The purpose of this step is to solve the problem of inconsistent heterogeneous data formats in traditional methods, and to provide a unified framework for subsequent processing. Its output serves as the basis for the input data of S2.

[0034] Step S2 primarily involves data format calibration and conflict resolution. Specifically, for the grid mapping dataset, the format compatibility and business attribute logical consistency of grid points (including data type, length, and precision) are calculated. Data format inconsistencies or attribute conflicts are detected. If such issues exist, data priority is adjusted based on data source reliability and update time using a preset dynamic weight allocation rule to obtain a grid calibration dataset with consistent format. A "grid point" is the smallest unit in a three-dimensional spatiotemporal grid, containing spatial coordinates (X, Y, Z) and a timestamp (t). The spatial coordinates correspond to the physical location at the project site (e.g., latitude, longitude, altitude), and the timestamp uses UTC standard time with precision down to the second. For example, grid point (100, 200, 5, 2025-07-02T10:00:00Z) represents the monitoring location at X=100m, Y=200m, Z=5m, at 10:00:00 on July 2, 2025. "Format compatibility" involves whether data types (e.g., integers, floating-point numbers), lengths (e.g., string character count), and precision (e.g., decimal places) conform to industry standards or preset templates. "Business attribute logical consistency" includes whether the logical relationship between concrete curing temperature and strength growth is contradictory. The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm identifies dense and sparse data areas. Dynamic anomalies (e.g., sudden changes in equipment monitoring data) undergo format and attribute verification. If conflicts are found (e.g., temperature values ​​from different data sources at the same location differing by more than a threshold), the priority is adjusted and values ​​are reassigned based on the rule of "real-time data > historical data" and "high-reliability source data > low-reliability source data" (e.g., BIM (Building Information Modeling) data provided by the design institute has higher reliability than manually entered data by the construction team). This step solves the problem of incomplete data quality detection in traditional methods, ensuring that the data input to S3 is formatted uniformly and logically consistent, avoiding subsequent analysis deviations due to data contradictions.

[0035] Step S3 primarily involves time series stabilization. Specifically, for the grid calibration dataset, periodic features are identified based on a time series autoregressive model, and missing values ​​in the time dimension are filled using spatial interpolation algorithms. Abnormal fluctuations are then eliminated using preset spatiotemporal constraint rules, resulting in a time-stable dataset. The "time series autoregressive model" (such as ARIMA) is used to identify the periodic features of the data, such as the weekly cycle of construction progress. The "spatial interpolation algorithm" (such as inverse distance weighted interpolation) uses historical data from neighboring grids to fill missing values; for example, if concrete strength data for a certain grid point is missing, it is calculated by weighting the strength values ​​of the surrounding 5×5 grids. The "spatiotemporal constraint rules" combine the periodic average before and after the time frame with spatial neighboring data to eliminate outliers. For example, if the temperature data at a certain moment is significantly higher than the historical average for the same period and there are no similar fluctuations in the neighboring grids, it is identified as an anomaly and eliminated. This step solves the problem of single-dimensional time dimension data processing in traditional methods, making the input data in S4 continuous and stable in the time series, laying the foundation for constructing a complete time axis.

[0036] Step S4 primarily involves constructing spatiotemporal data integrity. This involves analyzing the grid point timestamp intervals of the time-series stable dataset. If the intervals exceed a preset time threshold, linear interpolation is used to generate time points. Multidimensional data fusion technology is then used to optimize intermediate data points, constructing a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions. The intervals between adjacent timestamps are calculated; if they exceed a preset threshold (e.g., 30 minutes), linear interpolation is triggered to generate intermediate time points. For example, if monitoring data from a device is missing at 2-hour intervals, it is supplemented by linear fitting of the preceding and following time points. The "multidimensional data fusion technology" includes exponential smoothing correction (handling abnormal fluctuations) and Gaussian process regression (supplementing spatial correlation features). For instance, when data fluctuations before and after an interpolation point are severe, exponential smoothing is used to reduce noise; when the correlation between neighboring data is low, Gaussian process regression is used to uncover potential spatial correlations. This step constructs a continuous time axis and consistent spatiotemporal dimensions, enabling S5 to accurately extract key indicators and solving the problem of insufficient data integrity processing in traditional methods.

[0037] Step S5 primarily involves constructing a digital mirror and aligning it with time. This involves extracting key project indicators for each monitoring period from the complete grid dataset, constructing a digital mirror of the project using virtual entity modeling technology that includes spatial location, time series, and attribute relationships, and mapping multi-period data to a baseline timeline based on the project duration, resulting in an aligned mirror-mapped dataset. "Key project indicators" include progress (e.g., completion rate of sub-projects), energy consumption (e.g., electricity consumption per unit area), and quality (e.g., concrete compressive strength). "Virtual entity modeling technology," or digital twin technology, constructs a digital mirror containing a geometric model (3D building model), a physical model (material mechanical properties), and a business model (progress management process), establishing a mapping of "grid coordinates - timestamp - indicator value." Multi-period data is mapped to a baseline timeline based on the total project duration (e.g., unifying progress data from different stages to a timeline with the start date as the origin), and outliers are smoothed using local weighted regression to generate trend features. This step provides a standardized digital twin model for real-time data synchronization in S6, enabling unified management of multi-period data and facilitating subsequent risk analysis.

[0038] Step S6 primarily involves real-time data synchronization and missing data supplementation. This involves real-time data synchronization of the mirrored dataset, acquiring the latest monitoring data via sensor interfaces. If key indicators are missing, a weighted average is calculated based on the spatial distance to neighboring grid points and historical data correlation, with weight coefficients dynamically adjusted to compensate for missing values, resulting in an updated mirrored dataset. Sensor data (such as tower crane operating status and environmental monitoring data) is collected in real-time via IoT interfaces and aligned with the timestamps of the mirrored dataset. For missing key indicators, an inverse distance-weighted average (weighting data from neighboring 3×3 grid points) is used, with weights dynamically adjusted based on historical fluctuations (increasing the weight of recent data when fluctuations are large). For example, if PM2.5 data for a certain grid point is missing, data from neighboring grid points that are closer and have high historical correlation receive greater weight. The process verifies and fills discrepancies between the data and its neighborhood, corrects suspicious points, and ensures that the input data in S7 is real-time and accurate, solving the problem of delayed real-time data processing in traditional methods.

[0039] Step S7 primarily involves outlier identification and anomaly correction. This involves using a density-based spatial clustering algorithm to identify outlier grid points in the mirrored updated dataset. If their spatial distribution deviates from the core cluster and their temporal fluctuations exceed a preset range, outliers are corrected using time-series smoothing techniques, including moving averages and exponential smoothing, resulting in a validated, stable mirrored dataset. A density-based spatial clustering algorithm (such as the Local Outlier Factor, LOF) calculates the outlier degree of each grid point, identifying outliers whose spatial distribution deviates from the core cluster and whose temporal fluctuations exceed the limit. The K-nearest neighbor algorithm calculates the neighborhood mean as a correction reference, and cubic spline interpolation is used to fit the time-series curve. For example, if the vibration data of a device suddenly becomes abnormal, smoothing correction is performed using the vibration mean of neighboring devices and historical curve trends. This secondary validation ensures spatial neighborhood consistency and temporal continuity, generating a stable dataset that provides reliable data support for risk prediction in Step S8, solving the problem of inaccurate outlier handling in traditional methods.

[0040] Step S8 primarily involves multi-scenario risk simulation and prediction. Based on the mirrored stable dataset, it constructs multi-scenario operating conditions to meet project risk prediction needs. Using methods including Monte Carlo simulation and finite element analysis, it simulates project behavior patterns under different combinations of key indicators, identifies potential risk points exceeding risk thresholds, and generates a structured simulation analysis dataset. Scenarios such as normal operation, equipment failure, and extreme weather are defined, and the fluctuation range of key indicators is set (e.g., ±30% fluctuation of motor current in an equipment failure scenario). Monte Carlo simulation simulates project behavior by randomly generating a large number of indicator sequences and inputting them into the digital mirror model. Finite element analysis is used to simulate physical characteristics such as structural safety, for example, the stress distribution of building structures under extreme weather conditions. The Mahalanobis distance between the simulation output and the normal scenario is calculated to identify abnormal behavior points. An LSTM model is used to predict the duration and impact range of anomalies, generating structured data containing risk probability and impact degree. This step addresses the lack of multi-scenario simulation in traditional risk prediction methods, providing a basis for dynamic adjustments in S9.

[0041] Step S9 primarily involves generating a dynamic adjustment plan. This involves quantifying the differences in operational indicators under different scenarios using techniques including principal component analysis and cluster comparison on the simulation analysis dataset. Combined with a regression model of project configuration parameters and indicator performance, a dynamic adjustment plan encompassing equipment scheduling, resource allocation, and process adjustments is generated. Principal component analysis reduces the dimensionality of multi-scenario simulation data and extracts key features; K-means clustering identifies abnormal scenarios, calculates the Euclidean distance to the baseline scenario, and determines key adjustment parameters (such as equipment power and personnel configuration). A "configuration parameter-indicator performance" model is established through multiple linear regression to predict the adjustment magnitude, for example, predicting the required increase in construction personnel based on the risk of schedule delays. Combined with resource constraints, a priority ranking algorithm is used to generate equipment scheduling, resource allocation, and process adjustment plans, which are then output after simulation verification. This step addresses the lack of dynamic adjustment capabilities in traditional methods, achieving a closed loop from risk prediction to decision support.

[0042] The data analysis method provided in this embodiment forms a complete data processing closed loop through core technologies such as spatiotemporal grid standardization, multi-dimensional data verification, time series optimization, digital mirror construction, multi-scenario simulation, and dynamic adjustment. Specifically: by dividing the data into three-dimensional spatiotemporal grids, heterogeneous data from multiple stages such as planning, design, and construction are uniformly mapped, solving the fusion difficulties caused by inconsistent data formats in traditional methods and improving data integration efficiency; dynamic weight allocation rules, spatiotemporal interpolation algorithms, and outlier correction techniques improve the accuracy of data quality detection and the outlier identification rate, providing a more reliable data foundation for risk prediction. Based on digital twin multi-scenario simulation (Monte Carlo, finite element analysis) and real-time data synchronization mechanisms, the system can quickly respond to abnormal changes in project operation, and the risk prediction lead time is longer than that of traditional methods, enhancing dynamic adaptability. Through principal component analysis, regression models, and priority ranking algorithms, the generated equipment scheduling and resource allocation schemes can improve project resource utilization and reduce construction costs, effectively solving the problem of one-sided decision support in traditional methods, and comprehensively improving the intelligence level and scientific nature of digital management in engineering consulting.

[0043] In one implementation, step S1, which maps heterogeneous data from different projects onto a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset, specifically includes:

[0044] S11: Collect raw data streams containing spatial coordinates, timestamps, and business attributes from project stages;

[0045] S12: The original data stream is normalized in terms of spatial coordinates and time units by using a three-dimensional grid partitioning algorithm. Data from different data sources is converted to a unified grid coordinate system through a mapping function to obtain a preliminary standardized mapping dataset.

[0046] S13: For the pre-normalized mapping dataset, based on the pre-set out value detection threshold for each monitoring scenario data feature, mark data points that exceed the outlier detection threshold as outlier data points to form an annotated outlier dataset.

[0047] S14: Perform spatiotemporal autocorrelation analysis on the abnormal dataset, calculate the spatial clustering degree and time series correlation of the abnormal data points, determine whether the abnormal distribution has spatial clustering or time periodicity characteristics, and obtain the abnormal distribution feature set.

[0048] S15: Use the support vector machine algorithm to classify the risk level of abnormal data points in the abnormal distribution feature set, and generate a risk level heat map by combining the grid spatial location. Mark the areas with risk levels greater than or equal to the preset risk level and overlapping with the key nodes of the project as key attention areas.

[0049] S16: For key areas of concern, the pattern matching degree between real-time data streams and historical abnormal data is calculated using a dynamic time warping algorithm. The risk evolution trend is judged based on the slope of the matching degree curve. The risk level and evolution direction are mapped to a preliminary standardized mapping dataset to form a preliminary standardized grid mapping dataset containing risk characteristics.

[0050] This embodiment refines step S1 above, aiming to improve the risk prediction capability in the data standardization process through more refined anomaly detection and risk feature embedding. Specifically:

[0051] Step S11 primarily involves collecting raw data streams. "Project phases" encompass the entire project lifecycle: planning and design (e.g., CAD drawings, BIM model data), construction (e.g., progress reports, equipment monitoring data), and operation (e.g., energy consumption monitoring data, maintenance records). The "raw data stream" includes spatial coordinates (e.g., the 3D coordinates of various building components), timestamps (e.g., the specific time of data collection), and business attributes (e.g., concrete strength, equipment operating parameters). Its sources include IoT sensors, manual data entry systems, and design documents, and the data format may be heterogeneous, such as XML, GeoJSON, or database tables. This step provides raw data input for subsequent processing; the completeness and accuracy of the collected data directly impact the subsequent analysis results.

[0052] Step S12 primarily involves 3D mesh generation and coordinate normalization. The "3D mesh generation algorithm" divides the space into grids of fixed granularity (e.g., 10m × 10m × 5m) and the time dimension into fixed time intervals (e.g., 1 hour / grid) based on project scale and accuracy requirements. "Spatial coordinate normalization" transforms different coordinate systems (e.g., WGS84, Beijing 54) to a unified grid coordinate system using a coordinate transformation matrix; "time unit normalization" converts timestamps of different formats (e.g., milliseconds, seconds) to a unified time unit (e.g., UTC seconds). Mapping functions are customized based on the data source type; for example, component coordinates from the BIM model are mapped to the grid coordinate system through projection transformation, and sensor time-series data is mapped to the time dimension grid according to the acquisition time. This step generates a "preliminary normalized mapping dataset," providing a unified data foundation for anomaly detection in S13.

[0053] Step S13 primarily involves outlier detection and labeling. The "outlier detection threshold" is preset based on the historical data statistical characteristics of each monitoring scenario. For example, the normal range for concrete curing temperature is 20±5℃; values ​​exceeding this range are labeled as outliers. A threshold comparison algorithm is used to detect each data point in the preliminary normalized dataset, marking points exceeding the threshold as "outlier data points," thus forming an independent "outlier dataset." For instance, if the rebar stress value at a certain grid point exceeds the design threshold by 20% at a given moment, it is labeled as an outlier. This step initially filters out clearly anomalous data, preparing for subsequent in-depth analysis of anomaly characteristics.

[0054] Step S14 primarily involves spatiotemporal autocorrelation analysis. This analysis uses Geary's C or Moran's I index to calculate the spatial clustering of anomalous data points (e.g., the probability of multiple adjacent grids exhibiting anomalies simultaneously) and their temporal correlation (e.g., whether the anomalies repeat within a fixed period). For example, if multiple grids in a region show PM2.5 levels exceeding standards within the same time period for three consecutive days, autocorrelation analysis can determine that they exhibit spatial clustering and temporal periodicity. The "anomaly distribution feature set" includes spatial clustering indices (e.g., cluster radius) and temporal correlation coefficients, used to characterize the distribution patterns of anomalies and provide feature input for risk level classification in S15.

[0055] Step S15 primarily involves risk level classification and heatmap generation. The Support Vector Machine (SVM) algorithm takes anomaly distribution feature sets (such as spatial clustering, temporal correlation, and anomaly magnitude) as input to train a classification model that categorizes anomalous data points into low, medium, and high risk levels. For example, anomalies with high spatial clustering and strong temporal periodicity are classified as high-risk. Combining grid spatial location data, a "risk level heatmap" is generated using GIS technology, with color intensity indicating the level of risk. "Key project nodes" refer to areas that significantly impact project progress or quality (such as the core tube construction area or key equipment installation procedures). When the risk level is greater than or equal to a preset value (e.g., medium) and coincides with a key node, it is marked as a "key concern area." This step transforms data anomalies into a visualized risk assessment, clearly identifying areas requiring focused monitoring.

[0056] Step S16 primarily involves determining the risk evolution trend and mapping features. The "Dynamic Time Warping (DTW) algorithm" is used to calculate the pattern matching degree between real-time data streams and historical anomaly data, such as the similarity between the current equipment vibration data sequence and the vibration patterns before historical failures. The "matching degree curve slope" reflects the rate of change of the pattern matching degree over time; an increasing slope indicates a worsening risk trend. For example, a rapid increase in the matching degree from 0.3 to 0.8, with a slope of 0.5 / hour, indicates that the risk is escalating. The risk level and evolution direction (deterioration / mitigation) are mapped to a pre-normalized mapping dataset, ensuring that each grid point contains risk characteristics. This step extends the process from anomaly detection to risk trend prediction, providing data with risk characteristics for subsequent steps and enhancing the risk warning capability during the data standardization process.

[0057] This embodiment significantly improves the risk prediction capability of engineering consulting data analysis by embedding spatiotemporal anomaly detection, risk level classification, and trend prediction during the data standardization stage. Combining threshold detection, spatiotemporal autocorrelation analysis, and SVM classification increases the accuracy of anomaly data identification, especially for potential anomalies with spatiotemporal correlation (such as progressive equipment failures), achieving precise identification of anomaly data. By analyzing risk evolution trends using the DTW algorithm, the risk warning lead time is extended beyond the immediate detection of traditional methods, providing more time for project management to respond, thus achieving proactive risk warning. For example, trend analysis of bridge monitoring data predicted support displacement anomalies two days in advance, preventing accidents. Based on risk heatmaps and key node correlations, precise location of high-risk areas is achieved, centralizing monitoring resources and improving monitoring efficiency. For example, in high-rise building construction, high-risk areas in the concrete pouring stage are accurately identified, strengthening on-site inspections and reducing the incidence of quality accidents. By embedding risk characteristics into standardized datasets, subsequent data analysis (such as format verification in S2 and risk prediction in S8) can directly utilize risk characteristics, thereby improving the overall relevance and accuracy of the analysis and providing more forward-looking data support for the digital management of engineering consulting.

[0058] In one embodiment, step S2, which involves detecting data format and attribute conflicts in the grid mapping dataset and adjusting the data priority in the grid mapping dataset using a preset dynamic weight allocation rule to obtain a grid calibration dataset with a consistent format, specifically includes:

[0059] S21: The DBSCAN clustering algorithm is used to perform spatial density clustering of grid points, identify high-density clustered areas and sparsely distributed areas, and obtain grid grouped datasets.

[0060] S22: For grid points within each cluster group, extract the mean, variance, and rate of change features of the time series, calculate the change amplitude of adjacent time points through a sliding window, and mark those exceeding the preset fluctuation threshold as dynamic outliers;

[0061] S23: Extract the format parameters of dynamic anomalies, including data type, data length, and precision, and compare the format parameters with industry standard data format templates. Mark the mismatches as format anomalies and record the mismatch type. Also, analyze the business attributes of dynamic anomalies, verify the logical consistency between attributes through a decision tree rule engine, mark the contradictions as attribute conflict points, and record the conflict rules.

[0062] S24: Perform principal component analysis on format outliers and attribute conflict points, extract the first N principal components as feature vectors, calculate the cosine similarity with the historical outlier feature library, and mark those with similarity less than the preset similarity threshold as new outlier patterns.

[0063] S25: Based on the Kriging interpolation method, predict the potential expansion range of new anomaly patterns in the grid, determine whether they affect key areas by calculating the neighborhood correlation degree, adjust the priority of data affecting key areas according to the rule of "real-time data > historical data" and "high reliability source data > low reliability source data", and generate a grid calibration dataset with a consistent format.

[0064] This embodiment refines step S2 above. Through spatial density clustering, dynamic anomaly detection, anomaly pattern recognition, and priority adjustment, it achieves accurate calibration of grid mapping data, solving the problems of incomplete data format conflict detection and unintelligent processing in traditional methods. Specifically:

[0065] Step S21 primarily involves spatial density clustering. The "DBSCAN (Density-Based Spatial Clustering for Noise) algorithm" divides the data into high-density clustered regions (such as densely packed sensor areas in construction sites) and sparsely distributed regions (such as surrounding environmental monitoring points) based on the spatial distribution density of grid points. By setting the neighborhood radius ε and the minimum number of points MinPts (e.g., ε=50m, MinPts=5), adjacent grid points are clustered together, while noise points are grouped separately. For example, monitoring points for equipment such as tower cranes and material hoists at a construction site form a high-density cluster, while surrounding environmental monitoring points form a sparse cluster. The "grid grouping dataset" provides a grouping basis for subsequent anomaly detection based on regional characteristics, avoiding misjudgments caused by globally uniform detection.

[0066] Step S22 primarily involves detecting dynamic outliers. For each cluster of grid points, the statistical characteristics of the time series are calculated: mean (e.g., the average operating current of a device over a week), variance (the degree of data fluctuation), and rate of change (e.g., the hourly current change). A "sliding window" (e.g., a window size of 24 hours with a step size of 1 hour) is used to calculate the change in adjacent time points. If the change exceeds a preset fluctuation threshold (e.g., mean ± 20%), it is marked as a "dynamic outlier." For example, if the vibration data of a wind turbine suddenly increases by 30% within 30 minutes, exceeding the preset threshold, it is marked as a dynamic anomaly. This step combines spatial grouping and time series characteristics to accurately locate data fluctuation anomalies, providing targets for subsequent format and attribute verification.

[0067] Step S23 primarily involves detecting format anomalies and attribute conflicts. "Format parameters" include data type (e.g., vibration data should be floating-point but is integer), length (e.g., temperature sensor data should have three decimal places but only one), and precision (e.g., GPS coordinate precision should be meter-level but is kilometer-level). The format parameters of dynamic anomalies are compared with industry standard templates (e.g., the format specified in the "Construction Engineering Data Exchange Standard"). Mismatches are marked as "format anomalies," and the type of mismatch (e.g., incorrect type, insufficient length) is recorded. "Business attribute logic consistency" verification is implemented through a decision tree rule engine. For example, if "concrete strength ≥ 28 MPa, curing time should ≥ 28 days," and a data point has a strength of 30 MPa but a curing time of only 10 days, it is determined to be an attribute conflict, marked as an "attribute conflict point," and the conflict rule is recorded. This step detects data problems from both format and business logic dimensions, ensuring the standardization and logical consistency of the data.

[0068] Step S24 primarily involves identifying newly added anomaly patterns. Principal Component Analysis (PCA) is performed on format anomalies and attribute conflict points to extract the top N principal components (e.g., N=3) that explain more than 80% of the variance as feature vectors. Cosine similarity is calculated between these features and a historical anomaly feature library (which stores previously detected anomaly pattern features). Patterns with a similarity score less than a preset threshold (e.g., 0.6) are identified as "newly added anomaly patterns." For example, if a certain type of sensor data simultaneously exhibits type errors and insufficient length, and its similarity to a historical anomaly pattern is only 0.4, it is identified as a newly added anomaly. This step utilizes machine learning methods to identify unknown anomaly patterns, improving the system's adaptability and anomaly detection capabilities.

[0069] Step S25 primarily involves anomaly expansion prediction and priority adjustment. The "Kriging interpolation method" utilizes spatial autocorrelation to predict the potential expansion range of newly detected anomalies within the grid (e.g., a sensor malfunction might affect similar sensors within a 100m radius). "Neighborhood correlation" calculates the spatial distance and impact weight between the anomaly expansion range and key project areas (e.g., core tube, equipment room). If a key area is affected, priority adjustment rules are activated: "Real-time data > Historical data" (e.g., current sensor-collected real-time data takes precedence over yesterday's historical data) and "High-reliability source data > Low-reliability source data" (e.g., automated sensor data takes precedence over manually entered data). Conflicting data is reassigned through weight allocation to generate a "consistently formatted grid calibration dataset." For example, if there are conflicts in concrete strength data for a key area, real-time data from high-precision sensors is prioritized to replace manually entered historical data. This step enables intelligent processing of anomaly data, ensuring the accuracy of data in key areas.

[0070] This embodiment significantly improves the accuracy and intelligence of data format calibration through spatial grouping detection, dual-dimensional anomaly verification, new pattern recognition, and intelligent priority adjustment. The comprehensiveness of anomaly detection is enhanced; combining spatial density clustering and dynamic temporal features significantly improves the detection coverage of format anomalies and attribute conflicts compared to traditional methods, especially enhancing the detection capability for anomalies in sparsely distributed areas (such as data anomalies from remote environmental monitoring points). New anomaly pattern discovery is achieved through PCA and cosine similarity analysis, enabling the identification of new anomaly patterns not previously observed. This gives the system self-learning capabilities and improves the anomaly pattern detection rate. For example, communication protocol anomalies of new sensors are promptly identified, preventing large-scale data errors. Key area data is guaranteed; based on Kriging interpolation-based anomaly propagation prediction and priority adjustment rules, the accuracy and reliability of key area data are ensured. The accuracy rate of key area data is significantly improved compared to traditional methods, providing solid data support for decision-making at key project nodes. Data processing intelligence is enhanced; dynamic weight allocation rules automate the handling of data conflicts, reducing manual intervention, improving processing efficiency, and avoiding errors caused by human judgment, making the data calibration process more scientific and intelligent.

[0071] In one embodiment, step S3, which involves using a time-series autoregressive model and smoothing techniques to process missing values ​​and abnormal fluctuations in the grid calibration dataset to obtain a time-stable dataset, specifically includes:

[0072] S31: The ARIMA model is used to fit the time series of the grid calibration dataset, extract the seasonal features with a period of T, and calculate the periodic change intensity of each grid point. Among them, the change intensity is marked as a high dynamic grid point.

[0073] S32: The spatial distribution density of high dynamic grid points is calculated by kernel density estimation method. The Delaunay triangulation algorithm is used to divide the boundary of high density region. Isolated grid points with neighborhood association strength less than the preset association strength threshold are classified as isolated points.

[0074] S33: Establish a spatial buffer zone centered on the isolated point, calculate the number of grid points and data influence weights within the buffer zone coverage area, and fuse spatial density and periodic change intensity through a weighted overlay model;

[0075] S34: For grid points with missing values ​​in the time series of the fused data, the inverse distance weighted interpolation method is used to fill in the missing values ​​using historical data from the neighboring N×N grid; for abnormal fluctuation points, spatiotemporal constraints are eliminated by combining the mean of the previous M periods to generate a time-stable dataset, where M and N are both positive integers.

[0076] This embodiment focuses on the stability processing of time series data. By combining the ARIMA model with spatial analysis algorithms, it addresses the problem of the single time dimension data processing in traditional methods. Specifically:

[0077] Step S31 primarily involves identifying periodic features and marking high-dynamic points. The "ARIMA (Autoregressive Integral Moving Average) model" identifies the periodic characteristics of the data by fitting the autoregressive term, differencing order, and moving average term of the time series. For example, construction progress data often exhibits a weekly cycle (weekend shutdowns cause progress to slow down), and the ARIMA model can extract seasonal patterns with a period of T=7 days. The "intensity of periodic variation" is determined by calculating the ratio of the data fluctuation amplitude within the period to the mean. For instance, if the concrete curing temperature of a certain grid point fluctuates by more than 30% of the mean within a 7-day period, it is marked as a "high-dynamic grid point." This step provides the temporal dimension of feature input for the spatial density analysis in S32, allowing subsequent processing to focus on data points with drastic changes.

[0078] Step S32 primarily involves spatial density analysis and outlier classification. "Kernel density estimation" identifies data-dense areas (such as monitoring points in the center of building complexes) by calculating the spatial distribution density of highly dynamic grid points. "Delaunay triangulation" divides high-density areas into continuous geometric shapes, determining boundary ranges. "Neighborhood correlation strength" is determined by calculating the spatial distance and data correlation (such as Euclidean distance and Pearson correlation coefficient) between grid points and their neighbors. If the strength is less than a threshold (e.g., correlation coefficient < 0.3 and distance > 100m), it is classified as an "outlier" (such as a remote environmental monitoring station). This step combines the highly dynamic characteristics of the temporal dimension with spatial distribution, providing spatial location data for the buffer analysis in S33.

[0079] Step S33 primarily involves buffer analysis and feature fusion. A spatial buffer (e.g., a circular area with a radius of 50m) is established centered on the isolated point. The number of grid points within the buffer and the influence weight of each point on the isolated point are calculated (the closer the point, the higher the weight). The "weighted superposition model" linearly fuses spatial density (high-density areas have higher weights) and periodic variation intensity (higher intensity has higher weights) to generate a comprehensive weight. For example, if there are three highly dynamic grid points within 50m of an isolated point, with distances of 20m, 30m, and 40m respectively, corresponding to weights of 0.5, 0.3, and 0.2, and combined with its periodic variation intensity weight of 0.4, the final comprehensive weight is 0.4 × (0.5 + 0.3 + 0.2) = 0.4. This step provides the weight calculation basis for missing value imputation in S34, making the interpolation more consistent with the spatiotemporal distribution characteristics.

[0080] Step S34: Missing Value Imputation and Anomaly Removal. The "Inverse Distance Weighted Interpolation Method" uses historical data from a neighborhood N×N grid (e.g., 5×5) to calculate missing values ​​by weighting them according to the inverse of distance. For example, if the humidity data for a grid point at time t is missing, the historical humidity values ​​of the 10 nearest points in the surrounding 5×5 grid are taken and summed according to distance weights to obtain the imputed value. The "Spatiotemporal Constraint Rule" combines the mean of the preceding and following M periods (e.g., M=3) with spatial neighborhood data to remove outliers: if the temperature data at a certain moment is 50% higher than the mean of the preceding and following 3 periods and there are no similar fluctuations in the neighborhood, it is judged as an anomaly and removed. The "Time-Stable Dataset" generated in this step serves as the input to S4, ensuring the time series is continuous and free of abnormal fluctuations, laying the foundation for subsequent timeline construction.

[0081] This embodiment significantly improves the stability and integrity of time series data through dual-dimensional temporal and spatial data processing. Specifically: The accuracy of periodic feature extraction is improved; the ARIMA model's accuracy in identifying periodic features is significantly higher than traditional simple statistical methods (accuracy approximately 60%), especially in capturing complex cycles (such as irregular fluctuations in construction progress during the rainy season). The accuracy of missing value imputation is improved; the combination of inverse distance weighted interpolation and spatiotemporal weights reduces the error in missing value imputation. For example, in isolated point regions, the error of traditional single-time interpolation is ±15%, while this method reduces the error to ±9%. The reliability of outlier removal is enhanced; the spatiotemporal constraint rules significantly improve the outlier identification rate compared to traditional methods, avoiding misjudgments caused by single-time dimension removal (such as misjudging seasonal fluctuations as anomalies). The temporal stability of the data is optimized; the processed data time series exhibits high continuity, providing high-quality data for subsequent timeline construction in S4 and digital mirror modeling in S5.

[0082] In one embodiment, step S4, which involves analyzing the grid point timestamp intervals of the time-series stable dataset and constructing a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions by combining linear interpolation and multidimensional data fusion, specifically includes:

[0083] S41: Calculate the time interval between adjacent timestamps of the same grid point in the time-series stable dataset. If the time interval is greater than the preset time threshold, trigger linear interpolation.

[0084] S42: The intermediate data points generated by interpolation are optimized using multidimensional data fusion technology, including:

[0085] S43: Check the data fluctuation before and after the interpolation point for a preset time length. If the fluctuation is greater than the preset fluctuation threshold, introduce exponential smoothing correction. Also, calculate the synchronous data correlation of a preset number of grid points in the neighborhood. If the correlation is lower than the preset correlation threshold, supplement the spatial correlation features through Gaussian process regression.

[0086] S44: Construct a spatiotemporal integrity assessment model, and supplement the missing parts of grid points with more than a preset number of time axis break points or more than a preset percentage of missing spatial neighborhood data by comparing with historical data from the same period.

[0087] S45: Generate a complete grid dataset that is time-continuous and has consistent spatiotemporal dimensions through spatiotemporal consistency verification.

[0088] This embodiment aims to construct a continuous time axis and a consistent spatiotemporal dimension. By using linear interpolation and multidimensional data fusion, it addresses the shortcomings of traditional methods in data integrity processing. Specifically:

[0089] Step S41 primarily involves time interval detection and interpolation triggering. The "timestamp interval" refers to the time difference between the acquisition of two adjacent data points at the same grid point. A preset threshold is determined based on the data acquisition frequency (e.g., a threshold of 30 minutes for high-frequency monitoring data). If the interval exceeds the threshold (e.g., a 2-hour interval for data from a certain sensor), linear interpolation is triggered, for example, generating n time points at equal intervals between timest1 and t2. This step provides the basic time points for the interpolation optimization in S42, ensuring the initial continuity of the time axis.

[0090] Steps S42-S43 primarily involve multi-dimensional data fusion and optimization. "Exponential smoothing correction" handles situations where data fluctuates drastically before and after the interpolation point by assigning higher weights to recent data (e.g., α=0.8) to smooth noise. For example, if the data fluctuation before and after an interpolation point exceeds 20% of the mean, exponential smoothing is used to adjust the interpolation result to a weighted average of the last three time points. "Gaussian process regression" supplements spatial correlation features: when the correlation of neighboring grid point data is below a threshold (e.g., Pearson correlation coefficient < 0.5), a Gaussian kernel function is used to fit the spatial distribution and predict the spatial correlation features of the interpolation point. For example, if the temperature interpolation point of a certain grid point has low correlation with neighboring points, Gaussian process regression is used to combine the temperature gradients of the surrounding five points to supplement the spatial temperature field features. This step optimizes the temporal and spatial characteristics of the interpolation point, providing more reliable data for the integrity assessment in S44.

[0091] Step S44 primarily involves assessing spatiotemporal integrity and supplementing historical data. The "Spatiotemporal Integrity Assessment Model" sets a timeline break threshold (e.g., more than 5 break points for the same grid point) and a spatial missing threshold (e.g., more than 40% missing neighboring data). For grid points exceeding the threshold, supplementation is achieved through "historical data comparison": for example, if July temperature data for a grid point is missing, the average of historical data from the same period last July is extracted as a supplementary value. This step addresses the issue of large-scale data loss, ensuring that the spatiotemporal consistency verification in S45 has a complete data foundation.

[0092] Step S45 primarily performs spatiotemporal consistency verification. This is done through the following methods: 1. Time axis continuity verification (interval between any two adjacent timestamps ≤ a threshold); 2. Spatial dimension consistency verification (continuous change in spatial feature gradients of neighboring grid points). For example, it checks whether the elevation of adjacent grid points conforms to the terrain trend, avoiding spatial abrupt changes caused by interpolation. After successful verification, a "complete grid dataset" is generated, serving as input for extracting key indicators in S5 to ensure the spatiotemporal consistency of the digital mirror modeling.

[0093] This embodiment significantly improves the spatiotemporal continuity and integrity of data through multi-dimensional data fusion and integrity assessment. Specifically: Timeline continuity is significantly improved, meaning the rate of exceeding time stamp intervals decreases, and the temporal continuity of high-frequency monitoring data is higher, providing a precise time reference for subsequent dynamic analysis (such as real-time synchronization in S6). The accuracy of spatial correlation features is improved; after Gaussian process regression supplements spatial features, the correlation of neighboring data increases. For example, the interpolation error of elevation data in complex terrain areas decreases from ±5m to ±2m, better reflecting the actual terrain distribution. The ability to handle large-scale missing data is enhanced; the spatiotemporal integrity assessment model improves the efficiency of handling large-scale missing data. Traditional methods typically discard data with a missing rate >30%, while this method can supplement data using historical data from the same period, greatly improving data utilization. Spatiotemporal consistency is guaranteed; the spatiotemporal dimension consistency of the verified data is higher, avoiding analytical errors caused by time jumps or spatial abrupt changes in traditional methods. This provides a high-quality data foundation for digital mirror modeling in S5, improving the reliability of subsequent risk prediction.

[0094] In one embodiment, step S5, which involves extracting key indicators from the complete grid dataset, constructing a digital mirror of the project including spatial location, time series, and attribute relationships, and mapping multi-period data uniformly to a baseline timeline based on the project duration to obtain an aligned mirrored dataset, specifically includes:

[0095] S51: Define the set of key performance indicators (K) for the project K={ (schedule), (energy consumption) (Quality)}, extract the spatiotemporal distribution data of each period K from the complete grid dataset;

[0096] S52: Use digital twin technology to construct a digital mirror containing geometric model, physical model and business model, and establish a three-dimensional mapping relationship of "grid coordinates-time stamp-indicator value";

[0097] S53: Map the data from each monitoring cycle to a baseline time axis based on the total project duration through time offset correction and frequency unification to form a preliminary aligned dataset;

[0098] S54: For outliers in the initially aligned dataset where the index fluctuation exceeds the index fluctuation threshold, local weighted regression is used for smoothing correction. Calculate the rate of change of indicators across cycles and generate a mirrored dataset containing trend features;

[0099] S55: The final mirror mapping dataset is obtained by verifying the consistency of the baseline time axis and the accuracy of the mirror mapping.

[0100] This embodiment uses digital twin technology to construct a digital mirror of the project, achieving timeline alignment of multi-period data and providing a standardized model for subsequent real-time data synchronization and risk prediction. Specifically:

[0101] Step S51 mainly involves extracting key performance indicators (KPIs). "Key project KPIs" are defined based on engineering consulting requirements, such as schedule indicators. Energy consumption indicators for the completion rate of sub-projects (such as the percentage of completion of basic engineering). Electricity consumption per unit area ( ), quality indicators The concrete compressive strength (MPa) is calculated. Spatiotemporal distribution data for each indicator is extracted from the "Complete Grid Dataset" according to monitoring periods (e.g., weeks, months), for example, extracting concrete strength data for each grid point in weeks 1-4. This step provides core data support for the digital twin modeling of S52.

[0102] Step S52 primarily involves constructing a digital mirror image. "Digital twin technology" constructs the mirror image using the following models: 1. Geometric model (BIM-based 3D building model); 2. Physical model (e.g., the heat conduction equation for concrete); 3. Business model (e.g., CPM network plan for schedule management). A 3D mapping of "grid coordinates - timestamp - index value" is established; for example, mapping the concrete strength value of a grid point (X, Y, Z) at time t to the corresponding component in the BIM model. The digital mirror image generated in this step serves as the carrier for time axis alignment in S53, ensuring unified spatial positioning of multi-period data.

[0103] Step S53 primarily involves time axis alignment. "Time offset correction" adjusts the time origin of data for each period (e.g., unifying progress data from different construction phases to the project start date as the origin), and "frequency unification" resamples data from different collection frequencies (e.g., daily progress, weekly energy consumption) to a unified frequency (e.g., weekly data). For example, daily progress data is aggregated into weekly progress percentages and mapped to a baseline time axis (e.g., week 1, week 2, etc.). This step forms a "preliminary aligned dataset," providing a time-unified data foundation for anomaly correction in S54.

[0104] Step S54 primarily involves anomaly correction and trend feature generation. "Locally Weighted Regression (LWLR)" smooths outliers where index fluctuations exceed thresholds (e.g., strength index fluctuations > 15%) by assigning higher weights to neighboring points to fit the curve. For example, if concrete strength data suddenly decreases by 20% in a given week, LWLR is used to correct this by combining data from the preceding and following two weeks. "Cross-period change rate ΔK" is used to quantify trends, such as the progress index changing from week 1 (progress...) ) to the second week (progress) rate of change This process generates upward / downward trend features. This step infuses the mirrored dataset with dynamic trend information, providing a trend reference for verifying the accuracy of S55 and for subsequent real-time synchronization in S6.

[0105] Step S55 primarily involves verifying the mirror mapping. "Baseline Timeline Consistency Verification" checks whether the data for each period is accurately mapped to the baseline time point (e.g., whether the data for week 3 corresponds to the 3rd node on the timeline). "Mirror Mapping Accuracy Verification" compares the deviation between the actual monitored data and the mirror prediction data (e.g., the error between the actual progress and the mirror prediction progress is <5%). Upon successful verification, a "Final Mirror Mapping Dataset" is generated, serving as the baseline model for real-time data synchronization in S6, ensuring the accuracy of subsequent data updates.

[0106] This embodiment significantly improves the integration and analytical value of project data through digital twin modeling and timeline alignment. Specifically: Unified management of multi-period data maps heterogeneous data from planning, construction, and operation phases to a baseline timeline, improving data integration efficiency and resolving the low analytical efficiency caused by the scattered storage of multi-period data in traditional methods. Improved accuracy of digital mirroring: The three-dimensional mapping of geometric, physical, and business models ensures a high degree of matching between the mirror and the actual project. For example, the stress distribution simulation error of building structures is less than 8%, providing a reliable digital carrier for subsequent risk prediction (such as multi-scenario simulation in S8). Enhanced trend feature quantification capability: Cross-period change rate calculation improves the accuracy of indicator trend quantification, making it more objective than traditional manual judgment. For example, the early identification rate of schedule delay trends is improved, providing data support for dynamic adjustments in S9. Guaranteed data consistency: Timeline and mirror mapping verification ensure high data consistency, avoiding analytical biases caused by time misalignment or spatial mapping errors in traditional methods, and improving the accuracy of subsequent real-time data synchronization (such as S6).

[0107] In one embodiment, step S6, which involves acquiring the latest monitoring data through a sensor interface, and if key indicators are missing, performing a weighted average based on the spatial distance of neighboring grid points and the correlation of historical data, and dynamically adjusting the weighting coefficients to supplement missing values ​​in conjunction with data fluctuation amplitude, to update the mirror mapping dataset and obtain the mirror updated dataset, specifically includes:

[0108] S61: Collect sensor data in real time through the IoT interface, align the timestamp with the mirrored dataset, and trigger the data synchronization mechanism;

[0109] S62: For grid points with missing key indicators, a two-step method is used to fill them in, including: taking the real-time data of the nearest A×A grid points and performing an inverse distance weighted average, where A is a positive integer; and automatically adjusting the weights based on the fluctuation range of historical data.

[0110] S63: Perform spatial distribution verification on the filled data. If the difference with the neighboring grid data is greater than the preset difference threshold, it is marked as a suspicious point and corrected by the average of the data before and after the preset time.

[0111] S64: Calculate the deviation between real-time data and the predicted values ​​of the mirror model. For regions where the deviation is greater than the deviation threshold, the distribution characteristics are refitted using a Gaussian mixture model to generate a mirror updated dataset.

[0112] This embodiment achieves real-time updates of the mirrored dataset. Through spatiotemporal weighted imputation and bias correction, it addresses the problems of lag in real-time data processing and coarse-grained missing value imputation in traditional methods. Specifically:

[0113] Step S61 primarily involves real-time data acquisition and time alignment. The "IoT interface" connects to various sensors (such as RFID and vibration sensors) to collect data in real time (such as equipment operating parameters and environmental indicators). Real-time data is aligned with the "mirror-mapped dataset" through timestamp matching (e.g., accurate to the second), triggering a synchronization mechanism (e.g., synchronizing every 10 minutes). For example, the real-time load data collected by tower crane sensors is mapped to the corresponding grid points in the mirror model according to the timestamp. This step provides real-time data input for missing value filling in S62.

[0114] Step S62 primarily involves a two-step missing value imputation method. The "inverse distance-weighted average" method takes real-time data from neighboring A×A grids (e.g., 3×3) and calculates the missing value by weighting it according to the inverse of the distance; the closer the distance, the higher the weight (e.g., 0.8 weight for a distance of 5m, 0.4 weight for a distance of 10m). The "dynamic weight adjustment" method automatically corrects for historical data fluctuations: when fluctuations are large (e.g., historical standard deviation of load data > 10%), the weight of neighboring points is increased (e.g., the weight coefficient is multiplied by 1.2), and the weight of historical data is decreased; conversely, when fluctuations are small, the opposite is true. For example, if wind speed data for a certain grid point is missing, the nearest point in the neighboring 3×3 grid has a weight of 0.6; due to large historical fluctuations, the weight is adjusted to 0.7. This step, combining spatiotemporal characteristics and historical patterns, imputs missing values, providing preliminary imputation results for the spatial verification in S63.

[0115] Step S63 primarily involves spatial distribution verification and correction of suspicious points. The difference between the filled data and the neighboring grid data (e.g., Euclidean distance) is calculated. If the difference exceeds a threshold (e.g., difference > 20%), it is marked as a "suspicious point." For example, if the filled temperature for a grid point is 35℃, while neighboring points are all around 25℃, the difference exceeds the threshold. By adjusting the average of the data over a preset time period (e.g., 2 hours), the suspicious point is adjusted to an average of 26℃. This step ensures that the filled data conforms to spatial distribution patterns, providing reliable real-time data for the deviation analysis in S64.

[0116] Step S64 primarily involves bias analysis and distribution refitting. "Bias" = |Real-time Data - Mirror Predicted Value| / Mirror Predicted Value. If the bias exceeds a threshold (e.g., 15%), it indicates a significant deviation between the mirror model and the actual data. The "Gaussian Mixture Model" refits the data distribution characteristics of the biased region using a linear combination of multiple Gaussian distributions. For example, if the energy consumption data bias in a certain region consistently exceeds 20%, a Gaussian Mixture Model is used to fit the high, medium, and low energy consumption distributions, updating the prediction parameters of the mirror model. This step generates a "Mirror Update Dataset," which serves as input to S7, ensuring the mirror model dynamically optimizes with real-time data.

[0117] This embodiment significantly improves the real-time performance and accuracy of the mirror model through real-time data synchronization and intelligent data completion. Specifically: Real-time data processing efficiency is improved, with the IoT interface enabling minute-level data synchronization, significantly more efficient than traditional manual import (once a day). This allows the mirror model to reflect the project's current status in real time, providing data support for immediate response to sudden risks (such as equipment failure). Missing value completion accuracy is improved, with a two-step completion method combining spatial distance, historical correlation, and dynamic weighting, reducing the error. For example, when wind speed data is missing, the traditional single-distance weighted error is ±8%, while this method reduces the error to ±5%. Spatial distribution rationality is enhanced, with spatial verification ensuring high neighborhood consistency of the completed data, avoiding spatial abrupt changes (such as a sudden abnormal rise in temperature data) caused by isolated completion in traditional methods, making the spatial characteristics of the mirror model more consistent with the actual scenario. The model's dynamic optimization capability is improved, with the Gaussian mixture model refitting biased regions, continuously optimizing the prediction accuracy of the mirror model. Initial prediction errors are high, but decrease after multiple updates, providing a more accurate model foundation for subsequent outlier detection in S7 and risk prediction in S8.

[0118] In one embodiment, step S7, which involves identifying outlier grid points in the mirror update dataset and correcting these outlier grid points to obtain a verified mirror stable dataset, specifically includes:

[0119] S71: The local outlier factor algorithm is used to calculate the outlier degree of each grid point in the mirror update dataset. Grid points with an outlier degree > 2 are marked as outlier grid points.

[0120] S72: Perform spatiotemporal dual correction on outlier grid points, including using the K-nearest neighbor algorithm to calculate the mean of neighboring grid points as a correction reference value; and using cubic spline interpolation to fit the data curves of the previous and subsequent 24 hours to generate a smoothed time series.

[0121] S73: Perform a second verification on the corrected data, calculate spatial neighborhood consistency and time series continuity, and further correct areas that do not meet the standards through manual review rules;

[0122] S74: Finally, a mirror-stable dataset is generated through spatiotemporal stability evaluation.

[0123] This embodiment addresses the inaccuracy of outlier handling in traditional methods by identifying outliers and performing spatiotemporal correction, ensuring the stability of mirrored data. Specifically:

[0124] Step S71 primarily involves outlier identification. The "Local Outlier Factor (LOF) algorithm" determines outlier status by calculating the density ratio between a grid point and its neighbors: if a point's density is significantly lower than its neighbors (e.g., LOF > 2), it is marked as an "outlier grid point." For example, if the LOF value of a device's vibration data is 2.5, it indicates that its spatial distribution deviates from the core cluster. This step provides target points for the correction in S72, focusing on truly outlier points that deviate from the overall distribution.

[0125] Step S72 mainly involves spatiotemporal dual correction. The "K-nearest neighbor algorithm" selects the K nearest neighbors of the outlier (e.g., K=5) and calculates their mean as the correction reference value. For example, the vibration value of a certain outlier is... The mean of the 5 neighboring points is Preliminary revision to Cubic spline interpolation is used to fit the 24-hour data curves before and after outliers, ensuring a smooth and continuous time series and avoiding abrupt changes in data after correction. For example, vibration data of outliers are embedded into the curves before and after the outliers to generate a smoothed sequence. This step combines spatial neighborhood features and time series trends for dual correction, providing preliminary stable data for the secondary verification of S73.

[0126] Step S73 primarily involves secondary verification and manual correction. "Spatial Neighborhood Consistency" verifies the difference between outliers and their neighbors (e.g., Euclidean distance < 15%), while "Time Series Continuity" verifies the smoothness of the first-order difference of the corrected data (e.g., absolute difference < 10%). For areas that do not meet the standards (e.g., difference > 20%), manual review rules are initiated. For example, engineers determine whether it is a genuine anomaly (e.g., equipment failure) or a data error (e.g., sensor failure). If it is a data error, forced correction using the neighborhood mean is applied. This step compensates for the limitations of algorithmic correction, ensuring the accuracy of critical data.

[0127] Step S74 primarily involves assessing spatiotemporal stability. This is evaluated using the following metrics: 1. Mean spatial outlier < 1.5; 2. Standard deviation of the time series < 10% of the historical mean; 3. Neighborhood consistency rate > 95%. Upon successful assessment, a "mirror-stabilized dataset" is generated, serving as input for risk prediction in step S8 to ensure the reliability of the simulated data.

[0128] This embodiment significantly improves the stability and reliability of mirrored data through precise outlier identification and spatiotemporal correction. Specifically: outlier identification accuracy is improved; the LOF algorithm's accuracy in identifying outliers is significantly higher than the traditional threshold method (accuracy of approximately 70%), especially enhancing its ability to identify "hidden outliers" (such as slowly deviating device parameters). The rationality of outlier correction is enhanced; i.e., spatiotemporal dual correction results in high neighborhood consistency and temporal continuity of the corrected data. For example, after anomaly correction, vibration data aligns with the vibration trend of neighboring devices, avoiding spatial logic contradictions caused by traditional single-time smoothing. Data stability is significantly optimized; i.e., the spatiotemporal fluctuation amplitude of the corrected data is reduced, and the standard deviation decreases significantly relative to the historical mean, providing a stable data foundation for S8's multi-scenario simulation and improving the credibility of risk prediction results. The efficiency of manual review is improved; i.e., secondary verification automatically locates non-compliant areas, reducing the workload of manual review, focusing on key outliers, and improving the overall efficiency of data processing.

[0129] In one embodiment, step S8, which involves constructing multi-scenario operating conditions based on the mirror-stabilized dataset to generate a structured simulation analysis dataset for project risk prediction, specifically includes:

[0130] S81: Define a collection of multiple scenarios Set the fluctuation range of key indicators for each scenario;

[0131] S82: Use the Monte Carlo simulation method to generate indicator sequences for each scenario, input them into the digital mirror model for behavioral simulation, and record the output parameters;

[0132] S83: Perform anomaly detection on the simulated output data, calculate the Mahalanobis distance between each indicator and the normal scene, and mark points with a distance > 2 as abnormal behavior points;

[0133] S84: Perform time series trend analysis on abnormal behavior points, use LSTM model to predict the duration and scope of impact of abnormalities, and generate risk feature vectors containing the probability of risk occurrence and the degree of impact.

[0134] S85: By using data fusion technology, risk characteristics are associated with grid spatial location and timestamps, and the simulation analysis dataset is output in a structured format of "risk level-occurrence time-affected area".

[0135] This embodiment addresses the lack of dynamism and multi-scenario coverage in traditional risk prediction methods through multi-scenario simulation and risk feature extraction. Specifically:

[0136] Step S81 mainly involves defining multiple scenarios and setting indicator fluctuations. The "multi-scenario set" is defined based on common risk scenarios in engineering consulting, such as: For normal operation (indicator fluctuation ±10%) This could be due to equipment malfunction (such as tower crane motor current fluctuations of +30% to +50%). For extreme weather conditions (e.g., wind speed > 25 m / s), the fluctuation range of key indicators is set for each scenario as input conditions for the S82 simulation.

[0137] Step S82 primarily involves Monte Carlo simulation and behavior recording. The "Monte Carlo simulation" generates a large number of sequences that conform to the fluctuation range of scenario indicators (e.g., generating 1000 sets of current sequences for equipment failure scenarios), and inputs them into a digital mirror model (e.g., a structural mechanics model, a schedule simulation model) for simulation. Output parameters are recorded, such as the stress distribution of the structural model under extreme weather conditions and the number of days of project delay in the schedule model due to equipment failure. This step provides a massive amount of simulation data for anomaly detection in S83.

[0138] Step S83 primarily involves identifying anomalous behavior points. The "Mahathano distance" measures the difference between the simulated output indicators and the normal scenario, considering the correlation between indicators (such as the covariance of stress and displacement). Points with a distance > 2 (i.e., exceeding twice the standard deviation) are marked as "abnormal behavior points." For example, if the Mahalano distance of structural stress under a certain simulated condition is 2.3, it indicates a significant deviation from the normal scenario. This step transforms the simulated data into risk signals, providing target points for the trend analysis in S84.

[0139] Step S84 primarily involves risk trend prediction and feature vector generation. An LSTM (Long Short-Term Memory) network models the time series of anomalous behavior points, predicting the duration and scope of the anomaly. For example, an anomaly in structural stress is predicted by LSTM to last 48 hours and affect three surrounding grid areas. The risk feature vector includes: 1. Probability of occurrence (e.g., frequency of anomaly occurrence / number of simulations); 2. Degree of impact (e.g., percentage of stress exceeding design value). This step expands single-point anomalies into trend predictions, providing quantitative risk characteristics for the structured output of S85.

[0140] Step S85 primarily involves establishing a spatiotemporal correlation between risk characteristics. Using GIS spatial analysis technology, risk characteristics (probability, severity) are correlated with grid coordinates and timestamps to generate structured data (e.g., JSON format): `{"Risk Level": "High", "Occurrence Time": "2025-07-05 14:00", "Affected Area": ​​"50m surrounding grid (10,20,5)"}`. The "Simulation Analysis Dataset" output from this step serves as the input for the dynamic adjustment in S9, ensuring that the adjustment plan has a clear risk target.

[0141] This embodiment significantly improves the comprehensiveness and foresight of risk prediction through multi-scenario simulation and intelligent prediction. Specifically: Enhanced multi-scenario coverage: More than three typical scenarios are defined, covering extreme conditions not considered by traditional methods (such as equipment combination failures and multiple natural disasters), resulting in higher risk scenario coverage. Improved risk prediction accuracy: Monte Carlo simulation combined with LSTM prediction reduces the prediction error of risk occurrence probability and impact range, making it more accurate than traditional experience-based judgments. For example, the lead time for structural risk prediction under extreme weather conditions is extended from 12 hours to 48 hours. Enhanced risk feature quantification: Mahalanobis distance and risk feature vectors enable quantitative expression of risk. The vague judgment of "potential risk" in traditional methods is transformed into precise descriptions such as "risk probability 75%, impact exceeding threshold 20%", providing a quantitative basis for decision-making. Standardized structured data output: A unified risk structured format allows data to be directly input into the S9 dynamic adjustment model, achieving seamless integration from risk prediction to decision support and improving decision-making efficiency.

[0142] In one embodiment, step S9, which generates a dynamic adjustment scheme including equipment scheduling, resource allocation, and process adjustment based on the simulation analysis dataset, specifically includes:

[0143] S91: Principal component analysis is used to reduce the dimensionality of multi-scene simulation data in the simulation analysis dataset. The top X principal components are extracted as scene feature vectors, and abnormal scenes are identified by K-means clustering.

[0144] S92: Quantify the differences in indicators under abnormal scenarios, calculate the Euclidean distance between each abnormal scenario and the benchmark scenario, and use indicators whose Euclidean distance is greater than the preset distance threshold as key adjustment parameters.

[0145] S93: Establish a "configuration parameter-indicator performance" model through multiple linear regression to predict the adjustment range of parameters including equipment power and personnel configuration.

[0146] S94: Based on the adjustment range and resource constraints, a priority sorting algorithm is used to determine the resource allocation order, generate a dynamic adjustment scheme that includes equipment scheduling plan, personnel configuration scheme and process parameter adjustment table, and output the final scheme after simulation verification.

[0147] This embodiment generates an operable dynamic adjustment plan through data dimensionality reduction, indicator difference quantification, and regression modeling, addressing the lack of scientific rigor in traditional decision support methods. Specifically:

[0148] Step S91 primarily involves scene dimensionality reduction and anomaly identification. Principal Component Analysis (PCA) reduces the dimensionality of multi-scene simulation data (e.g., 1000 sets of simulation data, each containing 20 indicators), extracting the top X principal components (e.g., X=3) that explain more than 85% of the variance, serving as the scene feature vectors. K-means clustering groups the feature vectors, identifying "abnormal scenes" that differ significantly from normal scenes (e.g., cluster centers are more than 2 standard deviations away from normal scenes). For example, a composite scene of equipment failure and extreme weather is clustered as a high-risk abnormal scene.

[0149] Step S92 primarily involves determining key adjustment parameters. The Euclidean distance between the abnormal scenario and the baseline scenario (normal operation) is calculated. Indicators where the distance is greater than a threshold (e.g., 2) are identified as "key adjustment parameters." For example, if the Euclidean distance between the structural stress in an abnormal scenario and the baseline scenario is 2.5, the stress index is listed as a key parameter. This step clarifies the adjustment target for regression modeling in S93.

[0150] Step S93 primarily involves predicting adjustments to configuration parameters. "Multiple linear regression" establishes a mapping model between configuration parameters (e.g., equipment power P, number of personnel N) and performance indicators (e.g., stress σ, construction period D). The model parameters are trained using historical data (e.g., σ = 0.5P + 0.3N + Z), where Z represents the random error term, indicating other influencing factors not explicitly expressed in the model (e.g., environmental variables not included in the model, measurement errors, model simplification errors, etc.). The predicted adjustment magnitude for key parameters is then calculated; for example, to reduce stress σ by 20%, it is predicted that equipment power P needs to increase by 15% and the number of personnel N needs to increase by 10%.

[0151] Step S94 primarily involves generating dynamic adjustment schemes. The "priority sorting algorithm" determines the allocation order based on the adjustment magnitude and resource constraints (such as the upper limit of available equipment power and the maximum number of personnel). For example, equipment power adjustment has a higher priority than personnel allocation. Specific schemes are generated: 1. Equipment scheduling plan (e.g., activating a backup generator); 2. Personnel allocation plan (e.g., adding 20 construction workers); 3. Process adjustment table (e.g., extending concrete curing time to 30 days). The effectiveness of the scheme is verified through simulation (e.g., whether the stress σ drops below the threshold after adjustment), and the final scheme is output.

[0152] This embodiment significantly improves the scientific rigor and operability of decision support through data-driven dynamic adjustments. Specifically: The accuracy of indicator difference quantification is improved; PCA dimensionality reduction combined with Euclidean distance calculation results in higher accuracy in quantifying scenario differences, transforming traditional experience-based indicator adjustments into data-driven, precise decision-making. The scientific rigor of the adjustment scheme is enhanced; the multiple linear regression model reduces the prediction error of configuration parameter adjustment magnitudes, for example, reducing the deviation between predicted equipment power adjustments and actual needs, avoiding over- or under-adjustment. Resource allocation efficiency is optimized; the priority ranking algorithm improves resource utilization. For example, in equipment failure scenarios, the resource allocation scheme can shorten the project delay from 7 days to 3 days, while keeping cost increases within 10%. The operability of the scheme is improved; the structured scheduling plan, configuration scheme, and process adjustment table directly guide on-site execution, transforming vague suggestions such as "strengthening monitoring" in traditional methods into concrete operational steps, improving execution efficiency.

[0153] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A data analysis method for intelligent management of engineering consulting, characterized in that, The method includes: Heterogeneous data from different projects are mapped onto a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset. The data format and attribute conflicts in the grid mapping dataset are detected, and the data priority in the grid mapping dataset is adjusted by a preset dynamic weight allocation rule to obtain a grid calibration dataset with a consistent format. For the aforementioned grid calibration dataset, a time series autoregressive model and smoothing techniques are used to handle missing values ​​and abnormal fluctuations, resulting in a time-stable dataset. The grid point timestamp intervals of the time-series stable dataset are analyzed, and a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions is constructed by combining linear interpolation and multidimensional data fusion. Key indicators are extracted from the complete grid dataset to construct a digital mirror of the project that includes spatial location, time series, and attribute associations. Multi-period data are uniformly mapped to a benchmark timeline based on the project duration to obtain an aligned mirror mapping dataset. The latest monitoring data is obtained through the sensor interface. If key indicators are missing, a weighted average is calculated based on the spatial distance of neighboring grid points and the correlation of historical data. The weight coefficients are dynamically adjusted in combination with the data fluctuation amplitude to supplement the missing values. The mirror mapping dataset is then updated to obtain the mirror update dataset. Outlier grid points in the mirror update dataset are identified and corrected to obtain a verified mirror stable dataset. Based on the aforementioned mirror-stabilized dataset, multi-scenario operating conditions are constructed to meet the project risk prediction needs, generating a structured simulation analysis dataset. Based on the simulation analysis dataset, a dynamic adjustment scheme including equipment scheduling, resource allocation, and process adjustment is generated.

2. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The process of mapping heterogeneous data from different projects onto a unified three-dimensional spatiotemporal grid to obtain a preliminarily standardized grid mapping dataset includes: Collect raw data streams containing spatial coordinates, timestamps, and business attributes from each stage of the project. A three-dimensional mesh partitioning algorithm is used to normalize the spatial coordinates and unify the time units of the original data stream. Data from different data sources is transformed to a unified mesh coordinate system through a mapping function to obtain a preliminary standardized mapping dataset. For the pre-standardized mapping dataset, based on the outlier detection thresholds pre-set according to the data characteristics of each monitoring scenario, data points exceeding the outlier detection thresholds are marked as outlier data points, forming an annotated outlier dataset; Spatiotemporal autocorrelation analysis is performed on the abnormal dataset to calculate the spatial clustering degree and time series correlation of the abnormal data points, determine whether the abnormal distribution has spatial clustering or time periodicity characteristics, and obtain the abnormal distribution feature set. The support vector machine algorithm is used to classify the risk level of abnormal data points in the abnormal distribution feature set. Combined with the grid spatial location, a risk level heat map is generated, and areas with a risk level greater than or equal to the preset risk level and overlapping with the key nodes of the project are marked as key areas of concern. For key areas of concern, the pattern matching degree between real-time data streams and historical anomaly data is calculated using a dynamic time warping algorithm. The risk evolution trend is judged based on the slope of the matching degree curve. The risk level and evolution direction are mapped to a preliminary standardized mapping dataset, forming a preliminary standardized grid mapping dataset containing risk characteristics.

3. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The process of detecting data format and attribute conflicts in the grid mapping dataset involves adjusting the data priority in the grid mapping dataset using a preset dynamic weight allocation rule to obtain a grid calibration dataset with a consistent format, including: The DBSCAN clustering algorithm is used to perform spatial density clustering of grid points to identify high-density clustered areas and sparsely distributed areas, thus obtaining a grid grouped dataset. For each cluster group of grid points, the mean, variance, and rate of change of the time series are extracted. The change amplitude of adjacent time points is calculated by a sliding window, and points that exceed the preset fluctuation threshold are marked as dynamic anomalies. Extracting dynamic anomalies includes format parameters such as data type, data length, and precision. The format parameters are then compared with industry standard data format templates. Mismatches are marked as format anomalies and the mismatch type is recorded. Additionally, the business attributes of dynamic anomalies are analyzed. The logical consistency between attributes is verified through a decision tree rule engine. Inconsistencies are marked as attribute conflict points and conflict rules are recorded. Principal component analysis is performed on format outliers and attribute conflict points. The top N principal components are extracted as feature vectors. The cosine similarity with the historical outlier feature library is calculated. If the similarity is less than the preset similarity threshold, it is marked as a new outlier pattern. Based on the Kriging interpolation method, the potential expansion range of new anomaly patterns in the grid is predicted. The impact on key areas is determined by the neighborhood correlation degree calculation. The priority of data affecting key areas is adjusted according to the rules of "real-time data > historical data" and "high reliability source data > low reliability source data", and a grid calibration dataset with consistent format is generated.

4. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The aforementioned grid calibration dataset is processed using a time-series autoregressive model and smoothing techniques to handle missing values ​​and abnormal fluctuations, resulting in a time-stable dataset, including: The ARIMA model is used to fit the time series of the grid calibration dataset, and seasonal features with a period of T are extracted. The periodic variation intensity of each grid point is calculated, and grid points with a variation intensity greater than a preset intensity threshold are marked as high dynamic grid points. The spatial distribution density of high dynamic grid points is calculated by kernel density estimation method. The Delaunay triangulation algorithm is used to divide the boundary of high density region. Isolated grid points with neighborhood association strength less than the preset association strength threshold are classified as isolated points. A spatial buffer is established with isolated points as the center. The number of grid points and the data influence weight within the buffer coverage area are calculated. The spatial density and the intensity of periodic changes are fused through a weighted overlay model. For grid points with missing values ​​in the time series of the fused data, the inverse distance weighted interpolation method is used to fill in the missing values ​​using historical data from the neighboring N×N grid. For abnormal fluctuation points, spatiotemporal constraints are eliminated by combining the mean of the previous M periods to generate a time-stable dataset, where M and N are both positive integers.

5. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The analysis of the grid point timestamp intervals of the time-series stable dataset, combined with linear interpolation and multidimensional data fusion, constructs a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions, including: Calculate the time interval between adjacent timestamps of the same grid point within the time-stable dataset. If the time interval is greater than a preset time threshold, trigger linear interpolation. The intermediate data points generated by interpolation are optimized using multidimensional data fusion technology, including: Check the data fluctuation before and after the interpolation point for a preset time length. If the fluctuation is greater than the preset fluctuation threshold, introduce exponential smoothing correction. Also, calculate the correlation of synchronous data of a preset number of grid points in the neighborhood. If the correlation is lower than the preset correlation threshold, supplement the spatial correlation features through Gaussian process regression. A spatiotemporal integrity assessment model is constructed. For grid points with more than a preset number of time axis break points or more than a preset percentage of missing spatial neighborhood data, the missing parts are supplemented by comparison with historical data from the same period. By verifying spatiotemporal consistency, a complete grid dataset that is both temporally continuous and spatiotemporally consistent is generated.

6. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The process involves extracting key indicators from the complete grid dataset, constructing a digital mirror of the project that includes spatial location, time series, and attribute relationships, and mapping multi-period data uniformly to a baseline timeline based on the project duration to obtain an aligned mirrored dataset, including: Define the set of key metrics for the project The spatiotemporal distribution data of each period K is extracted from the complete grid dataset, where, For the sake of progress, For energy consumption, For quality; Digital twin technology is used to construct a digital mirror containing a geometric model, a physical model, and a business model, and to establish a three-dimensional mapping relationship of "grid coordinates - timestamp - indicator value"; Data from each monitoring cycle are mapped to a baseline time axis based on the total project duration through time offset correction and frequency unification to form a preliminary aligned dataset. For outliers in the initially aligned dataset where the index fluctuation exceeds the index fluctuation threshold, local weighted regression is used for smoothing correction. Calculate the rate of change of indicators across cycles and generate a mirrored dataset containing trend features; The final mirrored mapping dataset is obtained by verifying the consistency of the baseline time axis and the accuracy of the mirrored mapping.

7. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The process involves acquiring the latest monitoring data through a sensor interface. If key indicators are missing, a weighted average is calculated based on the spatial distance of neighboring grid points and the correlation with historical data. The weighting coefficients are dynamically adjusted in conjunction with data fluctuation amplitude to supplement missing values. This process updates the mirrored mapping dataset, resulting in a mirrored updated dataset, which includes: Sensor data is collected in real time through the IoT interface, and timestamps are aligned with the mirrored dataset to trigger a data synchronization mechanism. For grid points with missing key indicators, a two-step method is used to fill them in: taking the real-time data of the nearest A×A grid points and performing an inverse distance weighted average, where A is a positive integer; and automatically adjusting the weights based on the fluctuation range of historical data. Spatial distribution verification is performed on the filled data. If the difference with the neighboring grid data is greater than the preset difference threshold, it is marked as a suspicious point and corrected by the average of the data before and after the preset time period. The deviation between real-time data and the predicted values ​​of the mirror model is calculated. For regions where the deviation is greater than the deviation threshold, the distribution characteristics are refitted using a Gaussian mixture model to generate a mirror updated dataset.

8. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The process of identifying outlier grid points in the mirror update dataset and correcting these outlier grid points to obtain a verified mirror stable dataset includes: The local outlier factor algorithm is used to calculate the outlier degree of each grid point in the mirror update dataset. Grid points with an outlier degree > 2 are marked as outlier grid points. Outlier grid points are corrected in both time and space. This includes using the K-nearest neighbor algorithm to calculate the mean of neighboring grid points as a correction reference value, and using cubic spline interpolation to fit the data curves of the previous and next 24 hours to generate a smoothed time series. The corrected data is then subjected to a second verification, which calculates spatial neighborhood consistency and temporal series continuity. Areas that do not meet the standards are further corrected through manual review rules. Finally, a mirror-stable dataset is generated through spatiotemporal stability evaluation.

9. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, Based on the mirror-stabilized dataset, multi-scenario operating conditions are constructed to meet the project risk prediction needs, generating a structured simulation analysis dataset, including: Define a collection of multiple scenarios For each scenario, a fluctuation range for key indicators is set, among which, For normal operation, Equipment malfunction. For extreme weather; The Monte Carlo simulation method is used to generate indicator sequences for each scenario, which are then input into a digital mirror model for behavioral simulation, and the output parameters are recorded. Anomaly detection is performed on the simulated output data, and the Mahalanobis distance between each indicator and the normal scene is calculated. Points with a distance > 2 are marked as abnormal behavior points. Time series trend analysis is performed on abnormal behavior points, and the LSTM model is used to predict the duration and scope of impact of the anomalies, generating a risk feature vector that includes the probability of risk occurrence and the degree of impact. By using data fusion technology, risk characteristics are associated with grid spatial location and timestamps, and simulation analysis datasets are output in a structured format of "risk level-occurrence time-affected area".

10. The data analysis method for intelligent management of engineering consulting according to claim 1, characterized in that, The process of generating a dynamic adjustment scheme based on the simulation analysis dataset, including equipment scheduling, resource allocation, and process adjustment, includes: Principal component analysis is used to reduce the dimensionality of multi-scenario simulation data in the simulation analysis dataset. The top X principal components are extracted as scene feature vectors, and abnormal scenes are identified by K-means clustering. The differences in indicators under abnormal scenarios are quantified, and the Euclidean distance between each abnormal scenario and the benchmark scenario is calculated. Indicators with an Euclidean distance greater than a preset distance threshold are used as key adjustment parameters. A "configuration parameter-indicator performance" model was established using multiple linear regression to predict the adjustment range of parameters including equipment power and personnel configuration. Based on the adjustment range and resource constraints, a priority sorting algorithm is used to determine the resource allocation order, and a dynamic adjustment scheme including equipment scheduling plan, personnel configuration plan and process parameter adjustment table is generated. After simulation verification, the final scheme is output.

Citation Information

Patent Citations

  • Hydraulic engineering data intelligent monitoring method and system based on digital twinning

    CN119783356A

  • Prefabricated building component stability monitoring and early warning method and system

    WO2025118307A1