Data analysis method for engineering consultation digital intelligent management
Through space-time grid standardization and multi-scenario simulation, the problems of low data integration efficiency and insufficient accuracy in engineering consulting data analysis have been solved, and the intelligent and scientific transformation of engineering consulting management has been achieved.
Patent Information
- Application Number
- CN202511159045.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing engineering consulting data analysis methods are insufficient in the comprehensiveness, accuracy and dynamic adaptability of data processing, and are unable to meet the needs of digital management of engineering consulting, especially in terms of low data integration efficiency, insufficient analysis accuracy, incomplete detection of inconsistent data formats, single time dimension processing, and lack of multi-scenario simulation and dynamic adjustment capabilities.
A spatiotemporal grid standardization method is used to map heterogeneous data into a unified framework, and format conflicts are eliminated through dynamic weight allocation rules. Time series autoregressive models and spatial interpolation algorithms are combined to handle missing values and outliers. A digital mirror is constructed and multi-scenario simulation is performed to generate dynamic adjustment plans.
It achieves efficient integration and format consistency of heterogeneous data, improves the accuracy and real-time performance of data analysis, can accurately identify potential risks and generate scientific resource scheduling plans, and promotes the transformation of engineering consulting management towards intelligence and scientific management.
Smart Images

Figure CN120653941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a data analysis method for digital management of engineering consulting. Background Art
[0002] In today's rapidly developing digital era, the engineering consulting industry is accelerating its transformation toward digital and intelligent management. Engineering consulting projects often involve multiple stages, including planning, design, construction, and operations. Each stage generates massive amounts of heterogeneous data, including geospatial data, project progress data, and equipment monitoring data. This data contains critical information about project operations. Efficient analysis can assist in project decision-making, optimize resource allocation, and predict potential risks. Therefore, data analysis has become a core component of digital and intelligent management in engineering consulting.
[0003] At present, traditional engineering consulting data analysis methods mostly use a combination of manual processing and simple statistical tools. They are difficult to deal with complex heterogeneous data and have problems such as low data integration efficiency and insufficient analysis accuracy. With the development of technology, some methods have tried to use grid division, time series analysis and other technologies for data processing, but there are still many defects. For example, in the data standardization stage, existing technologies have difficulty in accurately mapping heterogeneous data from different project sources and different format standards into a unified framework, resulting in difficulties in data fusion; in the data quality detection link, the detection of inconsistent data formats or attribute conflicts is not comprehensive and accurate enough, and potential data contradictions cannot be effectively identified; in terms of data integrity processing, the method of handling missing data values and abnormal fluctuation points in the time dimension is relatively simple, which is difficult to meet the needs of complex engineering scenarios; in addition, existing technologies lack the ability to simulate and adjust based on multiple scenarios and dynamic adjustments in project risk prediction and configuration optimization, and cannot provide comprehensive and accurate decision support for engineering consulting.
[0004] To sum up, the existing engineering consulting data analysis methods have obvious deficiencies in the comprehensiveness, accuracy and dynamic adaptability of data processing, and are unable to meet the growing demand for digital management of engineering consulting. There is an urgent need for a more efficient, accurate and dynamically adjustable data analysis method to improve the intelligence level of engineering consulting project management and the scientific nature of decision-making. Summary of the Invention
[0005] The main purpose of this invention is to provide a data analysis method for digital management of engineering consulting, aiming to solve the technical problem that the existing engineering consulting data analysis methods have obvious deficiencies in the comprehensiveness, accuracy and dynamic adaptability of data processing, and are difficult to meet the growing demand for digital management of engineering consulting.
[0006] In order to achieve the above-mentioned object of the invention, the first aspect of the present invention proposes a data analysis method for digital management of engineering consulting, the method comprising: Obtain raw business data streams from multiple project links, and use a pre-established spatiotemporal grid partitioning method to map heterogeneous data from different projects into a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset; For the grid mapping dataset, by calculating the format compatibility of grid points including data type, length, precision and logical consistency of business attributes, data format inconsistency or attribute conflict problems are detected. If data format inconsistency or attribute conflict problems exist, the data priority is adjusted according to the reliability of the data source and the update time through the preset dynamic weight allocation rules to obtain a grid calibration dataset with consistent format; For the grid calibration dataset, periodic features are identified based on the time series autoregressive model, missing values in the time dimension are filled in with a spatial interpolation algorithm, and abnormal fluctuation points are eliminated through preset spatiotemporal constraint rules to obtain a time series stable dataset; Analyze the grid point timestamp intervals of the time-stable dataset. If the intervals exceed a preset time threshold, use linear interpolation to generate time points. Combined with multidimensional data fusion technology, optimize the intermediate data points to construct a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions. Extracting key project indicators for each monitoring period from the complete grid data set, constructing a digital mirror of the project containing spatial location, time series, and attribute associations using virtual entity modeling technology, and uniformly mapping multi-period data to a benchmark timeline based on the project duration to obtain an aligned mirror mapping data set; Implement real-time data synchronization on the mirror mapping dataset, obtain the latest monitoring data through the sensor interface, and if key indicators are missing, perform weighted averaging based on the spatial distance of adjacent grid points and the correlation of historical data, and dynamically adjust the weight coefficient based on the data fluctuation amplitude to supplement the missing values to obtain a mirror updated dataset; For the mirror update dataset, a density-based spatial clustering algorithm is used to identify outlier grid points. If their spatial distribution deviates from the core cluster and the time series fluctuation exceeds a preset range, the outliers are corrected using time series smoothing techniques including moving average and exponential smoothing to obtain a verified mirror stable dataset. Based on the mirrored stable dataset, multiple scenario operating conditions are constructed to meet the project risk prediction requirements. Project behavior patterns under different key indicator combinations are simulated using methods including Monte Carlo simulation and finite element analysis to identify potential risk points that exceed risk thresholds and generate a structured simulation analysis dataset. For the simulation analysis data set, the differences in operating indicators under different scenarios are quantified through techniques including principal component analysis and cluster comparison. Combined with the regression model of project configuration parameters and indicator performance, a dynamic adjustment plan including equipment scheduling, resource allocation, and process adjustment is generated.
[0007] Beneficial effects: The data analysis method for digital management of engineering consulting of the present invention maps heterogeneous data in multiple links to a unified framework by constructing a standardized space-time grid system, thereby solving the problem of low efficiency of traditional integration; with the help of format compatibility verification and dynamic weight allocation, it eliminates data format conflicts and attribute contradictions, and improves analysis accuracy; relying on time series optimization and digital mirror modeling, it achieves continuity and consistency in the space-time dimensions of data, laying the foundation for real-time monitoring; through multi-scenario simulation and dynamic adjustment algorithms, it accurately identifies potential risks and generates resource scheduling plans, so that engineering consulting management forms a complete closed loop from data collection to decision support, effectively enhancing the system's dynamic adaptability to complex working conditions, promoting the transformation of project management to intelligent and scientific methods, and realizing full-cycle management of engineering data and intelligent risk decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 This is a flow chart of a data analysis method for digital management of engineering consulting according to an embodiment of the invention; the realization of the purpose, functional features and advantages of the invention will be further explained in conjunction with the embodiment and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0009] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0010] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "above", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements, modules, modules and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, modules, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any module and all combinations of one or more associated listed items.
[0011] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which this invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.
[0012] Reference Figure 1 , an embodiment of the present invention provides a data analysis method for digital management of engineering consulting, including: S1: Map heterogeneous data from different projects into a unified 3D spatiotemporal grid to obtain a preliminary standardized grid mapping dataset; S2: Detecting data format and attribute conflicts in the grid mapping dataset, adjusting data priorities in the grid mapping dataset using a preset dynamic weight allocation rule, and obtaining a grid calibration dataset with consistent format; S3: For the grid calibration dataset, using a time series autoregressive model and smoothing technology to process missing values and abnormal fluctuations, to obtain a time series stable dataset; S4: Analyze the grid point timestamp intervals of the time-series stable dataset, and construct a complete grid dataset with a continuous time axis and consistent spatiotemporal dimensions by combining linear interpolation and multidimensional data fusion; S5: extracting key indicators from the complete grid dataset, constructing a digital mirror of the project including spatial location, time series, and attribute associations, and uniformly mapping multi-period data to a benchmark timeline based on the project duration to obtain an aligned mirror mapping dataset; S6: Obtain the latest monitoring data through the sensor interface. If key indicators are missing, weighted average is performed based on the spatial distance between adjacent grid points and the correlation of historical data. The weight coefficient is dynamically adjusted based on the data fluctuation amplitude to supplement the missing values. The mirror mapping dataset is updated to obtain a mirror update dataset. S7: Identify outlier grid points in the mirrored updated data set, and correct the outlier grid points to obtain a verified mirrored stable data set; S8: Based on the mirror stable data set, construct multi-scenario operating conditions according to project risk prediction requirements and generate a structured simulation analysis data set; S9: Based on the simulation analysis data set, a dynamic adjustment plan including equipment scheduling, resource allocation, and process adjustment is generated.
[0013] This embodiment relates to a data analysis method for digital and intelligent management of engineering consulting. Its core is to solve the problems of low efficiency in heterogeneous data integration, insufficient analysis accuracy, and lack of dynamic adaptability in traditional methods through multi-dimensional data processing and dynamic modeling. Specifically: Step S1 primarily involves spatiotemporal grid standardization. This involves acquiring raw business data streams from multiple project stages and mapping the heterogeneous data from these different projects onto a unified three-dimensional spatiotemporal grid using a pre-established spatiotemporal grid partitioning method, thereby generating a preliminarily standardized grid-mapped dataset. "Raw business data streams" refer to heterogeneous data collected from the planning, design, construction, and operation stages of engineering consulting projects, containing spatial coordinates (e.g., latitude, longitude, and altitude), timestamps (e.g., UTC time), and business attributes (e.g., progress percentage, equipment operating parameters). These data may be in formats such as CSV, JSON, or GIS vector data. The "spatiotemporal grid partitioning method" divides three-dimensional space (X, Y, and Z axes) into grids of fixed or dynamic granularity, adding time as the fourth dimension to form a unified spatiotemporal coordinate system. Data from different data sources is converted to this coordinate system using mapping functions (e.g., projection transformation and time unit normalization). For example, the spatial coordinates of a building BIM model and time series data from meteorological monitoring can be mapped onto the same grid, thereby achieving a "preliminary standardized grid-mapped dataset." The purpose of this step is to solve the problem of inconsistent formats of heterogeneous data in traditional methods and provide a unified framework for subsequent processing. Its output serves as the input data basis of S2.
[0014] Step S2 primarily involves data format calibration and conflict resolution. Specifically, the grid mapping dataset is inspected for format inconsistencies or attribute conflicts by calculating the format compatibility of grid points, including data type, length, and precision, as well as the logical consistency of business attributes. If any data format inconsistencies or attribute conflicts exist, data priority is adjusted based on data source reliability and update time using a pre-set dynamic weighting rule to obtain a consistent grid calibration dataset. A "grid point" is the smallest unit in a three-dimensional spatiotemporal grid and consists of spatial coordinates (X, Y, Z) and a timestamp (t). The spatial coordinates correspond to the physical location of the project site (e.g., latitude, longitude, and altitude). The timestamp is in UTC with second-level accuracy. For example, the grid point (100,200,5,2025-07-02T10:00:00Z) represents the monitoring location at 10:00:00 on July 2, 2025, at X = 100m, Y = 200m, and Z = 5m. "Format compatibility" involves whether the data type (e.g., integer, floating point), length (e.g., number of string characters), and precision (e.g., number of decimal places) conform to industry standards or preset templates. "Business attribute logical consistency" also involves, for example, whether there are inconsistencies in the logical relationship between concrete curing temperature and strength growth. The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm is used to identify data-dense and data-sparse areas. Dynamic anomalies (e.g., sudden changes in equipment monitoring data) are then format and attribute verified. If conflicts are detected (e.g., temperature differences between different data sources at the same location exceeding a threshold), the priority is adjusted and the values are reassigned based on the principles of "real-time data > historical data" and "high-reliability source data > low-reliability source data" (e.g., BIM (Building Information Modeling) data provided by a design institute is more reliable than manually entered data by construction teams). This step addresses the incomplete data quality checks inherent in traditional methods, ensuring that the data input to S3 is uniform in format and logically consistent, thus avoiding discrepancies in subsequent analysis.
[0015] Step S3 primarily performs time series stabilization. Specifically, the grid calibration dataset is subjected to periodicity analysis using a time series autoregressive model. Spatial interpolation algorithms are then used to fill missing values in the time dimension. Pre-set spatiotemporal constraints are then used to eliminate anomalous fluctuations, resulting in a time-stabilized dataset. Time series autoregressive models (such as ARIMA) are used to identify periodic data, such as the weekly cycle of construction progress. Spatial interpolation algorithms (such as inverse distance weighted interpolation) use historical data from neighboring grids to fill missing values. For example, if concrete strength data for a particular grid point is missing, it is calculated by weighting the strength values of the surrounding 5×5 grids. Spatiotemporal constraints combine the mean of the preceding and subsequent periods with spatial neighborhood data to eliminate anomalous points. For example, if temperature data at a given moment is significantly higher than the historical average for the same period and no similar fluctuations are observed in neighboring grids, it is considered an anomaly and removed. This step addresses the issue of single-dimensional data processing in traditional methods, ensuring that the input data for S4 is continuous and stable in the time series, laying the foundation for constructing a complete timeline.
[0016] Step S4 primarily establishes spatiotemporal data integrity. This involves analyzing the grid point timestamp intervals of the time-stable dataset. If these exceed a preset time threshold, linear interpolation is used to generate time points. Multidimensional data fusion techniques are then used to optimize intermediate data points, constructing a grid-complete dataset with a continuous time axis and consistent spatiotemporal dimensions. The intervals between adjacent timestamps are calculated. If these exceed a preset threshold (e.g., 30 minutes), linear interpolation is triggered to generate intermediate time points. For example, if a device's monitoring data is missing two hours apart, linear fitting of the preceding and following time points is used to supplement the data. Multidimensional data fusion techniques include exponential smoothing (to address fluctuation anomalies) and Gaussian process regression (to supplement spatial correlation features). For example, when data before and after the interpolation point fluctuate significantly, exponential smoothing is used to reduce noise. When correlation between neighboring data is low, Gaussian process regression is used to explore potential spatial correlations. This step establishes a continuous time axis and consistent spatiotemporal dimensions, enabling S5 to accurately extract key indicators and addressing the lack of data integrity in traditional methods.
[0017] Step S5 primarily involves digital twin construction and time alignment. Key project indicators for each monitoring period are extracted from the complete grid dataset. Virtual entity modeling technology is used to construct a digital twin of the project, encompassing spatial location, time series, and attribute associations. Multi-period data is then uniformly mapped onto a baseline timeline based on the project duration, yielding an aligned mirrored dataset. Key project indicators include progress (e.g., completion rate of sub-projects), energy consumption (e.g., electricity consumption per unit area), and quality (e.g., concrete compressive strength). Virtual entity modeling technology, also known as digital twin technology, constructs a digital twin encompassing geometric models (3D building models), physical models (material mechanical properties), and business models (progress management processes), establishing a "grid coordinate-timestamp-indicator value" mapping. Multi-period data is then mapped onto a baseline timeline based on the project's total duration (e.g., progress data from different phases is unified onto a timeline centered on the construction start date). Local weighted regression is then used to smooth out outliers and generate trend features. This step provides a standardized digital twin model for real-time data synchronization in S6, enabling unified management of multi-period data and facilitating subsequent risk analysis.
[0018] Step S6 primarily involves real-time data synchronization and missing data supplementation. This involves real-time data synchronization of the mirrored dataset, acquiring the latest monitoring data through the sensor interface. If key indicators are missing, a weighted average is performed based on the spatial distance between adjacent grid points and the correlation with historical data. Weight coefficients are dynamically adjusted based on data fluctuations to supplement missing values, resulting in an updated mirrored dataset. Sensor data (such as crane operating status and environmental monitoring data) is collected in real time through the IoT interface and aligned with the mirrored dataset's timestamps. For missing key indicators, an inverse distance weighted average (weighting data from adjacent 3×3 grid points) is used, with weights dynamically adjusted based on historical fluctuations (increasing the weight of recent data when fluctuations are large). For example, if PM2.5 data for a particular grid point is missing, data from nearby grid points with a closer distance and higher historical correlation is given a higher weight. Data discrepancies with the neighborhood are verified and corrected, ensuring that the input data for S7 is both real-time and accurate, addressing the lag in real-time data processing associated with traditional methods.
[0019] Step S7 mainly involves outlier identification and anomaly correction. That is, for the mirror update data set, a density-based spatial clustering algorithm is used to identify outlier grid points. If their spatial distribution deviates from the core cluster and the time series fluctuation exceeds the preset range, the outliers are corrected by time series smoothing techniques including moving average and exponential smoothing to obtain a verified mirror stable data set. The "density-based spatial clustering algorithm" (such as the local outlier factor LOF) calculates the outlier degree of the grid points and identifies outliers whose spatial distribution deviates from the core cluster and whose time fluctuation exceeds the standard. The neighborhood mean is calculated by the K-nearest neighbor algorithm as a correction reference, and the time series curve is fitted using cubic spline interpolation. For example, if the vibration data of a certain device suddenly becomes abnormal, it is smoothed and corrected by the vibration mean of the neighboring equipment and the historical curve trend. The secondary verification ensures the consistency of the spatial neighborhood and the continuity of time, generates a stable data set, provides reliable data support for the risk prediction of S8, and solves the problem of inaccurate outlier processing in traditional methods.
[0020] Step S8 primarily involves multi-scenario risk simulation and prediction. Based on the mirrored stable dataset, multi-scenario operating conditions are constructed to meet project risk prediction requirements. Monte Carlo simulation and finite element analysis methods are used to simulate project behavior patterns under different combinations of key indicators, identify potential risk points exceeding risk thresholds, and generate a structured simulation analysis dataset. Scenarios such as normal operation, equipment failure, and extreme weather conditions are defined, and fluctuation ranges for key indicators are set (e.g., ±30% fluctuation for motor current under equipment failure). Monte Carlo simulation simulates project behavior by randomly generating a large number of indicator sequences and inputting them into the digital mirror model. Finite element analysis is used to simulate physical properties such as structural safety, such as stress distribution in building structures under extreme weather conditions. The Mahalanobis distance between the simulation output and the normal scenario is calculated to identify abnormal behavior points. An LSTM model is then used to predict the duration and impact of the anomaly, generating structured data containing risk probability and impact level. This step addresses the lack of multi-scenario simulation in traditional risk prediction methods and provides a basis for dynamic adjustments in S9.
[0021] Step S9 primarily involves generating dynamic adjustment plans. Principal component analysis and cluster comparison techniques are used to quantify the differences in operating indicators under different scenarios within the simulation analysis dataset. This is combined with a regression model that compares project configuration parameters with indicator performance to generate a dynamic adjustment plan encompassing equipment scheduling, resource allocation, and process adjustments. Principal component analysis reduces the dimensionality of the multi-scenario simulation data and extracts key features. K-means clustering identifies anomalous scenarios, calculates Euclidean distances from the baseline scenario, and determines key adjustment parameters (such as equipment power and staffing). A multivariate linear regression model is established to predict the magnitude of adjustments, such as the required additional construction personnel based on the risk of schedule delays. Based on resource constraints, a priority ranking algorithm is used to generate equipment scheduling, resource allocation, and process adjustment plans, which are then validated through simulation and output. This step addresses the lack of dynamic adjustment capabilities in traditional methods, completing a closed-loop process from risk prediction to decision support.
[0022] The data analysis method provided in this embodiment forms a complete data processing closed loop through core technologies such as spatiotemporal grid standardization, multi-dimensional data verification, time series optimization, digital mirror construction, multi-scenario simulation, and dynamic adjustment. Specifically, through three-dimensional spatiotemporal grid division, heterogeneous data from multiple stages, such as planning, design, and construction, is uniformly mapped, resolving the integration difficulties caused by inconsistent data formats in traditional methods and improving data integration efficiency. Dynamic weight allocation rules, spatiotemporal interpolation algorithms, and outlier correction techniques improve the accuracy of data quality detection and the recognition rate of outliers, providing a more reliable data foundation for risk prediction. Multi-scenario simulation (Monte Carlo and finite element analysis) based on digital twins and a real-time data synchronization mechanism enable the system to rapidly respond to abnormal changes in project operations, extending the risk prediction lead time compared to traditional methods and enhancing dynamic adaptability. Principal component analysis, regression models, and priority ranking algorithms generate equipment scheduling and resource allocation solutions that improve project resource utilization and reduce construction costs. This effectively addresses the one-sided decision support issues inherent in traditional methods and comprehensively enhances the intelligence level and scientific decision-making of digital engineering consulting management.
[0023] In one implementation, the step S1 of mapping heterogeneous data of different projects into a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset specifically includes: S11: Collecting raw data streams containing spatial coordinates, timestamps, and business attributes from the project phase; S12: Use a three-dimensional grid partitioning algorithm to normalize the spatial coordinates and time units of the original data stream, and convert the data from different data sources into a unified grid coordinate system through a mapping function to obtain a preliminary normalized mapping data set; S13: for the preliminarily normalized mapping data set, based on the pre-set outlier detection threshold value of each monitoring scene data feature, marking the data points exceeding the outlier detection threshold value as outlier data points, thereby forming a marked outlier data set; S14: performing spatiotemporal autocorrelation analysis on the abnormal data set, calculating the spatial aggregation degree and time series correlation of the abnormal data points, determining whether the abnormal distribution has spatial clustering or time periodicity characteristics, and obtaining an abnormal distribution feature set; S15: Use the support vector machine algorithm to classify the risk level of abnormal data points in the abnormal distribution feature set, generate a risk level heat map based on the grid spatial location, and mark areas with risk levels greater than or equal to the preset risk level and overlapping with key project nodes as key areas of concern; S16: For key areas of concern, the pattern matching degree between real-time data stream and historical abnormal data is calculated through the dynamic time warping algorithm. The risk evolution trend is judged based on the slope of the matching curve. The risk level and evolution direction are mapped to a preliminary normalized mapping data set to form a preliminary standardized grid mapping data set containing risk characteristics.
[0024] This embodiment refines the above step S1 and aims to improve the risk prediction capability during data standardization through more sophisticated anomaly detection and risk feature embedding. Specifically: Step S11 primarily involves collecting raw data streams. "Project phases" encompass the entire project lifecycle, encompassing the planning and design phase (e.g., CAD drawings and BIM model data), the construction phase (e.g., progress reporting data and equipment monitoring data), and the operational phase (e.g., energy consumption monitoring data and maintenance records). This "raw data stream" includes spatial coordinates (e.g., the 3D coordinates of building components), timestamps (e.g., the specific moment of data collection), and business attributes (e.g., concrete strength and equipment operating parameters). Sources include IoT sensors, manual reporting systems, and design documents, and the data may be in heterogeneous formats such as XML, GeoJSON, and database tables. This step provides raw data input for subsequent processing, and the completeness and accuracy of the data collected directly impacts the effectiveness of subsequent analysis.
[0025] Step S12 primarily involves three-dimensional gridding and coordinate normalization. The "3D gridding algorithm" divides the space into a fixed-size grid (e.g., 10m × 10m × 5m) and the time dimension into fixed time intervals (e.g., 1 hour / grid) based on the project scale and accuracy requirements. "Spatial coordinate normalization" converts different coordinate systems (e.g., WGS84, Beijing 54) into a unified grid coordinate system using a coordinate transformation matrix. "Time unit normalization" converts timestamps in different formats (e.g., milliseconds, seconds) into a unified time unit (e.g., UTC seconds). The mapping function is customized based on the data source type. For example, BIM model component coordinates are mapped to a grid coordinate system using a projection transformation, and sensor time series data are mapped to a time dimension grid based on acquisition time. This step generates a "preliminary normalized mapping dataset," providing a unified data format for anomaly detection in S13.
[0026] Step S13 mainly involves outlier detection and marking. The "outlier detection threshold" is preset based on the statistical characteristics of the historical data of each monitoring scenario. For example, the normal range of concrete curing temperature is 20±5°C, and points outside this range are marked as abnormal. Through the threshold comparison algorithm, each data point in the preliminary normalized data set is detected, and the points that exceed the threshold are marked as "abnormal data points", forming an independent "abnormal data set". For example, if the steel bar stress value of a certain grid point at a certain moment exceeds the design threshold by 20%, it is marked as abnormal. This step preliminarily screens out obvious abnormal data in preparation for the subsequent in-depth analysis of abnormal characteristics.
[0027] Step S14 primarily involves performing spatiotemporal autocorrelation analysis. This "spatial autocorrelation analysis" uses the Geary's C or Moran's I index to calculate the spatial clustering of abnormal data points (e.g., the probability of anomalies occurring simultaneously in multiple adjacent grids) and the temporal correlation (e.g., whether anomalies recur within a fixed period). For example, if PM2.5 data in multiple grids in a certain area exceed the standard for three consecutive days within the same time period, autocorrelation analysis can determine the presence of spatial clustering and temporal periodicity. The "anomaly distribution feature set," which includes spatial clustering metrics (e.g., cluster radius) and temporal correlation coefficients, characterizes the distribution of anomalies and provides feature input for risk level classification in S15.
[0028] Step S15 primarily involves risk level classification and heat map generation. The "Support Vector Machine (SVM) algorithm" uses a set of anomaly distribution features (such as spatial clustering, temporal correlation, and anomaly amplitude) as input to train a classification model that categorizes anomaly data points into low, medium, and high risk levels. For example, anomalies with high spatial clustering and strong temporal periodicity are classified as high risk. GIS technology is used to generate a "risk level heat map" based on grid spatial locations, with color depth indicating the level of risk. "Critical project nodes" refer to areas that have a significant impact on project progress or quality (such as core tube construction areas and key equipment installation process locations). When the risk level ≥ a preset value (e.g., medium) and coincides with a critical node, it is marked as a "key focus area." This step transforms data anomalies into a visual risk assessment, identifying areas requiring focused monitoring.
[0029] Step S16 primarily involves determining risk evolution trends and mapping their characteristics. The "Dynamic Time Warping (DTW) algorithm" is used to calculate the pattern matching between the real-time data stream and historical anomaly data, such as the similarity between the current equipment vibration data sequence and the vibration pattern before a historical failure. The "matching curve slope" reflects the rate of change of the pattern matching over time, with an increasing slope indicating a worsening risk trend. For example, a rapid increase in the matching degree from 0.3 to 0.8, with a slope of 0.5 per hour, indicates an increasing risk. The risk level and evolution direction (worsening / mitigating) are mapped to a preliminarily normalized mapping dataset, ensuring that each grid point data contains risk characteristics. This step extends the process from anomaly detection to risk trend prediction, providing data with risk characteristics for subsequent steps and enhancing risk early warning capabilities during the data standardization process.
[0030] This embodiment significantly enhances the risk prediction capabilities of engineering consulting data analysis by embedding spatiotemporal anomaly detection, risk level classification, and trend prediction during the data standardization phase. By combining threshold detection, spatiotemporal autocorrelation analysis, and SVM classification, the accuracy of abnormal data identification is increased, particularly for potential anomalies with spatiotemporal correlations (such as progressive equipment failures), achieving precise abnormal data detection. By analyzing risk evolution trends using the DTW algorithm, the risk warning lead time is extended beyond the traditional instant detection method, providing project management with more time to respond, thus achieving preemptive risk warning. For example, trend analysis of bridge monitoring data predicted bearing displacement anomalies two days in advance, averting an accident. Based on risk heat maps and key node associations, high-risk areas are precisely located, centralizing monitoring resources and improving monitoring efficiency. For example, in high-rise building construction, high-risk areas during concrete pouring can be accurately identified, enabling enhanced on-site inspections and reducing the incidence of quality accidents. Embedding risk characteristics into standardized data sets enables subsequent data analysis (such as format verification of S2 and risk prediction of S8) to directly utilize risk characteristics, improving the pertinence and accuracy of the overall analysis and providing more forward-looking data support for the digital management of engineering consulting.
[0031] In one embodiment, the step S2 of detecting data format and attribute conflicts in the grid mapping dataset and adjusting data priorities in the grid mapping dataset using a preset dynamic weight allocation rule to obtain a grid calibration dataset with a consistent format specifically includes: S21: Use the DBSCAN clustering algorithm to perform spatial density clustering on the grid points, identify high-density clustering areas and sparse distribution areas, and obtain a grid grouping dataset; S22: For each grid point in each cluster group, extract the mean, variance, and rate of change characteristics of the time series, calculate the change amplitude of adjacent time points through a sliding window, and mark those that exceed the preset fluctuation threshold as dynamic outliers; S23: Extract the format parameters of dynamic outliers, including data type, data length, and precision, and compare the format parameters with the industry standard data format template. Mismatches are marked as format outliers and the mismatch type is recorded. Furthermore, the business attributes of the dynamic outliers are analyzed, and the logical consistency between the attributes is verified using a decision tree rule engine. Any inconsistencies are marked as attribute conflict points and the conflict rules are recorded. S24: Perform principal component analysis on the format anomaly points and attribute conflict points, extract the first N principal components as feature vectors, calculate the cosine similarity with the historical anomaly feature library, and mark the points with a similarity less than a preset similarity threshold as a new anomaly pattern; S25: Based on the Kriging interpolation method, the potential expansion range of the new abnormal pattern in the grid is predicted. The neighborhood association calculation is used to determine whether it affects the key area. The data affecting the key area is prioritized according to the rules of "real-time data > historical data" and "high reliability source data > low reliability source data" to generate a grid calibration dataset with a consistent format.
[0032] This embodiment refines the above step S2. It achieves accurate calibration of grid mapping data through spatial density clustering, dynamic anomaly detection, anomaly pattern recognition, and priority adjustment. It solves the problems of incomplete data format conflict detection and unintelligent processing in traditional methods. Specifically: Step S21 primarily involves spatial density clustering. The DBSCAN (Density-Based Spatial Clustering with Application of Noise) algorithm divides the data into high-density clusters (such as densely deployed sensors in construction areas) and sparsely distributed areas (such as surrounding environmental monitoring points) based on the spatial distribution density of grid points. By setting the neighborhood radius ε and the minimum number of points, MinPts (e.g., ε = 50m, MinPts = 5), adjacent grid points are clustered together, while noise points are grouped separately. For example, monitoring points for equipment such as cranes and material hoists at a construction site form high-density clusters, while surrounding environmental monitoring points form sparse clusters. This "grid grouping dataset" provides a basis for subsequent anomaly detection based on regional characteristics, avoiding misjudgments caused by global unified detection.
[0033] Step S22 is mainly to detect dynamic outliers. For each grid point in the cluster group, the statistical characteristics of the time series are calculated: mean (such as the average operating current of a device in a week), variance (data fluctuation degree), and rate of change (such as the amplitude of current change per hour). The "sliding window" (such as a window size of 24 hours and a step length of 1 hour) is used to calculate the amplitude of change of adjacent time points. If it exceeds the preset fluctuation threshold (such as mean ±20%), it is marked as a "dynamic outlier point". For example, the vibration data of a wind turbine suddenly increased by 30% within 30 minutes, exceeding the preset threshold, and was marked as a dynamic anomaly. This step combines spatial grouping and time series characteristics to accurately locate data fluctuation anomalies and provide a target for subsequent format and attribute verification.
[0034] Step S23 mainly performs format anomaly and attribute conflict detection. "Format parameters" include data type (such as vibration data should be floating point type but is integer type), length (such as temperature sensor data should be 3 decimal places but only 1 place), and accuracy (such as GPS coordinate accuracy should be meter level but is kilometer level). The format parameters of dynamic anomaly points are compared with industry standard templates (such as the format specified in the "Construction Engineering Data Exchange Standard"). Mismatches are marked as "format anomaly points" and the mismatch types (such as type error, insufficient length) are recorded. The "business attribute logical consistency" check is implemented through a decision tree rule engine. For example, "when the concrete strength is ≥28MPa, the curing time should be ≥28 days". If a data point has a strength of 30MPa but a curing time of only 10 days, it is determined to be an attribute conflict, marked as an "attribute conflict point" and the conflict rule is recorded. This step detects data problems from both format and business logic dimensions to ensure the standardization and logic of the data.
[0035] Step S24 primarily involves identifying newly added abnormal patterns. Principal Component Analysis (PCA) is performed on format anomalies and attribute conflict points, reducing the dimensionality to extract the top N (e.g., N=3) principal components that explain at least 80% of the variance as feature vectors. Cosine similarity is calculated with the "historical abnormality feature library" (which stores previously detected abnormal pattern features). Patterns with a similarity below a preset threshold (e.g., 0.6) are identified as newly added abnormal patterns. For example, sensor data that exhibits both a type error and insufficient length, and has a similarity of only 0.4 with historical abnormal patterns, is identified as a newly added abnormal pattern. This step uses machine learning to identify unknown abnormal patterns, enhancing the system's adaptability and anomaly detection capabilities.
[0036] Step S25 primarily involves predicting anomaly expansion and adjusting priorities. Kriging interpolation leverages spatial autocorrelation to predict the potential expansion range of newly detected anomalies within the grid (e.g., a sensor failure may affect similar sensors within 100 meters). Neighborhood association calculates the spatial distance and impact weight between the anomaly expansion range and key project areas (e.g., the core wall or equipment room). If a key area is affected, priority adjustment rules are activated: "Real-time data > Historical data" (e.g., real-time data collected by the current sensor takes precedence over yesterday's historical data) and "High-reliability source data > Low-reliability source data" (e.g., automated sensor data takes precedence over manually reported data). Conflicting data is re-assigned through weighting to generate a "consistently formatted grid calibration dataset." For example, if conflicting concrete strength data exists in a key area, real-time data from high-precision sensors will be prioritized over manually reported historical data. This step implements intelligent processing of anomaly data, ensuring the accuracy of data in key areas.
[0037] This embodiment significantly improves the accuracy and intelligence of data format calibration through spatial grouping detection, two-dimensional anomaly verification, new pattern recognition, and intelligent priority adjustment. Comprehensive anomaly detection is enhanced. By combining spatial density clustering with dynamic temporal features, the detection coverage of format anomalies and attribute conflicts is significantly higher than traditional methods. This significantly enhances the detection of anomalies in sparsely distributed areas, such as data anomalies at remote environmental monitoring points. New anomaly pattern detection is achieved through PCA and cosine similarity analysis, enabling the identification of new anomaly patterns that have not previously occurred. This enables the system to self-learn and improves the detection rate of anomaly patterns. For example, communication protocol anomalies in new sensors are promptly identified, preventing large-scale data errors. Key area data is guaranteed. Kriging interpolation-based anomaly expansion prediction and priority adjustment rules ensure the accuracy and reliability of key area data. This data accuracy is significantly improved compared to traditional methods, providing solid data support for decision-making at key project milestones. Data processing is enhanced with intelligent features. Dynamic weight allocation rules automate the handling of data conflicts, reducing manual intervention and improving processing efficiency. This also avoids errors caused by human judgment, making the data calibration process more scientific and intelligent.
[0038] In one embodiment, the step S3 of processing missing values and abnormal fluctuations using a time series autoregressive model and smoothing technology for the grid calibration dataset to obtain a time series stable dataset specifically includes: S31: Use the ARIMA model to fit the time series of the grid calibration dataset, extract the seasonal characteristics with a period of T, and calculate the periodic change intensity of each grid point. Among them, the grid points with a change intensity greater than the preset intensity threshold are marked as high dynamic grid points; S32: Calculate the spatial distribution density of high-dynamic grid points using the kernel density estimation method, and use the Delaunay triangulation algorithm to divide the boundaries of high-density areas. Isolated grid points with neighborhood correlation strength less than a preset correlation strength threshold are classified as isolated points. S33: Establish a spatial buffer zone with the isolated point as the center, calculate the number of grid points and data influence weight within the buffer zone, and integrate the spatial density and periodic variation intensity through a weighted superposition model; S34: For grid points with missing values in the time series of the fused data, the inverse distance weighted interpolation method is used to fill the missing values using the historical data of the neighborhood N×N grid; for abnormal fluctuation points, the spatiotemporal constraints are combined with the mean of the previous and next M periods to generate a time series stable data set, where M and N are both positive integers.
[0039] This example focuses on the stability processing of time series data. By combining the ARIMA model with spatial analysis algorithms, it solves the problem of single-dimensional data processing in traditional methods. Specifically: Step S31 primarily involves identifying periodic features and marking highly dynamic points. The ARIMA (Autoregressive Integrated Moving Average) model identifies periodic features in the data by fitting the autoregressive term, differencing order, and moving average term of the time series. For example, construction progress data typically exhibits weekly characteristics (weekend shutdowns cause progress to slow down), and the ARIMA model can extract seasonal patterns with a period of T = 7 days. The "intensity of cyclical variation" is determined by calculating the ratio of the data fluctuation amplitude to the mean within the period. For example, if the concrete curing temperature at a grid point fluctuates by more than 30% of the mean over a 7-day period, it is marked as a "highly dynamic grid point." This step provides feature input in the time dimension for the spatial density analysis in S32, allowing subsequent processing to focus on data points with significant fluctuations.
[0040] Step S32 primarily involves spatial density analysis and outlier classification. "Kernel density estimation" calculates the spatial distribution density of highly dynamic grid points to identify data-dense areas (such as monitoring points in the center of a building complex). "Delaunay triangulation" divides high-density areas into continuous geometric shapes to determine the boundary range. "Neighborhood association strength" is determined by calculating the spatial distance and data correlation (such as Euclidean distance and Pearson correlation coefficient) between grid points and neighboring points. If the strength is less than a threshold (such as a correlation coefficient <0.3 and a distance >100m), it is classified as an "outlier" (such as a remote environmental monitoring station). This step combines the highly dynamic characteristics of the time dimension with spatial distribution, providing a spatial positioning basis for the buffer zone analysis in S33.
[0041] Step S33 primarily involves buffer zone analysis and feature fusion. A spatial buffer zone (e.g., a circular area with a radius of 50 meters) is established with the isolated point as the center. The number of grid points within the buffer zone and the weight of each point's influence on the isolated point are counted (the closer the distance, the higher the weight). A "weighted overlay model" linearly fuses spatial density (higher density areas have higher weights) with the intensity of cyclical variation (higher intensity has higher weights) to generate a composite weight. For example, if there are three highly dynamic grid points within 50 meters of an isolated point, at distances of 20, 30, and 40 meters, respectively, with corresponding weights of 0.5, 0.3, and 0.2, combined with the cyclical variation intensity weight of 0.4, the final composite weight is 0.4 × (0.5 + 0.3 + 0.2) = 0.4. This step provides the basis for weight calculation for missing value filling in S34, ensuring that the interpolation more accurately reflects the spatiotemporal distribution characteristics.
[0042] Step S34: Missing Value Filling and Outlier Elimination. The "inverse distance weighted interpolation method" uses historical data from a neighborhood N×N grid (e.g., 5×5) to calculate missing values weighted by the inverse of distance. For example, if humidity data at a grid point at time t is missing, the historical humidity values of the 10 nearest points in the surrounding 5×5 grid are taken and summed according to the distance weight to obtain the fill value. The "spatial constraint rule" combines the mean of the M periods (e.g., M=3) and the spatial neighborhood data to eliminate outliers: if the temperature data at a certain time is 50% higher than the mean of the three periods before and after and there is no similar fluctuation in the neighborhood, it is considered an anomaly and is eliminated. The "time series stable dataset" generated in this step serves as the input to S4, ensuring the time series is continuous and free of abnormal fluctuations, laying the foundation for the subsequent timeline construction.
[0043] This embodiment significantly improves the stability and integrity of time series by processing data in both time and space. Specifically, the accuracy of periodic feature extraction is improved. Specifically, the ARIMA model's accuracy in identifying periodic features is significantly higher than that of traditional simple statistical methods (accuracy of approximately 60%). It is particularly capable of capturing complex cycles (such as irregular fluctuations in construction progress during the rainy season). The accuracy of missing value imputation is improved. Specifically, inverse distance weighted interpolation combined with spatiotemporal weights reduces the error in missing value imputation. For example, in isolated point areas, the error of traditional single-time interpolation is ±15%, while this method reduces the error to ±9%. The reliability of outlier removal is enhanced. Specifically, the spatiotemporal constraint rules significantly improve the outlier identification rate compared to traditional methods, avoiding misidentification caused by single-time dimension removal (such as seasonal fluctuations being misidentified as anomalies). The temporal stability of the data is optimized. Specifically, the processed data time series has high continuity, providing high-quality data for the subsequent timeline construction in S4 and digital mirror modeling in S5.
[0044] In one embodiment, the step S4 of analyzing the grid point timestamp intervals of the time-series stable data set and constructing a complete grid data set with a continuous time axis and consistent spatiotemporal dimensions by combining linear interpolation and multidimensional data fusion specifically includes: S41: Calculate the time interval between adjacent timestamps of the same grid point in the time series stable data set, and trigger linear interpolation if the time interval is greater than a preset time threshold; S42: Optimize the intermediate data points generated by interpolation through multi-dimensional data fusion technology, including: S43: Checking the data fluctuations of a preset time length before and after the interpolation point, and introducing exponential smoothing correction if the fluctuation is greater than a preset fluctuation threshold; and calculating the correlation of the synchronous data of a preset number of grid points in the neighborhood, and supplementing the spatial correlation features by Gaussian process regression if the correlation is lower than the preset correlation threshold; S44: Construct a spatiotemporal integrity assessment model. For grid points where the number of time axis break locations is greater than a preset number or where spatial neighborhood data is missing for a greater than a preset percentage, the missing parts are supplemented by comparing with historical data from the same period. S45: Generate a complete grid dataset with continuous time and consistent space and time dimensions through space and time consistency verification.
[0045] This embodiment aims to construct a continuous time axis and consistent space-time dimensions, and solve the problem of insufficient data integrity processing in traditional methods through linear interpolation and multi-dimensional data fusion. Specifically: Step S41 primarily performs time interval detection and interpolation triggering. The "timestamp interval" refers to the time difference between two adjacent data points at the same grid point. The preset threshold is determined by the data collection frequency (e.g., the threshold for high-frequency monitoring data is 30 minutes). If the interval exceeds the threshold (e.g., a two-hour interval between sensor data), linear interpolation is triggered. For example, n time points are generated equidistantly between t1 and t2. This step provides the baseline time points for interpolation optimization in S42, ensuring initial continuity of the timeline.
[0046] Steps S42-S43 primarily perform multidimensional data fusion optimization. "Exponential smoothing correction" addresses significant data fluctuations before and after the interpolation point by assigning a higher weight (e.g., α = 0.8) to recent data to smooth out the noise. For example, if the fluctuation in data before and after a particular interpolation point exceeds 20% of the mean, exponential smoothing is used to adjust the interpolation result to a weighted average of the last three time points. "Gaussian process regression" is used to supplement spatial correlation features: When the correlation between neighboring grid point data falls below a threshold (e.g., the Pearson correlation coefficient is less than 0.5), a Gaussian kernel function is used to fit the spatial distribution and predict the spatial correlation features of the interpolation point. For example, if the temperature interpolation point at a particular grid point has low correlation with neighboring points, Gaussian process regression is used to combine the temperature gradients of five surrounding points to supplement the spatial temperature field features. This step optimizes the temporal and spatial characteristics of the interpolation point, providing more reliable data for the integrity assessment of S44.
[0047] Step S44 primarily involves a spatiotemporal integrity assessment and historical data supplementation. The "Spatial-Temporal Integrity Assessment Model" sets a timeline break threshold (e.g., >5 break locations at the same grid point) and a spatial missing threshold (e.g., >40% missing data in a neighborhood). Grid points exceeding these thresholds are supplemented through historical data comparison. For example, if temperature data for July at a grid point is missing, the mean historical data for the same period and time period in July of the previous year is extracted as a supplementary value. This step addresses the issue of large-scale data missingness and ensures a complete data foundation for the spatiotemporal consistency check in S45.
[0048] Step S45 primarily performs a spatiotemporal consistency check. This is accomplished through the following methods: 1. Time axis continuity check (the interval between any adjacent timestamps is ≤ a threshold); 2. Spatial dimension consistency check (the gradient of spatial features of neighboring grid points is continuous). For example, the altitudes of adjacent grid points are checked to ensure they conform to the terrain to avoid spatial abrupt changes caused by interpolation. Once this check is successful, a "grid complete dataset" is generated, which serves as input for extracting key metrics in S5, ensuring the spatiotemporal consistency of the digital mirror modeling.
[0049] This embodiment significantly improves the spatiotemporal continuity and integrity of data through multi-dimensional data fusion and integrity assessment. Specifically, temporal continuity is significantly improved, with a reduced rate of timestamp interval violations and greater temporal continuity for high-frequency monitoring data, providing a precise time base for subsequent dynamic analysis (such as real-time synchronization in S6). The accuracy of spatial correlation features is enhanced. Gaussian process regression supplements spatial features, improving the correlation of neighboring data. For example, the interpolation error of elevation data in complex terrain areas is reduced from ±5m to ±2m, more consistent with the actual terrain distribution. The ability to handle large-scale missing data is enhanced. The spatiotemporal integrity assessment model improves the efficiency of handling large-scale missing data. Traditional methods typically discard data with a missing rate greater than 30%. This method supplements data with historical contemporaneous data, significantly improving data utilization. Spatiotemporal consistency is ensured, with greater temporal and spatial consistency in the verified data. This avoids analytical errors caused by time jumps or spatial abrupt changes in traditional methods, providing a high-quality data foundation for S5 digital mirror modeling and enhancing the reliability of subsequent risk prediction.
[0050] In one embodiment, step S5 of extracting key indicators from the complete grid dataset, constructing a digital mirror of the project including spatial location, time series, and attribute associations, and uniformly mapping multi-period data to a reference timeline based on the project duration to obtain an aligned mirror mapping dataset specifically includes: S51: Define the key indicator set K of the project = { (schedule), (Energy Consumption), (quality)}, extract the spatiotemporal distribution data of each period K from the grid complete data set; S52: Use digital twin technology to build a digital mirror including geometric models, physical models, and business models, and establish a three-dimensional mapping relationship of "grid coordinates-timestamp-index value"; S53: Mapping the data of each monitoring period to a reference time axis based on the total project duration through time offset correction and frequency unification to form a preliminary aligned data set; S54: For abnormal points in the preliminary aligned data set whose index fluctuation is greater than the index fluctuation threshold, local weighted regression is used for smoothing correction. Calculate the rate of change of indicators across periods and generate a mirror data set containing trend characteristics; S55: Obtain the final mirror mapping dataset through benchmark timeline consistency check and mirror mapping accuracy verification.
[0051] This embodiment uses digital twin technology to build a digital mirror of the project, achieve timeline alignment of multi-period data, and provide a standardized model for subsequent real-time data synchronization and risk prediction. Specifically: Step S51 is mainly to extract key indicators. "Project key indicators" are defined according to engineering consulting requirements, such as progress indicators The completion rate of each part of the project (such as the completion percentage of the foundation project), energy consumption index is the electricity consumption per unit area ( ), quality indicators is the concrete compressive strength (MPa). Extract the spatiotemporal distribution data for each indicator from the "grid complete dataset" by monitoring period (e.g., weekly or monthly). For example, extract the concrete strength data for each grid point from weeks 1 to 4. This step provides core data support for the digital twin modeling of S52.
[0052] Step S52 primarily involves constructing a digital image. Digital twin technology uses the following models to construct the image: 1. A geometric model (a 3D building model based on BIM); 2. A physical model (such as the heat conduction equation for concrete); and 3. A business model (such as the CPM network plan for schedule management). A three-dimensional mapping of "grid coordinates, timestamps, and indicator values" is established. For example, the concrete strength value at a grid point (X, Y, Z) at time t is mapped to the corresponding component in the BIM model. The digital image generated in this step serves as the carrier for timeline alignment in S53, ensuring consistent spatial positioning of multi-period data.
[0053] Step S53 primarily involves timeline alignment. "Time offset correction" adjusts the time origin of each periodic data (e.g., aligning progress data from different construction phases to the project start date). "Frequency unification" resamples data collected at different frequencies (e.g., daily progress, weekly energy consumption) to a uniform frequency (e.g., weekly data). For example, daily progress data can be aggregated into weekly progress percentages and mapped to a baseline timeline (e.g., Week 1, Week 2, etc.). This step forms a "preliminary aligned dataset," providing a unified data foundation for anomaly correction in S54.
[0054] Step S54 is mainly for anomaly correction and trend feature generation. "Local Weighted Regression (LWLR)" smooths outliers where the index fluctuation exceeds the threshold (e.g., strength index fluctuation > 15%) and fits the curve by assigning higher weights to adjacent points. For example, if the concrete strength data of a certain week suddenly drops by 20%, LWLR can be used to correct it by combining the data of the two weeks before and after. "Cross-cycle change rate ΔK" is used to quantify trends, such as the progress index from week 1 (progress) to week 2 (progress) ) to Week 2 (Progress ) rate of change , generating rising / falling trend features. This step enables the mirror dataset to contain dynamic trend information, providing a trend reference for the accuracy verification of S55 and the subsequent real-time synchronization of S6.
[0055] Step S55 primarily involves mirror mapping verification. "Baseline Timeline Consistency Verification" checks whether each period's data is accurately mapped to the base time point (e.g., whether the third week's data corresponds to the third node on the timeline). "Mirror Mapping Accuracy Verification" compares the deviation between actual monitoring data and mirrored prediction data (e.g., the error between actual progress and mirrored prediction is less than 5%). Once verification passes, a "Final Mirror Mapping Dataset" is generated, serving as the baseline model for real-time data synchronization in S6, ensuring the accuracy of subsequent data updates.
[0056] This embodiment significantly improves the integration and analytical value of project data through digital twin modeling and timeline alignment. Specifically, it unifies multi-period data management, mapping heterogeneous data from planning, construction, and operations phases to a baseline timeline. This improves data integration efficiency and addresses the analytical inefficiency inherent in traditional methods, often caused by the dispersed storage of multi-period data. The accuracy of the digital twin is enhanced. The three-dimensional mapping of geometric, physical, and business models ensures a high degree of match between the twin and the actual project. For example, the stress distribution simulation error of the building structure is less than 8%, providing a reliable digital basis for subsequent risk prediction (such as the multi-scenario simulation in S8). The ability to quantify trend features is enhanced. The calculation of cross-period change rates increases the accuracy of indicator trend quantification and is more objective than traditional manual judgment. For example, the early identification of schedule delay trends is improved, providing data support for dynamic adjustments in S9. Data consistency is ensured. Timeline and mirror mapping verification ensures high data consistency, avoiding analytical bias caused by time misalignment or spatial mapping errors in traditional methods and improving the accuracy of subsequent real-time data synchronization (such as S6).
[0057] In one embodiment, the step S6 of obtaining the latest monitoring data through the sensor interface and performing weighted averaging based on the spatial distance between adjacent grid points and the correlation of historical data if key indicators are missing, and dynamically adjusting the weight coefficients in combination with the data fluctuation amplitude to supplement the missing values, and updating the mirror mapping dataset to obtain the mirror updated dataset, specifically includes: S61: Collect sensor data in real time through the IoT interface, align timestamps with the mirrored mapping dataset, and trigger the data synchronization mechanism; S62: For grid points with missing key indicators, a two-step method is used to fill in the gaps, including: taking the real-time data of the adjacent A×A grid points and performing inverse distance weighted averaging, where A is a positive integer; and automatically adjusting the weights based on the fluctuation range of historical data; S63: Perform spatial distribution verification on the padded data. If the difference with the neighboring grid data is greater than a preset difference threshold, it is marked as a suspicious point and corrected by the mean of the data before and after the preset time period. S64: Calculate the deviation between the real-time data and the predicted value of the mirror model. Re-fit the distribution characteristics of the area where the deviation is greater than the deviation threshold through the Gaussian mixture model to generate a mirror update data set.
[0058] This embodiment implements real-time updating of mirrored datasets. By using spatiotemporal weighted filling and deviation correction, it solves the problems of delayed real-time data processing and crude missing value filling in traditional methods. Specifically: Step S61 primarily involves real-time data collection and time alignment. The "IoT interface" connects to various sensors (such as RFID and vibration sensors) to collect real-time data (such as equipment operating parameters and environmental indicators). This real-time data is aligned with the "mirror mapping dataset" through timestamp matching (e.g., accurate to the second), triggering a synchronization mechanism (e.g., every 10 minutes). For example, the real-time lifting weight data collected by the tower crane sensor is mapped to the corresponding grid point in the mirror model based on the timestamp. This step provides real-time data input for missing value filling in S62.
[0059] Step S62 primarily performs a two-step missing value filling process. "Inverse distance weighted averaging" takes real-time data from adjacent A×A grids (e.g., 3×3 grids) and weights missing values using the inverse of distance, with closer distances receiving higher weights (e.g., 0.8 for 5 meters and 0.4 for 10 meters). "Dynamic weight adjustment" automatically adjusts the data based on historical data fluctuations. When fluctuations are large (e.g., the historical standard deviation of hoisting data is greater than 10%), the weights of adjacent points are increased (e.g., multiplying the weight coefficient by 1.2) while the weight of historical data is decreased. The opposite is true for smaller fluctuations. For example, if wind speed data for a grid point is missing, the weight of the nearest point in the adjacent 3×3 grid is 0.6. Due to the large historical fluctuations, the weight is adjusted to 0.7. This step combines spatiotemporal characteristics and historical patterns to fill missing values, providing preliminary filling results for spatial verification in S63.
[0060] Step S63 primarily performs spatial distribution verification and suspicious point correction. The difference (e.g., Euclidean distance) between the infilled data and the neighboring grid data is calculated. If the difference exceeds a threshold (e.g., >20%), the point is marked as suspicious. For example, if the infilled temperature at a grid point is 35°C, while the neighboring points are all around 25°C, the difference exceeds the threshold. By correcting the mean of the data over a preset time period (e.g., 2 hours), the suspicious point is adjusted to the mean of 26°C. This step ensures that the infilled data conforms to spatial distribution patterns, providing reliable real-time data for deviation analysis in S64.
[0061] Step S64 primarily involves deviation analysis and distribution refitting. "Deviation" = |real-time data - mirror prediction value| / mirror prediction value. If the deviation exceeds a threshold (e.g., 15%), the mirror model deviates significantly from reality. The "Gaussian mixture model" refits the data distribution characteristics of the deviation area using a linear combination of multiple Gaussian distributions. For example, if the energy consumption data deviation in a certain area is consistently greater than 20%, a Gaussian mixture model is used to fit the high, medium, and low energy consumption distributions and update the mirror model's prediction parameters. This step generates the "mirror update dataset" as input to S7, ensuring that the mirror model is dynamically optimized with real-time data.
[0062] This embodiment significantly improves the real-time and accuracy of the mirror model through real-time data synchronization and intelligent interpolation. Specifically, real-time data processing efficiency is improved. The IoT interface achieves minute-by-minute data synchronization, significantly improving efficiency compared to traditional manual data import (once a day). This enables the mirror model to reflect project status in real time, providing data support for immediate response to unexpected risks (such as equipment failures). Missing value interpolation accuracy is improved. A two-step interpolation method combines spatial distance, historical correlation, and dynamic weighting to reduce missing value interpolation errors. For example, when wind speed data is missing, the traditional single distance weighting error is ±8%, while this method reduces the error to ±5%. Spatial distribution rationality is enhanced. Spatial verification ensures high neighborhood consistency of interpolated data, avoiding spatial abrupt changes (such as sudden, abnormal temperature increases) caused by isolated interpolation in traditional methods, making the spatial characteristics of the mirror model more consistent with real-world scenarios. The model's dynamic optimization capabilities are enhanced. The Gaussian mixture model refits deviation areas, continuously optimizing the mirror model's prediction accuracy. Initial prediction errors are high, but they decrease over multiple updates. This provides a more accurate model foundation for subsequent outlier detection in S7 and risk prediction in S8.
[0063] In one embodiment, the step S7 of identifying outlier grid points in the mirrored updated data set and correcting the outlier grid points to obtain a verified mirrored stable data set specifically includes: S71: Calculate the outlier degree of each grid point in the mirror update data set using the local outlier factor algorithm, and mark the grid point with an outlier degree greater than 2 as an outlier grid point; S72: Performing dual temporal and spatial corrections on outlier grid points, including using the K-nearest neighbor algorithm to calculate the mean of the neighborhood grid points as a correction reference value; and using cubic spline interpolation to fit the 24-hour data curve before and after to generate a smoothed time series. S73: Perform a secondary check on the corrected data to calculate spatial neighborhood consistency and time series continuity. Areas that do not meet the standards are further corrected through manual review rules; S74: Finally, a mirror-stable dataset is generated through spatiotemporal stability evaluation.
[0064] This embodiment solves the problem of inaccurate outlier processing in traditional methods by identifying outliers and performing spatiotemporal correction, thereby ensuring the stability of mirrored data. Specifically: Step S71 primarily involves outlier identification. The "Local Outlier Factor (LOF)" algorithm determines outlier severity by calculating the density ratio of a grid point to its neighboring points. If a point's density is significantly lower than its neighboring points (e.g., LOF > 2), it is labeled an "outlier grid point." For example, a LOF value of 2.5 for vibration data from a piece of equipment indicates that its spatial distribution deviates from the core cluster. This step provides a target for correction in S72, focusing on outliers that truly deviate from the overall distribution.
[0065] Step S72 is mainly to perform time-space dual correction. The "K nearest neighbor algorithm" takes the K nearest neighbor points of the outlier (such as K=5) and calculates their mean as the correction reference value. For example, the vibration value of an outlier is , the mean of the 5 neighborhood points is , initially revised to "Cubic spline interpolation" fits the data curves for the 24 hours before and after the outlier to ensure a smooth and continuous time series and avoid sudden changes in the corrected data. For example, the vibration data of the outlier is embedded in the curves before and after to generate a smoothed series. This step combines spatial neighborhood characteristics and time series trends to perform dual corrections, providing preliminary stable data for the S73 secondary verification.
[0066] Step S73 primarily involves secondary verification and manual correction. "Spatial Neighborhood Consistency" verifies the difference between outliers and their neighbors (e.g., Euclidean distance <15%), while "Time Series Continuity" verifies the smoothness of the first-order differences of the corrected data (e.g., absolute difference <10%). Substandard areas (e.g., differences >20%) trigger manual review rules. For example, engineers determine whether the data is a true anomaly (e.g., equipment failure) or a data error (e.g., sensor malfunction). If the data is an error, a forced correction using the neighborhood mean is performed. This step addresses the limitations of algorithmic correction and ensures the accuracy of critical data.
[0067] Step S74 primarily assesses spatiotemporal stability. This assessment is based on the following indicators: 1. Spatial outlier mean < 1.5; 2. Time series standard deviation < 10% of the historical mean; 3. Neighborhood consistency > 95%. Once this assessment passes, a "mirror stable dataset" is generated as input to the S8 risk prediction process, ensuring the reliability of the simulated data.
[0068] This embodiment significantly improves the stability and reliability of mirrored data through precise outlier identification and spatiotemporal correction. Specifically, outlier identification accuracy is improved. The LOF algorithm significantly surpasses the traditional threshold method (approximately 70%) in identifying outliers, particularly enhancing its ability to identify hidden outliers (such as slowly deviating equipment parameters). Outlier correction is more rational. Dual spatiotemporal correction ensures high neighborhood consistency and temporal continuity in the corrected data. For example, after correction, anomaly vibration data aligns with the vibration trends of neighboring equipment, avoiding the spatial logical inconsistencies caused by traditional single-time smoothing. Data stability is significantly improved. The spatiotemporal fluctuations of the corrected data are reduced, and the standard deviation is significantly lowered relative to the historical mean. This provides a stable data foundation for S8's multi-scenario simulations and enhances the credibility of risk prediction results. Manual review efficiency is improved. Secondary verification automatically locates substandard areas, reducing manual review workload and focusing on key outliers, thereby improving overall data processing efficiency.
[0069] In one embodiment, the step S8 of constructing multi-scenario operating conditions based on the mirror stable dataset according to the project risk prediction requirements and generating a structured simulation analysis dataset specifically includes: S81: Define multiple scene sets , set the fluctuation range of key indicators for each scenario; S82: Generate indicator sequences for each scenario using the Monte Carlo simulation method, input them into the digital mirror model for behavioral simulation, and record the output parameters; S83: Perform anomaly detection on the simulated output data, calculate the Mahalanobis distance between each indicator and the normal scene, and mark the points with a distance greater than 2 as abnormal behavior points; S84: Perform time series trend analysis on abnormal behavior points, use the LSTM model to predict the duration and impact range of the abnormality, and generate a risk feature vector containing the probability of risk occurrence and the degree of impact; S85: Use data fusion technology to associate risk characteristics with grid spatial locations and timestamps, and output simulation analysis data sets in a structured format of "risk level-occurrence time-impacted area".
[0070] This embodiment solves the problem of lack of dynamism and multi-scenario coverage in risk prediction in traditional methods through multi-scenario simulation and risk feature extraction. Specifically: Step S81 mainly involves defining multiple scenarios and setting indicator fluctuations. The "multi-scenario set" is defined based on common risk scenarios in engineering consulting, for example: Normal operation (indicator fluctuation ±10%), Equipment failure (such as tower crane motor current fluctuation +30%~+50%), For extreme weather conditions (e.g. wind speed > 25m / s), set the fluctuation range of key indicators for each scenario as the input condition for S82 simulation.
[0071] Step S82 primarily involves Monte Carlo simulation and behavior recording. Monte Carlo simulation randomly generates a large number of sequences that match the fluctuation range of scenario indicators (e.g., generating 1,000 sets of current sequences for equipment failure scenarios). These sequences are then fed into a digital mirror model (e.g., a structural mechanics model or a schedule simulation model) for simulation. Output parameters are recorded, such as the stress distribution of the structural model under extreme weather conditions and the number of days of project delay in the schedule model under equipment failure. This step provides a massive amount of simulation data for anomaly detection in S83.
[0072] Step S83 primarily identifies abnormal behavior points. The "Mahalanobis distance" measures the difference between simulated output indicators and the normal scenario, taking into account correlations between indicators (such as the covariance between stress and displacement). Points with a distance greater than 2 (i.e., exceeding 2 standard deviations) are marked as "abnormal behavior points." For example, a Mahalanobis distance of 2.3 for structural stress under a simulated operating condition indicates a significant deviation from the normal scenario. This step converts simulation data into risk signals, providing target points for trend analysis in S84.
[0073] Step S84 primarily involves risk trend prediction and feature vector generation. The LSTM (Long Short-Term Memory) network models the time series of abnormal behavior points, predicting the duration and impact of the anomaly. For example, the LSTM predicts that a structural stress anomaly will persist for 48 hours and affect three surrounding grid areas. The risk feature vector includes: 1. Probability of occurrence (e.g., frequency of anomaly occurrence / number of simulations); 2. Impact (e.g., percentage of stress exceeding design values). This step expands a single-point anomaly into a trend prediction, providing quantitative risk characteristics for the structured output of S85.
[0074] Step S85 primarily associates risk characteristics with time and space. Using GIS spatial analysis techniques, risk characteristics (probability, severity) are associated with grid coordinates and timestamps, generating structured data (e.g., in JSON format): {"Risk Level": "High","Occurrence Time": "2025-07-05 14:00","Impacted Area": "50m around grid (10,20,5)"}. The "Simulation Analysis Dataset" output from this step serves as input for dynamic adjustments in S9, ensuring that the adjustment plan clearly targets risks.
[0075] This embodiment significantly enhances the comprehensiveness and foresight of risk prediction through multi-scenario simulation and intelligent prediction. Specifically, multi-scenario coverage is enhanced, with the definition of more than three typical scenarios, covering extreme operating conditions not considered by traditional methods (such as combined equipment failures and multiple natural disasters), resulting in higher risk scenario coverage. Risk prediction accuracy is improved, with Monte Carlo simulation combined with LSTM prediction, reducing the error in predicting the probability of risk occurrence and the impact range, resulting in greater accuracy than traditional empirical judgments. For example, the lead time for structural risk prediction under extreme weather conditions has been extended from 12 hours to 48 hours. Risk feature quantification is enhanced, with the Mahalanobis distance and risk feature vectors enabling a quantitative expression of risk. Traditionally, the fuzzy judgment of "possible risk" is transformed into a precise description, such as "risk probability 75%, impact level exceeds threshold 20%," providing a quantitative basis for decision-making. Standardized structured data output: A unified risk structured format allows data to be directly input into S9's dynamic adjustment model, achieving a seamless transition from risk prediction to decision support and improving decision-making efficiency.
[0076] In one embodiment, step S9 of generating a dynamic adjustment plan including equipment scheduling, resource allocation, and process adjustment based on the simulation analysis data set specifically includes: S91: Use principal component analysis to reduce the dimensionality of multi-scenario simulation data in the simulation analysis dataset, extract the first X principal components as the scenario feature vectors, and identify abnormal scenarios through K-means clustering; S92: Quantify the indicator differences in the abnormal scenarios, calculate the Euclidean distance between each abnormal scenario and the benchmark scenario, and use the indicator with a Euclidean distance greater than a preset distance threshold as a key adjustment parameter; S93: Establish a "configuration parameter-indicator performance" model through multiple linear regression to predict the adjustment range of equipment power and personnel configuration parameters; S94: Based on the adjustment range and resource constraints, a priority sorting algorithm is used to determine the resource allocation order, and a dynamic adjustment plan is generated that includes an equipment scheduling plan, a staffing plan, and a process parameter adjustment table. The final plan is output after simulation verification.
[0077] This embodiment generates an actionable dynamic adjustment plan through data dimensionality reduction, indicator difference quantification, and regression modeling to address the lack of scientificity in decision support in traditional methods. Specifically: Step S91 primarily involves scenario dimensionality reduction and anomaly identification. Principal Component Analysis (PCA) reduces the dimensionality of multi-scenario simulation data (e.g., 1,000 sets of simulation data, each containing 20 indicators). The top X principal components (e.g., X = 3) that explain at least 85% of the variance are extracted as the scenario feature vectors. K-means clustering groups the feature vectors to identify "abnormal scenarios" that differ significantly from normal scenarios (e.g., cluster centers >2 standard deviations from normal scenarios). For example, a combination of equipment failure and extreme weather conditions is clustered as a high-risk abnormal scenario.
[0078] Step S92 primarily determines key adjustment parameters. The Euclidean distance between the abnormal scenario and the baseline scenario (normal operation) is calculated. Indicators with a distance greater than a threshold (e.g., 2) are identified as "key adjustment parameters." For example, if the Euclidean distance between the structural stress in an abnormal scenario and the baseline scenario is 2.5, the stress indicator is designated as a key parameter. This step clarifies the adjustment target for the regression modeling in S93.
[0079] Step S93 primarily involves predicting configuration parameter adjustments. A "multivariate linear regression" algorithm establishes a mapping model between configuration parameters (e.g., equipment power P, number of personnel N) and performance indicators (e.g., stress σ, construction period D). The model parameters are trained using historical data (e.g., σ = 0.5P + 0.3N + Z), where Z represents the random error term, representing other influencing factors not explicitly expressed in the model (e.g., environmental variables not included in the model, measurement errors, model simplification errors, etc.). The adjustment range for key parameters is predicted. For example, to reduce stress σ by 20%, the predicted adjustment range for equipment power P is 15%, and the number of personnel N is 10%.
[0080] Step S94 primarily involves generating dynamic adjustment plans. A "priority sorting algorithm" determines the allocation order based on the adjustment range and resource constraints (such as the upper limit of available equipment power and the maximum number of personnel). For example, equipment power adjustment takes precedence over personnel allocation. The generated plan includes: 1. Equipment scheduling plan (such as activating a backup generator); 2. Personnel allocation plan (such as adding 20 construction workers); and 3. Process adjustment table (such as extending the concrete curing time to 30 days). The plan's effectiveness is verified through simulation (e.g., whether the adjusted stress σ falls within the threshold), and the final plan is output.
[0081] This embodiment significantly enhances the scientific nature and operability of decision support through data-driven dynamic adjustments. Specifically, the quantification accuracy of indicator differences is improved. PCA dimensionality reduction combined with Euclidean distance calculations achieves higher accuracy in quantifying scenario differences, transforming the traditional empirically based indicator adjustment approach into a data-driven, precise decision. The scientific nature of the adjustment plan is enhanced. The multivariate linear regression model reduces the prediction error of configuration parameter adjustment ranges. For example, the deviation between predicted equipment power adjustments and actual demand is reduced, avoiding over- or under-adjustment. Resource allocation efficiency is optimized. The priority ranking algorithm improves resource utilization. For example, in the case of equipment failure, the resource allocation plan can reduce project delays from 7 days to 3 days while keeping cost increases within 10%. The operability of the plan is enhanced. Structured scheduling plans, configuration plans, and process adjustment tables directly guide on-site execution. Vague recommendations such as "strengthen monitoring" in traditional approaches are transformed into concrete operational steps, improving execution efficiency.
[0082] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A data analysis method for digital management of engineering consulting, characterized by: The method comprises: Map heterogeneous data from different projects into a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset; Detecting data format and attribute conflicts in the grid mapping dataset, adjusting data priorities in the grid mapping dataset using a preset dynamic weight allocation rule, and obtaining a grid calibration dataset with a consistent format; For the grid calibration data set, a time series autoregressive model and smoothing technology are used to process missing values and abnormal fluctuations to obtain a time series stable data set; Analyze the grid point timestamp intervals of the time-series stable data set, and construct a complete grid data set with a continuous time axis and consistent spatiotemporal dimensions by combining linear interpolation with multidimensional data fusion; Extracting key indicators from the complete grid data set, constructing a digital mirror of the project including spatial location, time series, and attribute associations, and uniformly mapping multi-period data to a benchmark timeline based on the project duration to obtain an aligned mirror mapping data set; The latest monitoring data is obtained through the sensor interface. If key indicators are missing, a weighted average is performed based on the spatial distance between adjacent grid points and the correlation of historical data. The weight coefficient is dynamically adjusted based on the data fluctuation amplitude to supplement the missing values, and the mirror mapping dataset is updated to obtain a mirror update dataset. Identifying outlier grid points in the mirrored updated data set and correcting the outlier grid points to obtain a verified mirrored stable data set; Based on the mirror stable data set, multiple scenario operating conditions are constructed according to the project risk prediction requirements to generate a structured simulation analysis data set; Based on the simulation analysis data set, a dynamic adjustment plan including equipment scheduling, resource allocation, and process adjustment is generated.
2. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The heterogeneous data of different projects are mapped into a unified three-dimensional spatiotemporal grid to obtain a preliminary standardized grid mapping dataset, including: Collect raw data streams containing spatial coordinates, timestamps, and business attributes from the project phase; A three-dimensional grid partitioning algorithm is used to normalize the spatial coordinates and unify the time units of the original data stream. The data from different data sources are converted to a unified grid coordinate system through a mapping function to obtain a preliminary normalized mapping data set. For the preliminarily normalized mapping data set, based on the pre-set outlier detection threshold value of each monitoring scene data feature, data points exceeding the outlier detection threshold value are marked as outlier data points, thereby forming a marked outlier data set; Performing spatiotemporal autocorrelation analysis on the abnormal data set, calculating the spatial aggregation degree and time series correlation of the abnormal data points, determining whether the abnormal distribution has spatial clustering or temporal periodicity characteristics, and obtaining an abnormal distribution feature set; Use the support vector machine algorithm to classify the risk level of abnormal data points in the abnormal distribution feature set, generate a risk level heat map based on the grid spatial location, and mark areas with risk levels greater than or equal to the preset risk level and overlapping with key project nodes as key areas of concern; For key areas of concern, the pattern matching degree between real-time data stream and historical abnormal data is calculated through dynamic time warping algorithm, the risk evolution trend is judged based on the slope of the matching curve, and the risk level and evolution direction are mapped to a preliminary normalized mapping data set to form a preliminary standardized grid mapping data set containing risk characteristics.
3. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The detecting of data format and attribute conflicts in the grid mapping dataset, adjusting the data priority in the grid mapping dataset by a preset dynamic weight allocation rule, and obtaining a grid calibration dataset with a consistent format, includes: The DBSCAN clustering algorithm is used to perform spatial density clustering on grid points, identify high-density clustering areas and sparse distribution areas, and obtain a grid grouping data set; For each grid point in the cluster group, the mean, variance, and rate of change characteristics of the time series are extracted, and the amplitude of changes between adjacent time points is calculated through a sliding window. Points exceeding the preset fluctuation threshold are marked as dynamic outliers. Extract dynamic outliers, including format parameters such as data type, data length, and precision, and compare the format parameters with industry standard data format templates. Mismatches are marked as format outliers and the mismatch type is recorded. Furthermore, the business attributes of the dynamic outliers are analyzed, and the logical consistency between the attributes is verified using a decision tree rule engine. Inconsistencies are marked as attribute conflicts and the conflict rules are recorded. Perform principal component analysis on format anomalies and attribute conflict points, extract the first N principal components as feature vectors, calculate the cosine similarity with the historical anomaly feature library, and mark those with similarity less than the preset similarity threshold as new anomaly patterns; The potential expansion range of the newly added abnormal pattern in the grid is predicted based on the Kriging interpolation method. The neighborhood association degree is calculated to determine whether it affects the key area. The data affecting the key area is prioritized according to the rules of "real-time data > historical data" and "high reliability source data > low reliability source data" to generate a grid calibration dataset with a consistent format.
4. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The grid calibration dataset is processed using a time series autoregressive model and smoothing technology to process missing values and abnormal fluctuations to obtain a time series stable dataset, including: The ARIMA model is used to fit the time series of the grid calibration dataset, and the seasonal characteristics with a period of T are extracted. The periodic variation intensity of each grid point is calculated, where the grid points with variation intensity greater than the preset intensity threshold are marked as high dynamic grid points. The spatial distribution density of high-dynamic grid points is calculated by the kernel density estimation method, and the Delaunay triangulation algorithm is used to divide the boundaries of high-density areas. Isolated grid points with neighborhood correlation strength less than the preset correlation strength threshold are classified as isolated points. A spatial buffer zone is established with the isolated point as the center, the number of grid points and data influence weight within the buffer zone are calculated, and the spatial density and periodic variation intensity are integrated through a weighted superposition model. For grid points with missing values in the time series of the fused data, the inverse distance weighted interpolation method is used to fill the missing values using the historical data of the neighborhood N×N grid; for abnormal fluctuation points, the spatiotemporal constraints are combined with the mean of the previous and next M periods to generate a time series stable data set, where M and N are both positive integers.
5. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The analyzing the grid point timestamp intervals of the time-series stable data set and combining linear interpolation with multidimensional data fusion to construct a complete grid data set with a continuous time axis and consistent spatiotemporal dimensions includes: Calculate the time interval between adjacent timestamps of the same grid point in a time-stable dataset. If the time interval is greater than a preset time threshold, linear interpolation is triggered. The intermediate data points generated by interpolation are optimized through multi-dimensional data fusion technology, including: Check the data fluctuations of the preset time length before and after the interpolation point. If the fluctuation is greater than the preset fluctuation threshold, exponential smoothing correction is introduced; and the correlation of the synchronous data of a preset number of grid points in the neighborhood is calculated. If the correlation is lower than the preset correlation threshold, Gaussian process regression is used to supplement the spatial correlation features; Construct a spatiotemporal integrity assessment model to supplement the missing parts of grid points with more than a preset number of time axis break locations or more than a preset percentage of missing spatial neighborhood data by comparing with historical data of the same period; Through spatiotemporal consistency verification, a complete grid dataset with continuous time and consistent spatiotemporal dimensions is generated.
6. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The key indicators are extracted from the complete grid dataset to construct a digital mirror of the project containing spatial location, time series, and attribute associations, and the multi-period data is uniformly mapped to a reference timeline based on the project duration to obtain an aligned mirror mapping dataset, including: Define a set of key indicators for the project , extract the spatiotemporal distribution data of each period K from the grid complete dataset, where, For progress, is energy consumption, For quality; Use digital twin technology to build a digital mirror of the geometric model, physical model, and business model, and establish a three-dimensional mapping relationship of "grid coordinates-timestamp-index value"; Map the data of each monitoring period to the benchmark time axis based on the total project duration through time offset correction and frequency unification to form a preliminary aligned data set; For abnormal points in the preliminary alignment data set whose index fluctuation is greater than the index fluctuation threshold, local weighted regression is used for smoothing correction. Calculate the rate of change of indicators across periods and generate a mirror data set containing trend characteristics; The final mirror mapping dataset is obtained through benchmark timeline consistency check and mirror mapping accuracy verification.
7. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The latest monitoring data is obtained through the sensor interface. If key indicators are missing, weighted average is performed based on the spatial distance of adjacent grid points and the correlation of historical data. The weight coefficient is dynamically adjusted in combination with the data fluctuation amplitude to supplement the missing values. The mirror mapping dataset is updated to obtain a mirror update dataset, including: Collect sensor data in real time through the IoT interface, align timestamps with the mirrored mapping dataset, and trigger the data synchronization mechanism; For grid points with missing key indicators, a two-step method is used to fill them, including: taking the real-time data of the adjacent A×A grid points and performing inverse distance weighted averaging, where A is a positive integer; and automatically adjusting the weights according to the fluctuation range of historical data; The spatial distribution of the filled data is checked. If the difference with the neighboring grid data is greater than the preset difference threshold, it is marked as a suspicious point and corrected by the mean of the data before and after the preset time length; The deviation between the real-time data and the predicted value of the mirror model is calculated. The distribution characteristics of the areas where the deviation is greater than the deviation threshold are refitted through the Gaussian mixture model to generate a mirror update dataset.
8. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: The identifying of outlier grid points in the mirrored updated data set and correcting the outlier grid points to obtain a verified mirrored stable data set includes: The local outlier factor algorithm is used to calculate the outlier degree of each grid point in the mirror update dataset, and grid points with outlier degrees greater than 2 are marked as outlier points. The outlier grid points were corrected in both time and space, including using the K-nearest neighbor algorithm to calculate the mean of the neighborhood grid points as a correction reference value; and using cubic spline interpolation to fit the data curves before and after 24 hours to generate a smoothed time series. The corrected data were rechecked to calculate spatial neighborhood consistency and time series continuity, and non-compliant areas were further corrected using manual review rules; Finally, a mirror-stable dataset is generated through spatiotemporal stability evaluation.
9. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: Based on the mirror stable data set, multi-scenario operating conditions are constructed according to the project risk prediction requirements to generate a structured simulation analysis data set, including: Defining multiple scene collections , set the key indicator fluctuation range for each scenario, where: For normal operation, For equipment failure, for extreme weather; The Monte Carlo simulation method is used to generate indicator sequences for each scenario, which are input into the digital mirror model for behavioral simulation and the output parameters are recorded. Perform anomaly detection on the simulated output data and calculate the Mahalanobis distance between each indicator and the normal scene. Points with a distance greater than 2 are marked as abnormal behavior points. Perform time series trend analysis on abnormal behavior points, use the LSTM model to predict the duration and impact range of the abnormality, and generate a risk feature vector that includes the probability of risk occurrence and the degree of impact; Data fusion technology is used to associate risk characteristics with grid spatial locations and timestamps, and simulation analysis data sets are output in a structured format of "risk level-occurrence time-impacted area." 10. The data analysis method for digital management of engineering consulting according to claim 1 is characterized in that: Generating a dynamic adjustment plan including equipment scheduling, resource allocation, and process adjustment based on the simulation analysis data set includes: Principal component analysis is used to reduce the dimensionality of multi-scenario simulation data in the simulation analysis dataset, and the first X principal components are extracted as scene feature vectors. K-means clustering is used to identify abnormal scenes. Quantify the indicator differences in abnormal scenarios and calculate the Euclidean distance between each abnormal scenario and the baseline scenario. The indicator with a Euclidean distance greater than the preset distance threshold is used as the key adjustment parameter; A "configuration parameter-indicator performance" model was established through multiple linear regression to predict the adjustment range of equipment power and personnel configuration parameters; Based on the adjustment range and resource constraints, a priority sorting algorithm is used to determine the resource allocation order, and a dynamic adjustment plan is generated that includes equipment scheduling plan, staffing plan, and process parameter adjustment table. The final plan is output after simulation verification.
Citation Information
Patent Citations
Hydraulic engineering data intelligent monitoring method and system based on digital twinning
CN119783356A
Prefabricated building component stability monitoring and early warning method and system
WO2025118307A1
Cited By
Project four-control digital monitoring optimization method and system based on dynamic model
CN120952718A
Real-time collaborative awareness method based on extreme climate risk
CN121144768A
Real-time collaborative sensing method based on extreme climate risk
CN121144768B
Electric energy meter metering abnormity analysis method and system
CN121234254A
Flexible freight bag automatic equipment virtual debugging system and method based on 3D model technology
CN121351362A