A multi-source construction data analysis and storage method and device
Patent Information
- Application Number
- CN202610865064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-16
AI Technical Summary
[0004]为了解决现有技术因数据抖动、空间漂移与事件重复导致的对齐质量差、边权分配不真实、无法提取施工主链以及缺乏冲突回溯与分层存储机制的问题
[0013]This invention addresses the issues of data arrival jitter and spatial drift by collecting multi-source construction data and adjusting the time window width and grid granularity based on reliability assessment. Simultaneously, it constructs a multi-dimensional association graph encompassing time, space, and event nodes, and adjusts edge weights based on node interaction states and topological stability parameters to represent the spatiotemporal relationships and change logic among construction elements. Furthermore, by compressing paths and folding composite nodes in the association graph to extract the longest construction main chain, isolated and redundant information is eliminated, reducing computational analysis and storage overhead. Combined with a calculated threshold to extract multi-dimensional difference fingerprints for consistency verification, it can trigger conflict source backtracking when deviating from the constraint template. By optimizing the data organization structure through local reorganization and hierarchical storage, it achieves rapid traceability of construction anomalies and reliable data storage management.
Smart Images

Figure CN122412893B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data analysis, and in particular relates to a method and apparatus for multi-source construction data analysis and storage. Background Technology
[0002] The construction site deploys numerous IoT sensors and business information systems, generating real-time multi-source construction data including personnel distribution, machinery operation, material flow, and environmental monitoring. Due to the harsh construction environment, inconsistent accuracy of data acquisition equipment, and inherent network latency, this multi-source data exhibits arrival jitter, spatial drift, and event duplication in both temporal and spatial dimensions, leading to data deviations in construction event correlation analysis, progress management, and status verification. By combining spatiotemporal correlation modeling with consistency verification techniques, a graph structure model containing spatiotemporal nodes and construction events is constructed to track and detect conflicts in the construction process. The analysis results are then written into a database for storage management.
[0003] However, existing technologies lack adjustment mechanisms based on the credibility of multi-source data, resulting in low fusion quality of spatiotemporal segment sequences. In the construction and computation of spatiotemporal correlation graphs, the allocation of temporal, spatial, and entity connection weights between nodes is rather coarse, failing to incorporate topological stability parameters of node divergent connections to correct edge weights, thus failing to accurately represent the dynamic construction interaction relationships. Faced with redundant graph nodes, traditional methods cannot extract the main construction chain, leading to a lack of basis for comprehensively calculating differences in temporal closure, spatial location difference, resource overlap, and process violation characteristics. Furthermore, when verification indicators deviate abnormally, existing technologies lack conflict source backtracking methods and hierarchical storage strategies for local reorganization, resulting in wasted storage resources, low query efficiency, and an inability to meet the actual business needs of high-frequency traceability and high consistency analysis of modern engineering data. Summary of the Invention
[0004] To address the problems of poor alignment quality, inaccurate edge weight allocation, inability to extract the main construction chain, and lack of conflict backtracking and hierarchical storage mechanisms caused by data jitter, spatial drift, and event duplication in existing technologies.
[0005] In a first aspect, the present invention provides a method for analyzing and storing multi-source construction data, the method comprising:
[0006] Multi-source construction data is collected, and the data is encoded based on time reference and grid division to obtain a spatiotemporal segment sequence. The reliability of each data source is calculated based on the arrival jitter value, spatial drift value and event repetition rate in adjacent time windows. The time window width and grid granularity are adjusted to align the original data based on the comprehensive value of the reliability of each data source, and interval segmentation is performed to generate an aligned spatiotemporal segment sequence.
[0007] Based on the alignment of spatiotemporal fragment sequences, an association graph containing time nodes, spatial nodes, and event nodes is constructed. Temporal edges, spatial edges, and entity edges between nodes are established in the association graph. Initial edge weights are set according to the distribution of the number of event conflicts and the number of stable events in node interactions. The node topology stability parameters are calculated by statistically analyzing the density of node divergent connections. The edge weight correction coefficient is calculated in combination with the node topology stability parameters to correct the initial edge weights and obtain the corrected edge weights.
[0008] The association graph is compressed by folding adjacent event nodes within the same process with a time interval less than a threshold into composite nodes. Unconnected isolated branches are removed to extract the longest construction main chain with key composite nodes. The threshold is determined by the modified edge weight statistics and node topology stability parameters. The time closure, spatial position difference, resource overlap and process violation are calculated based on the longest construction main chain and merged to generate a difference fingerprint. When the difference fingerprint deviates from the constraint template, the conflict source backtracking is triggered. The association graph is locally reorganized according to the backtracking results to obtain hierarchical storage blocks, and the generated tags are written to the repository.
[0009] In another aspect, the present invention provides a multi-source construction data analysis and storage device, the device comprising:
[0010] The generation module is used to collect multi-source construction data, encode the data based on time reference and grid division to obtain a spatiotemporal segment sequence, calculate the reliability of each data source based on arrival jitter value, spatial drift value and event repetition rate in adjacent time windows, and adjust the time window width and grid granularity to align the original data based on the comprehensive value of the reliability of each data source, and perform interval segmentation to generate an aligned spatiotemporal segment sequence.
[0011] The correction module is used to construct an association graph containing time nodes, spatial nodes and event nodes based on the aligned spatiotemporal segment sequence. In the association graph, temporal edges, spatial edges and entity edges between each node are established. The initial edge weights are set according to the distribution of the number of event conflicts and the number of stable events in node interactions. The node topology stability parameters are calculated by statistically analyzing the density of node divergent connections. The edge weight correction coefficient is calculated in combination with the node topology stability parameters to correct the initial edge weights and obtain the corrected edge weights.
[0012] The reorganization module is used to perform path compression processing on the association graph. It folds adjacent event nodes within the same process and with a time interval of less than a threshold into composite nodes, removes unconnected isolated branches, and extracts the longest construction main chain with key composite nodes. The threshold is determined by the modified edge weight statistics and node topology stability parameters. Based on the longest construction main chain, it calculates time closure, spatial location difference, resource overlap, and process violation and merges them to generate difference fingerprints. When the difference fingerprint deviates from the constraint template, it triggers conflict source backtracking. Based on the backtracking results, it locally reorganizes the association graph to obtain hierarchical storage blocks, generates tags, and writes them to the repository.
[0013] This invention addresses the issues of data arrival jitter and spatial drift by collecting multi-source construction data and adjusting the time window width and grid granularity based on reliability assessment. Simultaneously, it constructs a multi-dimensional association graph encompassing time, space, and event nodes, and adjusts edge weights based on node interaction states and topological stability parameters to represent the spatiotemporal relationships and change logic among construction elements. Furthermore, by compressing paths and folding composite nodes in the association graph to extract the longest construction main chain, isolated and redundant information is eliminated, reducing computational analysis and storage overhead. Combined with a calculated threshold to extract multi-dimensional difference fingerprints for consistency verification, it can trigger conflict source backtracking when deviating from the constraint template. By optimizing the data organization structure through local reorganization and hierarchical storage, it achieves rapid traceability of construction anomalies and reliable data storage management. Attached Figure Description
[0014] Figure 1 A flowchart for multi-source construction data analysis and storage methods;
[0015] Figure 2 A schematic diagram of the statistical distribution of topological stability under the core event connection density triggering;
[0016] Figure 3 A schematic diagram summarizing the four categories of indicators for multidimensional deviation fingerprint features;
[0017] Figure 4 A line graph illustrating the optimization of daily computing power consumption performance using the composite node time scale threshold strategy. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] In this application, a method for multi-source construction data analysis and storage is provided, such as... Figure 1 As shown, it includes:
[0020] S1 collects multi-source construction data, encodes the data based on time reference and grid division to obtain a spatiotemporal segment sequence, calculates the reliability of each data source based on arrival jitter value, spatial drift value and event repetition rate in adjacent time windows, and adjusts the time window width and grid granularity to align the original data based on the comprehensive value of the reliability of each data source, and performs interval segmentation to generate an aligned spatiotemporal segment sequence.
[0021] The consumer application programming interface of the Apache Kafka distributed stream processing platform is called to collect multi-source construction data in real time and extract timestamp and latitude / longitude coordinate attributes.
[0022] The network time protocol is used to set the time base for global time synchronization. The formatting function of the standard time processing library is called to convert the timestamp into a time series with a unified standard time zone. The spatial indexing algorithm in the geospatial data processing library GeoPandas is called, and the GeoHash algorithm is used to map the geographic coordinates into a spatial grid code of a specific level. The time series and the spatial grid code are spliced together to obtain a spatiotemporal segment sequence.
[0023] A sliding time window is established, and the variance calculation function in the NumPy statistical library is used to calculate the variance of the arrival time interval of data packets within adjacent time windows. After normalization, this variance is used as the arrival jitter value. The theoretical coordinates of entities are predicted using the Kalman filter algorithm, and the absolute value of the distance deviation between the actual displacement and the theoretically predicted displacement of the same entity is calculated using the semi-versus formula. After normalization, this deviation is used as the spatial drift value. The number of times construction tasks with identical feature codes are repeatedly uploaded within a set period is counted, and after normalization, this number is used as the event repetition rate. The normalized arrival jitter value, spatial drift value, and event repetition rate are each assigned a preset weight and then summed. The weighted summation result is subtracted from the constant 1 to obtain the credibility of each data source, which is not lower than the preset lower limit.
[0024] Based on the comprehensive credibility values of various data sources, the time window width and grid granularity are adjusted using an adaptive adjustment strategy. When the credibility is lower than a preset threshold, the time window width is increased by a linear interpolation algorithm, and the encoding level of the GeoHash algorithm is reduced to coarsen the grid granularity. Conversely, the time window width is reduced and the encoding level is increased. The time resampling and aggregation functions of the Pandas data analysis library are called to resample, aggregate, and align the original data according to the adjusted time window width and grid granularity. The binary interval tree algorithm is used to slice the data stream according to the aligned time boundaries to generate an aligned spatiotemporal segment sequence.
[0025] In an optional embodiment, the step of encoding the data based on a time base and grid partitioning to obtain a spatiotemporal segment sequence includes:
[0026] Subtract the starting zero point corresponding to the unified time base from the timestamp of the multi-source construction data, divide by the preset time segment duration and round down to obtain the time segment code;
[0027] The physical coordinates of the multi-source construction data are obtained. The longitudinal distance offset of the physical coordinates relative to the longitude starting point of the construction area is calculated and divided by a preset longitudinal grid unit length and rounded down. The latitudinal distance offset of the physical coordinates relative to the latitude starting point is calculated and divided by a preset latitudinal grid unit length and rounded down. The combined calculation results are used to obtain the spatial segmentation code.
[0028] The time segment code, spatial segment code, and original input data text structure items are concatenated to form a standard combined string, generating the spatiotemporal segment sequence.
[0029] Set a unified time reference starting point, such as setting the project start time as 2023-01-01, with 00:00:00 as 0s. Convert the timestamps contained in the multi-source construction data into second-level Unix timestamps and subtract them from the starting point. Then divide by the preset time segment duration, preferably in the range of 300s to 3600s, for example, a value of 900s. Round the result down to obtain the time segment code, for example, the code T0042.
[0030] Obtain the GPS coordinates of the construction data. Based on the preset longitude and latitude starting points of the construction area, calculate the absolute distance offset of these starting points in the longitude and latitude directions using the spherical distance formula or a plane projection coordinate system. For example, the longitude offset is 598.5m and the latitude offset is 888.0m. Divide the longitude and latitude distance offsets by the preset grid unit length. For example, the preset grid unit length for both longitude and latitude is 10m, which is within the preferred range of 10m to 50m. Round down to obtain the longitude index 59 and the latitude index 88. Concatenate them with underscores to obtain the spatial segment code, for example, S_59_88.
[0031] The generated time-segmented and spatial-segmented codes are concatenated with the original input data text structure items, such as device IDEXcavator-01 and operation command Dig, using the delimiter "|", to form a standard combined string similar to T0042|S_59_88|Excavator-01|Dig. All real-time incoming strings are then appended sequentially to the in-memory time-series database according to their sequence structure, thus generating a complete spatiotemporal segment sequence for high-frequency retrieval and matching in the next step.
[0032] In an optional embodiment, the step of calculating the reliability of each data source based on arrival jitter, spatial drift, and event repetition rate within adjacent time windows, adjusting the time window width and grid granularity to align the original data based on the comprehensive reliability values of each data source, and performing interval segmentation to generate an aligned spatiotemporal segment sequence includes:
[0033] Calculate the variance of the time interval between two consecutive data arrivals in the data source, and use the normalized value as the arrival jitter value;
[0034] Calculate the absolute value of the spatial distance deviation between the actual location of the current sampling point and the theoretical location predicted based on historical data for the same entity among adjacent sampling points, and use it as the spatial drift value after normalization.
[0035] The number of times that construction task records with identical statistical feature codes are repeatedly uploaded within a set period is normalized and used as the event repetition rate.
[0036] The arrival jitter value, spatial drift value and event repetition rate are each assigned a preset weight and summed by weight. The weighted sum is then subtracted from the constant to obtain the credibility of each data source, which is not lower than a preset lower limit. The mean of the credibility of each data source is then calculated as the comprehensive credibility value of each data source.
[0037] Extract the baseline credibility threshold and baseline time window size issued during the initialization phase. Multiply the ratio of the baseline credibility threshold to the comprehensive credibility value of each data source by the baseline time window size to obtain a new duration value, and replace the time window width.
[0038] Extract the baseline confidence threshold and standard grid side length size issued during the initialization phase. Multiply the arithmetic square root of the ratio of the baseline confidence threshold to the comprehensive confidence value of each data source by the standard grid side length size to obtain a new side length value, and replace the grid granularity.
[0039] The aligned spatiotemporal segment sequence is generated by performing interval segmentation in the full sample data space using the adjusted time window width and grid granularity as truncation boundaries.
[0040] Within the current calculation cycle, the time intervals of adjacent data packets from each sensor data source are extracted and the statistical variance is calculated. The variance is mapped to the [0,1] interval using the minimax normalization formula to obtain the arrival jitter value. For example, if the variance is 12.5, the jitter value after normalization is 0.15. Simultaneously, the current theoretical coordinates of the entity are predicted using the Kalman filter algorithm. The absolute value of the Euclidean distance deviation between the actual GPS feedback coordinates and the theoretical coordinates is calculated, and a similar normalization operation is performed to obtain the spatial drift value. For example, if the drift distance is 1.2m, the normalized value is 0.08. The statistics are set for a period of time, preferably 1 to 5 minutes, for example, the number of log entries with identical but duplicated feature codes, such as the specific pouring task ID Task_A1, within 3 minutes. After normalization, the event duplication rate is obtained, for example, the duplication rate value is 0.05.
[0041] Preset empirical weights are assigned to arrival jitter value, spatial drift value, and event repetition rate, with a preferred weight allocation ratio of 0.4:0.4:0.2. A weighted sum is performed, and the sum is subtracted from the constant to obtain the credibility of each data source. When the credibility is lower than a preset lower limit, preferably 0.3, it is forced to be 0.3. The arithmetic mean of the credibility of each source is calculated to obtain a comprehensive value, for example, a comprehensive score of 0.85. When adjusting parameters, the baseline credibility threshold (preferably 0.8) and the baseline time window size (e.g., 1000s) in the initialization configuration file are retrieved and multiplied by the ratio of the baseline credibility threshold to the comprehensive value, resulting in a new duration of approximately 941s, which replaces the original time window width. The standard grid side length (e.g., 20m) is retrieved and multiplied by the square root of the ratio of the baseline credibility threshold to the comprehensive value, approximately 0.969, resulting in a new side length of approximately 19.38m, which covers the current grid granularity. Based on the real-time generated boundaries, approximately 941s and 19.38m, the cached samples in the global database are re-divided. The starting endpoints are defined by cutting cursors to obtain a set of aligned spatiotemporal segment sequence files with strictly uniform scale and free from spatiotemporal misalignment noise.
[0042] S2: Based on the aligned spatiotemporal segment sequence, construct an association graph containing time nodes, spatial nodes, and event nodes. Establish temporal edges, spatial edges, and entity edges between nodes in the association graph. Set initial edge weights according to the distribution of event conflict times and stability times of node interactions. Calculate node topology stability parameters by statistically analyzing the density of node divergent connections. Combine the node topology stability parameters to calculate edge weight correction coefficients and correct the initial edge weights to obtain corrected edge weights.
[0043] The graph structure is initialized by calling the multidirected graph instantiation function in the NetworkX graph computation and analysis library. The time nodes are extracted from the timestamps in the aligned spatiotemporal segment sequence, the spatial nodes are extracted from the grid code, and the event nodes are extracted from the business semantics. The three types of nodes are loaded into the in-memory graph object through the graph structure's built-in node addition function.
[0044] The function to add edges is called to create directed temporal edges between time nodes ordered by time, undirected spatial edges between spatial nodes corresponding to adjacent spatial grids, and undirected entity edges between event nodes and the time and spatial nodes corresponding to the event, thus obtaining the association graph.
[0045] Traverse the entity edges of the association graph. Define event conflict as the same spatial node being associated with an event node belonging to a mutually exclusive process at the same time node, and define stability as no mutually exclusive process association occurring. Set a memory counter to traverse and accumulate the number of event conflicts and stability times of each node interaction. For entity edges containing event nodes, calculate the ratio of the stability times to the total number of interactions to obtain the basic weight. Combine this with a preset penalty coefficient to generate a decay penalty term. Subtract the decay penalty term from the basic weight and map it to the positive value range through the Sigmoid activation function as the initial edge weight. For temporal edges and spatial edges, set the number of event conflicts to 0 and the stability times to the total number of interactions. Calculate the initial edge weight using the same rules and write the initial edge weight into the attribute dictionary of the corresponding edge.
[0046] The total number of entity edges extending outward from the target node is used as the divergent connection density. The mean and variance of the initial edge weights of all edges connected to the target node are calculated. The divergent connection density is multiplied by the mean of the initial edge weights and then divided by the sum of the variance and the smoothing constant to obtain the node topological stability parameters.
[0047] Extract the topological stability parameters of the starting and ending nodes of the edges, calculate the geometric mean of the two to obtain the original correction coefficient, constrain the original correction coefficient to the standard range of 0.5-1.5 by linear scaling to obtain the final edge weight correction coefficient, multiply the initial edge weight by the edge weight correction coefficient, and update the edge attribute dictionary of the association graph to obtain the corrected edge weight.
[0048] In an optional embodiment, the construction of an association graph containing time nodes, spatial nodes, and event nodes based on the aligned spatiotemporal segment sequence, establishing temporal edges, spatial edges, and entity edges between nodes in the association graph, and setting initial edge weights according to the distribution of the number of event conflicts and stable events in node interactions, includes:
[0049] The aligned spatiotemporal segment sequence is deduplicated and extracted, and the independent timestamps are mapped to time nodes, the two-dimensional coordinate grid numbers are mapped to spatial nodes, and the business operation process types are mapped to event nodes;
[0050] Connecting adjacent time nodes in ascending order of time is defined as temporal edges, and connecting spatial nodes that conform to the Cartesian adjacency association rule is defined as spatial edges.
[0051] Within the same data packet directory identifier, all time nodes, space nodes, and event nodes belonging to the same identifier are connected, and a network structure is built by pairwise interactions, which are set as entity edges to form the association graph.
[0052] For edges containing event nodes, the total number of reverse mutual exclusion error uploads of the actual actions of the physical objects associated with the two nodes is the event conflict count, and the number of correct matching and flow counts is the stable count;
[0053] For temporal and spatial edges, the event conflict count is set to zero and the stability count is set to the total number of interactions for the corresponding edge.
[0054] The basic weight is obtained by calculating the ratio of the stable number of times to the total number of interactions. A decay penalty term is generated using the number of event conflicts. The basic weight is subtracted from the decay penalty term and mapped to a positive value range through an activation function, which is then set as the initial edge weight of the corresponding connection edge.
[0055] During the graph construction phase, the aligned spatiotemporal segment sequence is deduplicated using a hash set to extract three types of entities: timestamps accurate to the second are instantiated as time nodes, such as Node_T_1001; two-dimensional coordinate grid numbers are instantiated as spatial nodes, such as Node_S_59_88; and parsed business operation process types, such as hoisting and welding, are instantiated as event nodes, such as Node_E_Weld.
[0056] In terms of connection establishment, the time node sequence is sorted according to the time axis order, and the connection between adjacent nodes in time sequence is used as time sequence edge; the spatial nodes are judged by octagonal adjacency determination to generate spatial edges for the corresponding nodes of adjacent grids; at the same time, through the directory identifier in the packet header, such as packet IDPK_001, a bridge is used to establish completely undirected pairwise interconnections between time, space and event nodes under this identifier in the graph database, which are defined as multimodal entity edges, thereby constructing a basic heterogeneous association graph containing topological and spatiotemporal features.
[0057] During the initial edge weight calculation and assignment process, the connection type is retrieved: if it is a business-related edge connecting event nodes, the interaction records of the two related node objects in the system log are retrieved, and the total number of reverse mutual exclusion error uploads caused by process reversal and resource deadlock is counted as the event conflict count. For example, if there are 5 errors due to device collision in the record, the number of correct matching flows that are successfully completed is counted as the stable count. For example, if there are 95 normal executions, the total number of interactions is 100. If it is a simple temporal edge or spatial edge, the default is that there is no conflict, that is, the event conflict count is zero, and the total frequency of edge interaction is directly assigned to the stable count.
[0058] Divide the number of stable events by the total number of interactions to obtain the basic weight of the edge, such as 0.95. Combine the number of event conflicts with the penalty coefficient, preferably within the range of 0.01 to 0.05, such as 0.02, to generate a decay penalty term of 0.1. Subtract the decay penalty term from the basic weight to obtain the median value of 0.85. Use the Sigmoid activation function to process this median value to ensure a smooth mapping and that it falls strictly within the range greater than zero, for example, the output is 0.701. Write this result as a parameter into the initial edge weight field of the connection attribute in the graph database.
[0059] In an optional embodiment, the step of calculating the node topology stability parameter by statistically analyzing the density of divergent connections, and then calculating the edge weight correction coefficient based on the node topology stability parameter to correct the initial edge weight and obtain the corrected edge weight includes:
[0060] The total number of edges extending outward from the target node to connect to the entity is used as the connection density.
[0061] Calculate the mean and variance of the initial edge weights of all edges connected to the target node;
[0062] The topological stability parameters of the target node are obtained by multiplying the connection density by the mean of the initial edge weights and dividing by the sum of the variance and the smoothing constant.
[0063] Extract the topological stability parameters of the starting node and the ending node of the initial edge weight to be corrected, calculate the geometric mean of the two, and then linearly scale them according to a preset scaling factor to obtain the edge weight correction coefficient.
[0064] Multiply the initial edge weight to be corrected by the edge weight correction coefficient, and set it as the corrected edge weight after edge update.
[0065] For any given target node in the association graph, such as a specific event node, call the graph traversal algorithm to retrieve the first-order neighborhood, count the total number of edges radiating outward from the target node and actually connecting other entities, and use this integer value as the connection density. In the example, the total number of diverging edges of a high-frequency scheduling node is 45.
[0066] The initial edge weight attribute values stored in the 45 connected edges are traversed, and the arithmetic mean of the initial edge weights is calculated using statistical formulas. Assuming the calculated mean is 0.72, and the sample variance representing the degree of weight fluctuation is assumed to be 0.04, to avoid calculation anomalies caused by a zero denominator, the arithmetic mean 0.72 is multiplied by the connection density 45 to obtain 32.4. This is divided by the sample variance 0.04 and a preset small positive real number, i.e., a smoothing constant, preferably between 1e-3 and 1e-2, for example, 0.01. The sum of these two is 0.05, and the topological stability parameter of the target node is calculated to be 648.0. A higher value indicates richer node connections and a highly concentrated and stable weight distribution. To visually present the overall distribution law of the node topological stability parameter calculated by this method in the full set of construction event samples, the statistical distribution of topological stability under the core event connection density trigger is shown below. Figure 2 As shown.
[0067] During the edge weight correction phase, each edge in the graph is processed sequentially to obtain the topological stability parameters corresponding to the starting and ending nodes of the edge to be corrected. For example, the starting stability parameter is extracted as 648.0, and the ending stability parameter is 800.0. The two values are multiplied and then squared to calculate the geometric mean of the two values as the edge weight correction coefficient, resulting in a correction coefficient of approximately 720.0. In practical applications, this can be proportionally constrained to a dimensionless range of 0.5 to 1.5 using a linear scaling factor, such as 1.15 after scaling.
[0068] The original edge weight, such as 0.75, is multiplied by the calculated edge weight correction coefficient of 1.15. The resulting product, 0.8625, is overwritten into the original weight field in the graph data model and established as the corrected edge weight of the connection edge.
[0069] S3. Perform path compression on the association graph, fold adjacent event nodes within the same process with a time interval less than a threshold into composite nodes, remove disconnected isolated branches and extract the longest construction main chain with key composite nodes. The threshold is determined by the modified edge weight statistics and node topology stability parameters. Calculate time closure, spatial position difference, resource overlap and process violation based on the longest construction main chain and merge them to generate difference fingerprints. When the difference fingerprint deviates from the constraint template, trigger conflict source backtracking. Based on the backtracking results, locally reorganize the association graph to obtain hierarchical storage blocks, generate tags and write them to the repository.
[0070] Traverse the event nodes of the association graph, obtain the cumulative average of the corrected edge weights of all edges on the candidate path as the corrected edge weight statistic, and sum the mean of the path node topology stability parameters after the corrected edge weight statistic is scaled in the same way as the correction coefficient, according to the preset weight, and then multiply by the preset time conversion coefficient to obtain the folding threshold.
[0071] Compare the absolute differences in the time attributes of adjacent event nodes within the same process. When the absolute difference is less than the threshold, call the node shrinking function of the NetworkX library to aggregate and fold the above adjacent event nodes into composite nodes, synchronously merge the associated edges and accumulate the corresponding corrected edge weights.
[0072] The connected component statistics algorithm is called to find isolated nodes and branches in the graph that cannot be connected to the start and end nodes of the main chain. The node removal function is called to remove them. The key composite nodes remaining in the graph are traversed based on the topology sorting algorithm and the planning algorithm. The weighted path length is calculated by accumulating the modified edge weights on the path. The directed path with the largest sum of modified edge weights is extracted as the longest construction main chain.
[0073] Based on the longest construction main chain, the actual total operating time is extracted, and the time closure is obtained by dividing the target total cycle time of the process by the actual total operating time. The Euclidean distance algorithm is used to calculate the geometric distance between the proposed spatial coordinates and the actual landing coordinates to obtain the spatial landing difference. The percentage of overlap of resource usage exceeding the maximum load threshold in the same time period and space is calculated to obtain the resource overlap. The actual process is compared with the standard process, and the number of reverse violation steps and the total number of violations counted by reverse sorting are calculated and divided by the total number of nodes in the main chain sequence to obtain the process violation degree.
[0074] The temporal closure, spatial location difference, resource overlap, and process violation are converted into equal-length hexadecimal feature code fields and concatenated to obtain a difference fingerprint string. This string is then mapped byte by byte to a fixed-dimensional feature vector.
[0075] The cosine similarity calculation function is invoked to evaluate the similarity between the difference fingerprint feature vector and the preset benchmark constraint template vector. When the similarity is lower than a predetermined threshold, the conflict source backtracking process is triggered. The benchmark constraint template vector is a standard feature vector generated by calculating and encoding four core indicators under ideal construction conditions based on construction specifications, BIM design drawings, and standard procedures.
[0076] Calculate the first-order difference of edge weights between adjacent nodes along the longest construction main chain. When the absolute value of the difference is greater than the mutation threshold, locate the node near that position as the conflict source. The mutation threshold can be set as the mean of all edge weight differences plus twice the standard deviation.
[0077] Based on the location of the conflict source, a subgraph local segmentation algorithm based on conflict anchor points is invoked to segment the association graph into multiple local subgraph clusters. These clusters are then reorganized hierarchically according to time and spatial levels. Finally, the columnar serialization function of the Apache Arrow data processing library is called to convert the reorganized subgraphs into hierarchical storage blocks in Parquet file format.
[0078] Generate metadata tags containing a globally unique identifier for the conflict source and a nanosecond-level timestamp. Call the distributed file system interface and use the transmission control protocol to persist the tags as index fields in the distributed object repository along with the hierarchical storage blocks.
[0079] In an optional embodiment, the path compression process on the association graph, which involves folding adjacent event nodes within the same process and with a time interval less than a threshold into composite nodes, and removing disconnected isolated branches to extract the longest construction main chain containing key composite nodes, includes:
[0080] The corrected edge weights of all edges on the candidate path are summed and divided by the total number of nodes passed through to obtain the arithmetic mean, thus yielding the corrected edge weight statistics.
[0081] The corrected edge weight statistics are weighted and summed together with the mean of the topological stability parameters of all nodes to which the path is attached. The result is then multiplied by a preset time conversion coefficient and set as the threshold.
[0082] Within the same specified code grid, find adjacent event node sequences that are completely duplicated, have a unified task flow name, and have a consecutive time interval between their first and last occurrences that is less than the threshold.
[0083] Collapse consecutive adjacent event nodes that meet the conditions into a composite node that includes a start and end time and a unique identifier object header;
[0084] Isolated branches without connected start and end nodes are removed, and the directed path with the largest sum of modified edge weights that retains all key composite nodes and the original connected nodes is retained as the longest construction main chain.
[0085] Before analyzing the construction sequence, a complete candidate path is extracted from the starting point to the ending point using a depth-first search strategy. The corrected edge weights of all connecting edges traversed on the candidate path are read and arithmetically summed. The sum is then divided by the total number of nodes covered by the current path. For example, if the sum of weights is 18.5, the path contains 20 nodes. The corrected edge weight statistics are calculated, such as an average value of 0.925.
[0086] The arithmetic mean of the corrected edge weight statistics and the topological stability parameters of all nodes on the scaled path is calculated. Assuming this mean is 1.25, a weighted summation is performed using preset proportions, preferably 0.4 and 0.6, yielding 1.12. This summation result is multiplied by a preset time conversion coefficient mapped to a time scale, in seconds per unit weight (e.g., 1000). The calculated threshold is then output, resulting in a threshold duration of 1120 seconds. When performing graph structure dimensionality reduction based on the calculated threshold, a specific spatial segmentation code is defined within the association graph. Regular expression matching is used to retrieve consecutive node sequences with completely identical business attributes (e.g., all event names are related to "secondary structure casting" and sorted by timestamp), where the time difference between the first and last events of adjacent nodes is strictly less than the calculated threshold of 1120 seconds. For the isomorphic subgraph exhibiting highly dense characteristics, a node folding algorithm is triggered at the graph database level to delete the entire subgraph and generate a composite node in place to replace it. The new composite node records the timestamps of the earliest to the latest occurrence nodes in the series as the start and end times, and synthesizes a unique identifier object header named after the main process and suffixed with _Complex.
[0087] Based on the Tarjan connectivity test algorithm, isolated branch structures that do not reach the two boundary endpoints of the graph and are in a fault-free state are removed. Only the directed path that successfully connects all key composite nodes and the remaining folded regular nodes and has the largest sum of modified edge weights is saved as the longest construction main chain for fingerprint evaluation in the next stage.
[0088] In an optional embodiment, the step of calculating time closure, spatial location difference, resource overlap, and process violation based on the longest construction main chain and merging them to generate a difference fingerprint includes:
[0089] Extract the total actual operating time of the longest construction main chain, divide the target total cycle time of the process by the total actual operating time to obtain the quotient, and use it as the time closure degree;
[0090] Obtain the positioning coordinate vector of the spatial node in the longest construction main chain, and calculate the absolute distance deviation on the plane between the positioning coordinate vector and the center placement vector of the construction specification drawing, as the spatial placement difference;
[0091] The total usage of object name resources within the same time period and space is calculated, and the percentage of overlap where the cumulative object items exceed the maximum resource load threshold is extracted as the resource overlap degree.
[0092] Find the sequence of steps in the prescribed process, compare the reverse violation steps of the longest monitored main construction chain with the total number of violations counted by the reverse count, and divide by the total number of sequence nodes to obtain the process violation degree.
[0093] The temporal closure, spatial location difference, resource overlap, and process violation are converted into equal-length feature encoding fields and merged to generate the differential fingerprint.
[0094] For the longest construction main chain, the timestamp of the node in the main chain is read and the timestamp of the initial node is subtracted to obtain the actual total running time. For example, if it actually takes 150 hours, the target total cycle time that the overall task should theoretically achieve is extracted from the local specification library. For example, the rated total cycle time is 120 hours. The rated total cycle time is divided by the actual time to obtain the quotient ratio of the two, which is 0.80. This dimensionless value is assigned as the time closure degree.
[0095] In the spatial dimension, the sequence of center coordinate points mapped by each associated composite spatial node on the main chain is extracted and weighted averaged to calculate a synthetic actual positioning coordinate vector, such as planar coordinates X:500, Y:600. This vector is then subtracted from the pre-stored center placement vector of the BIM component specification drawing, such as X:502, Y:598, and the L2 norm modulus is taken to calculate the geometric absolute distance deviation between the two on the same projection plane, which is approximately 2.82m, and recorded as the spatial positioning difference. Simultaneously, a joint query is performed using the main chain time window and spatial range to sum the cumulative allocation scheduling hours or number of cranes or pump trucks within the same local area and time interval, such as a foundation pit area within a day. If the total number of shifts, such as 12 shifts, exceeds the predetermined maximum safe load threshold for equipment in that area in the engineering management database, such as a safety threshold of 10 shifts, the excess is divided by the load threshold to calculate a percentage, i.e., 20% overlap, which is extracted as the resource overlap.
[0096] A directed graph topology sequence comparison algorithm is used to compare the timing of the main chain nodes with a pre-defined sequence of execution steps, such as the standard process from A to B to C. Reverse violations, such as B occurring before A, are detected. The algorithm also counts the number of consecutive steps that are completely opposite to the standard process. These two types of violations are mutually exclusive and not counted repeatedly. If the total number of anomalies is 3, it is divided by the total number of sequence nodes in the longest main chain being analyzed (e.g., 30 nodes), and the corresponding ratio of 0.10 is calculated as the process violation degree. The reverse counting refers to the number of consecutive steps in the construction process that are completely opposite to the standard process. The reverse violation steps refer to the number of violations in the construction process where the standard order is violated, or where a single or adjacent step is reversed.
[0097] The four calculated indicators were formatted and aligned, uniformly scaled, and mapped into equal-width 8-bit hexadecimal feature code fields. For example, 0.80 was mapped to CC, 2.82 to 2D, 0.20 to 33, and 0.10 to 19. These fields were then concatenated into a continuous data string such as CC2D3319. This string was then mapped to a numerical dimension for every two bits, generating a fixed-dimensional difference fingerprint feature vector. This string is the unique difference fingerprint representing the deviation of the current process from the multidimensional benchmark, and is written to the repository for use as a judgment benchmark for conflict source backtracking. Based on the full construction data sample of this test, the summary statistics of the four core indicators corresponding to the multidimensional deviation difference fingerprint are as follows: Figure 3 As shown.
[0098] The experiment used real data from multi-source construction sensors generated over three months in a large-scale building transportation hub project, covering location stream video analysis events and equipment operation logs. The control group, employing basic fixed spatiotemporal mesh matching and static graph analysis, achieved a data alignment accuracy of 72.4% and a conflict / anomaly alarm accuracy of 68.2%, with an average daily data processing time of 420 seconds. Ablation group one, utilizing only time window width and mesh granularity adjustments, improved the data alignment accuracy to 86.5%. Ablation group two, adding node topology stability parameters and edge weight correction mechanisms, further improved the key process identification rate from 75.6% to 89.3%. The complete solution group, including threshold path compression and composite node folding, achieved a data alignment accuracy of 88.1%, a key process identification rate of 94.7%, and a conflict detection accuracy of 91.5%, while reducing the daily data processing time to 315 seconds. The optimization of daily computing power and time consumption through the use of composite node time scale threshold strategies resulted in improved monitoring results. Figure 4 As shown, the topology stability correction coefficient, constructed based on connection density and weight variance, enhances the weight of core edges in real business flows, avoiding interference from abnormal edges on the relational graph structure. By implementing compound folding and isolated branch removal on duplicate adjacent nodes according to thresholds, not only is the longest construction main chain purified to generate multidimensional difference fingerprints for anomaly tracing, but the dimensionality reduction of the graph structure also reduces computational power consumption and processing latency.
[0099] In this application, a multi-source construction data analysis and storage device includes the following modules:
[0100] The generation module is used to collect multi-source construction data, encode the data based on time reference and grid division to obtain a spatiotemporal segment sequence, calculate the reliability of each data source based on arrival jitter value, spatial drift value and event repetition rate in adjacent time windows, and adjust the time window width and grid granularity to align the original data based on the comprehensive value of the reliability of each data source, and perform interval segmentation to generate an aligned spatiotemporal segment sequence.
[0101] The correction module is used to construct an association graph containing time nodes, spatial nodes and event nodes based on the aligned spatiotemporal segment sequence. In the association graph, temporal edges, spatial edges and entity edges between each node are established. The initial edge weights are set according to the distribution of the number of event conflicts and the number of stable events in node interactions. The node topology stability parameters are calculated by statistically analyzing the density of node divergent connections. The edge weight correction coefficient is calculated in combination with the node topology stability parameters to correct the initial edge weights and obtain the corrected edge weights.
[0102] The reorganization module is used to perform path compression processing on the association graph. It folds adjacent event nodes within the same process and with a time interval of less than a threshold into composite nodes, removes unconnected isolated branches, and extracts the longest construction main chain with key composite nodes. The threshold is determined by the modified edge weight statistics and node topology stability parameters. Based on the longest construction main chain, it calculates time closure, spatial location difference, resource overlap, and process violation and merges them to generate difference fingerprints. When the difference fingerprint deviates from the constraint template, it triggers conflict source backtracking. Based on the backtracking results, it locally reorganizes the association graph to obtain hierarchical storage blocks, generates tags, and writes them to the repository.
[0103] In an optional embodiment, the step of encoding the data based on a time base and grid partitioning to obtain a spatiotemporal segment sequence includes:
[0104] Subtract the starting zero point corresponding to the unified time base from the timestamp of the multi-source construction data, divide by the preset time segment duration and round down to obtain the time segment code;
[0105] The physical coordinates of the multi-source construction data are obtained. The longitudinal distance offset of the physical coordinates relative to the longitude starting point of the construction area is calculated and divided by a preset longitudinal grid unit length and rounded down. The latitudinal distance offset of the physical coordinates relative to the latitude starting point is calculated and divided by a preset latitudinal grid unit length and rounded down. The combined calculation results are used to obtain the spatial segmentation code.
[0106] The time segment code, spatial segment code, and original input data text structure items are concatenated to form a standard combined string, generating the spatiotemporal segment sequence.
[0107] In an optional embodiment, the step of calculating the reliability of each data source based on arrival jitter, spatial drift, and event repetition rate within adjacent time windows, adjusting the time window width and grid granularity to align the original data based on the comprehensive reliability values of each data source, and performing interval segmentation to generate an aligned spatiotemporal segment sequence includes:
[0108] Calculate the variance of the time interval between two consecutive data arrivals in the data source, and use the normalized value as the arrival jitter value;
[0109] Calculate the absolute value of the spatial distance deviation between the actual location of the current sampling point and the theoretical location predicted based on historical data for the same entity among adjacent sampling points, and use it as the spatial drift value after normalization.
[0110] The number of times that construction task records with identical statistical feature codes are repeatedly uploaded within a set period is normalized and used as the event repetition rate.
[0111] The arrival jitter value, spatial drift value and event repetition rate are each assigned a preset weight and summed by weight. The weighted sum is then subtracted from the constant to obtain the credibility of each data source, which is not lower than a preset lower limit. The mean of the credibility of each data source is then calculated as the comprehensive credibility value of each data source.
[0112] Extract the baseline credibility threshold and baseline time window size issued during the initialization phase. Multiply the ratio of the baseline credibility threshold to the comprehensive credibility value of each data source by the baseline time window size to obtain a new duration value, and replace the time window width.
[0113] Extract the baseline confidence threshold and standard grid side length size issued during the initialization phase. Multiply the arithmetic square root of the ratio of the baseline confidence threshold to the comprehensive confidence value of each data source by the standard grid side length size to obtain a new side length value, and replace the grid granularity.
[0114] The aligned spatiotemporal segment sequence is generated by performing interval segmentation in the full sample data space using the adjusted time window width and grid granularity as truncation boundaries.
[0115] In an optional embodiment, the construction of an association graph containing time nodes, spatial nodes, and event nodes based on the aligned spatiotemporal segment sequence, establishing temporal edges, spatial edges, and entity edges between nodes in the association graph, and setting initial edge weights according to the distribution of the number of event conflicts and stable events in node interactions, includes:
[0116] The aligned spatiotemporal segment sequence is deduplicated and extracted, and the independent timestamps are mapped to time nodes, the two-dimensional coordinate grid numbers are mapped to spatial nodes, and the business operation process types are mapped to event nodes;
[0117] Connecting adjacent time nodes in ascending order of time is defined as temporal edges, and connecting spatial nodes that conform to the Cartesian adjacency association rule is defined as spatial edges.
[0118] Within the same data packet directory identifier, all time nodes, space nodes, and event nodes belonging to the same identifier are connected, and a network structure is built by pairwise interactions, which are set as entity edges to form the association graph.
[0119] For edges containing event nodes, the total number of reverse mutual exclusion error uploads of the actual actions of the physical objects associated with the two nodes is the event conflict count, and the number of correct matching and flow counts is the stable count;
[0120] For temporal and spatial edges, the event conflict count is set to zero and the stability count is set to the total number of interactions for the corresponding edge.
[0121] The basic weight is obtained by calculating the ratio of the stable number of times to the total number of interactions. A decay penalty term is generated using the number of event conflicts. The basic weight is subtracted from the decay penalty term and mapped to a positive value range through an activation function, which is then set as the initial edge weight of the corresponding connection edge.
[0122] In an optional embodiment, the step of calculating the node topology stability parameter by statistically analyzing the density of divergent connections, and then calculating the edge weight correction coefficient based on the node topology stability parameter to correct the initial edge weight and obtain the corrected edge weight includes:
[0123] The total number of edges extending outward from the target node to connect to the entity is used as the connection density.
[0124] Calculate the mean and variance of the initial edge weights of all edges connected to the target node;
[0125] The topological stability parameters of the target node are obtained by multiplying the connection density by the mean of the initial edge weights and dividing by the sum of the variance and the smoothing constant.
[0126] Extract the topological stability parameters of the starting node and the ending node of the initial edge weight to be corrected, calculate the geometric mean of the two, and then linearly scale them according to a preset scaling factor to obtain the edge weight correction coefficient.
[0127] Multiply the initial edge weight to be corrected by the edge weight correction coefficient, and set it as the corrected edge weight after edge update.
[0128] In an optional embodiment, the path compression process on the association graph, which involves folding adjacent event nodes within the same process and with a time interval less than a threshold into composite nodes, and removing disconnected isolated branches to extract the longest construction main chain containing key composite nodes, includes:
[0129] The corrected edge weights of all edges on the candidate path are summed and divided by the total number of nodes passed through to obtain the arithmetic mean, thus yielding the corrected edge weight statistics.
[0130] The corrected edge weight statistics are weighted and summed together with the mean of the topological stability parameters of all nodes to which the path is attached. The result is then multiplied by a preset time conversion coefficient and set as the threshold.
[0131] Within the same specified code grid, find adjacent event node sequences that are completely duplicated, have a unified task flow name, and have a consecutive time interval between their first and last occurrences that is less than the threshold.
[0132] Collapse consecutive adjacent event nodes that meet the conditions into a composite node that includes a start and end time and a unique identifier object header;
[0133] Isolated branches without connected start and end nodes are removed, and the directed path with the largest sum of modified edge weights that retains all key composite nodes and the original connected nodes is retained as the longest construction main chain.
[0134] In an optional embodiment, the step of calculating time closure, spatial location difference, resource overlap, and process violation based on the longest construction main chain and merging them to generate a difference fingerprint includes:
[0135] Extract the total actual operating time of the longest construction main chain, divide the target total cycle time of the process by the total actual operating time to obtain the quotient, and use it as the time closure degree;
[0136] Obtain the positioning coordinate vector of the spatial node in the longest construction main chain, and calculate the absolute distance deviation on the plane between the positioning coordinate vector and the center placement vector of the construction specification drawing, as the spatial placement difference;
[0137] The total usage of object name resources within the same time period and space is calculated, and the percentage of overlap where the cumulative object items exceed the maximum resource load threshold is extracted as the resource overlap degree.
[0138] Find the sequence of steps in the prescribed process, compare the reverse violation steps of the longest monitored main construction chain with the total number of violations counted by the reverse count, and divide by the total number of sequence nodes to obtain the process violation degree.
[0139] The temporal closure, spatial location difference, resource overlap, and process violation are converted into equal-length feature encoding fields and merged to generate the differential fingerprint.
[0140] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0141] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-source construction data analysis and storage method, characterized in that, include: Multi-source construction data is collected, and the data is encoded based on time reference and grid division to obtain a spatiotemporal segment sequence. The reliability of each data source is calculated based on the arrival jitter value, spatial drift value and event repetition rate in adjacent time windows. The time window width and grid granularity are adjusted to align the original data based on the comprehensive value of the reliability of each data source, and interval segmentation is performed to generate an aligned spatiotemporal segment sequence. Based on the alignment of spatiotemporal fragment sequences, an association graph containing time nodes, spatial nodes, and event nodes is constructed. Temporal edges, spatial edges, and entity edges between nodes are established in the association graph. Initial edge weights are set according to the distribution of the number of event conflicts and the number of stable events in node interactions. The node topology stability parameters are calculated by statistically analyzing the density of node divergent connections. The edge weight correction coefficient is calculated in combination with the node topology stability parameters to correct the initial edge weights and obtain the corrected edge weights. The association graph is compressed by folding adjacent event nodes within the same process with a time interval less than a threshold into composite nodes. Unconnected isolated branches are removed to extract the longest construction main chain with key composite nodes. The threshold is determined by the modified edge weight statistics and node topology stability parameters. The time closure, spatial position difference, resource overlap and process violation are calculated based on the longest construction main chain and merged to generate a difference fingerprint. When the difference fingerprint deviates from the constraint template, the conflict source backtracking is triggered. The association graph is locally reorganized according to the backtracking results to obtain a hierarchical storage block and generate a tag to write to the repository. The calculation of time closure, spatial location difference, resource overlap, and process violation based on the longest construction main chain, and the merging of these factors to generate a difference fingerprint, includes: Extract the total actual operating time of the longest construction main chain, divide the target total cycle time of the process by the total actual operating time to obtain the quotient, and use it as the time closure degree; Obtain the positioning coordinate vector of the spatial node in the longest construction main chain, and calculate the absolute distance deviation on the plane between the positioning coordinate vector and the center placement vector of the construction specification drawing, as the spatial placement difference; The total usage of object name resources within the same time period and space is calculated, and the percentage of overlap where the cumulative object items exceed the maximum resource load threshold is extracted as the resource overlap degree. Find the sequence of steps in the prescribed process, compare the reverse violation steps of the longest monitored main construction chain with the total number of violations counted by the reverse count, and divide by the total number of sequence nodes to obtain the process violation degree. The temporal closure, spatial location difference, resource overlap, and process violation are converted into equal-length feature encoding fields and merged to generate the differential fingerprint.
2. The method of claim 1, wherein, The process of encoding data based on a time base and grid partitioning to obtain a spatiotemporal segment sequence includes: Subtract the starting zero point corresponding to the unified time base from the timestamp of the multi-source construction data, divide by the preset time segment duration and round down to obtain the time segment code; The physical coordinates of the multi-source construction data are obtained. The longitudinal distance offset of the physical coordinates relative to the longitude starting point of the construction area is calculated and divided by a preset longitudinal grid unit length and rounded down. The latitudinal distance offset of the physical coordinates relative to the latitude starting point is calculated and divided by a preset latitudinal grid unit length and rounded down. The combined calculation results are used to obtain the spatial segmentation code. The time segment code, spatial segment code, and original input data text structure items are concatenated to form a standard combined string, generating the spatiotemporal segment sequence.
3. The method according to claim 1, characterized in that, The process involves calculating the reliability of each data source based on arrival jitter, spatial drift, and event repetition rate within adjacent time windows, adjusting the time window width and grid granularity to align the original data based on the comprehensive reliability values of each data source, and performing interval segmentation to generate an aligned spatiotemporal segment sequence, including: Calculate the variance of the time interval between two consecutive data arrivals in the data source, and use the normalized value as the arrival jitter value; Calculate the absolute value of the spatial distance deviation between the actual location of the current sampling point and the theoretical location predicted based on historical data for the same entity among adjacent sampling points, and use it as the spatial drift value after normalization. The number of times that construction task records with identical statistical feature codes are repeatedly uploaded within a set period is normalized and used as the event repetition rate. The arrival jitter value, spatial drift value and event repetition rate are each assigned a preset weight and summed by weight. The weighted sum is then subtracted from the constant to obtain the credibility of each data source, which is not lower than a preset lower limit. The mean of the credibility of each data source is then calculated as the comprehensive credibility value of each data source. Extract the baseline credibility threshold and baseline time window size issued during the initialization phase. Multiply the ratio of the baseline credibility threshold to the comprehensive credibility value of each data source by the baseline time window size to obtain a new duration value, and replace the time window width. Extract the baseline confidence threshold and standard grid side length size issued during the initialization phase. Multiply the arithmetic square root of the ratio of the baseline confidence threshold to the comprehensive confidence value of each data source by the standard grid side length size to obtain a new side length value, and replace the grid granularity. The aligned spatiotemporal segment sequence is generated by performing interval segmentation in the full sample data space using the adjusted time window width and grid granularity as truncation boundaries.
4. The method according to claim 3, characterized in that, The method involves constructing an association graph containing time nodes, spatial nodes, and event nodes based on aligned spatiotemporal segment sequences. Temporal edges, spatial edges, and entity edges are established between nodes in the association graph. Initial edge weights are set according to the distribution of event conflict and stability frequency in node interactions, including: The aligned spatiotemporal segment sequence is deduplicated and extracted, and the independent timestamps are mapped to time nodes, the two-dimensional coordinate grid numbers are mapped to spatial nodes, and the business operation process types are mapped to event nodes; Connecting adjacent time nodes in ascending order of time is set as temporal edges, and connecting spatial nodes that conform to the Cartesian adjacency association rule is set as spatial edges. Within the same data packet directory identifier, all time nodes, space nodes, and event nodes belonging to the same identifier are connected, and a network structure is built by pairwise interactions, which are set as entity edges to form the association graph. For edges containing event nodes, the total number of reverse mutual exclusion error uploads of the actual actions of the physical objects associated with the two nodes is the event conflict count, and the number of correct matching and flow counts is the stable count; For temporal and spatial edges, the event conflict count is set to zero and the stability count is set to the total number of interactions for the corresponding edge. The basic weight is obtained by calculating the ratio of the stable number of times to the total number of interactions. A decay penalty term is generated using the number of event conflicts. The basic weight is subtracted from the decay penalty term and mapped to a positive value range through an activation function, which is then set as the initial edge weight of the corresponding connection edge.
5. The method according to claim 1, characterized in that, The process of calculating node topology stability parameters by statistically analyzing the density of divergent connections, and then using these parameters to calculate edge weight correction coefficients to correct the initial edge weights and obtain corrected edge weights includes: The total number of edges extending outward from the target node to connect to the entity is used as the connection density. Calculate the mean and variance of the initial edge weights of all edges connected to the target node; The topological stability parameters of the target node are obtained by multiplying the connection density by the mean of the initial edge weights and dividing by the sum of the variance and the smoothing constant. Extract the topological stability parameters of the starting node and the ending node of the initial edge weight to be corrected, calculate the geometric mean of the two, and then linearly scale them according to a preset scaling factor to obtain the edge weight correction coefficient. Multiply the initial edge weight to be corrected by the edge weight correction coefficient, and set it as the corrected edge weight after edge update.
6. The method according to claim 1, characterized in that, The path compression process of the association graph, which folds adjacent event nodes within the same process with a time interval less than a threshold into composite nodes, and removes unconnected isolated branches to extract the longest construction main chain containing key composite nodes, includes: The corrected edge weights of all edges on the candidate path are summed and divided by the total number of nodes passed through to obtain the arithmetic mean, thus yielding the corrected edge weight statistics. The corrected edge weight statistics are weighted and summed together with the mean of the topological stability parameters of all nodes to which the path is attached. The result is then multiplied by a preset time conversion coefficient and set as the threshold. Within the same specified code grid, find adjacent event node sequences that are completely duplicated, have a unified task flow name, and have a consecutive time interval between their first and last occurrences that is less than the threshold. Collapse consecutive adjacent event nodes that meet the conditions into a composite node that includes a start and end time and a unique identifier object header; Isolated branches without connected start and end nodes are removed, and the directed path with the largest sum of modified edge weights that retains all key composite nodes and the original connected nodes is retained as the longest construction main chain.
7. A multi-source construction data analysis and storage device, characterized in that, Includes the following modules: The generation module is used to collect multi-source construction data, encode the data based on time reference and grid division to obtain a spatiotemporal segment sequence, calculate the reliability of each data source based on arrival jitter value, spatial drift value and event repetition rate in adjacent time windows, and adjust the time window width and grid granularity to align the original data based on the comprehensive value of the reliability of each data source, and perform interval segmentation to generate an aligned spatiotemporal segment sequence. The correction module is used to construct an association graph containing time nodes, spatial nodes and event nodes based on the aligned spatiotemporal segment sequence. In the association graph, temporal edges, spatial edges and entity edges between each node are established. The initial edge weights are set according to the distribution of the number of event conflicts and the number of stable events in node interactions. The node topology stability parameters are calculated by statistically analyzing the density of node divergent connections. The edge weight correction coefficient is calculated in combination with the node topology stability parameters to correct the initial edge weights and obtain the corrected edge weights. The reorganization module is used to perform path compression processing on the association graph. It folds adjacent event nodes within the same process and with a time interval less than a threshold into composite nodes, removes unconnected isolated branches, and extracts the longest construction main chain with key composite nodes. The threshold is determined by the modified edge weight statistics and node topology stability parameters. Based on the longest construction main chain, it calculates time closure, spatial position difference, resource overlap, and process violation and merges them to generate difference fingerprints. When the difference fingerprint deviates from the constraint template, it triggers conflict source backtracking. Based on the backtracking results, it locally reorganizes the association graph to obtain hierarchical storage blocks and generates tags to write to the repository. The calculation of time closure, spatial location difference, resource overlap, and process violation based on the longest construction main chain, and the merging of these factors to generate a difference fingerprint, includes: Extract the total actual operating time of the longest construction main chain, divide the target total cycle time of the process by the total actual operating time to obtain the quotient, and use it as the time closure degree; Obtain the positioning coordinate vector of the spatial node in the longest construction main chain, and calculate the absolute distance deviation on the plane between the positioning coordinate vector and the center placement vector of the construction specification drawing, as the spatial placement difference; The total usage of object name resources within the same time period and space is calculated, and the percentage of overlap where the cumulative object items exceed the maximum resource load threshold is extracted as the resource overlap degree. Find the sequence of steps in the prescribed process, compare the reverse violation steps of the longest monitored main construction chain with the total number of violations counted by the reverse count, and divide by the total number of sequence nodes to obtain the process violation degree. The temporal closure, spatial location difference, resource overlap, and process violation are converted into equal-length feature encoding fields and merged to generate the differential fingerprint.
8. The apparatus according to claim 7, characterized in that, The process of encoding data based on a time base and grid partitioning to obtain a spatiotemporal segment sequence includes: Subtract the starting zero point corresponding to the unified time base from the timestamp of the multi-source construction data, divide by the preset time segment duration and round down to obtain the time segment code; The physical coordinates of the multi-source construction data are obtained. The longitudinal distance offset of the physical coordinates relative to the longitude starting point of the construction area is calculated and divided by a preset longitudinal grid unit length and rounded down. The latitudinal distance offset of the physical coordinates relative to the latitude starting point is calculated and divided by a preset latitudinal grid unit length and rounded down. The combined calculation results are used to obtain the spatial segmentation code. The time segment code, spatial segment code, and original input data text structure items are concatenated to form a standard combined string, generating the spatiotemporal segment sequence.
9. The apparatus according to claim 7, characterized in that, The process involves calculating the reliability of each data source based on arrival jitter, spatial drift, and event repetition rate within adjacent time windows, adjusting the time window width and grid granularity to align the original data based on the comprehensive reliability values of each data source, and performing interval segmentation to generate an aligned spatiotemporal segment sequence, including: Calculate the variance of the time interval between two consecutive data arrivals in the data source, and use the normalized value as the arrival jitter value; Calculate the absolute value of the spatial distance deviation between the actual location of the current sampling point and the theoretical location predicted based on historical data for the same entity among adjacent sampling points, and use it as the spatial drift value after normalization. The number of times that construction task records with identical statistical feature codes are repeatedly uploaded within a set period is normalized and used as the event repetition rate. The arrival jitter value, spatial drift value and event repetition rate are each assigned a preset weight and summed by weight. The weighted sum is then subtracted from the constant to obtain the credibility of each data source, which is not lower than a preset lower limit. The mean of the credibility of each data source is then calculated as the comprehensive credibility value of each data source. Extract the baseline credibility threshold and baseline time window size issued during the initialization phase. Multiply the ratio of the baseline credibility threshold to the comprehensive credibility value of each data source by the baseline time window size to obtain a new duration value, and replace the time window width. Extract the baseline confidence threshold and standard grid side length size issued during the initialization phase. Multiply the arithmetic square root of the ratio of the baseline confidence threshold to the comprehensive confidence value of each data source by the standard grid side length size to obtain a new side length value, and replace the grid granularity. The aligned spatiotemporal segment sequence is generated by performing interval segmentation in the full sample data space using the adjusted time window width and grid granularity as truncation boundaries.
Citation Information
Patent Citations
Constructional engineering risk assessment method and system for multi-source anomaly monitoring
CN120525331A
Intelligent logistics log anomaly detection method based on block chain and space-time compression traceability graph
CN121907548A