Engineering data management method and device, equipment and storage medium

By introducing a quality verification layer based on knowledge graphs and a strategy-driven mechanism, the problem of data and entity disconnect in engineering data management is solved. Semantic encapsulation of data and automatic correction of abnormal data are achieved, improving the accuracy and response speed of data processing and providing complete data support for engineering projects.

CN121563291AInactive Publication Date: 2026-02-24SHENZHEN PIN HIGH-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511670101.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing engineering data management methods are ill-suited to the complex and rapidly changing conditions of engineering sites, leading to a disconnect between data and physical entities, delayed anomaly feedback, and difficulties in historical tracing, which in turn affects the accuracy and timeliness of data-driven decision-making.

Method used

By introducing a knowledge graph-based quality verification layer and a strategy-driven anomaly data correction mechanism, the system realizes the unit construction of engineering entity identification and monitoring data, entity information detection, anomaly data marking and correction, and generates a full-cycle data view of the project.

Benefits of technology

It achieves semantic encapsulation and dynamic association of engineering entity data, improves the accuracy of data processing and the timeliness of anomaly response, provides a data foundation with complete historical traceability for project decision-making, and supports pre-event early warning, in-event control and post-event analysis of engineering management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121563291A_ABST
    Figure CN121563291A_ABST
Patent Text Reader

Abstract

The invention relates to an engineering data management method, device and equipment and a storage medium, and the method comprises the steps: collecting an engineering entity identifier and monitoring data of an engineering site for unit construction, and forming a basic data unit; performing entity information detection on the component identifier in the basic data unit and a pre-constructed engineering knowledge graph, and when it is detected that the component identifier is not matched with information of the engineering knowledge graph, extracting entity information corresponding to the component identifier, and performing engineering node construction and relation updating on the engineering knowledge graph according to the entity information, generating an updated knowledge graph, performing detection again until matching succeeds, and generating an engineering data unit. According to the method, the data unit is verified in real time through the quality rule based on the knowledge graph, and the abnormal data is automatically corrected by adopting a strategy polling mechanism, so that the accuracy of data processing and the timeliness of abnormal response are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of engineering data management, and in particular to an engineering data management method, apparatus, equipment, and storage medium. Background Technology

[0002] Engineering data management, as a core component of modern large-scale engineering projects, plays a crucial role in ensuring construction quality, controlling project costs, and improving operation and maintenance efficiency. With the deepening of smart construction sites and digital construction, how to achieve automated collection, semantic integration, and full lifecycle tracking of multi-source, heterogeneous monitoring data from engineering sites has become a key challenge in this field. Existing data management methods focus on unidirectional data collection and storage, or lack a deep understanding of the dynamic changes of engineering entities and the inherent semantic relationships of data during the data processing stage. This relatively static and isolated processing model is difficult to adapt to the complex and rapidly changing realities of engineering site entities, leading to problems such as data disconnect from entities, delayed anomaly feedback, and difficulties in historical tracing, thereby affecting the accuracy and timeliness of data-driven decision-making. Summary of the Invention

[0003] The main objective of this invention is to provide a collaborative control method and system for air source heat pump aggregation prediction and peak shaving task instructions. By introducing a quality verification layer based on knowledge graphs and a strategy-driven abnormal data correction mechanism, the traditional data inspection is upgraded from a single value domain judgment to consistency verification in the engineering semantic context.

[0004] To achieve the above objectives, the present invention provides an engineering data management method, comprising: Collect engineering entity identification and monitoring data from the engineering site to construct basic data units; The component identifier in the basic data unit is compared with the pre-built engineering knowledge graph for entity information detection. When the information of the component identifier in the engineering knowledge graph does not match, the entity information corresponding to the component identifier is extracted. Based on the entity information, the engineering knowledge graph is used to construct engineering nodes and update relationships. An updated knowledge graph is generated and the detection is repeated until a match is successful, and an engineering data unit is generated. The engineering data unit is subjected to quality verification according to the preset engineering quality rules, and the engineering data unit that fails the quality verification is marked as an abnormal data unit. The abnormal data units are polled and corrected according to the correction strategy table, and the corrected data units are periodically constructed with the updated knowledge graph to output a full-cycle data view.

[0005] Furthermore, the engineering entity identification and monitoring data collected at the engineering site are used to construct basic data units, including: Receive multi-source data streams from the engineering site, identify and parse the multi-source data streams to obtain the engineering entity identifier and the monitoring data; Each of the engineering entity identifiers is associated and bound with at least one of the monitoring data to form a binding combination; The binding combination is subjected to basic verification according to the predefined identifier specification. When the engineering entity identifier in the binding combination passes the format verification and the monitoring data passes the threshold pre-detection, a time stamp with a unified format is added to the binding combination to generate the basic data unit.

[0006] Further, the step of performing entity information detection between the component identifier in the basic data unit and the pre-constructed engineering knowledge graph, and when a mismatch is detected between the component identifier and the information in the engineering knowledge graph, extracting the entity information corresponding to the component identifier, includes: Based on the analysis of the engineering entity identifier, the component identifier in the basic data unit is extracted. All entity information nodes in the engineering knowledge graph are traversed. The component identifier is compared with each entity information node in string form to obtain the comparison information. When all the comparison information does not match, the node mismatch result is determined and output; In response to the node mismatch result, the data source type, collection timestamp, monitoring parameter sequence and data collection location information associated with the component identifier are located and extracted from the basic data unit; The extracted data source type, the collection timestamp, the monitoring parameter sequence, and the data collection location information are combined to form the entity information.

[0007] Furthermore, the process of constructing engineering nodes and updating relationships in the engineering knowledge graph based on the entity information, generating an updated knowledge graph, and re-detecting it until a match is successful, thereby generating engineering data units, includes: Based on the data source type and data collection location information in the entity information, node attributes are constructed to generate a set of attribute fields. A new engineering node is created in the engineering knowledge graph, the attribute field set is written into the attribute list of the new engineering node, and the relationship edge of the new engineering node in the engineering knowledge graph is updated according to the monitoring parameter sequence to generate the updated knowledge graph; The component identifier in the basic data unit is re-detected with the updated knowledge graph. If the match is successful, the attribute field set is associated and attached to the basic data unit to generate the engineering data unit. If a match fails, repeat the node construction and relationship update steps described above until a match is found.

[0008] Further, the step of creating a corresponding new project node in the project knowledge graph, writing the attribute field set into the attribute list of the new project node, and updating the relationship edges of the new project node in the project knowledge graph according to the monitoring parameter sequence to generate the updated knowledge graph includes: Extract the node identifier and the field name and field value of each attribute field from the set of attribute fields; Based on the node identifier, a new project node is created in the project knowledge graph, and the field name of each attribute field is matched with the node attribute template of the project knowledge graph. When the field name matches successfully, the field value is written to the corresponding attribute item in the attribute list of the new project node; When the field name fails to match, the field name and the field value are combined into a new attribute item and added to the attribute list; After completing the matching of all the attribute fields, the relationship edge between the new project node and the project knowledge graph is located based on the monitoring parameter sequence, and the association attribute of the relationship edge is set. After the relationship is updated, the updated knowledge graph is generated.

[0009] Further, the step of performing quality verification on the engineering data unit according to preset engineering quality rules, and marking the engineering data unit that fails the quality verification as an abnormal data unit, includes: Extract the specification indicators to be verified from the engineering data unit, and select a rule from the preset engineering quality rule set as the current verification rule in sequence; The specification indicators are compared with the specification requirements in the current verification rules. When the specification indicators do not meet the specification requirements, the rule deviation information is recorded and the current verification is determined to have failed. Based on the rule deviation information and the preset anomaly type mapping table, type matching is performed to obtain the anomaly type code; Based on the anomaly type code, the corresponding quality status identifier is extracted from the engineering quality rule set, and an anomaly identification timestamp is generated; The anomaly type code is associated with the corresponding engineering data unit, and the anomaly identification timestamp and the quality status identifier are added to form the anomaly data unit.

[0010] Furthermore, the step of performing policy polling and data correction on the abnormal data units according to the correction strategy table, and periodically constructing the obtained corrected data units with the updated knowledge graph to output a full-cycle data view includes: The abnormal data unit is matched sequentially with each correction strategy in the correction strategy table. When the abnormal data unit meets the triggering condition of any of the correction strategies, the correction strategy is marked as the current execution strategy. Based on the current execution strategy, abnormal data correction is performed on the abnormal data unit to generate an intermediate corrected data unit; The intermediate correction data unit is mapped to the updated knowledge graph and the nodes are verified. If the verification is successful, the intermediate correction data unit is marked as a correction data unit. The corrected data unit is integrated with the data unit node in the updated knowledge graph in a time sequence to construct the full-cycle data view of the project.

[0011] The present invention also provides an engineering data management device, applied to the engineering data management method described in any one of the above claims, comprising: The data acquisition module is used to collect engineering entity identification and monitoring data from the engineering site and construct them into basic data units. The analysis module is used to perform entity information detection between the component identifier in the basic data unit and the pre-built engineering knowledge graph. When the information of the component identifier in the engineering knowledge graph does not match, the entity information corresponding to the component identifier is extracted. Based on the entity information, the engineering knowledge graph is used to construct engineering nodes and update relationships. An updated knowledge graph is generated and the detection is performed again until a match is successful, and an engineering data unit is generated. The association module is used to perform quality verification on the engineering data unit according to the preset engineering quality rules, and mark the engineering data unit that fails the quality verification as an abnormal data unit. The processing module is used to perform policy polling and data correction on the abnormal data units according to the correction strategy table, and periodically construct the obtained corrected data units with the updated knowledge graph to output a full-process cycle data view.

[0012] The present invention also provides an engineering data management device, comprising: Memory, used to store programs; A processor is used to execute the program to implement the various steps of the engineering data management method described in any of the above-mentioned embodiments.

[0013] The present invention also provides a storage medium storing computer instructions for causing a computer to perform any of the methods described above.

[0014] The present invention provides an engineering data management method, apparatus, device, and storage medium, which have the following beneficial effects: By constructing integrated units for on-site entity identification and monitoring data, and dynamically matching and expanding the engineering knowledge graph for each unit, semantic encapsulation and dynamic association of engineering entity data are achieved, effectively solving the problem of data and entity disconnect. Real-time verification of data units is performed using quality rules based on the knowledge graph, and an automated correction of abnormal data is achieved through a strategy polling mechanism, significantly improving the accuracy of data processing and the timeliness of anomaly response. By periodically associating the corrected data units with the knowledge graph, a complete and reliable full-cycle data view of the entire project is generated, providing a data foundation with complete historical traceability for project decision-making, supporting pre-event warning, in-event control, and post-event analysis in project management. Attached Figure Description

[0015] Figure 1 This is a flowchart of an engineering data management method provided by the present invention; Figure 2 This is a structural diagram of an engineering data management device provided for the present invention; Figure 3 This is a structural diagram of an engineering data management device provided for the present invention.

[0016] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0018] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.

[0019] Reference Figure 1 This invention provides an engineering data management method, comprising: Step S10: Collect engineering entity identification and monitoring data from the engineering site to construct basic data units; Step S20: Perform entity information detection between the component identifier in the basic data unit and the pre-built engineering knowledge graph. When a mismatch is detected between the component identifier and the engineering knowledge graph, extract the entity information corresponding to the component identifier. Based on the entity information, construct engineering nodes and update relationships in the engineering knowledge graph. Generate an updated knowledge graph and re-detect until a match is successful, and generate an engineering data unit. Step S30: Perform quality verification on the engineering data units according to the preset engineering quality rules, and mark the engineering data units that fail the quality verification as abnormal data units; Step S40: Perform policy polling and data correction on abnormal data units according to the correction strategy table, and periodically construct the obtained corrected data units with the updated knowledge graph to output a full-cycle data view of the entire project.

[0020] Based on the steps described above, the detailed process is as follows: Step S10: This involves receiving and parsing data streams generated by various data acquisition terminals on-site (such as sensors, RFID readers, cameras, or manual data entry terminals). These data streams follow different communication protocols and data formats. The parsing operation involves unpacking and decoding them to identify two types of key information: First, there is the unique engineering entity identifier (such as a specific component, equipment, or structural part). Second, there are the state parameters of that entity measured or recorded at a specific moment, i.e., monitoring data. After parsing, the data enters the association and binding stage. Based on the inherent temporal proximity, spatial homogeneity, or logical correlation in the data, an engineering entity identifier is combined with one or more monitoring data generated within the same time period to form a binding combination.

[0021] The binding combination is then subjected to basic verification, which is based on predefined identifier specifications (such as specific encoding rules, length, and character sets) and threshold pre-detection ranges of monitoring data (such as reasonable value ranges for physical quantities). Only binding combinations with correct identifier formats and monitoring data that pass the pre-detection are considered valid, and a unified and accurate timestamp is attached to them, encapsulating them into basic data units with a standard structure that include three elements: entity identifier, monitoring value, and timestamp.

[0022] Step S20: Key component identifiers are extracted from the basic data units. Then, all existing entity information nodes in the pre-built engineering knowledge graph are traversed, and the extracted component identifiers are precisely compared with the identifier information of each node. This comparison process aims to confirm whether the current entity is a known entity in the knowledge graph.

[0023] If a node that matches exactly is found, it means that the entity has been defined by the knowledge graph, and its semantic information, such as attributes and relationships, is known. This predefined semantic information is then extracted from the successfully matched node and associated as additional attributes with the original basic data unit, thereby generating a semantically rich engineering data unit.

[0024] Conversely, if no matching node is found after the traversal, the component identifier is determined to correspond to an unknown new entity. At this point, detailed information related to this new entity is further extracted from the basic data units, such as data source type, collection location, and monitoring parameter characteristics. Using this information, a dynamic expansion operation is performed in the engineering knowledge graph: an engineering node representing this new entity is created, initial attributes are set for it, and the relationship edges in the knowledge graph are updated based on its potential relationships with other known entities (such as spatial relationships and functional affiliations, which can be inferred from the collection location and parameter characteristics or established according to domain rules), thereby generating an updated knowledge graph.

[0025] Based on the updated knowledge graph, the original basic data units are re-matched. This cycle continues until a match is successful, ultimately generating engineering data units. This process ensures that regardless of whether the entity is known or not, it can ultimately be semantically integrated into the knowledge system.

[0026] Step S30: The pre-defined set of engineering quality rules constitutes the verification criteria. These rules go beyond simple numerical range checks and cover multiple dimensions such as data logical consistency, physical law compliance, and business rule compliance.

[0027] Specific specifications that need to be verified are extracted from engineering data units, such as the monitoring values ​​themselves, the trends of value changes, or the calculation relationships between multiple related parameters. Rules are selected from the engineering quality rule set in a predetermined order, and the extracted specifications are compared with the standard requirements defined by the current rule. This comparison is not a simple equality or range judgment, but involves calculations, reasoning based on rule logic, or cross-validation with related node information in a knowledge graph.

[0028] When a deviation from the specified requirements is detected, the specific information of the deviation is recorded, including the type of deviation, the degree of deviation, and the specific rule entry violated. Based on the recorded deviation information, the standardized classification code corresponding to this anomaly is determined by referring to a predefined anomaly type mapping table.

[0029] Generate a tag block containing information such as exception type code, exception trigger timestamp, and severity indicator, and associate this tag block with this project data unit to formally identify it as an exception data unit to be processed.

[0030] Step S40: The correction strategy table contains a series of correction strategies organized by priority or applicable conditions, such as data smoothing, interpolation completion, inference correction based on associated sensor data, or replacement rules for specific anomaly patterns. The characteristics of the anomalous data unit (such as anomaly type encoding and data source characteristics) are sequentially matched with the triggering conditions of each correction strategy. When a matching correction strategy is found, it is used as the current execution strategy, and the anomalous portion of the anomalous data unit is corrected according to the specific calculation or processing logic defined by that strategy, generating an intermediate correction result.

[0031] Intermediate correction results are not immediately effective; they require verification through a validation mechanism. Validation involves re-checking the consistency of the corrected data with the relevant rules and associated data in the updated engineering knowledge graph to ensure the correction results are semantically sound in engineering terms. If validation passes, the intermediate result is confirmed as the final corrected data unit; if validation fails, the next applicable strategy in the strategy table is activated, and the correction and validation cycle is repeated until success or all strategies have been tried.

[0032] For successful corrective data units, they are associated with the corresponding engineering entity nodes, and integrated into the entity's temporal state sequence in the knowledge graph based on their timestamp information. By systematically organizing and linking all normal and corrective data units along the timeline, a complete, coherent, and logically consistent full-cycle data view is finally constructed. This view enables complete tracing and visualization of the historical state evolution of engineering entities.

[0033] This invention provides an engineering data management method that integrates on-site entity identification and monitoring data into a single unit. Each unit is dynamically matched with and expanded using an engineering knowledge graph, achieving semantic encapsulation and dynamic association of engineering entity data, effectively solving the problem of data-entity disconnect. Real-time verification of data units is performed using knowledge graph-based quality rules, and an automated correction mechanism for abnormal data is employed through a policy polling mechanism, significantly improving the accuracy of data processing and the timeliness of anomaly response. By periodically associating the corrected data units with the knowledge graph, a complete and reliable full-cycle data view is generated, providing a data foundation with complete historical traceability for project decision-making, supporting pre-event warning, in-event control, and post-event analysis in engineering management.

[0034] In one embodiment, the identification and monitoring data of engineering entities collected at the engineering site are used to construct basic data units, including: Establish communication connections with various data acquisition terminals at the engineering site. These terminals are highly heterogeneous, including embedded sensors, RFID readers, image acquisition devices, and manual data entry terminals. They continuously or intermittently generate raw data streams and are configured with various communication interface adapters to support different industrial protocols, such as Modbus, OPC UA, and MQTT, in order to receive these raw, unprocessed byte streams or data packets.

[0035] After the data stream is received, it enters the identification and parsing stage. The core task of this stage is to accurately separate semantic information with clear engineering significance from the messy raw data. According to predefined or dynamically negotiated data format specifications, the data stream is deframed, decrypted, or deserialized, and converted into a structured or semi-structured intermediate data representation.

[0036] Based on pre-registered device metadata templates or identification information embedded in the data stream, intermediate data is subjected to content identification. The identification process focuses on two key data elements: engineering entity identifiers and monitoring data. Engineering entity identifiers are codes that uniquely represent a physical or functional entity. They are represented as a string of characters with specific encoding rules, such as RFID tag codes conforming to ISO standards, or component serial numbers conforming to project coding specifications.

[0037] Monitoring data consists of quantitative or qualitative values ​​describing specific attributes or states of the entity, such as stress values, temperature readings, vibration frequencies, or equipment on / off status. The parsing engine uses pattern matching, keyword extraction, or rule-based content analysis techniques to locate and extract these two types of information from intermediate data. The extracted engineering entity identifiers and monitoring data are temporarily stored separately and tagged with metadata such as data source, collection time point (if included in the original data), and data confidence level, laying the foundation for subsequent association and binding.

[0038] The system correctly pairs identification information representing "what it is" with monitoring information representing "what its status is," restoring their inherent connections in real-world engineering scenarios. The logic of this association and binding is not a simple random combination, but rather based on the spatiotemporal and logical correlations inherent in the data. Temporal correlation is the primary binding criterion. The system maintains a configurable time window, treating engineering entity identifiers and monitoring data belonging to the same data source and whose timestamps fall within this window as potentially belonging to the same monitoring event, thus establishing a correlation. For example, identifiers and readings reported from the same sensor node within a millisecond time difference will be preferentially bound.

[0039] Spatial or logical correlation is another important criterion. When the data itself does not contain precise timestamps, or when cross-data source binding is required, spatial location information (such as GPS coordinates, area codes) or logical hierarchical relationships (such as device affiliation, functional grouping) are relied upon for matching. For example, monitoring data reported by all sensors deployed within a specific structural area can be bound to engineering entity identifiers representing that area. During implementation, the system traverses and performs matching calculations on the data set output from the parsing phase according to predefined association rules. Successfully matched engineering entity identifiers and at least one monitoring data point are combined into a temporary data set, i.e., a binding combination. Each binding combination inherently represents one or more state observations of a specific engineering entity within a specific context (time, space). This binding combination constitutes a data entity with preliminary engineering semantics, but its completeness and accuracy still need to be verified.

[0040] The system uses a predefined identifier specification library and a monitoring parameter threshold library as verification standards. The verification process first performs format verification on the engineering entity identifiers within the bound combination. The identifier specification library defines the structured rules that various types of entity identifiers must follow, such as encoding length, character set limitations, checksum algorithms, and prefix / suffix specifications.

[0041] The identifier to be validated is matched against all standard templates in the library. Only identifiers that fully conform to a valid template rule are considered to have passed format validation. Garbled or invalid identifiers caused by transmission errors, device malfunctions, or unauthorized access are filtered out. After an identifier passes validation, the validation focus shifts to the monitoring data within the bound combination.

[0042] Threshold pre-detection does not involve complex business logic judgments, but rather performs basic rationality checks based on the inherent physical or engineering limits of the monitored parameters themselves. For example, concrete temperature monitoring values ​​should not be negative or far exceed the hydration heat limit of cement.

[0043] Based on the type of monitoring data, the corresponding reasonable value range is extracted from the threshold database for comparison. All monitoring data within the preset reasonable value range passes the threshold pre-check; if any monitoring data exceeds the reasonable range, the entire binding combination is temporarily suspended or marked as suspicious.

[0044] Only binding combinations that pass both of the above checks are considered valid data and enter the final encapsulation stage. A unified and precise time stamp is added to these valid binding combinations. This time stamp uses the internationally standardized Coordinated Universal Time (UTC) format, and the priority of the timestamp source is clearly defined. For example, the time signal from the built-in clock of the data acquisition terminal is given priority; if it is unavailable or unreliable, the system time of the data receiving server is used.

[0045] After adding timestamps, the previously loosely bound combinations are formally encapsulated into a basic data unit with a standard structure. This unit, as a complete data object, contains a verified engineering entity identifier, one or more pre-checked monitoring data points, and an authoritative timestamp, providing standardized input for subsequent deep integration and semantic processing with the engineering knowledge graph. Bound combinations that fail verification are logged and can be discarded, trigger alerts, or transferred to a manual review queue according to predefined policies.

[0046] This embodiment achieves unified access and semantic recognition of data from various acquisition terminals by standardizing the reception and parsing of multi-source heterogeneous data streams from the engineering site, effectively solving the integration challenges of complex data sources and inconsistent formats. By intelligently binding and verifying entity identifiers and monitoring data based on spatiotemporal correlation, high-quality basic data units are constructed, ensuring the validity and inherent consistency of the initial data and providing reliable input for subsequent processing. The use of a unified time stamp to encapsulate valid data gives each data unit a standardized time reference, laying a solid foundation for building an accurate full-cycle data view of the project and improving the traceability and comparability of data in the time dimension.

[0047] In one embodiment, the component identifiers in the basic data unit are compared with the pre-built engineering knowledge graph to perform entity information detection. When a mismatch is detected between the component identifier and the engineering knowledge graph information, the entity information corresponding to the component identifier is extracted, including: The generated basic data units are parsed. Each basic data unit is a structured data container containing a key field that uniquely identifies an engineering entity: the engineering entity identifier. The identifier is a composite field containing prefixes such as version information and project code. From this engineering entity identifier string, according to predefined syntax rules, the most crucial part used for entity matching in the knowledge graph is analyzed and extracted: the component identifier. After obtaining the clean component identifier, the process enters the knowledge graph traversal and matching phase. The pre-built engineering knowledge graph is treated here as a graph-structured database containing a large number of entity information nodes.

[0048] Each entity information node possesses one or more attribute values ​​to identify itself, which are standardized component codes or names. The matching engine initiates a traversal query, sequentially accessing each entity information node in the knowledge graph and reading its identification attributes. Subsequently, it performs a rigorous string comparison between the component identifier extracted from the basic data unit and the identification attributes of the currently accessed node.

[0049] This involves precise or fuzzy matching based on specific matching rules, such as requiring strings to be exactly equal, or allowing matching while ignoring case or specific delimiters. Each comparison produces a definite result, either "match" or "not match," which is recorded as comparison information. The traversal and comparison process continues until all candidate nodes have been traversed or a match is found ahead of time.

[0050] All comparison results are summarized and analyzed, and a definitive judgment is made to determine the next direction of the data flow. After traversing and comparing all relevant entity information nodes in the knowledge graph, the system obtains a complete set of comparison information. This set contains the results of every comparison attempt for the current component identifier.

[0051] The comparison information set is then evaluated holistically. The evaluation logic is deterministic: if no "match" result exists in the comparison information set, it means that no entity definition corresponding to the input component identifier has been found within the knowledge scope of the current knowledge graph. The entity matching operation will then be formally deemed a failure, and a structured node mismatch result will be generated. The node mismatch result is a signal containing contextual information, including the original component identifier that triggered the mismatch, the ID of the underlying data unit that initiated the matching operation, and metadata such as the timestamp of the matching failure.

[0052] The output of a node mismatch result indicates that the entity represented by the basic data unit is a "new entity" or an "unknown entity" in the knowledge graph, thereby triggering subsequent processing branches aimed at expanding the knowledge graph. This determination mechanism ensures that the process can clearly distinguish between known and unknown entities, providing clear triggering conditions for the adaptive evolution of the knowledge graph.

[0053] This operation maximizes the extraction of feature information from existing data carriers that can be used to define and describe the new entity, providing factual evidence for the expansion of the knowledge graph. It is a direct response to node mismatch results; once the determination result is generated, the corresponding information extraction routine is immediately initiated.

[0054] The implementation process focuses on the in-depth analysis of the original basic data units. These units, as the initial encapsulation of the data, not only contain core component identifiers but also rich contextual metadata. The goal of the extraction operation is to locate and obtain four types of key information: data source type, collection timestamp, monitoring parameter sequence, and data collection location information.

[0055] Extracting the data source type is achieved by accessing the data source identifier recorded in the basic data unit. This identifier indicates the type of acquisition device or system module that generated the data, such as a strain gauge, temperature sensor, GPS positioning module, or manual inspection terminal. This information helps infer the basic attribute category of the new entity. Extracting the acquisition timestamp involves reading the uniform timestamp assigned to the basic data unit in previous processes. This timestamp records the moment the entity was first perceived by the system and is crucial for constructing the entity's time-dimensional lifecycle.

[0056] Extracting the monitoring parameter sequence involves acquiring all monitoring data values ​​and their type names bound to the component identifier within the data unit. These parameter sequences intuitively reflect the observable state characteristics of the entity and serve as the direct basis for defining the entity's monitoring dimensions. Extracting data acquisition location information involves parsing location data from different sources, such as latitude and longitude coordinates from embedded GPS modules, area codes based on base stations or RFID, or construction zone numbers specified during manual entry. This information is used to establish the entity's coordinates in its spatial context. The extraction process accurately reads the above information from specific fields of the basic data unit by calling a predefined metadata access interface, ensuring data integrity and consistency.

[0057] Dispersed data items are integrated into a structured information object with complete semantic description capabilities, namely entity information. This entity information will serve as the authoritative data source for creating new nodes in the knowledge graph. The combination operation is not a simple data stacking, but a logical and structured assembly based on a predefined entity information model. This model defines the information fields required to describe an engineering entity and their interrelationships.

[0058] During implementation, the four extracted elements—data source type, collection timestamp, monitoring parameter sequence, and data collection location information—are mapped to corresponding fields in the entity information model. The data source type is assigned a "collection attribute" field to describe the entity's data source characteristics. The collection timestamp is recorded as the entity's "first identification time," serving as the starting point of the entity's lifecycle in the digital system. The monitoring parameter sequence is systematically parsed, and each parameter type and its initial reading are compiled into an "initial state characteristics" list, which clarifies which states can be monitored and recorded for this type of entity.

[0059] After standardization, the data collection location information is entered into the "Spatial Location" field to spatially locate the entity. All these fields are encapsulated in a unified data structure instance, either as key-value pairs or nested documents, thus forming a rich and clearly structured entity information object. This entity information object not only contains the core identifiers needed to identify the entity, but more importantly, it integrates the contextual features of its perception, preparing the data for the creation of a new engineering knowledge graph node with practical data support.

[0060] This embodiment, through rigorous string comparison between component identifiers and knowledge graph nodes, can clearly distinguish between known and unknown entities, thus providing an accurate basis for subsequent differential processing. By responding to mismatch results and automatically extracting multi-dimensional information such as data source, timestamp, parameter sequence, and location, a complete information set describing new entities can be constructed, providing sufficient data support for the adaptive expansion of the knowledge graph. By combining the extracted multi-dimensional information into structured entity information objects, standardized, machine-readable entity definitions can be formed, significantly improving the knowledge graph's ability to absorb newly discovered entities and the completeness of the knowledge system.

[0061] In one embodiment, the engineering knowledge graph is constructed and its relationships are updated based on entity information. The updated knowledge graph is then re-detected until a match is found, generating engineering data units, including: The system takes an entity information object as input. This entity information is a structured data container containing the data source type, collection timestamp, monitoring parameter sequence, and data collection location information. The node attribute construction operation first focuses on the data source type and data collection location information, as these two types of information most directly reveal the entity's basic static characteristics and context. The data source type value, such as "GPS displacement monitoring point" or "reinforcing bar stress gauge," is fed into a predefined attribute mapping rule base.

[0062] The rule base maintains a mapping between different data source types and a set of standard engineering attributes. The mapping operation automatically derives an initial set of attribute key-value pairs based on the detected data source type. For example, a data source type identified as "temperature sensor" is mapped to basic attribute fields such as "monitored physical quantity: temperature", "unit: degrees Celsius", and "expected range: -20 to 100".

[0063] Simultaneously, the data acquisition location information undergoes a standardization process. The raw location information exists in various formats, such as latitude and longitude coordinates, relative location descriptions, or regional codes. This process converts these heterogeneous location descriptions into coordinates or standard location identifiers under a standardized spatial reference system uniformly adopted within the knowledge graph. The processed standard location information is assigned attribute fields such as "spatial location" or "installation coordinates." All the attribute key-value pairs obtained through mapping and standardization are aggregated and organized into a structured set of attribute fields. This set constitutes an attribute blueprint for the initial definition of new entities, providing specific attribute content for creating corresponding nodes in the knowledge graph.

[0064] Formal integration of new entities into the core operations of the engineering knowledge graph involves node instantiation, attribute assignment, and the integration of the relationship network. It begins with the creation of a new, empty node object—the new engineering node—in the graph database of the engineering knowledge graph. This node is assigned a globally unique node identifier upon creation, but its semantic content is currently empty.

[0065] The attribute field set generated in the previous step is written into the attribute list of this new node. This is a field-by-field assignment process; each key-value pair in the attribute field set is read, and its value is filled into the corresponding attribute item in the new node's attribute list. For new attribute keys that are not yet predefined in the knowledge graph, the graph schema can be dynamically expanded to accommodate the new attribute. After the attribute assignment is completed, the new project node has basic self-descriptive information.

[0066] The process shifts to a more semantically deep relationship-building phase, based on the sequence of monitoring parameters contained in the entity information. This sequence not only includes readings but also implicitly reveals potential associations between the entity and other concepts or entities. During implementation, the type, name, or metadata of the monitoring parameters is analyzed, and based on predefined business rules, it is inferred which existing nodes or concept nodes in the graph should establish relationship edges with the new node. For example, a sensor node monitoring "main beam deflection" needs to establish a "monitoring" relationship with the engineering node representing "main beam"; its monitoring parameter type, "deflection," needs to establish a "belongs to" classification relationship with a standardized "mechanical index" concept node.

[0067] The relationship update operation involves creating new directed edges in the knowledge graph and adding attributes to them, such as relationship strength and initial creation time. Once the new nodes are created, the attribute list is populated, and the relationship edges are updated, the state of the engineering knowledge graph changes, generating an updated knowledge graph containing the new knowledge elements. This updated graph provides a more complete knowledge foundation for all subsequent graph-based operations.

[0068] The new entities have been confirmed to be correctly included in the knowledge graph, and structured semantic information is injected back into the original data units to achieve deep integration of data and knowledge. Using the updated knowledge graph as a new context, the entity information detection process is re-executed. The input for this detection remains the component identifiers carried in the original basic data units, but the database environment it searches now includes the newly created project nodes.

[0069] The detection mechanism traverses the updated graph space. Since the identifier attributes of newly added nodes have been officially stored in the database, this traversal can accurately hit the corresponding entity information nodes. A successful match triggers a semantic appending operation.

[0070] The complete set of attribute fields previously built based on entity information is associated with the original basic data unit as a semantic enhancement package. This association is not a simple data splicing, but rather, while retaining the original observations, timestamps, and other core data of the basic data unit, an independent attribute extension layer rich in engineering semantics is added to it.

[0071] The attribute extension layer transforms basic data, which originally only contained identifiers and numerical values, into rich contextual information defined by the knowledge graph, including types, spatial locations, and monitoring parameter dictionaries. After this process, the basic data unit is transformed into an engineering data unit. This unit not only contains raw, unvarnished monitoring facts but also integrates standardized semantic descriptions that are understandable to machines, becoming an information entity with a complete engineering context. This provides a data foundation that is both authentic and interpretable for subsequent quality verification, trend analysis, and decision support.

[0072] Matching failures are theoretically considered abnormal paths, meaning that the previous node creation or relationship update operation failed to bring the new node to a state where it could be successfully retrieved. Reasons for this include node identifier mapping rule conflicts, graph transaction commit delays, or interference from concurrent access. During implementation, if a re-check still results in a mismatch, retry logic is immediately triggered. This logic is not a simple, indiscriminate loop, but rather it re-invokes the complete node construction and relationship update sequence, starting from "building node attributes based on entity information." During repeated execution, the system carries previous execution context information, including re-validation of entity information, fine-tuning of attribute mapping rules, or attempts to establish nodes and relationships in the graph in a more explicit manner.

[0073] Each iteration is a correction to the knowledge graph state and a retry of the integration conditions. This "build-update-detect" loop constitutes a closed-loop process with a clear termination condition: the entity information detection returns a successful match. This design ensures the eventual consistency of the data processing logic, fundamentally preventing the accidental discarding of data units due to temporary technical failures or incomplete initial information, thus guaranteeing 100% coverage and extremely high robustness of the data pipeline processing capabilities.

[0074] This embodiment achieves adaptive expansion of the engineering knowledge graph by constructing structured attributes for newly identified entity information and dynamically creating knowledge graph nodes. This allows for the automatic absorption of new engineering objects appearing on-site without the need for manual pre-definition of all entities. By performing correlation analysis between new nodes and monitoring parameter sequences and updating the relationship edges in the graph, semantic connections between new entities and the existing knowledge system can be established, ensuring the consistency and integrity of the knowledge network. Through a closed-loop "creation-verification" mechanism until successful matching, each basic data unit is successfully transformed into a semantically rich engineering data unit, fundamentally avoiding the loss of effective data and improving the robustness and processing efficiency of the data pipeline.

[0075] In one embodiment, a new engineering node is created in the engineering knowledge graph, the set of attribute fields is written into the attribute list of the new engineering node, and the relationship edges of the new engineering node in the engineering knowledge graph are updated according to the monitoring parameter sequence to generate an updated knowledge graph, including: The input is a set of attribute fields generated through preprocessing. This set itself is a structured data container containing multiple attribute fields, each with its name and corresponding value. The primary goal of the extraction operation is to locate and obtain the node identifier.

[0076] The node identifier is a special attribute field whose name conforms to predefined naming conventions for identifying fields, such as "node ID," "unique identifier," or "component code." The system precisely locates the specific field carrying the node identifier by traversing the names of all fields in the attribute field set and matching them against a set of predefined identifier field naming rules.

[0077] Once the location is successful, the field value is extracted as the official node identifier and subjected to necessary format validation, such as checking whether it conforms to specific encoding rules or length requirements, to ensure it can serve as a unique identifier for a node in the knowledge graph. After extracting the node identifier, the process moves to the general processing of all attribute fields. At this point, the set of attribute fields is traversed, and each attribute field is accessed sequentially. For each accessed field, the operation deconstructs it into two basic components: the field name and the field value.

[0078] The field name is a string describing the semantic meaning of the attribute, such as "data source type" or "spatial location". The field value is the specific value that the attribute can take, and its data type can be string, number, boolean, or a more complex structure. This extraction process ensures that each pair of field name and field value is independently and completely separated, forming a data unit that can be processed independently by subsequent steps. All these extracted node identifiers and the set of field name-field value pairs provide standardized, atomic input data for subsequent node creation and attribute matching operations.

[0079] The extracted node identifier is used. This identifier serves as a unique code for the new entity in the knowledge graph and is used to initialize an empty node object. The operation of creating a new project node is performed in the knowledge graph's graph database environment. The result is the allocation of a unique internal node ID in the database, and the setting of the externally provided node identifier as the core identification attribute of the node. This newly created node currently only has the most basic identification information, and its attribute list is empty, awaiting filling. After the node is created, the process immediately moves to the attribute field matching stage. The goal of this stage is to match the attribute descriptions obtained from the raw data with the existing semantic specifications of the knowledge graph.

[0080] Each previously extracted field name-value pair is processed sequentially. For each pair, the field name is compared with the predefined node attribute template of the engineering knowledge graph. The node attribute template can be understood as a standardized attribute dictionary or pattern definition, which specifies which attribute items a certain type of engineering node in the knowledge graph should have, as well as the name, expected data type, and constraints of each attribute item.

[0081] The matching operation involves a precise string comparison between the current field name and the names of all defined attribute items in the template, searching for an attribute item whose name is exactly the same as the current field name. This matching process is performed field by field, aiming to determine whether each attribute field from the raw data can be mapped to the existing standardized attribute semantic framework of the knowledge graph. The matching result will directly determine the specific way to process the field value subsequently: whether to incorporate it into the standardized attribute system or to process it dynamically as an extended attribute.

[0082] A successful field name match indicates that the attribute field can be included in the standard semantic framework of the knowledge graph. When the name of a field exactly matches the name of a predefined attribute item in the node attribute template, a successful match is confirmed, and that attribute item is locked as the target for writing.

[0083] Before executing the write command, a verification process is initiated. This process performs a compliance review on the field values ​​to be written, based on the specifications defined in the template for the target attribute items. These specifications include, but are not limited to, data type constraints, value range limitations, format requirements, and business logic rules.

[0084] For example, if the target attribute is defined as an integer, the string-based number needs to be parsed and converted; if the attribute defines an enumerated value range, the field value needs to be checked to see if it belongs to one of the valid enumerated values; if the attribute represents geographic coordinates, its format needs to be verified to conform to the standard. For values ​​that do not conform to the specifications, a data cleaning routine can be triggered or the data can be marked as requiring review. After verification, the field value is officially assigned to the corresponding attribute in the new project node's attribute list. This write operation is atomic, ensuring that the assignment of each attribute is independent and accurate.

[0085] Once the writing is complete, the standardized attribute item carries specific information from the engineering site, enabling the newly created engineering node to possess semantic features that can be understood globally by the knowledge graph. This process is performed field by field, gradually building a standardized description of the new node using standardized attribute values.

[0086] When a field name cannot find a corresponding predefined attribute item in the node attribute template, a match failure event is triggered, indicating that a new descriptive dimension has been introduced into the current data. This pair of unmatched field names and values ​​is identified as a complete unit of information. This new attribute unit is then normalized to ensure it integrates harmoniously into the existing attribute list. Normalization includes standardizing and cleaning field names, such as unifying naming styles, eliminating ambiguity, and ensuring the uniqueness of names within the attribute list.

[0087] The normalized field names and their corresponding field values ​​are integrated into a new attribute item. This attribute item, as a complete key-value pair, is added to the attribute list of the new project node. This addition operation is accompanied by the recording of metadata about the new attribute item, such as its creation source and creation timestamp, to trace its origin. In this way, the knowledge graph pattern can incrementally learn during the instantiation process, continuously enriching its vocabulary describing project entities. This mechanism effectively solves the challenge of exhaustively enumerating all attributes in the early stages of system design due to the complexity and innovation of engineering sites, significantly improving the adaptability and foresight of the entire data management method.

[0088] New nodes are transformed from isolated data points into interconnected, organic components within a knowledge graph. The implementation process relies heavily on monitoring parameter sequences, which implicitly define the node's function, behavior, and interactions with other entities in the environment. Relationship localization based on these monitoring parameter sequences first requires in-depth semantic parsing of the sequences.

[0089] Metadata such as the type name, unit, and measurement object of the monitoring parameters are matched and reasoned against a pre-defined relational mapping rule base. These rules define the semantic relationship type implied by a specific type of monitoring parameter. For example, the parameter sequence of a sensor monitoring "bridge pier settlement" will trigger a rule indicating that a "monitors" relationship edge needs to be established between the sensor node and the engineering node representing the "bridge pier".

[0090] Relationship localization involves using rule-based reasoning to determine which one or more existing nodes in the graph a new node should connect to, and the specific semantic type of the connection. Once the relationship target and type are determined, a corresponding directed relationship edge is created in the graph, pointing from the new node to the target node.

[0091] Assignment attributes are set for relation edges, further characterizing the specific features of the relation. These attributes include the basis for relation establishment, initial establishment time, relation strength or confidence level, and other parameters characterizing relation properties. Setting association attributes makes the relation itself an entity carrying rich information. Once all inferred relation edges have been created and configured, the new project node is deeply embedded into the complex network of the project knowledge graph through multiple relation edges with clear semantics and attributes. At this point, the knowledge graph completes this incremental update, generating an updated knowledge graph containing new nodes, new attributes, and new relations, providing a more complete knowledge foundation for subsequent queries, analysis, and reasoning.

[0092] This embodiment ensures precise alignment between external data and the internal semantic model of the knowledge graph by standardizing and verifying attribute fields before writing them into node attributes, thereby improving the standardization and consistency of engineering data description. Dynamic integration and attribute list expansion of unmatched fields enable the knowledge graph to adaptively absorb new entity feature descriptions, enhancing the system's inclusiveness and flexibility in handling emerging attribute information from the engineering site. Semantic reasoning based on monitoring parameter sequences to locate and set relation edges automatically associates new entities with the existing knowledge network context, achieving a deep transformation from isolated data points to interconnected knowledge elements and strengthening the integrity and interpretability of the data.

[0093] In one embodiment, a quality check is performed on the engineering data unit according to a preset engineering quality rule, and engineering data units that fail the quality check are marked as abnormal data units, including: As a composite structure integrating raw monitoring data and semantic attributes of a knowledge graph, an engineering data unit contains a variety of information. The goal of the extraction operation is to locate and obtain specific data items that need to be subject to rule-based verification, i.e., specification indicators. These specification indicators include the specific numerical values ​​of the monitoring data, derived values ​​obtained through simple calculations, and logical states obtained by comparing them with the attributes of associated nodes in the knowledge graph. The extraction logic is executed according to a predefined specification indicator mapping table, which specifies which data fields need to be extracted or which indicators need to be calculated as verification objects for different types of engineering data units. For example, for a data unit representing a concrete strength sensor, the specification indicator to be verified is the raw pressure value it monitors; while for a data unit representing structural displacement, the indicator to be verified is the calculated displacement change rate.

[0094] After successfully extracting the specifications to be verified, the process enters the rule selection stage. The preset engineering quality rule set is an ordered collection of rules, where each rule defines a specific quality dimension and acceptance standard. The rules are ordered based on the verification priority; for example, basic integrity checks are performed first, followed by numerical rationality checks, and then complex logical consistency verifications. Sequential selection means starting with the first rule in the rule set and setting each rule sequentially as the current verification rule effective within the current verification cycle. This sequential execution mechanism ensures that the verification process is orderly, avoids cross-interference between rules, and provides a clear context for subsequent step-by-step processing.

[0095] The current verification rules not only include rule identification information, but more importantly, they define specific specifications. These specifications can take various forms, including numerical ranges, lists of enumerated values, reference curves, or logical conditions. The verification comparison operation involves checking the actual values ​​of the specifications against the requirements defined by the rules. This check is not always a simple equivalence or range judgment; it involves calculations based on rule logic. For example, if the rule requires that "the rate of change of monitored value A must not exceed threshold B," then the comparison operation needs to first calculate the rate of change and then compare it with the threshold. Similarly, if the rule requires that "the equipment status must be consistent with the status of the associated power node," then cross-data unit association queries and logical judgments are required.

[0096] When the comparison results reveal that the specifications do not meet the requirements, the exception handling process is triggered. First, detailed rule deviation information is recorded. This information not only records the fact of "non-compliance," but also includes specific details of the deviation, such as the actual value, the expected value, the absolute or relative amount of the deviation, and the specific rule clause violated. This detailed information is recorded in a structured manner, forming a complete deviation report. Based on the existence of this deviation, the system formally determines that the current quality verification for this rule has failed. This determination indicates that the engineering data unit has a defect in the quality dimension defined by the current rule, providing clear input for subsequent exception classification and labeling.

[0097] Specific, personalized rule deviation information is mapped to a unified, limited anomaly classification system, providing a precise classification basis for subsequent differentiated processing. The implementation process uses recorded rule deviation information as input. This information includes the technical details of the deviation, such as the violated rule ID, the calculation method of the deviation, and the difference between the actual and expected values. A pre-defined anomaly type mapping table is a key configuration, defining the correspondence between various rule deviation patterns and standard anomaly type codes. The mapping table structure includes matching conditions and corresponding anomaly type codes.

[0098] Matching conditions are defined based on a combination of multiple dimensions, including rule ID, the nature of the deviation (such as exceeding the upper or lower limit, missing data, or logical conflict), and the severity threshold of the deviation. The type matching operation compares the features in the rule deviation information with each matching condition in the mapping table to find a completely matching condition. For example, a rule deviation regarding "concrete temperature exceeding 70 degrees Celsius" matches the condition "monitoring value exceeding the process upper limit" in the mapping table and is thus categorized into the corresponding anomaly type code, such as "TEMP_OVER_LIMIT". Similarly, a logical error such as "data record timestamp is later than the received timestamp" is mapped to the "time-series logical anomaly" type code "TIME_LOGICAL_ERR".

[0099] The matching process needs to handle overlaps and priority issues. The mapping table defines explicit matching priorities or mutual exclusion rules. Upon successful matching, a standardized exception type code is output. This code, as a machine-readable identifier, concisely represents the essential category of the current exception, shielding it from the complex details of specific rule deviations. This allows subsequent processes to select and handle strategies based on type rather than specific details. This classification mechanism greatly enhances the efficiency and consistency of handling diverse quality issues.

[0100] Based on the identified anomaly type codes, a correlation query is performed within the engineering quality rule set. The engineering quality rule set not only defines inspection rules and specifications but also maintains metadata associated with various anomaly type codes, a crucial element of which is the quality status identifier. This identifier is a predefined symbol or code used to indicate the overall quality status that a data unit in this anomaly state should be assigned. For example, for minor anomalies caused by transient interference, the corresponding quality status identifier is "suspicious"; for serious anomalies affecting data reliability, the identifier is "failure"; and for anomalies requiring immediate alerts, the identifier is "dangerous."

[0101] The extraction operation retrieves the corresponding quality status identifier value from the rule set configuration based on the exception type code key. This step ensures that the severity or urgency of the exception can be passed down through an explicit status value. Next, the system generates an exception identification timestamp. This timestamp records the exact moment the exception was formally identified by the system. The timestamp generation follows a unified clock source and time format specification, ensuring its uniqueness and comparability within the distributed system. The exception identification timestamp is crucial; it marks the point in time when the data quality issue was discovered, and is of key value for subsequent problem tracing, performance statistics, and time-based correlation analysis. The extracted quality status identifier and the generated exception identification timestamp together constitute important status attributes of the exception data unit.

[0102] Taking the original engineering data unit containing the detected problem as the core, three types of metadata—anomaly type encoding, quality status identifier, and anomaly identification timestamp—are associated with it. This association is not a simple data appending, but is achieved through a specific tagging mechanism. A typical implementation is to dedicate a field area in the engineering data unit's data structure to record the quality status, called the quality header or tag block.

[0103] The associated tagging operation involves writing the anomaly type code, quality status identifier, and anomaly identification timestamp into this area. The anomaly type code indicates "what the problem is," the quality status identifier indicates "the current overall status of this data," and the anomaly identification timestamp indicates "when the problem was discovered." These tags coexist with the original content of the engineering data unit (such as monitoring data, entity identifiers, semantic attributes, and original timestamps) to form an enhanced data structure.

[0104] This tagged data entity, containing both the original data and complete quality information, is called an anomalous data unit. The formation of an anomalous data unit signifies that it is no longer merely an observation record, but an information entity carrying its own quality assessment results. It clearly informs downstream data users or processing flows that this data contains a specific type of identified quality problem, its quality status, and when the problem was identified. This provides a clear and consistent input basis for subsequent differentiated processing such as data correction, ignoring, review, or statistical analysis, achieving effective transmission and management of data quality information throughout the process.

[0105] This embodiment achieves standardized identification and classification of data quality issues by performing regularized quality verification and anomaly marking on engineering data units, thereby improving the standardization and consistency of quality issue handling. By mapping rule deviations to standardized anomaly codes and associating them with quality status identifiers, differentiated processing criteria can be provided for anomaly data of different natures, thus achieving precision and refinement in anomaly handling. By generating anomaly identification timestamps and associating them with original data units, the temporal information of quality problem occurrence and identification can be accurately recorded, providing a reliable data foundation for quality traceability and process analysis.

[0106] In one embodiment, abnormal data units are subjected to policy polling and data correction based on a correction strategy table, and the obtained corrected data units are periodically constructed with the updated knowledge graph to output a full-cycle data view, including: The system takes anomaly data units and a predefined correction policy table as input. The correction policy table is an ordered collection, where each correction policy explicitly specifies its applicable preconditions, i.e., triggering conditions. These triggering conditions are defined based on the metadata carried by the anomaly data unit, such as the anomaly type code, data source type, device identifier that generated the data, or characteristics of the anomaly value (such as deviation magnitude). Matching operations are performed sequentially according to the order of the policies in the table.

[0107] For the currently checked correction strategy, the system parses its trigger conditions and then extracts the corresponding feature attributes from the anomalous data unit for comparison. For example, one strategy's trigger condition is defined as "triggered when the anomaly type is 'SENSOR_SPIKE' and the data source is 'vibration sensor'"; another strategy's trigger condition is "triggered when the anomaly type is 'DATA_GAP' and the data missing duration is less than a set threshold". The comparison process is a logical judgment that checks whether the features of the anomalous data unit fully meet the trigger conditions of the strategy. Once met, the matching process immediately terminates, and the strategy is selected and marked as the current execution strategy.

[0108] This sequential matching and "first-match-first-hit" mechanism essentially establishes a priority system for strategies, with strategies listed earlier having higher execution priority. This ensures that for complex anomalies that can be handled by multiple strategies, the most appropriate processing method will be adopted according to the preset priority order. If no strategy that meets the triggering conditions is found after traversing the entire strategy table, it means that there is currently no automatic correction solution for the abnormal data unit, and it is transferred to a special processing flow, such as a manual review queue. Successfully marking the current execution strategy provides a clear instruction basis for subsequent targeted data correction operations.

[0109] Based on the specific rules and methods defined by the selected strategy, the abnormal parts in the abnormal data units are corrected, producing a preliminary correction result. This process revolves around the current execution strategy and the abnormal data units. The current execution strategy not only includes triggering conditions but, more importantly, defines the specific correction logic. These correction logics are diverse, including but not limited to: numerical replacement (e.g., replacing abnormal values ​​with the previous valid value, the average value of adjacent sensors during the same period, or values ​​calculated by interpolation), data smoothing (filtering data sequences containing spikes), null value imputation (impacting missing data segments), and even reasoning correction based on related data in a knowledge graph (e.g., calculating a reasonable replacement value based on equipment status and operating parameters).

[0110] Performing outlier correction involves applying these predefined logics to outlier data units. The operation locates the specific data point or segment within the unit that is marked as outlier, and then performs calculations or replacements according to the policy logic. For example, for a transient spike outlier, the policy logic is to "replace the spike peak with a linear interpolation of the two preceding and following normal data points"; for a continuous missing data segment, the policy logic is to "generate data for the missing segment using a time series forecasting model".

[0111] The entire correction process must ensure the structural integrity of the original data units. This means the correction operation primarily targets the abnormal monitoring values ​​themselves, while preserving the core identifier, timestamp, and other normal metadata of the data unit. After the correction operation is complete, a new data entity is generated: the intermediate corrected data unit. The intermediate corrected data unit contains the corrected monitoring data, but also retains traces of its origin from the abnormal data unit. For example, the original abnormal marker is converted to a "corrected" state, and the correction strategy ID used is recorded.

[0112] The process takes intermediate corrected data units and an updated engineering knowledge graph as input. Node mapping is the primary step, aiming to confirm that the engineering entity represented by the corrected unit has a clear and unique corresponding node in the knowledge graph. By extracting the engineering entity identifier from the corrected unit, a node query is performed in the knowledge graph. Successful mapping signifies that the entity is a defined and managed object in the knowledge graph, laying the foundation for subsequent relationship verification.

[0113] The relation verification performed is a deeper level of semantic checking. The verification logic is based on the business rules implied by the relation edges connected to the entity node in the knowledge graph. For example, the knowledge graph defines a rule: "The vibration sensor readings of the pump station should have a specific covariant relationship with the inlet and outlet pressure values." Relation verification will check whether the corrected vibration data still satisfies this covariant relationship with the associated pressure data in the same time context.

[0114] For the correction data representing "beam deflection," its value is verified to ensure it falls within the reasonable deformation range for this type of beam under a known load spectrum. This range information is derived from the material properties and design parameters associated with the beam nodes in the spectrum. The verification process involves complex logical judgments or correlation analysis of multiple data sources. If an intermediate correction data unit passes all applicable relational verifications, proving that its correction result is reasonable and consistent within the engineering context, then the unit is formally designated as a correction data unit.

[0115] The process takes all corrected data units (along with previously normal engineering data units) and an updated knowledge graph as input. The core of the temporal integration operation is to systematically associate the data units with their corresponding engineering entity nodes in the knowledge graph based on their timestamp attributes. For each engineering entity node in the knowledge graph, the operation aggregates all data units that are temporally associated with that node. These data units are arranged in chronological order according to their timestamps, forming a state sequence unfolding along the timeline. This sequence clearly shows the historical record of all key state observations of the entity from its first perception to the current time point, including its normal state, abnormal events, and corrected state. The integration process is not just simple temporal sorting; it also involves associating and aligning the temporal state data of different entities based on the relationships between entities defined in the knowledge graph (such as spatial relationships, logical relationships, and parent-child relationships). For example, aligning different types of sensor data from the same structural part on the timeline allows for the analysis of multi-parameter coupling effects; associating the operating state sequences of upstream and downstream equipment allows for the analysis of fault propagation paths.

[0116] Through this cross-entity temporal correlation, the final constructed full-cycle data view is a multi-dimensional, interconnected temporal data cube. This view not only provides the independent history of each entity, but also reveals the interactions between entities in the time dimension and the behavioral patterns of the overall engineering system, providing a unique and authoritative data foundation for engineering diagnosis, predictive maintenance, and global optimization.

[0117] This embodiment sequentially matches abnormal data units with a correction strategy table and performs targeted corrections, enabling the allocation of appropriate processing strategies for different types of anomalies. This allows for precise and flexible handling of diverse data anomaly issues. By semantically verifying intermediate correction results with a knowledge graph, it ensures that the corrected data conforms to engineering logic and business rules, thereby improving data usability while guaranteeing data rationality and credibility. By temporally integrating the verified corrected data with knowledge graph nodes, the complete state evolution sequence of engineering entities can be systematically reconstructed, generating a full-cycle data view that is both accurate and comprehensive, supporting in-depth analysis and traceability.

[0118] Reference Figure 2 The present invention also provides an engineering data management device, which is applied to the engineering data management method described in any of the above-mentioned embodiments, comprising: The data acquisition module is used to collect engineering entity identification and monitoring data from the engineering site and construct them into basic data units. The analysis module is used to perform entity information detection between the component identifiers in the basic data units and the pre-built engineering knowledge graph. When a mismatch is detected between the component identifier and the engineering knowledge graph, the entity information corresponding to the component identifier is extracted. Based on the entity information, the engineering knowledge graph is used to construct engineering nodes and update relationships. An updated knowledge graph is generated and the detection is performed again until a match is successful, and an engineering data unit is generated. The association module is used to perform quality checks on engineering data units according to preset engineering quality rules, and to mark engineering data units that fail the quality check as abnormal data units. The processing module is used to perform policy polling and data correction on abnormal data units according to the correction strategy table, and periodically construct the obtained corrected data units with the updated knowledge graph to output a full-cycle data view of the entire project.

[0119] This invention provides an engineering data management device that integrates on-site entity identification and monitoring data into a single unit. By dynamically matching and expanding an engineering knowledge graph for each unit, it achieves semantic encapsulation and dynamic association of engineering entity data, effectively solving the problem of data-entity disconnect. Real-time verification of data units is performed using knowledge graph-based quality rules, and an automated correction mechanism for abnormal data is employed through a policy polling mechanism, significantly improving the accuracy of data processing and the timeliness of anomaly response. By periodically associating the corrected data units with the knowledge graph, a complete and reliable full-cycle data view is generated, providing a data foundation with complete historical traceability for project decision-making and supporting pre-event warning, in-event control, and post-event analysis in engineering management.

[0120] Reference Figure 3 As shown, the present invention also provides an engineering data management device, comprising: Memory, used to store programs; A processor is used to execute programs to implement the various steps of an engineering data management method that includes any of the above-mentioned features.

[0121] In this embodiment, the processor and memory can be connected via a bus or other means. The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive. The processor may be a general-purpose processor, such as a central processing unit, digital signal processor, application-specific integrated circuit, or one or more integrated circuits configured to implement embodiments of the present invention.

[0122] The present invention also provides a storage medium storing computer instructions for causing a computer to perform any of the methods described above.

[0123] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the system and each module described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for managing engineering data, characterized in that, include: Collect engineering entity identification and monitoring data from the engineering site to construct basic data units; The component identifier in the basic data unit is compared with the pre-built engineering knowledge graph for entity information detection. When the information of the component identifier in the engineering knowledge graph does not match, the entity information corresponding to the component identifier is extracted. Based on the entity information, the engineering knowledge graph is used to construct engineering nodes and update relationships. An updated knowledge graph is generated and the detection is repeated until a match is successful, and an engineering data unit is generated. The engineering data unit is subjected to quality verification according to the preset engineering quality rules, and the engineering data unit that fails the quality verification is marked as an abnormal data unit. The abnormal data units are polled and corrected according to the correction strategy table, and the corrected data units are periodically constructed with the updated knowledge graph to output a full-cycle data view.

2. The engineering data management method according to claim 1, characterized in that, The collected engineering entity identification and monitoring data from the engineering site are used to construct basic data units, including: Receive multi-source data streams from the engineering site, identify and parse the multi-source data streams to obtain the engineering entity identifier and the monitoring data; Each of the engineering entity identifiers is associated and bound with at least one of the monitoring data to form a binding combination; The binding combination is subjected to basic verification according to the predefined identifier specification. When the engineering entity identifier in the binding combination passes the format verification and the monitoring data passes the threshold pre-detection, a time stamp with a unified format is added to the binding combination to generate the basic data unit.

3. The engineering data management method according to claim 1, characterized in that, The step involves performing entity information detection between the component identifier in the basic data unit and the pre-constructed engineering knowledge graph. When a mismatch is detected between the component identifier and the information in the engineering knowledge graph, the entity information corresponding to the component identifier is extracted, including: Based on the analysis of the engineering entity identifier, the component identifier in the basic data unit is extracted. All entity information nodes in the engineering knowledge graph are traversed. The component identifier is compared with each entity information node in string form to obtain the comparison information. When all the comparison information does not match, the node mismatch result is determined and output; In response to the node mismatch result, the data source type, collection timestamp, monitoring parameter sequence and data collection location information associated with the component identifier are located and extracted from the basic data unit; The extracted data source type, the collection timestamp, the monitoring parameter sequence, and the data collection location information are combined to form the entity information.

4. The engineering data management method according to claim 3, characterized in that, The process of constructing engineering nodes and updating relationships in the engineering knowledge graph based on the entity information, generating an updated knowledge graph, re-detecting it until a match is successful, and generating engineering data units includes: Based on the data source type and data collection location information in the entity information, node attributes are constructed to generate a set of attribute fields. A new engineering node is created in the engineering knowledge graph, the attribute field set is written into the attribute list of the new engineering node, and the relationship edge of the new engineering node in the engineering knowledge graph is updated according to the monitoring parameter sequence to generate the updated knowledge graph; The component identifier in the basic data unit is re-detected with the updated knowledge graph. If the match is successful, the attribute field set is associated and attached to the basic data unit to generate the engineering data unit. If a match fails, repeat the node construction and relationship update steps described above until a match is found.

5. The engineering data management method according to claim 4, characterized in that, The process of creating a corresponding new project node in the project knowledge graph, writing the attribute field set into the attribute list of the new project node, and updating the relationship edges of the new project node in the project knowledge graph according to the monitoring parameter sequence to generate the updated knowledge graph includes: Extract the node identifier and the field name and field value of each attribute field from the set of attribute fields; Based on the node identifier, a new project node is created in the project knowledge graph, and the field name of each attribute field is matched with the node attribute template of the project knowledge graph. When the field name matches successfully, the field value is written to the corresponding attribute item in the attribute list of the new project node; When the field name fails to match, the field name and the field value are combined into a new attribute item and added to the attribute list; After completing the matching of all the attribute fields, the relationship edge between the new project node and the project knowledge graph is located based on the monitoring parameter sequence, and the association attribute of the relationship edge is set. After the relationship is updated, the updated knowledge graph is generated.

6. The engineering data management method according to claim 1, characterized in that, The step of performing quality verification on the engineering data units according to preset engineering quality rules, and marking the engineering data units that fail the quality verification as abnormal data units, includes: Extract the specification indicators to be verified from the engineering data unit, and select a rule from the preset engineering quality rule set as the current verification rule in sequence; The specification indicators are compared with the specification requirements in the current verification rules. When the specification indicators do not meet the specification requirements, the rule deviation information is recorded and the current verification is determined to have failed. Based on the rule deviation information and the preset anomaly type mapping table, type matching is performed to obtain the anomaly type code; Based on the anomaly type code, the corresponding quality status identifier is extracted from the engineering quality rule set, and an anomaly identification timestamp is generated; The anomaly type code is associated with the corresponding engineering data unit, and the anomaly identification timestamp and the quality status identifier are added to form the anomaly data unit.

7. The engineering data management method according to claim 1, characterized in that, The process of polling and correcting the abnormal data units according to the correction strategy table, and periodically constructing the corrected data units with the updated knowledge graph to output a full-cycle data view includes: The abnormal data unit is matched sequentially with each correction strategy in the correction strategy table. When the abnormal data unit meets the triggering condition of any of the correction strategies, the correction strategy is marked as the current execution strategy. Based on the current execution strategy, abnormal data correction is performed on the abnormal data unit to generate an intermediate corrected data unit; The intermediate correction data unit is mapped to the updated knowledge graph and the nodes are verified. If the verification is successful, the intermediate correction data unit is marked as a correction data unit. The corrected data unit is integrated with the data unit node in the updated knowledge graph in a time sequence to construct the full-cycle data view of the project.

8. An engineering data management device, characterized in that, The engineering data management method applied to any one of claims 1-7 includes: The data acquisition module is used to collect engineering entity identification and monitoring data from the engineering site and construct them into basic data units. The analysis module is used to perform entity information detection between the component identifier in the basic data unit and the pre-built engineering knowledge graph. When the information of the component identifier in the engineering knowledge graph does not match, the entity information corresponding to the component identifier is extracted. Based on the entity information, the engineering knowledge graph is used to construct engineering nodes and update relationships. An updated knowledge graph is generated and the detection is performed again until a match is successful, and an engineering data unit is generated. The association module is used to perform quality verification on the engineering data unit according to the preset engineering quality rules, and mark the engineering data unit that fails the quality verification as an abnormal data unit. The processing module is used to perform policy polling and data correction on the abnormal data units according to the correction strategy table, and periodically construct the obtained corrected data units with the updated knowledge graph to output a full-process cycle data view.

9. An engineering data management device, characterized in that, include: Memory, used to store programs; A processor is configured to execute the program to implement the various steps of the engineering data management method as described in any one of claims 1-7.

10. A storage medium, characterized in that, The computer contains computer instructions for causing the computer to perform the method according to any one of claims 1 to 7.