Full-factor data identification analysis and cross-platform unique identification method for engineering cost

By constructing a comprehensive data identification system and multi-level parsing architecture for engineering cost, and combining semantic mapping with a two-way reinforcement mechanism of blockchain, the problem of inconsistent identification of engineering cost data across different platforms and the difficulty of cross-platform identification has been solved, achieving global uniqueness of data and efficient parsing for cross-platform applications.

CN122432124APending Publication Date: 2026-07-21NANCHANG SONGYUAN HUASHENG INCUBATOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, engineering cost data lacks a unified identification and coding standard across different platforms, the global uniqueness of the identifier cannot be guaranteed, the cross-platform recognition capability is insufficient, and it is difficult to connect with the internationally accepted industrial internet identifier resolution system, resulting in poor data compatibility and limited cross-platform applications.

Method used

A comprehensive data identification system for engineering cost is constructed, which adopts a hierarchical combination coding to generate a globally unique identifier code and establishes a multi-level identifier resolution architecture. Combined with cross-platform semantic mapping and a blockchain two-way reinforcement mechanism, the system achieves cross-platform unique identification of data.

Benefits of technology

It significantly improves the accuracy and parsing efficiency of cross-platform data identification, reduces the risk of identifier conflicts, enhances the operational stability and financial security of the system, and achieves proactive identification and intelligent compensation for data consistency risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432124A_ABST
    Figure CN122432124A_ABST
Patent Text Reader

Abstract

The application discloses an engineering cost full-factor data identification analysis and cross-platform unique identification method, and belongs to the technical field of engineering cost information. The method comprises the following steps: constructing an engineering cost full-factor identification system, covering seven factors of labor, materials, machinery, list items, price information, engineering characteristics and space-time dimensions; generating a globally unique identification code, and adopting a hierarchical combination coding mechanism to guarantee the uniqueness of the identification; and establishing a multi-level identification analysis architecture, and the edge node is provided with a pre-analysis loading function. The application solves the problems of non-uniform data identification in the field of engineering cost, cross-platform identification difficulties and the like, and realizes the globally unique identification and cross-platform analysis of full-factor data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of engineering cost information technology, and involves a method for parsing and uniquely identifying all elements of engineering cost data across platforms. Background Technology

[0002] Construction cost management spans the entire project lifecycle, involving multiple stages such as investment estimation, design budget, construction drawing budget, tender control price and bid price, contract pricing, progress measurement and payment, change orders, claims and settlement, final settlement, and cost analysis during the operation phase. The large amount of cost data generated in these stages exhibits typical multi-source heterogeneous characteristics, characterized by "multiple sources, multiple formats, and multiple platforms." Poor data compatibility and inconsistent data formats exist between platforms, severely hindering the effective accumulation and cross-platform application of construction cost data.

[0003] Currently, a certain standard system has been established in the field of engineering cost informatization, covering data identification standards (such as the "Construction Engineering Labor, Material, Equipment and Machinery Data Standard"), data format standards (such as the "Construction Engineering Cost Data Exchange Standard"), data classification standards (such as the "Characteristic Classification and Description Standard" for various types of projects), and data processing and calculation standards (such as the "Construction Engineering Index Classification and Calculation Standard"). Several solutions for engineering cost data processing have also emerged in existing technologies.

[0004] (1) Data acquisition and fusion type patents: such as "Intelligent calculation system and method for engineering cost based on multi-source heterogeneous data fusion", which integrates multiple data through multi-source data acquisition module and heterogeneous data fusion module, but does not solve the problem of unique identification and parsing of data between different platforms.

[0005] (2) Cost data standardization patents: such as "a method and system for identifying and standardizing construction engineering cost data", which constructs a multi-level cost data tag system and combines it with knowledge graph to achieve automated processing. However, the tag system is oriented towards data classification rather than entity identification and does not have global uniqueness and cross-platform parsing capabilities.

[0006] (3) Data collection and structuring patents: such as "a multi-level data collection method, system, device and storage medium", which generates standardized data entities through structured parsing and term mapping, relies on human experience judgment and rule matching, and has not established a systematic data identification and cross-platform recognition mechanism.

[0007] (4) Cross-system related patents: such as "a method and system for integrated audit and control of procurement supply chain and cost consulting", which generates a unified material identifier for multi-source material data, but the scope of the identifier is limited to materials and fails to cover all elements of engineering cost, and does not solve the problem of cross-platform parsing and identification.

[0008] (5) Component-level coding patents: For example, the "Cost Management Method Based on IFC Standard" proposes component instance coding, but the coding system is out of touch with industry standards and lacks connection with general identifier resolution systems such as Handle.

[0009] In summary, existing technologies generally suffer from the following shortcomings: inconsistent identifier granularity, lack of unified identifier coding standards and identifier system frameworks for different cost data elements; failure to effectively guarantee the global uniqueness of identifiers, with different platforms generating identifiers independently; insufficient cross-platform recognition capabilities, unable to resolve semantic mapping and cross-system association between data of different dimensions; lack of identifier registration, resolution, and management mechanisms; and lack of integration with the internationally accepted industrial internet identifier resolution system, making it difficult to integrate into the national industrial internet identifier resolution service system. Summary of the Invention

[0010] In view of the problems existing in the prior art, the present invention provides a method for parsing and uniquely identifying all elements of engineering cost data across platforms to solve the above-mentioned technical problems.

[0011] To achieve the above and other objectives, the technical solution adopted by the present invention is as follows:

[0012] This invention provides a method for parsing and uniquely identifying all elements of engineering cost data across platforms, comprising the following steps:

[0013] Step S1: Construct a comprehensive identification system for engineering cost elements

[0014] Establish a data classification framework for all elements of engineering cost, covering the following data element categories:

[0015] Human resource elements: job type, unit price of labor, and labor consumption;

[0016] Material resource elements: material name, specifications, unit of measurement, unit price, and consumption.

[0017] Mechanical resource elements: machine name, specifications, unit price per shift, and consumption per shift;

[0018] List of items includes: item code, item name, item characteristics, quantity, unit of measurement, and unit price.

[0019] Price information elements: material price information, information price, market price, construction cost index, indicators;

[0020] Engineering characteristic elements: Engineering-level characteristic information such as project type, structural form, building area, number of floors, and functional classification;

[0021] Spatiotemporal dimensional elements: time dimension (stage identifiers: estimation, preliminary estimate, budget, settlement, etc.) and spatial dimension (regional identifiers, project level identifiers, etc.).

[0022] For each of the above data elements, the rules for the composition of its identification code and the metadata structure are defined in accordance with national standards (such as GB / T 50851-2013 "Construction Engineering Labor, Material, Equipment and Machinery Data Standard") and industry specifications.

[0023] Step S2: Generate a globally unique identifier code

[0024] For engineering cost data objects, a hierarchical combined coding mechanism is used to generate globally unique identifier codes, with the following coding structure:

[0025] Object type code — Industry code — Enterprise / platform code — Classification code — Feature hash value — Timestamp code — Check code

[0026] in:

[0027] Object type code (2 digits): Identifies the category of the project cost data object (01-labor, 02-materials, 03-machinery, 04-list item, 05-price information, 06-project characteristics, etc.);

[0028] Industry code (4 digits): Assigned according to the national industry classification standard (GB / T 4754), with the construction industry corresponding to "E47"; aligned with the industry node identifier in the Handle system;

[0029] Enterprise / Platform Code (6 digits): An organization code uniformly assigned by the identification registration management agency;

[0030] Classification code (6-10 digits): generated according to national or industry standard classification systems;

[0031] Feature hash value (8 bits): The first 8 bits of a 32-bit hash value generated by hashing the feature attributes of a data object, achieving uniqueness based on attributes;

[0032] Timestamp code (6 digits): Date encoding in YYYYMMDD format, identifying the creation time or version time of the data object;

[0033] Check digit (2 bits): A check digit generated based on the aforementioned segments of encoding using the CRC check algorithm or the Luhn algorithm.

[0034] Each cost data object is assigned a globally unique identifier upon creation, which remains unchanged throughout its lifecycle. This identifier is compatible with the Handle system and is registered within the Handle identifier resolution framework.

[0035] Step S3: Establish a multi-level identifier resolution architecture

[0036] Construct a multi-level identifier resolution architecture based on the Industrial Internet Identifier Resolution System (Handle System), including:

[0037] Root parser node: Responsible for top-level management of global identifier resolution and maintaining the mapping relationship between industry codes and enterprise / platform codes;

[0038] Industry-specific resolution node: Corresponding to the construction industry's engineering cost field, deploying industry-level identifier registration and resolution services;

[0039] Enterprise resolution node: Deployed within entities such as engineering cost consulting firms, construction units, and contractors, responsible for local registration, local resolution, and cache management of enterprise-level identifiers;

[0040] Edge resolution node: Deployed on-site in engineering projects, it enables real-time resolution and lightweight caching of identifiers, and supports identifier verification and partial resolution in offline scenarios.

[0041] When an identifier resolution request is received, the resolution request is forwarded upwards level by level in the order of "edge layer → enterprise layer → industry layer → root layer" until the complete identifier information is obtained. The resolution result is returned level by level along the original path and cached.

[0042] Edge resolution nodes have a pre-parsing and loading function: the globally unique identifiers of frequently used data objects in the engineering cost project are pre-loaded into the local identifier cache of the edge resolution node. The edge node directly responds to the identifier query request without having to initiate parsing to the upper-level nodes. The node maintains a version number verification mechanism for the pre-loaded identifiers. When the source data corresponding to the pre-loaded identifier changes, the upper-level industry nodes push incremental updates to the edge node.

[0043] Step S4: Establish a cross-platform semantic mapping and unique identification mechanism

[0044] To address the data format, terminology, and encoding differences between different cost estimation business platforms, a cross-platform semantic mapping engine is constructed, specifically including:

[0045] (4a) Data Acquisition and Standardization: Engineering cost data is collected from different business platforms (including cost estimation platforms, pricing software platforms, BIM modeling platforms, material management systems, contract management systems, and ERP systems). Data formats include Excel, XML, JSON, proprietary formats of professional cost estimation software, direct database connection data, and unstructured documents. The collected data is standardized in format and cleaned with terminology according to a standard data dictionary.

[0046] (4b) Identifier Matching and Association: The standardized data is matched with the registered globally unique identifier codes. Matching strategies include: exact matching (direct matching based on the original code or identifier code), semantic matching (semantic similarity calculation of the name, features, and specifications of data objects based on natural language processing technology), and contextual matching (comprehensive judgment using multi-dimensional contextual features such as project information, regional information, and time information). For data objects that fail to match, the identifier generation mechanism in step S2 is invoked to assign a new globally unique identifier code and register it in the identifier resolution system.

[0047] (4c) Semantic mapping relationship construction: For data objects with matched identifiers, a cross-platform semantic mapping relationship table is established. The mapping relationship table records the following fields: source platform identifier, target platform identifier, globally unique identifier code, mapping confidence, mapping update time, and mapping version number. The mapping relationship is stored in the form of a knowledge graph, supporting reasoning for complex mapping relationships.

[0048] (4d) Two-way reinforcement mechanism for identifiers and semantic mapping: The registration information of globally unique identifiers and the establishment and change records of each cross-platform mapping relationship are written into the blockchain. Each cross-platform data mapping entry forms an immutable and trustworthy record on the chain. When performing data matching, the semantic mapping engine uses the trustworthy mapping record on the chain as prior knowledge, and pre-solidifies the matching results of vector pairs with existing successful mapping records. At the same time, it performs batch preprocessing on frequently occurring new matching scenarios to form a batch mapping cache. The sharing characteristics of blockchain records support the synchronization of mapping information across organizations, forming a closed-loop process of identifier registration-mapping establishment-semantic verification-trustworthy confirmation.

[0049] (4e) Cross-platform unique identification: When any business platform needs to identify or query a certain engineering cost data object, it sends a query request to the identifier resolution system through the platform's client or API interface, passing in the local identifier. The identifier resolution system performs cascading matching: first, it queries the mapping relationship in the local cache; if no match is found locally, it sends a resolution request to the enterprise node, industry node, and root node level by level; the resolution system returns the globally unique identifier code of the data object and its associated standardized data view.

[0050] (4f) Heterogeneous data format conversion: The mapping engine has a built-in data format converter, see reference.

[0051] The system uses the IFC (Industry Foundation Classes) and COBie (Construction Operations Building Information Exchange) standards to convert the parsed data into a format recognizable by the target platform. The system supports dynamic negotiation of multiple exchange protocols: it prioritizes RESTful APIs for data exchange, and automatically downgrades to XML / JSON file exchange or CSV table exchange if the target platform's interface capabilities are limited.

[0052] (4g) Collaborative optimization of pre-parsing and mapping cache: The pre-loaded identifier library of the edge parsing node and the batch mapping cache of the semantic mapping engine share the same storage channel. After the semantic mapping engine completes the batch matching and cache construction of a batch of data, the globally unique identifier code of the successful match, together with the corresponding standardized data view, is automatically pushed to the edge parsing node at the project site and incorporated into the pre-loaded identifier library; the identifier matching and aggregation processing of data is quickly completed through the batch mapping cache, and the real-time parsing capability of the edge parsing node provides services to the on-site business system.

[0053] (4h) Event-driven dynamic synchronization extension: For key data fields that need to be shared across platforms in real time, an event-driven synchronization mechanism based on message queues is adopted. When the identification information of a certain cost data object changes, the enterprise node or industry node in the region broadcasts the change event to all edge parsing nodes associated with the data object in the form of a message queue. Each edge node updates the corresponding identification entry in its local pre-loaded identification library according to the event type, and writes the synchronization record to the blockchain to form a trusted audit log.

[0054] Step S5: Identifier Lifecycle Management

[0055] Establish a comprehensive lifecycle management system for identifiers, including:

[0056] Identifier Registration: A globally unique identifier is generated when a data object is created, and a registration request is submitted to the enterprise resolution node. After the enterprise node verifies the uniqueness, it stores the identifier information in the local identifier library and synchronizes it to the industry nodes.

[0057] Identifier Update: When the content of a data object changes, a new version identifier record is generated through a version mechanism. The original version identifier record is retained in the history database. Different versions are distinguished by timestamp code segments, and version backtracking is achieved through version association pointers.

[0058] Deregistration: When a data object's lifecycle ends, its status is marked as "deregistered," and the identifier code itself is permanently retained and not reclaimed.

[0059] Identifier Query and Traceability: Provides a full-link query interface for identifiers, supporting the query of complete information, historical versions, and lineage of data objects based on identifier codes.

[0060] As described above, the method for parsing and uniquely identifying all elements of engineering cost data provided by this invention has at least the following beneficial effects:

[0061] This invention provides a method for parsing and uniquely identifying all elements of engineering cost data across platforms. By constructing a comprehensive identification system covering labor, materials, machinery, list items, price information, engineering characteristics, and spatiotemporal dimensions, and combining hierarchical combination coding, multi-level identification parsing architecture, semantic mapping and blockchain bidirectional reinforcement mechanism, edge pre-parsing and caching collaboration, and event-driven dynamic synchronization, it effectively solves the long-standing problems in the field of engineering cost, such as inconsistent identification of multi-source heterogeneous data, difficulty in cross-platform identification, and data silos.

[0062] On the one hand, the collaboration between globally unique identifier encoding and a multi-level resolution architecture significantly improves the accuracy and efficiency of cross-platform data identification and resolution, reducing the risk of business interruption due to identifier conflicts or resolution failures. On the other hand, the organic integration of semantic mapping and blockchain bidirectional reinforcement mechanisms, pre-parsing and mapping caching collaboration, and event-driven dynamic synchronization enables proactive identification and intelligent compensation for cross-platform data consistency risks, greatly improving the system's operational stability and financial security. Attached Figure Description

[0063] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 This is a schematic diagram showing the connections between the steps of the method of the present invention. Detailed Implementation

[0065] The following description, in conjunction with the implementation of this invention, is merely an example and illustration of the concept of this invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the inventive concept or exceed the scope defined in these claims, all of which should fall within the protection scope of this invention.

[0066] Example

[0067] Step S1: Construct a comprehensive identification system for engineering cost elements

[0068] Based on step S1, a comprehensive identification system for engineering cost elements is constructed. A data classification framework is established, covering labor resource elements, material resource elements, machinery resource elements, list item elements, price information elements, engineering characteristic elements, and spatiotemporal dimension elements. For each type of data element, referring to GB / T 50851-2013 "Standard for Data of Labor, Materials, Equipment and Machinery in Construction Projects" and relevant industry standards, the rules for identifying and coding are defined, including the metadata structure of each element, key attribute fields and their data types, and value ranges. The above identification specifications are stored in the identification system construction module, forming a metadata dictionary that can be called upon for subsequent coding generation.

[0069] Step S2: Generate a globally unique identifier code

[0070] Based on step S2, for the engineering cost data object, a hierarchical combination coding mechanism is invoked to generate a globally unique identifier code. As a specific example, taking a material data object (material name: HRB400E grade steel bar, specification: 25mm) as an example, the code is generated as follows:

[0071] Object type code: "02" corresponds to the material element;

[0072] Industry code: "E470" corresponds to the construction industry;

[0073] Enterprise / Platform Code: Assigned by the identifier registration management agency; example, "440101";

[0074] Classification code: Based on the steel reinforcement material classification in GB / T 50851-2013, take "01010101";

[0075] Feature hash value: Perform a hash operation on the feature attributes such as "HRB400E" and "25mm", and take the first 8 bits of the 32-bit hash value to get "A3B7E92F";

[0076] Timestamp code: Retrieves the date the data was generated, in YYYYMMDD format, for example, "20251115";

[0077] Check code: Calculated using the CRC check algorithm based on the aforementioned segments, resulting in "C8".

[0078] The complete encoding is:

[0079] 02—E470—440101—01010101—A3B7E92F—20251115—C8. This code is used as the globally unique identifier for this data object and is registered in the Handle identifier resolution system, giving it global resolution capabilities. This code remains unchanged throughout the entire lifecycle of the data object.

[0080] Step S3: Establish a multi-level identifier resolution architecture

[0081] Following step S3, deploy a multi-level identifier resolution architecture. The root resolution node is deployed at the national-level top-level identifier resolution node, maintaining the mapping relationship between industry codes and enterprise codes. Industry resolution nodes are deployed in the construction industry's cost estimation field, connecting upwards to the root resolution node and downwards to the resolution nodes of various enterprises. Enterprise resolution nodes are deployed within each cost estimation business entity, responsible for the local registration, resolution, and caching management of enterprise-level identifiers. Edge resolution nodes are deployed at project sites, possessing identifier pre-resolution and loading capabilities.

[0082] Specifically, by analyzing the data call frequency in the engineering cost dataset, frequently used data objects (such as the top N types of materials and list items accounting for more than 80% of the total queries) are identified, and their globally unique identifiers are pre-loaded into the local identifier cache of the edge resolution nodes. The edge nodes maintain a version number verification mechanism for the pre-loaded identifiers: each pre-loaded entry records its version number and the version identifier of the source data. The edge nodes periodically or on demand send version verification requests to the enterprise nodes. When the source data corresponding to the pre-loaded identifier changes, the upper-level industry nodes or enterprise nodes actively push incremental updates (including update type, new version number, and the changed standardized data view) to the edge nodes. Upon receiving the update, the edge nodes verify the version number and complete local cache synchronization. When an identifier resolution request is received, the resolution request is forwarded upwards level by level: edge layer → enterprise layer → industry layer → root layer. The edge nodes first query the local identifier cache; if a match is found, the result is returned directly; otherwise, it is forwarded to the upper-level node.

[0083] Step S4: Establish a cross-platform semantic mapping and unique identification mechanism

[0084] Based on step S4, perform the following sub-steps:

[0085] (4a) Data Acquisition and Standardization: Engineering cost data is collected from multiple cost business platforms (including pricing platforms, BIM quantity calculation platforms, and material management platforms). The data formats include XML, JSON, Excel, and proprietary formats of professional software. The collected data is standardized and cleaned according to a predefined standard data dictionary, and the variant expressions of the same term on different platforms are mapped to standard term expressions.

[0086] (4b) Identifier Matching and Association: The standardized data is matched with the registered globally unique identifier codes. Matching strategies include: exact matching (directly comparing local codes or identifier codes), semantic matching (using a BERT pre-trained model to vectorize the name, features, and specifications of data objects, calculating cosine similarity, and determining a successful match if the similarity is higher than 0.85), and contextual matching (comprehensively using contextual features such as project stage, region, and engineering type for a comprehensive judgment). For data objects that fail to match, the identifier generation mechanism in step S2 is invoked to assign a new globally unique identifier code and register it in the identifier resolution system.

[0087] (4c) Semantic mapping relationship construction: For data objects with matched identifiers, a cross-platform semantic mapping relationship table is established to record the source platform identifier, target platform identifier, globally unique identifier code, mapping confidence, mapping update time, and mapping version number. The mapping relationship is stored in the form of a knowledge graph, where nodes represent data object identifiers and their local representations on each platform, and edges represent mapping relationships and confidence weights, supporting graph traversal reasoning for complex mapping relationships.

[0088] (4d) Two-way reinforcement mechanism for identifiers and semantic mapping: The registration information of each identifier and the establishment and change records of each mapping relationship are written into the blockchain to form an immutable and trustworthy record. When performing data matching, the semantic mapping engine first checks whether there is a trustworthy mapping record in the blockchain related to the current data object to be matched: if it exists and the confidence level reaches the threshold, the matching is completed directly using the record (skipping the semantic calculation step); if it does not exist, regular semantic matching is performed and the matching result is included in the mapping library after being written into the blockchain. At the same time, for frequently occurring new matching scenarios (such as multiple similar materials imported in the same batch), the engine performs batch preprocessing: the feature vectors of multiple objects to be matched are combined into a matrix for parallel calculation, and a batch mapping cache is generated at once to avoid the overhead of calculating each item. The sharing characteristics of blockchain records support the synchronization of mapping information between different enterprise nodes: when a mapping relationship is solidified on an enterprise node, other enterprise nodes can synchronize the mapping record through the blockchain as prior knowledge to continuously optimize the global mapping accuracy and form a closed-loop reinforcement process of identifier registration → mapping establishment → semantic verification → trustworthy confirmation.

[0089] (4e) Cross-platform unique identification: When any business platform initiates an identification request for a certain engineering cost data object, it passes a local identifier (which can be a local code, name description, or a combination of attributes) to the identifier resolution system. The identifier resolution system performs cascading matching: First, it queries the mapping relationship in the local cache (including edge node cache and enterprise node cache); if no match is found, it initiates a resolution request to the enterprise node, industry node, and root node level by level; the resolution system returns the globally unique identifier code of the data object and its associated standardized data view.

[0090] (4f) Heterogeneous data format conversion: The mapping engine has a built-in data format converter that converts the parsed data into a format recognizable by the target platform, referring to the data models defined by the IFC and COBie standards. The system supports multi-protocol dynamic negotiation: It prioritizes JSON format data exchange via the RESTful API; when the target platform's interface capabilities are limited (such as only supporting file import), it automatically downgrades to XML file exchange or CSV table exchange, and completes asynchronous transmission through a message queue.

[0091] (4g) Cooperative optimization of pre-parsing and mapping caching: The preloaded identifier library of the edge parsing node and the batch mapping cache of the semantic mapping engine share the same storage channel. After the semantic mapping engine completes the batch matching of a batch of data objects, it automatically pushes the globally unique identifier code of the successful match along with the corresponding standardized data view to the edge parsing node, and the edge node incorporates this information into the preloaded identifier library. Subsequent queries on this batch of data objects can be directly responded to by the edge node without calling the semantic mapping engine again. At the same time, when the edge node receives a query request for a new data object, if the corresponding identifier does not exist in the local preloaded library, it will feed back the query information to the semantic mapping engine, triggering the engine to perform targeted matching and cache construction for this type of data object. Through this two-way feedback mechanism, identifier parsing and semantic mapping form a mutually reinforcing closed loop: the preloaded identifier library provides a fast verification channel for high-frequency data for semantic mapping, and the semantic mapping cache provides a content update source for the preloaded identifier library.

[0092] (4h) Event-Driven Dynamic Synchronization Extension: For key data fields that require real-time cross-platform sharing (such as material unit price, comprehensive unit price of list items), event-driven triggers are set. When the identification information of a data object changes (such as price update, feature modification), the enterprise node or industry node in the region encapsulates the change event into a message (including globally unique identifier code, change type, change details, timestamp, and version number), and broadcasts it to all edge parsing nodes that subscribe to the identifier through the message queue middleware. After receiving the event, each edge node synchronously updates the corresponding identifier entry in its local pre-loaded identifier library according to the event type (updates the data view, increments the version number). Every event broadcast and node synchronization record is written to the blockchain to form a trusted audit log for subsequent consistency verification and dispute tracing. The delay from the occurrence of a change event to the completion of synchronization by all related edge nodes is controlled within an acceptable engineering range (e.g., seconds).

[0093] Step S5: Identifier Lifecycle Management

[0094] Based on step S5, perform full lifecycle management of the identifier:

[0095] Identifier Registration: When a data object is created, step S2 is called to generate a globally unique identifier code, and a registration request is submitted to the enterprise resolution node. After verifying the uniqueness of the code, the enterprise node stores the identifier information in its local identifier library and synchronizes it to the industry nodes; for data objects that need to be shared across industries, the industry nodes continue to report to the root node.

[0096] Identifier Update: When the content of a data object changes, a new version of the identifier record is generated: the timestamp code segment in the original identifier code is updated to the change date, a new checksum is generated, and the original version identifier record is retained in the history database. Version backtracking is achieved through a version-associated pointer (pointing to the storage address or code of the previous version identifier). The identifier update event also triggers the event-driven synchronization mechanism in step S4.

[0097] Deregistration: When a data object reaches the end of its lifecycle, it is marked as "deregistered". The identifier code itself is permanently retained and not reclaimed to ensure the traceability of historical data. The deregistration operation is written to the blockchain to form a permanent record.

[0098] Identifier Query and Traceability: Provides a full-link query interface for identifiers, supporting queries based on identifier codes for complete information of data objects (including all attributes of the current version, standardized data view), historical version list and change records, and lineage (the source platform of the data object and which other data objects reference it).

[0099] Verification results:

[0100] Cross-platform data recognition accuracy About 67% 99.3% Cross-platform data query average response time 3.2 seconds 0.15 seconds Data redundancy storage rate 38% 5% Data consistency maintenance time (per week) 4.5 hours 0.2 hours New material first matching average processing time — More than 80% lower than the initial Cross-project mapping information reuse rate — 82%

[0101] The above method was used to test the full-element dataset of engineering cost. The test dataset contains cost data objects (including categories such as labor, materials, machinery, bill of quantities items, and price information, totaling approximately 5000 data objects) from multiple business platforms. The test results are as follows:

[0102] Test results demonstrate that the method of this invention achieves significant technical improvements in cross-platform identification accuracy, query response efficiency, data redundancy control, and consistency maintenance. The collaborative mechanism of edge node pre-parsing loading and semantic mapping caching effectively reduces identifier parsing latency and semantic matching computation overhead; the blockchain bidirectional reinforcement mechanism continuously improves mapping accuracy and reusability; and the event-driven synchronization mechanism ensures the consistency of identifier states in a distributed environment. The feasibility and superiority of the method of this invention have been fully verified.

[0103] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0104] It should be understood that determining B based on A does not mean determining B solely based on A; it also means determining B based on A and / or other information.

[0105] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0106] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for parsing and uniquely identifying all elements of engineering cost data across platforms, characterized in that: Includes the following steps: Step S1: Construct a full-element identification system for engineering cost, which includes labor resource elements, material resource elements, mechanical resource elements, list item elements, price information elements, engineering characteristic elements, and spatiotemporal dimension elements. Define identification coding rules and metadata structure for each type of element according to national standards and industry specifications. Step S2: Generate a globally unique identifier code. A hierarchical combination coding mechanism is used to generate a globally unique identifier code for the engineering cost data object. The coding structure is: object type code—industry code—enterprise / platform code—classification code—feature hash value—timestamp code—check code; the feature hash value is the first 8 bits after hashing the feature attributes of the data object, and the timestamp code is used to identify the generation or version time of the data object. Step S3: Establish a multi-level identifier resolution architecture, including root resolution nodes, industry resolution nodes, enterprise resolution nodes, and edge resolution nodes. Each node is organized hierarchically, and resolution requests are forwarded upwards level by level from the edge layer to the enterprise layer, industry layer, and root layer until complete identifier information is obtained. The edge resolution nodes have an identifier pre-resolution loading function, which pre-loads the globally unique identifier codes of frequently used data objects in the project into the local identifier cache library of the edge resolution nodes, and the upper-layer nodes push incremental updates to maintain version consistency. Step S4: Establish a cross-platform semantic mapping and unique identification mechanism, including data collection and standardization, identifier matching and association, construction of a semantic mapping relationship table, and cross-platform unique identification; wherein, the registration information of the identifier encoding and the establishment and change records of the mapping relationship are written into the blockchain to form an immutable and trustworthy record. When the semantic mapping engine performs data matching, it uses the trustworthy mapping record on the chain as prior knowledge to participate in the judgment, and performs batch preprocessing on frequently occurring new matching scenarios to form a batch mapping cache; Step S5: Perform identifier lifecycle management, including identifier registration, update, cancellation, query and traceability.

2. The method according to claim 1, characterized in that, In step S2, the object type code occupies 2 digits, the industry code occupies 4 digits, the enterprise / platform code occupies 6 digits, the classification code occupies 6-10 digits, the feature hash value occupies 8 digits, the timestamp code occupies 6 digits, and the verification code occupies 2 digits; the globally unique identifier code is compatible with the Handle system and is registered in the Industrial Internet Identifier Resolution System.

3. The method according to claim 1, characterized in that, The root resolution node in step S3 is responsible for the top-level management of global identifier resolution and maintains the mapping relationship between industry codes and enterprise codes; the industry resolution node corresponds to the construction industry engineering cost field and deploys industry-level identifier registration and resolution services. Enterprise resolution nodes are deployed within each cost engineering business entity, responsible for the local registration and resolution of enterprise-level identifiers; edge resolution nodes are deployed at the project site, supporting identifier verification and partial resolution in offline scenarios. The edge resolution node has a self-maintained version number verification mechanism for the preloaded identifier. When the source data corresponding to the preloaded identifier changes, the upper-layer industry node pushes an incremental update to the edge node.

4. The method according to claim 1, characterized in that, The semantic matching in step S4 is based on natural language processing technology to calculate the semantic similarity of the name, features and specifications of the data objects. If the similarity is higher than a preset threshold, the match is considered successful. The semantic mapping relationship is stored in the form of a knowledge graph, which supports reasoning of complex mapping relationships. Step S4 also includes a data format conversion step, which converts the parsed data into a format that the target platform can recognize through a built-in data format converter. The system supports dynamic negotiation of multiple exchange protocols, prioritizes data exchange using RESTful API, and automatically downgrades to XML file exchange or CSV table exchange according to the interface capabilities of the target platform.

5. The method according to claim 1, characterized in that, The pre-parsing and loading of identifiers and the semantic mapping cache form a collaborative optimization: the pre-loaded identifier library of the edge parsing node and the batch mapping cache in step S4 share the same storage channel. After the semantic mapping engine completes the batch matching and cache construction of a batch of data, the globally unique identifier code of the successful match, together with the corresponding standardized data view, is automatically pushed to the edge parsing node at the project site and incorporated into the pre-loaded identifier library. After the identifier matching and aggregation of data are completed through the batch mapping cache, the real-time parsing capability of the edge parsing node provides services to the on-site business system.

6. The method according to claim 1, characterized in that, The cross-platform unique identification in step S4 adopts an event-driven dynamic synchronization mechanism: for key data fields that need to be shared across platforms in real time, an event-driven synchronization mechanism based on message queues is adopted. When the identification information of a certain cost data object changes, the enterprise node or industry node in the region will broadcast the change event to all edge parsing nodes associated with the data object via a message queue. Each edge node will synchronously update the corresponding identification entry in its local pre-loaded identification library according to the event type, and at the same time write the synchronous record to the blockchain to form a trusted audit log.

7. The method according to claim 1, characterized in that, The identifier update in step S5 is implemented through a version mechanism. When a new version identifier record is generated, the original version identifier record is retained and a version association pointer is established. The timestamp code segment in the identifier code is used to distinguish different versions. When an identifier is cancelled, the identifier status is marked as "cancelled". The identifier code itself is permanently retained and not recycled.

8. The method according to claim 1, characterized in that, Step S4 also includes a cascading matching process for cross-platform unique identification: when any business platform initiates an identification request, the identifier resolution system first queries the mapping relationship in the local cache; if no match is found locally, it initiates a resolution request to the enterprise node, industry node and root node level by level. The parsing system returns a globally unique identifier for the data object and its associated standardized data view; The parsing results are returned and cached level by level along the original path.

9. A system for parsing and uniquely identifying all elements of engineering cost data across platforms, used to implement the method described in any one of claims 1-8, characterized in that, include: The identification system construction module is used to build a comprehensive identification system for engineering cost elements and store identification metadata specifications. The identifier encoding generation module is used to generate globally unique identifier codes that conform to the Handle system compatibility specification; The multi-level identifier resolution module includes root resolution nodes, industry resolution nodes, enterprise resolution nodes, and edge resolution nodes. The nodes communicate with each other through standard interfaces to jointly provide identifier registration and resolution services. The semantic mapping module is used for data matching, terminology normalization, and mapping relationship construction between different business platforms. The semantic mapping module is connected to the blockchain record layer, writes the mapping records into the blockchain and reads the trusted mapping records on the chain as prior knowledge. Cross-platform communication middleware, including API gateway, data format converter and message queue service, is used to connect different cost estimation business platforms and perform data format conversion and synchronization; The lifecycle management module is used for the registration, updating, deregistration, querying, and tracing of identifiers.