Chemical engineering cost data cleaning method and system

By constructing a multimodal database and integrating knowledge graphs and deep learning models, the chemical engineering cost data cleaning method and system solve the problems of multi-source heterogeneity and semantic fragmentation in chemical engineering cost data. It achieves accurate alignment and dynamic updates of cross-professional terms, improves the accuracy and timeliness of data cleaning, and provides high-quality data support for intelligent cost analysis.

CN120910031APending Publication Date: 2025-11-07CHINA TIANCHEN ENGINEERING CORPORATION LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510782865.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Chemical cost data suffers from differences in terminology across different professional fields, scattered data sources, heterogeneous formats, and frequent semantic conflicts. Existing technologies struggle to achieve cross-professional terminology mapping, dynamic updates, and complex semantic scenario parsing, resulting in insufficient timeliness and accuracy of data cleaning, weak unstructured data processing capabilities, and frequent implicit logical errors.

Method used

A multimodal database is constructed, integrating domain knowledge graphs and deep learning models. Through a cross-professional terminology mapping library and a dual verification mechanism, the standardization and dynamic updating of chemical cost data are achieved, including the transformation of unstructured data, terminology feature extraction and intelligent cleaning. A cloud-edge collaborative architecture is adopted to ensure data privacy and cross-enterprise collaborative optimization.

Benefits of technology

It has achieved unified collection and standardized integration of fragmented features of chemical cost data across disciplines, media, and time periods, improving the accuracy and timeliness of data identification, reducing hidden errors, and providing a high-quality foundation for intelligent cost analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910031A_ABST
    Figure CN120910031A_ABST
Patent Text Reader

Abstract

The invention discloses a chemical engineering cost data cleaning method and system, and the method comprises the steps: S1, gathering heterogeneous data sources of a plurality of specialties in the chemical engineering field, and forming a multi-modal database, S2, converting unstructured data and real-time streaming data in the multi-modal database into structured data, S3, extracting terms of the multi-modal database in the S2, and S4, extracting the terms of the multi-modal database in the S2, constructing a dynamically updated cross-terminology mapping library; s4, intelligently cleaning terms in the cross-terminology mapping library and outputting a standardized term library; and S5, adopting a dual verification mechanism for the standardized term library to improve the recognition accuracy. According to the chemical cost data cleaning method and system provided by the invention, the multi-modal database is constructed, accurate alignment and dynamic updating of cross-terminology are realized by fusing the domain knowledge graph and the deep learning model, and the standardized database is generated, so that the industrial problems of multi-source isomerism and semantic segmentation of the chemical cost data are effectively solved; and a high-quality data basis is provided for intelligent cost analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a chemical engineering cost data cleaning method and system. BACKGROUND

[0002] In the field of chemical engineering cost, data cleaning is the core basic work to realize accurate cost control and resource optimization. With the transformation of the chemical industry to intelligent and intensive, enterprises need to integrate massive cost data from multiple professional fields such as process design, equipment procurement, construction installation, and financial accounting. However, due to the differences in terminology expression of the same concept by different professionals (for example, "anti-corrosion engineering quantity" may be expressed as "surface treatment engineering quantity" in the civil engineering profession, and may be associated with "anti-corrosion layer thickness parameter" in the equipment profession), combined with scattered data sources (bidding documents, completion settlement reports, equipment manuals, etc.), heterogeneous formats (structured tables, unstructured texts, image scans, etc.), the data homogeneity is low and semantic conflicts occur frequently. According to industry statistics, about 37-50% of the cost deviation is caused by data misreading due to insufficient terminology standardization. In addition, the traditional method has significant defects in dealing with chemical specific scenarios: it lacks deep analysis ability for key elements such as dimension unit conversion (such as the implicit association of "ton·kilometer" and "cubic meter / hour") and process parameter semantics (such as the difference in grading standards of "pressure rating" in different equipment), resulting in more than 30% of implicit errors in the cleaned data. At the same time, the chemical industry data updates and iterates quickly (such as new carbon capture device cost indicators), and the existing static knowledge base is difficult to dynamically adapt to the evolution of terminology, while emerging technologies such as AI-professional terminology library learning models have not been deeply integrated with domain knowledge, which restricts the timeliness and accuracy of data cleaning.

[0003] The current technical solution has the following significant defects: 1. Insufficient cross-professional term mapping capability: different professional fields have different expressions for the same concept (such as "anti-corrosion engineering quantity" in civil engineering and equipment), which lacks dynamic semantic association, leading to data misreading. 2. Reliance on static rule base and template: using fixed templates or user-defined rules makes it difficult to adapt to dynamically changing terms and scenarios. The pricing items of new devices (such as electrochemical adsorption equipment) cannot be covered by existing templates, leading to data omission or mislabeling. The static rule base has a long update cycle (usually annual update), which cannot match the average iteration speed of new chemical device cost indicators. 3. Lack of complex semantic scene analysis: unable to analyze the implicit association between cross-modal parameters (such as the implicit association between "tons-kilometers" and "cubic meters / hour" related to equipment energy efficiency). Traditional models only handle surface data formats (such as table structure), but cannot understand process parameter classification standards (such as the difference in "pressure rating" definition in reaction kettles and pipes). 4. Conflict between dynamic knowledge update and privacy protection: the lack of deep integration of AI-professional term library learning model and domain knowledge through collaborative technologies leads to privacy leakage risk when updating cross-enterprise term libraries. The knowledge base has significant lag, and the error rate of new device cost analysis is 8-12% higher than the international benchmark (6%). 5. Weak unstructured data processing capability: there are blind spots in semantic analysis of unstructured text such as drawing annotations and handwritten corrections, relying on manual supplementation. Existing OCR and NLP technologies cannot combine with chemical industry knowledge (such as the influence of material density variables on dimension conversion), leading to implicit logic errors (such as "tons" directly converted to "cubic meters"). SUMMARY

[0004] The present application aims to solve the problems in the prior art and discloses a chemical cost data cleaning method and system, which constructs a multi-modal database, realizes accurate alignment and dynamic update of cross-professional terms by fusing domain knowledge graph and deep learning model, and generates a standardized database, effectively solving the industry problems of "multi-source heterogeneous, semantic fragmentation" of chemical cost data.

[0005] The present application is realized by the following technical solutions:

[0006] The present application first provides a chemical cost data cleaning method, comprising the following steps:

[0007] S1. Collecting and forming a multi-modal database from multiple professional heterogeneous data sources in the chemical field,

[0008] S2. Converting unstructured data and real-time stream data in the multi-modal database into structured data,

[0009] S3. Extracting terms from the multi-modal database in S2 and constructing a dynamically updated cross-professional term mapping library;

[0010] S4. Intelligent cleaning of terms in the cross-professional term mapping library and output of a standardized term library;

[0011] S5. Using a double verification mechanism on the standardized term library to improve recognition accuracy.

[0012] As a further solution, the multi-modal database comprises structured data, unstructured data and real-time streaming data,

[0013] The structured data is data with fixed format and explicit organization, and is convenient for computer program processing and query;

[0014] The unstructured data includes handwritten characters in electronic documents, paper documents or paper files, and image data;

[0015] The real-time streaming data includes audio data and video data.

[0016] As a further solution, the method for converting unstructured data and real-time streaming data into structured data in S2 is:

[0017] S21. Cleaning and converting electronic documents in the multi-modal database into structured text;

[0018] S22. Parsing and converting image data in the multi-modal database into structured text;

[0019] S23. Transcribing and converting real-time streaming data in the multi-modal database into structured text.

[0020] As a further solution, the specific method for extracting term features of the multi-modal database in S2 in S3 is:

[0021] S311. Identifying physical quantity units in the multi-modal database terms, and mapping and converting the physical quantity units with the international standard unit library to form a term mapping library;

[0022] S312. Based on the term mapping library, constructing a knowledge graph in the chemical field based on the chemical process flow, and extracting process parameters of the terms in the knowledge graph;

[0023] S313. Analyzing the association relationship of the terms in the knowledge graph in the engineering quantity list, technical specification book and other documents through a graph neural network, establishing a cross-document term network, and forming a cross-document term mapping library.

[0024] As a further solution, the method for constructing a dynamically updated cross-professional term mapping library in S3 is:

[0025] S321. Extract historical term data from the cross-document term mapping library for determining whether the newly emerged term is a new term;

[0026] S322. Automatically incorporate the newly emerged term and update the mapping conversion result of the term and the international standard unit library through the incremental learning framework, and simultaneously update the cross-document term mapping library in real time;

[0027] S323. When the definitions of the same term by different professions conflict, the priority of the term definition is determined according to the conflict resolution rule to form a cross-professional term mapping library.

[0028] The conflict resolution rule is based on industry priority and weight coefficient.

[0029] As a further scheme, the specific method of S4 of intelligently cleaning and outputting the standardized term library is:

[0030] S41. Adopting Levenshtein distance algorithm and Jaccard similarity algorithm to perform spelling correction and preliminary deduplication on the term;

[0031] S42. Combining the ontology relationship in the knowledge graph, the context meaning of the polysemous word is analyzed to eliminate ambiguity.

[0032] S43. According to the “Construction Engineering Quantity List Valuation Specification” and the chemical industry quota standard, a quasi-standardized term library with version identification is generated, and the quasi-standardized term library contains standardized term codes.

[0033] As a further scheme, the double verification mechanism includes manual verification and AI verification,

[0034] The method of manual verification includes: setting up a multi-level auditing process, sampling and reviewing the cleaning results by cost engineers, and labeling controversial term cases, and finally adding the term name recognized by everyone to the standardized term library;

[0035] The method of AI verification includes: cross-comparing the semantic analysis result of the AI model with the standardized term library to form effective feedback, and confirming the final name of the term in the standardized term library.

[0036] As a further scheme, the chemical cost data cleaning method further includes filtering and cleaning sensitive words in the standardized term library, automatically shielding terms related to business secrets, and only retaining the cleaned data after desensitization.

[0037] The present application also provides a chemical cost data cleaning system which adopts the chemical cost data cleaning method, and the system comprises:

[0038] A data collection module is configured to collect multi-professional heterogeneous data sources and form a multi-modal database;

[0039] A multi-modal data processing module is configured to convert different modal data in the multi-modal database into structured data for facilitating computer program processing and query;

[0040] A term feature extraction module is configured to extract context semantic features, dimensional attributes and industry correlation of terms in the structured data, and construct a dynamically updated cross-professional term mapping library;

[0041] An intelligent cleaning engine module is configured to perform fuzzy matching, synonym disambiguation and context correlation analysis on the terms in the cross-professional term mapping library, and output a standardized term library;

[0042] A transfer learning and dynamic evaluation module is configured to train a transfer learning model using historical project cleaning data, quickly adapt to the term system of new engineering types, and automatically evaluate the influence of term cleaning by the intelligent cleaning engine module on cost indicators, and generate an optimization suggestion report;

[0043] A feedback optimization module is configured to improve the recognition accuracy of the standardized term library through a double verification mechanism;

[0044] A permission management module is configured to establish a three-level user permission system and perform hierarchical verification on the permissions of the standardized term library;

[0045] A traceability module is configured to record the modification log of the standardized term library, and support version rollback and audit tracking;

[0046] A visualization verification module is configured to display the relationship graph of terms in the standardized term library, the cleaning effect heat map and abnormal term early warning, and support correction of the hierarchical standard deviation of core parameters through an interactive interface;

[0047] As a further scheme, the system is deployed in a cloud-edge collaborative architecture, including

[0048] An edge node is configured to complete real-time data cleaning on a local server, and support low-latency on-site data verification;

[0049] A cloud center is configured to gather cost data of the whole company, realize joint optimization of cross-enterprise term libraries through the chemical cost data cleaning method, and ensure data privacy;

[0050] A blockchain storage is configured to hash and chain the cleaning records of key terms in the standardized term library, and ensure data traceability and compliance.

[0051] Compared with the prior art, the chemical cost data cleaning method and system have the following advantages:

[0052] In the aspect of chemical cost data cleaning method:

[0053] (1) The chemical cost data cleaning method described in the application constructs a multi-modal database, realizes accurate alignment and dynamic update of cross-specialty terms by fusing a domain knowledge graph and a deep learning model, and generates a standardized database, effectively solving the industry problem of "multi-source heterogeneous, semantic fragmentation" of chemical cost data, and providing a high-quality data basis for intelligent cost analysis.

[0054] (2) The chemical cost data cleaning method described in the application simultaneously collects the fragmented characteristics of cross-specialty, cross-carrier and cross-time-effect of chemical cost data, and establishes a unified multi-modal database, effectively solving the standardization integration problem of multi-source heterogeneous data.

[0055] (3) In the intelligent cleaning process of the chemical cost data cleaning method described in the application, errors are corrected, duplicates are removed, and sensitive terms are desensitized, which is beneficial to the lightweight standardized term library.

[0056] (4) The chemical cost data cleaning method described in the application uses a double verification mechanism to improve the recognition accuracy of the standardized term library.

[0057] In the aspect of chemical cost data cleaning system:

[0058] (1) The chemical cost data cleaning system described in the application realizes cross-enterprise term library collaborative optimization under the premise of protecting data privacy through the chemical cost data cleaning method and blockchain technology.

[0059] (2) The chemical cost data cleaning system described in the application can verify the permission classification of the standardized term library, and can correct the classification standard deviation of the core parameters in the standardized term library through a man-machine interaction interface. BRIEF DESCRIPTION OF DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0061] Figure 1 The flowchart of the chemical cost data cleaning method described in the embodiments of the application;

[0062] Figure 2 The principle block diagram of the chemical cost data cleaning system described in the embodiments of the application. DETAILED DESCRIPTION

[0063] For the purpose of understanding the present application, the present application will be described below more fully with the embodiments of the present application, but the scope of the present application is not limited thereby.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the specification herein is for describing particular embodiments only and is not intended to be limiting of the application. Unless otherwise defined, all terms used in disclosing the application, including technical and scientific terms, terms used in describing the application, terms commonly used, and terms specifically defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the application pertains. It will be understood that the terms used herein are intended to describe the present application and not to limit the application, unless otherwise defined in the specification and claims. In describing and claiming the present application, the following terminology will be used in accordance with the definitions set out below. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. "Set" means one or more.

[0065] In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, can be fixedly connected, can be detachably connected, or integrally connected; can be mechanically connected, can be electrically connected; can be directly connected, can be indirectly connected through an intermediate medium, and can be the communication between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0066] Noun explanation:

[0067] Levenshtein distance: Levenshtein distance is a kind of edit distance. It refers to the minimum number of editing operations required to convert one string into another. The allowed editing operations include replacing one character with another character, inserting a character, and deleting a character.

[0068] Jaccard similarity: Jaccard Similarity is a statistical measure used to compare the similarity and diversity between limited sample sets.

[0069] A chemical cost data cleaning method, as shown in Figure 1 includes the following steps:

[0070] S1. A plurality of professional heterogeneous data sources in the chemical field are combined to establish a multi-modal database,

[0071] In one embodiment, the plurality of professionals includes civil engineering, installation, process, electrical, etc.

[0072] The multi-modal database includes structured data, unstructured data, and real-time streaming data.

[0073] The structured data refers to data with fixed format and explicit organization. This type of data is highly organized and easy to process and query by computer programs.

[0074] The unstructured data is a non-structured document, specifically including Excel, Word, etc. Electronic documents, handwritten text in paper bid documents or paper files, image data (such as PDF signed process data sheet), etc.

[0075] The real-time streaming data includes audio data and video data, such as conference recordings, on-site audio and video related to cost evidence files, etc.

[0076] This method simultaneously collects the fragmented characteristics of chemical cost data across disciplines (process / equipment / finance), across carriers (tables / texts / images), and across time (historical data / real-time indicators), and establishes a unified multi-modal database, effectively solving the problem of standardized integration of multi-source heterogeneous data.

[0077] S2. Identify, read, write, and convert the unstructured data and real-time streaming data in the multi-modal database into structured data;

[0078] The S2 conversion to structured data specifically includes the following methods:

[0079] S21. Clean and convert electronic documents in the multi-modal database to structured text;

[0080] In one embodiment, the cost table in Excel is aligned by rows and columns, the merged cells are parsed and the formula is checked, and it is converted to structured text and cleaned;

[0081] S22. Analyze and convert image data in the multi-modal database to structured text;

[0082] In one embodiment, the handwritten terms in the paper bid are extracted by OCR technology, and similarity matching is performed with the multi-modal database terms;

[0083] S23. Transcribe and convert real-time streaming data in the multi-modal database into structured text;

[0084] In one embodiment, the colloquial terms in the conference recording, such as "pipe main material price", are converted into structured text and cleaned.

[0085] The method converts unstructured data and real-time streaming data in the multi-modal database into unified structured text, unifies the standards, and facilitates the subsequent unified processing of data in the multi-modal database.

[0086] S3. Extract the terms of the multi-modal database in S2, and construct a dynamically updated cross-professional term mapping library;

[0087] Specifically, the specific method for extracting the term features of the multi-modal database in S2 in S3 is:

[0088] S311. Identify the physical quantity units in the multi-modal database terms, and map and convert the physical quantity units with the international standard unit library to form a term mapping library;

[0089] In one embodiment, the non-international unit (SI) "yuan / m 3 " is mapped and converted with the international standard unit library as follows:

[0090] A1. Term analysis

[0091] "Yuan / m 3 ": cost / cubic meter → monetary unit (yuan) divided by volume unit (m 3 ).

[0092] A2. International unit (SI) reference

[0093] SI base units:

[0094] Length: meter (m)

[0095] Currency: no international standard unit (usually use local currency code, such as CNY Renminbi)

[0096] A3. Conversion of composite units

[0097] "Yuan / m 3 ":

[0098] Currency: yuan (non-SI unit, need to mark the currency type, such as CNY).

[0099] Volume: cubic meter (m 3 , SI derived unit).

[0100] Mapping result: CNY / m 3 (Money unit cannot be converted, but volume is SI unit).

[0101] A4. Non-SI unit processing

[0102] The currency is a non-convertible unit, and the original unit is retained, but it must be clearly marked as currency: yuan, CNY.

[0103] A5. Complete mapping table, as shown in Table 1

[0104] Table 1

[0105] Original physical quantity unit Type of physical quantity SI unit mapping Remarks m2 / m2 3 ]] Cost density CNY / m 3 ]] Currency unit needs to be specified type

[0106] In another embodiment, the method for mapping and converting the international unit (SI) "ton-kilometer" with the International System of Units library is:

[0107] B1. Term resolution

[0108] "Ton-kilometer": mass unit (ton) multiplied by distance unit (kilometer) → measure of freight volume or transportation work.

[0109] B2. International unit (SI) comparison

[0110] SI base units:

[0111] Length: meter (m)

[0112] Mass: kilogram (kg)

[0113] B3. Compound unit conversion

[0114] "Ton-kilometer":

[0115] Mass: 1 ton = 1000 kg (SI compatible).

[0116] Distance: 1 kilometer = 1000 m (SI compatible).

[0117] Mapping result: 1000 kg × 1000 m = 10 6 kg·m (convertible to SI units).

[0118] B4. SI unit processing

[0119] Convert to SI units such as kilogram (kg), meter (m), etc., and retain the conversion coefficient (such as 10 6 ).

[0120] B5. Complete mapping table, as shown in Table 2

[0121] Table 2

[0122] Original physical quantity unit Type of physical quantity SI unit mapping Remarks Ton·kilometer Transportation workload / freight volume 10 6 kg·m Mass distance product, fully SI-able

[0123] S312. Based on the term mapping library in S311, a knowledge graph in the chemical field is constructed based on the chemical process flow, and process parameters of terms in the knowledge graph are extracted;

[0124] The knowledge graph is a semantic knowledge network constructed for the cost field in the chemical industry. The entities (such as materials, equipment, process flow, etc.), attributes (such as price, performance parameters, etc.), and relationships between entities (such as "material A is used in process B") in this field are described in a structured manner. The main purpose is to organize the professional knowledge in this field in the form of a graph, to lay the foundation for complex tasks such as intelligent search, reasoning, decision-making, etc.

[0125] In one embodiment, in the chemical process flow "reactor installation" process, the process parameter "pressure rating" is extracted;

[0126] In another embodiment, in the chemical process flow "pipeline corrosion prevention" process, the process parameter "corrosion resistance rating" is extracted;

[0127] S313. By analyzing the association of terms in the knowledge graph in the documents such as the bill of quantities, technical specification book, etc. through a graph neural network (GNN), a cross-document term network is established, and a cross-document term mapping library is formed.

[0128] Specifically, the method for constructing a dynamically updated cross-specialty term mapping library in S3 is:

[0129] S321. Extract historical term data from the cross-document term mapping library for judging whether a newly appeared term is a new term;

[0130] The sources of the historical term data include bidding documents, completion settlement reports, and equipment manuals, etc.

[0131] S322. Through an incremental learning framework, newly appeared terms are automatically included and the mapping and conversion results of the terms and the international standard unit library are updated, and the cross-document term mapping library is updated in real time;

[0132] In one embodiment, the term is carbon capture device installation fee.

[0133] S323. When the definitions of the same term by different specialties conflict, the priority of the term definition is determined according to the conflict resolution rules, and a cross-specialty term mapping library is formed.

[0134] The conflict resolution rules are based on industry priority and weight coefficient for judgment.

[0135] S4. The terms in the cross-specialty term mapping library are intelligently cleaned and a standardized term library is output; it is conducive to lightweight cross-specialty term mapping library, while improving the accuracy of data.

[0136] Preferably, intelligent cleaning includes fuzzy matching, semantic disambiguation, and contextual analysis of terms;

[0137] The specific method is as follows:

[0138] S41. In the fuzzy matching stage, the Levenshtein distance algorithm and the Jaccard similarity algorithm are used to perform spelling correction and preliminary deduplication of terms;

[0139] Levenshtein distance primarily works by calculating the minimum number of edit operations to match potentially correct spellings, such as miswriting "flange connection" as "flange joint" or "seamless" as "seemless." When such spelling errors occur, it can provide a certain degree of error correction feedback.

[0140] Jaccard primarily uses keyword overlap rates for semantic deduplication. For example, "seamless carbon steel pipe DN25 class150" and "Class150 DN25 carbon steel seamless pipe" have extremely high similarity or the same meaning. Therefore, one of them can be deleted.

[0141] In one embodiment, if both "flange connection" and "flange joint" appear in the terminology, the Levenshtein distance algorithm will correct the spelling of the incorrect term "flange joint" and change it to "flange connection". At this point, there are two "flange connections". The Jaccard similarity algorithm will then perform preliminary deduplication of the duplicate terms and delete the redundant term "flange connection".

[0142] S42. In the semantic disambiguation and context association analysis stage, combine the ontology relations in the knowledge graph to analyze the contextual meaning of polysemous words;

[0143] The ontological relationship refers to the relationship between two concepts, mainly including four basic relationships:

[0144] (Part-whole) relationship: indicates that one concept is a component of another concept. For example, equipment costs are a part of direct engineering costs.

[0145] (Inheritance) Relationship: This indicates that one concept is a subclass of another concept. For example, Pipe Materials is a subclass of Pipe Specialty.

[0146] (Instance-Class) Relationship: This indicates that a specific instance belongs to a certain class. For example, a sewage pump is an instance that belongs to the class "pump".

[0147] (Attribute) Relationship: Indicates that a concept possesses a certain attribute. For example, an electrolytic cell has the attribute of "weight".

[0148] The following is an example of process equipment, Table 3 is a professional classification of part of the process equipment, such as large category is static equipment, equipment filling, catalyst filling, and the following is further divided into various small categories, which is the feature of chemical project process equipment.

[0149] Table 3

[0150] Static equipment Equipment filling Catalyst loading Tower Porcelain ring random pile Φ15_50 Ammonia catalyst Tray Porcelain ring random pile Φ80_150 CO conversion catalyst Reactor Porcelain ring arrangement Φ25 inside Urea catalyst Container Porcelain ring arrangement Φ80 inside Methanol catalyst Heat exchanger Porcelain ring arrangement Φ150 inside Glycol catalyst Air cooler Carbon steel random pile Adipic acid catalyst Process skid Stainless steel random pile Butanol catalyst Electrolytic cell Aluminum random pile Vulcanization medium Demister Plastic ring random pile Activated carbon Electric dust collector Porcelain ring random pile 50*50*4.5 Acrylonitrile catalyst Porcelain ring random pile 100*100*10 Silica gel catalyst

[0151] Table 4 shows the main categories of process equipment, and further classifies the contents of static equipment-reactor to facilitate the classification and summary of cost data.

[0152] Table 4

[0153]

[0154]

[0155] S43. In the standardization output stage: according to the “Construction Engineering Quantity List Valuation Specification” and the chemical industry quota standard, the standardized terminology code with version identification is generated.

[0156] For example, “TCC-2025-PVC-pipe-installation-001” represents the cost of polyvinyl chloride pipe installation in 2025.

[0157] S5. Adopting double verification mechanism to improve the recognition accuracy of standardized terminology library in S4.

[0158] Specifically, through artificial verification and automatic test to generate adversarial samples, iteratively optimize the standardized terminology library, and improve the recognition accuracy of the terminology specific to the field of chemical cost.

[0159] In this embodiment, the terminology specific to the field of chemical cost includes “main material cost”, “comprehensive unit price”, “measure project cost”

[0160] The double verification mechanism includes artificial verification and AI verification, specifically:

[0161] The method of artificial verification includes: setting up a multi-level audit process, sampling and reviewing the cleaning results by cost engineers, and marking controversial terminology cases, while finally adding the terminology name recognized by everyone to the standardized terminology library, so as to improve the recognition accuracy of the standardized terminology library;

[0162] The method of AI verification includes: cross-comparing the semantic analysis results of AI model with the standardized terminology library to form effective feedback, and confirming the final name of the terminology in the standardized terminology library.

[0163] The chemical cost data cleaning method further comprises sensitive word filtering and cleaning, and the data in the standardized term library automatically shields terms related to business secrets (such as the name of a patented process), and only the desensitized cleaning results are retained.

[0164] A chemical cost data cleaning system, as shown in Figure 2 comprises

[0165] a data acquisition module for acquiring heterogeneous data sources of multiple specialties in the chemical field and forming a multi-modal database;

[0166] a multi-modal data processing module for identifying, reading, and writing different modal data in the multi-modal database and forming structured data;

[0167] a term feature extraction module for extracting the context semantic features, dimensional attributes, and industry association relationships of terms in the structured data and constructing a dynamically updated cross-specialty term mapping library;

[0168] an intelligent cleaning engine module for fuzzy matching, synonym disambiguation, and context association analysis of terms in the cross-specialty term mapping library and outputting a standardized term library;

[0169] a transfer learning and dynamic evaluation module that uses historical project cleaning data to train a transfer learning model, automatically adjusts the similarity threshold according to the project size, balances cleaning accuracy and efficiency, quickly adapts to the terminology system of new engineering types, automatically evaluates the impact of term cleaning on cost indicators, and generates an optimization recommendation report;

[0170] a feedback optimization module for improving the recognition accuracy of the standardized term library through a double verification mechanism;

[0171] a permission management module for establishing a three-level user permission system and implementing permission hierarchical verification of the standardized term library through asymmetric encryption technology;

[0172] a traceability module for recording modification logs of the standardized term library and supporting version rollback and audit tracking;

[0173] a visual verification module for displaying the relationship graph of terms in the standardized term library, the cleaning effect heat map, and abnormal term early warning and supporting the correction of the hierarchical standard deviation of core parameters through an interactive interface.

[0174] The permission management module includes a data access layer, a model training layer, and a core parameter layer, and implements permission hierarchical verification through RSA asymmetric encryption technology; the data access layer is used to view the content of the standardized term library; the model training layer is used to manage the standardized term library; and the core parameter layer is used to modify the new rules of the standardized term library, and the generation method of the standardized term library can be changed according to actual needs.

[0175] The term modification log includes modifier, timestamp, difference between pre / post modification terms, etc.

[0176] In one embodiment, the data access layer is enterprise-level read-only permission, the model training layer is standardized term library administrator permission, and the core parameter layer is knowledge base administrator permission.

[0177] The visualization verification module is used to generate a three-dimensional visualization board, including

[0178] The term relationship graph is used to display the mapping path between different professional terms in a Sankey diagram.

[0179] The cleaning effect heat map is used to calculate the term standardization rate and error rate according to the project type.

[0180] The abnormal term early warning board is used to alarm high-frequency controversial terms in real time.

[0181] In one embodiment, the term relationship graph can show the correlation of the anti-corrosion coating construction cost in the civil engineering and anti-corrosion professional; the cleaning effect heat map is according to the project type of EPC general contracting and construction general contracting; and the high-frequency controversial term is the difference of the measure project cost in different provincial quotas.

[0182] A chemical cost data cleaning system is deployed in a cloud edge collaborative architecture, including

[0183] The edge node is used to complete real-time data cleaning on the local server, and supports low-latency on-site data verification.

[0184] The cloud center is used to gather the cost data of the whole company, realize joint optimization of cross-enterprise term libraries through the chemical cost data cleaning method, and ensure data privacy.

[0185] The blockchain storage is used to hash and chain the cleaning records of key terms in the standardized term library, to ensure data traceability and compliance.

[0186] The chemical cost data cleaning system provided by the present application realizes collaborative optimization of cross-enterprise term libraries under the premise of ensuring data privacy through the chemical cost data cleaning method and blockchain technology.

[0187] In summary, the chemical cost data cleaning method and system provided by the present application constructs a multi-modal database, realizes accurate alignment and dynamic update of cross-professional terms by fusing field knowledge graph and deep learning model, effectively solves the industry problems of multi-source heterogeneous and semantic fragmentation of chemical cost data, and provides high-quality data basis for intelligent cost analysis.

[0188] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the application. Accordingly, the legal scope of the application is defined only by the appended claims.

Claims

1. A chemical cost data cleaning method, characterized in that, It comprises the following steps: S1. Collecting multiple professional heterogeneous data sources in the chemical field and forming a multi-modal database; S2. Converting unstructured data and real-time stream data in the multi-modal database into structured data; S3. Extracting the terms of the multi-modal database in S2 and constructing a dynamically updated cross-professional term mapping library; S4. Intelligent cleaning of the terms in the cross-professional term mapping library and output of a standardized term library; S5. Using a double verification mechanism to improve the recognition accuracy of the standardized term library.

2. The chemical cost data cleaning method according to claim 1, wherein, The multi-modal database comprises structured data, unstructured data, and real-time stream data, The structured data is data with fixed format and explicit organization, and is convenient for computer program processing and query; The unstructured data includes handwritten characters in electronic documents, paper bid documents, or paper files, and image data; The real-time stream data includes audio data and video data.

3. The chemical cost data cleaning method according to claim 2, wherein, The method for converting unstructured data and real-time stream data into structured data in S2 is: S21. Cleaning and converting electronic documents in the multi-modal database into structured text; S22. Parsing and converting image data in the multi-modal database into structured text; S23. Transcribing and converting real-time stream data in the multi-modal database into structured text.

4. The chemical cost data cleaning method according to claim 1, wherein, The specific method for extracting term features of the multi-modal database in S3 is: S311. Identifying physical quantity units in the multi-modal database terms, and mapping and converting the physical quantity units with the international standard unit library to form a term mapping library; S312. Based on the term mapping library, constructing a knowledge graph of the chemical field based on the chemical process flow, and extracting the process parameters of the terms in the knowledge graph; S313. Analyzing the association of the terms in the knowledge graph in the engineering quantity list, technical specification book, etc. through a graph neural network, establishing a cross-document term network, and forming a cross-document term mapping library.

5. The chemical cost data cleaning method according to claim 4, wherein, The method for constructing a dynamically updated cross-professional term mapping library in S3 is: S321. Extracting historical term data from the cross-document term mapping library to determine whether a newly appeared term is a new term; S322. Automatically incorporating newly appeared terms and updating the mapping and conversion results of the terms with the international standard unit library through an incremental learning framework, while updating the cross-document term mapping library in real time; S323. When the definitions of the same term by different professionals conflict, the priority of the term definition is determined according to the conflict resolution rules to form a cross-professional term mapping library; The conflict resolution rules are based on industry priority and weight coefficient.

6. The chemical cost data cleaning method according to claim 1, wherein, The specific method for intelligent cleaning and output of the standardized term library in S4 is: S41. Using the Levenshtein distance algorithm and the Jaccard similarity algorithm to correct spelling errors and preliminarily remove duplicates of the terms; S42. Analyzing the context meaning of polysemous words based on the ontology relationship in the knowledge graph to eliminate ambiguity; S43. Generating a quasi-standardized term library with version identification according to the "Construction Engineering Bill of Quantities Pricing Specification" and the chemical industry quota standard, which contains standardized term codes.

7. The chemical cost data cleaning method of claim 1, wherein, The double verification mechanism includes manual verification and AI verification, The method of manual verification includes setting up a multi-level review process, sampling and reviewing the cleaning results by cost engineers, and labeling controversial term cases, while finally adding the standardized term library with the term name unanimously recognized by everyone. The method of AI verification includes cross-comparing the semantic analysis results of the AI model with the standardized term library to form effective feedback and confirm the final name of the term in the standardized term library.

8. The chemical cost data cleaning method of claim 1, wherein, It also includes filtering and cleaning sensitive terms in the standardized term library, automatically shielding terms related to business secrets, and only retaining cleaned data after desensitization.

9. A chemical cost data cleaning system, which adopts the chemical cost data cleaning method of any one of claims 1-8, the system comprising: a data acquisition module for acquiring multi-specialty heterogeneous data sources and forming a multi-modal database; a multi-modal data processing module for converting different modal data in the multi-modal database into structured data for easy processing and querying by computer programs; a term feature extraction module for extracting the context semantic features, dimensional attributes, and industry association relationships of terms in the structured data, and constructing a dynamically updated cross-specialty term mapping library; an intelligent cleaning engine module for fuzzy matching, synonym disambiguation, and context association analysis of terms in the cross-specialty term mapping library, and outputting a standardized term library; a transfer learning and dynamic evaluation module that uses historical project cleaning data to train a transfer learning model, quickly adapts to new engineering type term systems, and automatically evaluates the impact of term cleaning by the intelligent cleaning engine module on cost indicators, generating an optimization recommendation report; a feedback optimization module for improving the recognition accuracy of the standardized term library through a double verification mechanism; a permission management module for establishing a three-level user permission system and verifying the permissions of the standardized term library by level; a traceability module for recording modification logs of the standardized term library, supporting version rollback and audit tracking; a visualization verification module for displaying the relationship graph of terms in the standardized term library, cleaning effect heat map, and abnormal term early warning, and supporting the correction of the hierarchical standard deviation of core parameters through an interactive interface.

10. The chemical cost data cleaning system of claim 9, wherein, The system is deployed in a cloud-edge collaborative architecture, including an edge node for completing real-time data cleaning on a local server, supporting low-latency on-site data verification; a cloud center for aggregating cost data from the entire company, implementing joint optimization of cross-enterprise term libraries through the chemical cost data cleaning method, while ensuring data privacy; blockchain storage for hashing and chaining the cleaning records of key terms in the standardized term library, ensuring data traceability and compliance.