Engineering budget checking method and system based on multi-source data fusion and rule engine
Through multi-source data fusion and rule engine technology, the power grid engineering design text is automatically processed, solving the problems of low efficiency, poor adaptability and insufficient accuracy in existing technologies, and achieving efficient and accurate engineering budget verification and investment management.
Patent Information
- Application Number
- CN202510815826.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-12
AI Technical Summary
During the feasibility study and preliminary design stages of power grid projects, existing technologies suffer from low matching efficiency, poor dynamic adaptability, insufficient intelligence, and high error and missed detection rates, which limits the accuracy and efficiency of project investment management.
By adopting a method based on multi-source data fusion and rule engine, through structured processing, natural language processing, machine learning and API interface technologies, key information in the design text is automatically extracted and matched, price information is updated in real time, and a dynamic knowledge base is built to verify the engineering budget.
It improves the efficiency and accuracy of engineering budget verification, ensures the timeliness and reliability of verification results, reduces manual intervention, and promotes data standardization management and intelligent decision-making.
Smart Images

Figure CN120634655A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method and system for verifying engineering budget estimates based on multi-source data fusion and a rule engine. Background Art
[0002] During the feasibility study and preliminary design phases of power grid projects, due to limited design depth, the technical parameters, project quantities, and other information required for budget estimates are primarily sourced from design documents such as design specifications, equipment and material inventories, and specialized funding proposals. Technical and economic personnel must manually review and extract key data, then verify this data against the budget estimate documents to ensure the design plan aligns with the project investment. However, design documents cover a wide range of disciplines, are highly technical, and contain fragmented information. This mechanical data extraction and verification process consumes considerable time for technical and economic personnel, and manual processing is prone to omissions and deviations, impacting the accuracy and efficiency of investment management.
[0003] The rapid development of artificial intelligence (AI) technology is providing new solutions for automated information extraction and intelligent analysis. Natural language processing (NLP) and large language models (LLMs) have demonstrated powerful capabilities in text understanding, information extraction, and structured data processing.
[0004] The technical and economic review of power transmission and transformation projects involves a vast amount of highly specialized design documents, including design specifications, equipment and materials lists, and design proposals. Currently, this information is primarily reviewed manually and then compared with the budget documents to ensure that the project investment aligns with the design plan.
[0005] However, current methods have the following major drawbacks: 1. Inefficient matching: Quotas and materials must be manually associated (e.g., the relationship between GD4-41 and copper busbars); 2. Poor dynamic adaptability: unable to automatically obtain the latest national grid information prices, and price updates lag behind; 3. Insufficient intelligence: Unable to process engineering unit conversions (e.g., 100m → 150m); unable to identify the aggregation relationship between dispersed materials and quotas (e.g., 13 cabinets scattered across two inventory records); 4. High error and missed detection rate: There is no fuzzy matching capability for the difference between the construction cost proposal "steel beam cast concrete roof panel" and the quota name. Summary of the Invention
[0006] The embodiments of the present application provide a method and system for verifying engineering budget estimates based on multi-source data fusion and a rule engine, so as to at least address the deficiencies in the above-mentioned related technologies.
[0007] In a first aspect, an embodiment of the present application provides a method for verifying an engineering budget estimate based on multi-source data fusion and a rule engine, comprising the following steps: Step 1: obtaining a number of design documents to be processed, and performing structured processing on each of the design documents to obtain corresponding structured data and unstructured data; Step 2: Using a natural language processing algorithm to extract key information from the unstructured data, and converting the extracted key information into a standardized data format to obtain corresponding standardized data; Step 3: Encoding the standardized data to obtain corresponding encoded data and non-encoded data, and calculating the non-encoded data based on an improved text similarity algorithm to identify the relevance between text descriptions; Step 4: Obtain price information through the API interface and automatically mark the data timeliness of the price information to achieve horizontal multi-source data consistency verification and vertical historical data rationality deviation check; Step 5: Record all manual correction operation data and expert decisions, and optimize matching rules and similarity thresholds through machine learning algorithms to build a dynamic knowledge base; Step 6: Construct a project budget verification model based on the dynamic knowledge base, and use the project budget verification model to implement budget verification of engineering data to generate a corresponding difference report.
[0008] Furthermore, the step 1 includes: Extracting the design specifications, equipment and material lists, and professional funding sheets from each of the design documents, and processing the design specifications, equipment and material lists, and professional funding sheets to obtain corresponding preliminary processed data; Key feature values are extracted from the preliminary processed data to obtain corresponding structured data and unstructured data.
[0009] Furthermore, the step 2 includes: Extracting key entity information from the unstructured data using a natural language processing algorithm, including equipment model, technical parameters, and engineering quantity data; The named entity recognition algorithm is used to automatically annotate technical and economic related entities in the text, and a standardized data format is established to obtain the corresponding standardized data.
[0010] Furthermore, the step three includes: Obtaining a coded unique identifier, and matching the standardized data according to the coded unique identifier to obtain corresponding coded data and non-coded data; The semantic similarity algorithm of the BERT model is used to process the non-coded data, and the matching degree between the quota items and the design description is calculated based on the dimensions of word order, word frequency and professional weight to obtain the corresponding similarity score.
[0011] Furthermore, the step 4 includes: Obtain price information through the API interface, mark the price information for timeliness, and compare the price information with the data in the design document, quota library, and price library to achieve horizontal verification of multi-source data; Acquire historical project data and verify the rationality deviation between the historical project data and the current project to achieve longitudinal verification of the data.
[0012] In a second aspect, the present invention further proposes a system for verifying engineering budget estimates based on multi-source data fusion and a rule engine, comprising: A structured processing module is used to obtain a number of design texts to be processed and perform structured processing on each of the design texts to obtain corresponding structured data and unstructured data; A data extraction module is used to extract key information from the unstructured data using a natural language processing algorithm, and convert the extracted key information into a standardized data format to obtain corresponding standardized data; A data processing module is used to encode the standardized data to obtain corresponding encoded data and non-encoded data, and calculate the non-encoded data based on an improved text similarity algorithm to identify the relevance between text descriptions; The data verification module is used to obtain price information through the API interface and automatically mark the data timeliness of the price information to achieve horizontal multi-source data consistency verification and vertical historical data rationality deviation check; The data optimization module is used to record all manual correction operation data and expert decisions, and optimize matching rules and similarity thresholds through machine learning algorithms to build a dynamic knowledge base; The estimate verification module is used to construct an engineering estimate verification model based on the dynamic knowledge base, and use the engineering estimate verification model to implement estimate verification of engineering data to generate a corresponding difference report.
[0013] Furthermore, the structured processing module is specifically used to: Extracting the design specifications, equipment and material lists, and professional funding sheets from each of the design documents, and processing the design specifications, equipment and material lists, and professional funding sheets to obtain corresponding preliminary processed data; Key feature values are extracted from the preliminary processed data to obtain corresponding structured data and unstructured data.
[0014] Furthermore, the data extraction module is specifically used to: Extracting key entity information from the unstructured data using a natural language processing algorithm, including equipment model, technical parameters, and engineering quantity data; The named entity recognition algorithm is used to automatically annotate technical and economic related entities in the text, and a standardized data format is established to obtain the corresponding standardized data.
[0015] Furthermore, the data processing module is specifically used to: Obtaining a coded unique identifier, and matching the standardized data according to the coded unique identifier to obtain corresponding coded data and non-coded data; The semantic similarity algorithm of the BERT model is used to process the non-coded data, and the matching degree between the quota items and the design description is calculated based on the dimensions of word order, word frequency and professional weight to obtain the corresponding similarity score.
[0016] Furthermore, the data verification module is specifically used to: Obtain price information through the API interface, mark the price information for timeliness, and compare the price information with the data in the design document, quota library, and price library to achieve horizontal verification of multi-source data; Acquire historical project data and verify the rationality deviation between the historical project data and the current project to achieve longitudinal verification of the data.
[0017] Compared with related technologies, the embodiment of the present application provides a method and system for engineering budget verification based on multi-source data fusion and rule engine. By extracting key information of the design text and intelligently matching it with the budget document, the traditional manual verification mode is changed, and the efficiency of budget verification is improved. Real-time access to industry data sources and combined with a multi-dimensional verification mechanism ensure the accuracy and timeliness of the verification results, providing reliable protection for engineering investment management. The application of natural language processing and knowledge graphs lays the foundation for building a large-scale model of the power profession, freeing technical and economic personnel from tedious data verification and turning them to creative work such as investment analysis and cost optimization. At the same time, it promotes standardized management of engineering data and provides support for intelligent decision-making.
[0018] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1Flowchart of a method for checking engineering budget estimates based on multi-source data fusion and rule engine in a first embodiment of the present invention; Figure 2 This is a structural block diagram of an engineering budget verification system based on multi-source data fusion and rule engine in the second embodiment of the present invention.
[0020] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts are within the scope of protection of this application.
[0022] Obviously, the drawings described below are merely examples or embodiments of the present application. Those skilled in the art can, without inventive effort, apply the present application to other similar scenarios based on these drawings. Furthermore, it is also understood that, although the effort involved in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, changes in design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as an insufficiency of the content disclosed in this application.
[0023] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.
[0024] Unless otherwise defined, the technical or scientific terms involved in this application should have the usual meaning understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "the" and the like involved in this application do not indicate a quantitative limitation and may indicate the singular or plural. The terms "include", "comprise", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units that are inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Example 1
[0025] See also Figure 1 , which shows a method for checking engineering budget estimates based on multi-source data fusion and rule engine in a first embodiment of the present invention. The method specifically includes steps S101 to S103: S101, obtaining a number of design documents to be processed, and performing structured processing on each of the design documents to obtain corresponding structured data and unstructured data; Furthermore, the step S101 specifically includes steps S1011 to S1012: S1011, extracting the design specifications, equipment and material lists, and professional funding proposals from each of the design documents, and processing the design specifications, equipment and material lists, and professional funding proposals to obtain corresponding preliminary processed data; S1012: Extract key feature values from the preliminarily processed data to obtain corresponding structured data and unstructured data.
[0026] During implementation, various input design documents (including design specifications, equipment and materials lists, and professional funding documents) are structured. Specifically, quota data from the design documents is received, including quota number, project name, unit, and quantity. The corresponding equipment parameters and engineering quantity data from the equipment and materials lists are also obtained. This raw data is standardized, with unified measurement units and key feature values extracted to facilitate subsequent intelligent matching. For example, a compound unit like "1.5 (100m)" is converted to a standard measurement representation.
[0027] Furthermore, the system automatically analyzes the correspondence between quota items and equipment and materials, identifying instances where unit conversion is required. A built-in conversion rule engine performs precise unit conversions and numerical calculations based on pre-set industry standards. For example, it converts the length of cable protection conduit from "100m" to "m" to ensure consistency with the measurement units in inventory data. The system also verifies the reasonableness of engineering quantity values.
[0028] After the match verification is complete, a verification report containing detailed comparison results is automatically generated. The report clearly marks the items that matched successfully or failed, and provides a detailed analysis of the differences. The report also records the conversion rules and matching logic used during the verification process, and this empirical data is fed back into the knowledge base to continuously optimize subsequent intelligent verification results.
[0029] S102, extracting key information from the unstructured data using a natural language processing algorithm, and converting the extracted key information into a standardized data format to obtain corresponding standardized data; Furthermore, the step S102 specifically includes steps S1021 and S1022: S1021, extracting key entity information from the unstructured data using a natural language processing algorithm, including equipment model, technical parameters, and engineering quantity data; S1022, automatically annotating technology-related entities in the text through a named entity recognition algorithm, and establishing a standardized data format to obtain corresponding standardized data.
[0030] In specific implementation, natural language processing technology is used to extract key information from unstructured text, including equipment models, technical parameters, engineering quantity data, etc. A named entity recognition algorithm is used to automatically annotate technical and economic entities in the text, and a standardized data format is established to provide unified data input for subsequent matching and verification.
[0031] Specifically, we conduct in-depth analysis of construction quota names and budget proposal descriptions. Using natural language processing technology, we remove redundant modifiers and extract core keywords. For example, we simplify "steel beam cast concrete roof slab" to "steel beam roof slab," retaining the key elements that best reflect the project's characteristics. We also establish a synonym library for professional terms to address compatibility issues between different expression methods.
[0032] S103, encoding the standardized data to obtain corresponding encoded data and non-encoded data, and calculating the non-encoded data based on an improved text similarity algorithm to identify the relevance between text descriptions; Furthermore, the step S103 specifically includes steps S1031 and S1032: S1031, obtaining a coding unique identifier, and matching the standardized data according to the coding unique identifier to obtain corresponding coded data and non-coded data; S1032: Process the non-encoded data using the semantic similarity algorithm of the BERT model, and calculate the matching degree between the quota item and the design description based on the dimensions of word order, word frequency, and professional weight to obtain a corresponding similarity score.
[0033] In specific implementation, the matching engine adopts a three-level progressive matching strategy: 1. Accurate code matching: Prioritize accurate matching through unique identifiers such as quota codes, equipment material codes, etc. 2. Semantic similarity matching: For data without explicit coding, semantic similarity calculation based on the BERT model is used to identify the relevance between text descriptions. 3. Expert rule matching: Apply preset industry knowledge rules to special scenarios, such as processing the professional correspondence between "GD4-41 quota" and "copper busbar TMY-125×10" To address the issue of inconsistent engineering quantity units, the system has a built-in unit conversion rule library that automatically performs standardized conversion calculations such as "1.5(100m)→150m".
[0034] Specifically, an improved text similarity algorithm is used to calculate the degree of match between quota items and design descriptions. The algorithm comprehensively considers multiple dimensions such as word order, word frequency, and professional weight to give a similarity score. For projects where the degree of match is near the critical value, a secondary verification process is automatically triggered to improve the accuracy of judgment by checking auxiliary information such as engineering quantity values and professional categories. The matching results are presented in a visual manner, allowing users to make final confirmation. For special cases, users can supplement their professional judgment through the interactive interface. These feedback data will be automatically collected by the system and used to optimize subsequent matching algorithms. At the same time, the system will record typical matching cases and continuously enrich the terminology knowledge base of the construction profession.
[0035] S104, obtaining price information through the API interface and automatically marking the data timeliness of the price information to achieve horizontal multi-source data consistency verification and vertical historical data rationality deviation check; Furthermore, the step S104 specifically includes steps S1041 and S1042: S1041, obtaining price information through the API interface, marking the price information for timeliness, and comparing the price information with the design document, the quota database, and the price database to achieve horizontal verification of multi-source data; S1042, obtaining historical project data, and performing rationality deviation verification between the historical project data and the current project to achieve longitudinal verification of the data.
[0036] In the specific implementation, it connects to external data sources such as the State Grid information price in real time, obtains the latest price information through the API interface, and automatically marks the timeliness of the data. The verification process adopts a dual-channel verification mechanism: 1. Horizontal verification: Compare the consistency of multiple sources of data such as design documents, quota database, price database, etc. 2. Longitudinal verification: Check the rationality deviation between historical project data and current project data For normative fees such as measures fees, a preset pre-regulation calculation model is used to automatically generate standard rates based on project type, region and other attributes, enabling automatic compliance review.
[0037] Specifically, a comprehensive timeliness monitoring mechanism has been established, with all price data labeled with expiration dates and sources. A regular scanning program automatically checks the timeliness of price data, identifying records that are about to expire or have already expired. Price data is categorized and managed by key attributes such as device type and voltage level to ensure targeted updates.
[0038] Furthermore, when a price data update is detected, the system automatically connects to official data sources, such as the State Grid price information, to obtain the latest price information through standardized APIs. New data undergoes multiple verification steps before being stored, including format checks, rationality verification, and historical data comparison. The system utilizes a progressive update strategy to ensure data transitions are completed without service interruption.
[0039] After a price update is completed, costs for affected items are automatically recalculated and differences are analyzed compared to pre-update calculations. A detailed price change impact report is generated, highlighting key equipment with significant price fluctuations, helping users quickly understand investment changes. Historical price data is also saved to support price trend analysis and forecasting.
[0040] S105, recording all manual correction operation data and expert decisions, and optimizing matching rules and similarity thresholds through machine learning algorithms to establish a dynamic knowledge base; During implementation, all manual corrections and expert decisions are recorded, and matching rules and similarity thresholds are continuously optimized through machine learning. A dynamic knowledge base is established to achieve the following functions: 1. Automatically expand the mapping relationship of professional terms 2. Adjust the weight of unit conversion rules 3. Update price fluctuation warning threshold Specifically, we monitor anomalies during the verification process in real time and identify typical matching failure patterns using a pre-set rule engine. For each anomaly, we automatically analyze the cause of the failure, classifying it into different types, such as unit mismatch, terminology discrepancy, and missing data. We also assess the severity of the anomaly to determine whether expert intervention is needed. When expert knowledge support is confirmed, the system searches the professional rule base to find the appropriate solution. The knowledge base adopts a multi-level organizational structure, encompassing general rules, specialized rules, and special rules. The system comprehensively considers factors such as the scope of application, frequency of use, and confidence level of the rules to select the optimal solution. For complex situations, it supports the combination of multiple rules.
[0041] S106: constructing a project budget verification model based on the dynamic knowledge base, and implementing budget verification of project data using the project budget verification model to generate a corresponding difference report.
[0042] In this embodiment, data interaction is achieved through a RESTful interface to ensure stable operation under large-scale engineering data. The verification results generate difference reports in real time and support multi-dimensional statistical analysis and visualization.
[0043] Specifically, a comprehensive feedback learning mechanism is established to record the handling process and user confirmation results of each exception case. Using machine learning algorithms, rule weights and application priorities are continuously adjusted. Furthermore, regular health checks are conducted on the knowledge base to identify knowledge points that require supplementation or revision, forming a continuously optimized closed-loop system. Experts can also directly maintain and expand the knowledge base content through a dedicated management interface.
[0044] In summary, the engineering budget verification method based on multi-source data fusion and rule engine in the above embodiments of the present invention changes the traditional manual verification mode by extracting key information of the design text and intelligently matching it with the budget document, thereby improving the efficiency of budget verification; real-time access to industry data sources, combined with a multi-dimensional verification mechanism, ensures the accuracy and timeliness of the verification results, and provides reliable protection for engineering investment management; the application of natural language processing and knowledge graphs lays the foundation for building a large model of the power profession, freeing technical and economic personnel from tedious data verification and turning them to creative work such as investment analysis and cost optimization, while promoting standardized management of engineering data and providing support for intelligent decision-making. Example 2
[0045] On the other hand, the present invention also proposes a system for checking engineering budget estimates based on multi-source data fusion and rule engine. Figure 2 , shown is a system for checking engineering budget estimates based on multi-source data fusion and rule engine in a second embodiment of the present invention, comprising: The structured processing module 11 is used to obtain a number of design documents to be processed and perform structured processing on each of the design documents to obtain corresponding structured data and unstructured data; A data extraction module 12 is configured to extract key information from the unstructured data using a natural language processing algorithm, and convert the extracted key information into a standardized data format to obtain corresponding standardized data; A data processing module 13 is configured to encode the standardized data to obtain corresponding encoded data and non-encoded data, and calculate the non-encoded data based on an improved text similarity algorithm to identify the relevance between text descriptions; The data verification module 14 is used to obtain price information through the API interface and automatically mark the data timeliness of the price information to achieve horizontal multi-source data consistency verification and vertical historical data rationality deviation check; Data optimization module 15, used to record all manual correction operation data and expert decisions, and optimize matching rules and similarity thresholds through machine learning algorithms to establish a dynamic knowledge base; The estimate verification module 16 is used to construct an engineering estimate verification model based on the dynamic knowledge base, and use the engineering estimate verification model to implement estimate verification of engineering data to generate a corresponding difference report.
[0046] Furthermore, the structured processing module 11 is specifically configured to: Extracting the design specifications, equipment and material lists, and professional funding sheets from each of the design documents, and processing the design specifications, equipment and material lists, and professional funding sheets to obtain corresponding preliminary processed data; Key feature values are extracted from the preliminary processed data to obtain corresponding structured data and unstructured data.
[0047] Furthermore, the data extraction module 12 is specifically configured to: Extracting key entity information from the unstructured data using a natural language processing algorithm, including equipment model, technical parameters, and engineering quantity data; The named entity recognition algorithm is used to automatically annotate technical and economic related entities in the text, and a standardized data format is established to obtain the corresponding standardized data.
[0048] Furthermore, the data processing module 13 is specifically configured to: Obtaining a coded unique identifier, and matching the standardized data according to the coded unique identifier to obtain corresponding coded data and non-coded data; The semantic similarity algorithm of the BERT model is used to process the non-coded data, and the matching degree between the quota items and the design description is calculated based on the dimensions of word order, word frequency and professional weight to obtain the corresponding similarity score.
[0049] Furthermore, the data verification module 14 is specifically used to: Obtain price information through the API interface, mark the price information for timeliness, and compare the price information with the data in the design document, quota library, and price library to achieve horizontal verification of multi-source data; Acquire historical project data and verify the rationality deviation between the historical project data and the current project to achieve longitudinal verification of the data.
[0050] The functions or operation steps implemented when the above modules are executed are substantially the same as those in the above method embodiments and will not be repeated here.
[0051] The embodiment of the present invention provides an engineering budget verification system based on multi-source data fusion and rule engine. Its implementation principle and technical effects are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the system embodiment, please refer to the corresponding content in the aforementioned method embodiment.
[0052] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0053] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for checking engineering budget estimates based on multi-source data fusion and rule engine, characterized in that: The following steps are involved: Step 1: obtaining a number of design documents to be processed, and performing structured processing on each of the design documents to obtain corresponding structured data and unstructured data; Step 2: Using a natural language processing algorithm to extract key information from the unstructured data, and converting the extracted key information into a standardized data format to obtain corresponding standardized data; Step 3: Encoding the standardized data to obtain corresponding encoded data and non-encoded data, and calculating the non-encoded data based on an improved text similarity algorithm to identify the relevance between text descriptions; Step 4: Obtain price information through the API interface and automatically mark the data timeliness of the price information to achieve horizontal multi-source data consistency verification and vertical historical data rationality deviation check; Step 5: Record all manual correction operation data and expert decisions, and optimize matching rules and similarity thresholds through machine learning algorithms to build a dynamic knowledge base; Step 6: Construct a project budget verification model based on the dynamic knowledge base, and use the project budget verification model to implement budget verification of engineering data to generate a corresponding difference report.
2. The engineering budget verification method based on multi-source data fusion and rule engine according to claim 1 is characterized in that: The step one comprises: Extracting the design specifications, equipment and material lists, and professional funding sheets from each of the design documents, and processing the design specifications, equipment and material lists, and professional funding sheets to obtain corresponding preliminary processed data; Key feature values are extracted from the preliminary processed data to obtain corresponding structured data and unstructured data.
3. The engineering budget verification method based on multi-source data fusion and rule engine according to claim 1 is characterized in that: The second step includes: Extracting key entity information from the unstructured data using a natural language processing algorithm, including equipment model, technical parameters, and engineering quantity data; The named entity recognition algorithm is used to automatically annotate technical and economic related entities in the text, and a standardized data format is established to obtain the corresponding standardized data.
4. The engineering budget verification method based on multi-source data fusion and rule engine according to claim 1 is characterized in that: The step three includes: Obtaining a coded unique identifier, and matching the standardized data according to the coded unique identifier to obtain corresponding coded data and non-coded data; The semantic similarity algorithm of the BERT model is used to process the non-coded data, and the matching degree between the quota items and the design description is calculated based on the dimensions of word order, word frequency and professional weight to obtain the corresponding similarity score.
5. The engineering budget verification method based on multi-source data fusion and rule engine according to claim 1 is characterized in that: The fourth step includes: Obtain price information through the API interface, mark the price information for timeliness, and compare the price information with the data in the design document, quota library, and price library to achieve horizontal verification of multi-source data; Acquire historical project data and verify the rationality deviation between the historical project data and the current project to achieve longitudinal verification of the data.
6. A project budget verification system based on multi-source data fusion and rule engine, characterized by: include: A structured processing module is used to obtain a number of design texts to be processed and perform structured processing on each of the design texts to obtain corresponding structured data and unstructured data; A data extraction module is used to extract key information from the unstructured data using a natural language processing algorithm, and convert the extracted key information into a standardized data format to obtain corresponding standardized data; A data processing module is used to encode the standardized data to obtain corresponding encoded data and non-encoded data, and calculate the non-encoded data based on an improved text similarity algorithm to identify the relevance between text descriptions; The data verification module is used to obtain price information through the API interface and automatically mark the data timeliness of the price information to achieve horizontal multi-source data consistency verification and vertical historical data rationality deviation check; The data optimization module is used to record all manual correction operation data and expert decisions, and optimize matching rules and similarity thresholds through machine learning algorithms to build a dynamic knowledge base; The estimate verification module is used to construct an engineering estimate verification model based on the dynamic knowledge base, and use the engineering estimate verification model to implement estimate verification of engineering data to generate a corresponding difference report.
7. The engineering budget verification system based on multi-source data fusion and rule engine according to claim 6 is characterized in that: The structured processing module is specifically used for: Extracting the design specifications, equipment and material lists, and professional funding sheets from each of the design documents, and processing the design specifications, equipment and material lists, and professional funding sheets to obtain corresponding preliminary processed data; Key feature values are extracted from the preliminary processed data to obtain corresponding structured data and unstructured data.
8. The engineering budget verification system based on multi-source data fusion and rule engine according to claim 6 is characterized in that: The data extraction module is specifically used for: Extracting key entity information from the unstructured data using a natural language processing algorithm, including equipment model, technical parameters, and engineering quantity data; The named entity recognition algorithm is used to automatically annotate technical and economic related entities in the text, and a standardized data format is established to obtain the corresponding standardized data.
9. The engineering budget verification system based on multi-source data fusion and rule engine according to claim 6 is characterized in that: The data processing module is specifically used for: Obtaining a coded unique identifier, and matching the standardized data according to the coded unique identifier to obtain corresponding coded data and non-coded data; The semantic similarity algorithm of the BERT model is used to process the non-coded data, and the matching degree between the quota items and the design description is calculated based on the dimensions of word order, word frequency and professional weight to obtain the corresponding similarity score.
10. The engineering budget verification system based on multi-source data fusion and rule engine according to claim 6 is characterized in that: The data verification module is specifically used for: Obtain price information through the API interface, mark the price information for timeliness, and compare the price information with the data in the design document, quota library, and price library to achieve horizontal verification of multi-source data; Acquire historical project data and verify the rationality deviation between the historical project data and the current project to achieve longitudinal verification of the data.
Citation Information
Cited By
Engineering quantity list automatic compiling and auditing method and system based on industrial big data
CN122022734A