Construction engineering cost data identification and standardization method and system

By constructing a five-level tagging system and knowledge graph, combined with a deep learning model, the problems of inconsistent data formats and low efficiency of manual processing in construction engineering cost data management have been solved. This has enabled efficient and accurate data identification and standardized processing, improving the comparability of data and the scientific nature of management decisions.

CN121579583APending Publication Date: 2026-02-27HANGZHOU RUICHENG INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202610123942.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

The management of construction project cost data suffers from problems such as inconsistent data formats, large differences in units and precision, lack of a unified labeling system, cumbersome and error-prone manual processing, and incomparability of cross-project data. These issues make it difficult to effectively integrate and process the data, and hinder the discovery of its potential value.

Method used

A five-level cost data labeling system and a construction engineering cost knowledge graph are constructed. A deep learning model is used for automated identification and standardization, including data cleaning, format normalization, field labeling, verification and normalization. The model is optimized through a closed-loop mechanism.

Benefits of technology

It enables structured representation and semantic association of multi-source heterogeneous cost data throughout the entire lifecycle of construction projects, improving data consistency, comparability, and reliability, reducing labor costs, enhancing intelligent data processing capabilities, and improving the scientific nature of management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579583A_ABST
    Figure CN121579583A_ABST
Patent Text Reader

Abstract

The invention discloses a building engineering cost data identification and standardization method and system. The method comprises the steps of collecting and preprocessing multi-source heterogeneous cost data to obtain an initial data set; constructing a cost knowledge graph of the five-level cost data label system and the integration field rule; labeling a historical wide table based on a label system to generate a training sample, and training a deep learning model to obtain a label matching model; mapping the initial data set into a standard wide table, and inputting the standard wide table into the model to complete field automatic labeling; performing field value verification and data normalization processing on the annotation wide table based on a knowledge graph rule; outputting a standardized cost wide table; and feeding back an audit result to a model training process to form closed-loop optimization. Through combination of the five-level label system and the knowledge graph, automatic and standardized processing of the cost data is realized, the data quality and cross-project comparability are remarkably improved, and reliable data support is provided for construction engineering cost management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of construction project cost data management technology, and in particular to a method and system for identifying and standardizing construction project cost data. Background Technology

[0002] Construction project cost data spans the entire project lifecycle, covering multiple stages including design, construction, and completion. The data comes from a wide range of sources, including various documents, business systems, and other information platforms, exhibiting a multi-source and heterogeneous nature. This data is the core basis for project cost control, budget calculation, risk warning, and cross-project comparative analysis. Its accuracy, consistency, and usability directly affect the scientific nature and effectiveness of project management decisions.

[0003] However, current construction project cost data management faces a series of challenges. First, data formats are often inconsistent across different projects and stages, with significant differences in units and precision, making effective data integration and processing difficult. Second, the lack of a unified labeling system leads to ambiguous meanings in data fields, making it impossible to ensure accurate correlation and matching between data. Furthermore, traditional data processing methods mainly rely on manual identification and organization, which is cumbersome, error-prone, inefficient, and costly. Finally, when comparing data across projects, the lack of standardized specifications makes data from different projects incomparable, hindering the extraction of potential value from the data.

[0004] Existing data processing technologies and methods have failed to fully address these issues, particularly when dealing with the large-scale application of construction project cost data, making it difficult to achieve efficient, accurate, and standardized management. Therefore, a new, systematic solution is urgently needed to automate the identification, precise labeling, and standardized processing of construction project cost data, thereby meeting the demands for data accuracy, consistency, and efficiency in construction project management. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for identifying and standardizing construction project cost data, which enables automated identification, accurate labeling, and standardized processing of construction project cost data.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for identifying and standardizing construction project cost data, comprising the following steps:

[0007] Step S1: Collect multi-source heterogeneous cost data throughout the entire life cycle of the building project, clean and normalize the data to obtain the initial cost dataset;

[0008] Step S2: Construct a five-level cost data labeling system and a construction engineering cost knowledge graph based on the initial cost dataset. The five-level cost data labeling system includes core projects, first-level processes, second-level processes, feature labels, and calculation indicators. The construction engineering cost knowledge graph integrates domain rules based on the labeling system.

[0009] Step S3: Based on the five-level cost data labeling system, label the historical building project cost wide table to form a training sample set, and divide the training sample set into a training set and a test set;

[0010] Step S4: Construct a deep learning model, train the deep learning model using the training set, and verify the model performance using the test set. When the label matching accuracy reaches a preset threshold, obtain the cost field label matching model.

[0011] Step S5: Map the initial cost dataset to a standard wide table structure, input the cost field label matching model, and output the labels corresponding to each field to complete the automatic labeling of the wide table fields;

[0012] Step S6: Based on the domain rules in the knowledge graph of construction engineering costs, perform field value validation on the labeled wide table and perform data normalization on the field values ​​with the same label;

[0013] Step S7: Output a standardized cost wide table, which includes field names, third-level labels, field values, data sources, and validation results;

[0014] Step S8: Review the standardized cost wide table. If the review does not meet the requirements, mark it as a negative sample. If the review meets the requirements, mark it as a positive sample and feed it back to step S4 to update the cost field label matching model. At the same time, output the standardized cost wide table that meets the requirements.

[0015] Furthermore, in step S1, the multi-source heterogeneous cost data includes structured and unstructured data from Excel, PDF, and Word documents; the data cleaning includes removing blank fields and duplicate records; and the format normalization process includes converting monetary data into a uniform unit of yuan and retaining two decimal places.

[0016] In step S2, the domain rules of the integrated construction engineering cost knowledge graph include the content calculation rules and the unit cost calculation rules.

[0017] The calculation rule for the content of the project is as follows: content of the project = total quantity of the project ÷ building area; the calculation rule for the unit cost is as follows: unit cost = total project cost ÷ total building area of ​​the project.

[0018] In step S3, the annotation information of the training sample set includes time, project, building, label, feature, total amount and unit; the ratio of the training set to the test set is 8:2.

[0019] Furthermore, in step S4, the deep learning model is obtained by fine-tuning the BERT model;

[0020] The input layer of the deep learning model includes text vectors of field names, field context information, and knowledge graph association rules;

[0021] The feature fusion layer of the deep learning model uses an attention mechanism to perform weighted fusion of the three types of vectors from the input layer. The formula for weighted fusion is as follows:

[0022] ;

[0023] in, Let represent the feature fusion vector, where α, β, and γ represent the preset first, second, and third learnable weight parameters, respectively. , , These represent text vectors, context vectors, and knowledge graph rule vectors, respectively.

[0024] The output layer of the deep learning model uses a step-Softmax function to normalize the hidden layer output to output the third-level label in the five-level cost data labeling system. The step-Softmax function is expressed as follows:

[0025] ;

[0026] in, Let Z represent the Softmax function in the step described above, and let Z represent the input vector. This represents the i-th vector in the input vector. Let represent an exponential function. The lower limit of accumulation j=1 indicates that the accumulation starts from the first element and continues to the kth element of the upper limit of accumulation, where k is the dimension of the input vector Z, i.e. the total number of categories.

[0027] Furthermore, in step S4, the training of the deep learning model uses the cross-entropy loss function to calculate the error between the predicted label and the manually labeled label, and iteratively adjusts the model parameters through the gradient descent algorithm;

[0028] The preset threshold is a label matching accuracy of ≥98%; the cross-entropy loss function is expressed as:

[0029] ;

[0030] Where L represents the cross-entropy loss function, and N represents the total number of categories in the three-level label. This represents the unique hot code value of the manually labeled tag. This represents the probability of the i-th type of label predicted by the model.

[0031] Furthermore, in step S6, the field value verification includes detecting whether the field value conforms to the domain rules and marking outliers;

[0032] The data normalization process includes converting the unit cost to "yuan / square meter" and the material usage to "cubic meter / square meter".

[0033] Furthermore, in the five-level cost data labeling system, the core projects include earthwork engineering, foundation pit support engineering, steel reinforcement engineering, pile foundation engineering, concrete engineering, and waterproofing engineering; the calculation indicators include the quantity of work, the comprehensive unit price, and the total price.

[0034] Furthermore, after step S7, an audit and optimization step is also included: the standardized cost wide table is audited by both the model and human review, label errors and data anomalies are corrected, and the corrected data is added to the training sample set as positive samples, and the erroneous data is added to the training sample set as negative samples.

[0035] A construction project cost data identification and standardization system, applied to the aforementioned construction project cost data identification and standardization method, includes:

[0036] The data acquisition module is used to collect multi-source heterogeneous cost data throughout the entire life cycle of a building project, clean and normalize the data to obtain an initial cost dataset.

[0037] The knowledge construction module, connected to the data acquisition module, is used to construct a five-level cost data labeling system and a construction engineering cost knowledge graph based on the initial cost dataset. The five-level cost data labeling system includes core projects, first-level processes, second-level processes, feature labels, and calculation indicators. The construction engineering cost knowledge graph integrates domain rules based on the labeling system.

[0038] The sample processing module, connected to the knowledge construction module, is used to label the historical building engineering cost wide table based on the five-level cost data labeling system, form a training sample set, and divide the training sample set into a training set and a test set.

[0039] The model training module, connected to the sample processing module, is used to construct a deep learning model, train the deep learning model using the training set, and verify the model performance using the test set. When the label matching accuracy reaches a preset threshold, the cost field label matching model is obtained.

[0040] The label matching module is connected to the data acquisition module and the model training module respectively. It is used to map the initial cost dataset into a standard wide table structure, input the cost field label matching model, and output the labels corresponding to each field to complete the automatic labeling of the wide table fields.

[0041] The verification and normalization module is connected to the tag matching module and the knowledge construction module respectively. It is used to verify the field values ​​of the labeled wide table based on the domain rules in the construction engineering cost knowledge graph, and to perform data normalization processing on the field values ​​of the same tag.

[0042] The audit and closed-loop update module is connected to the verification and normalization module. It is used to output a standardized cost wide table. The standardized cost wide table includes field names, three-level labels, field values, data sources and verification results. The module audits the standardized cost wide table. If the audit does not meet the requirements, it is marked as a negative sample. If the audit meets the requirements, it is marked as a positive sample. The module feeds back to step S4 to update the cost field label matching model. At the same time, it outputs the standardized cost wide table that meets the requirements.

[0043] The processor is used to control the operation of the above modules;

[0044] A computing card, connected to the processor, the model training module and the label matching module, is used to accelerate model training and label matching calculations.

[0045] A memory, connected to the processor, is used to store the knowledge graph, training sample set, model parameters, and standardized cost wide table.

[0046] The beneficial effects of this invention are:

[0047] The proposed method for identifying and standardizing construction project cost data, by constructing a unified and complete five-level cost data labeling system and combining it with a construction project cost knowledge graph, achieves structured expression and semantic association of multi-source heterogeneous cost data throughout the entire life cycle of construction projects. This fundamentally solves the problems of inconsistent field expression, missing labels, semantic ambiguity, and difficulty in cross-project reuse in traditional cost data.

[0048] By constructing a training sample set based on a large number of labeled cost wide tables and utilizing a deep learning model for automatic field label recognition, this invention enables automated, batch, and low-human-intervention-required data label matching of cost fields. Compared with traditional methods that rely on manual rules, this invention significantly improves the accuracy and processing scale of field recognition, reduces labor costs, and achieves intelligent processing of large-scale engineering cost data.

[0049] The process of this invention introduces multi-dimensional data cleaning, format normalization, field value validation, and aggregation and normalization of fields with the same label. This ensures that the final standardized cost wide table achieves a unified standard in field names, third-level labels, field values, and source information, significantly enhancing the consistency, comparability, traceability, and reliability of the data, and laying the foundation for enterprises to build high-quality cost data assets.

[0050] Furthermore, by feeding back the review results (positive / negative samples) to the model training process, this invention innovatively constructs a sustainable, self-evolving closed-loop model optimization mechanism. This mechanism enables the cost field label matching model to continuously iterate, automatically adapting to changes in different construction projects, different cost systems, and different data formats, thereby ensuring the model's stability and generalization ability in long-term operation.

[0051] The final output of the standardized cost spreadsheet can be directly used for various business scenarios such as cross-project cost comparison, budget calculation, cost risk assessment, and cost prediction model training. This greatly improves the informatization and intelligence level of engineering management, significantly enhances the scientific nature of cost management decisions and data support capabilities, and fully taps the business value of construction engineering cost data. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the steps of the construction engineering cost data identification and standardization method in this invention; Figure 2 This is a schematic diagram of the structure of the construction engineering cost data identification and standardization system in this invention. Attached reference numerals: 1. Data acquisition module; 2. Knowledge construction module; 3. Sample processing module; 4. Model training module; 6. Label matching module; 7. Verification and normalization module; 8. Review and closed-loop update module; 9. Processor; 10. Computing power card; 11. Memory. Detailed Implementation

[0053] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.

[0054] Example 1, refer to Figure 1This is the first embodiment of the present invention, which provides a method for identifying and standardizing construction project cost data. This method is applicable to the automated processing, intelligent annotation, and standardized output of multi-source, heterogeneous cost data throughout the entire lifecycle of a construction project. By integrating deep learning models and domain knowledge graphs, this method achieves accurate identification, labeling, rule validation, and normalization of cost data fields. Ultimately, it outputs a standardized wide cost table with a unified structure and clear semantics, providing a high-quality data foundation for cost analysis, cost control, and decision support.

[0055] Detailed steps and working principle of the method in Example 1:

[0056] Step S1: Collect multi-source heterogeneous cost data throughout the entire life cycle of the building project, clean and normalize the data to obtain the initial cost dataset;

[0057] Working principle: Cost data from different departments and in different formats is collected throughout the entire lifecycle of a construction project, including design, bidding, construction, and settlement. This data typically exists in the form of Excel spreadsheets, PDF reports, and Word documents, containing both structured data (such as numbers and text in tables) and unstructured data (such as cost information in paragraph descriptions). To achieve unified processing later, the raw data needs to be cleaned and formatted. Detailed implementation method:

[0059] Data Acquisition: Extract cost-related data from various documents through database interfaces, document parsing tools (such as Apache POI, PDFBox), or OCR technology.

[0060] Data cleaning: Automatically identify and remove completely blank fields and delete completely duplicate data records to reduce the interference of noisy data on model training. Specifically, during the data cleaning process, blank fields are removed by checking for non-empty fields, and duplicate records are removed by comparing data hash values.

[0061] Format normalization: All fields involving monetary amounts are uniformly processed, converting them to values ​​in "yuan" and retaining two decimal places. For example, "1,256,000 yuan" is converted to "1,256,000.00 yuan". This step ensures the consistency of data units, laying the foundation for subsequent calculations and comparisons. Specifically, during format normalization, amounts in different currencies are uniformly converted to RMB (yuan), all monetary data retain two decimal places, and the date format is uniformly "YYYY-MM-DD", ensuring the initial cost dataset has a standardized format.

[0062] Technical effect: This step solves the pain points of cost data being from diverse sources, with messy formats and poor quality. Through automated cleaning and normalization, a high-quality initial cost dataset is generated, providing clean and consistent data input for subsequent knowledge construction and model training.

[0063] Step S2: Construct a five-level cost data labeling system and a construction engineering cost knowledge graph based on the initial cost dataset. The five-level cost data labeling system includes core projects, first-level processes, second-level processes, feature labels, and calculation indicators. The construction engineering cost knowledge graph integrates domain rules based on the labeling system.

[0064] How it works: To assign unified semantic labels to cost data and embed domain knowledge, this step constructs a hierarchical labeling system and a knowledge graph that integrates domain rules. The labeling system provides a framework for data classification, while the knowledge graph encapsulates the industry's computational rules and logical relationships. Detailed implementation method:

[0066] Five-level cost data labeling system: This system consists of five levels, as shown in Table 1 below:

[0067]

[0068] Table 1

[0069] Table 1 is an example of a five-level cost data labeling system. The interpretation of the five-level cost data in Table 1 is as follows:

[0070] Core projects: These represent the largest cost category, such as earthwork, foundation pit support, steel reinforcement, pile foundation, concrete, and waterproofing.

[0071] Primary construction process: The main construction stage under the core project, such as the pile foundation project, which can be divided into concrete pipe pile drilling; precast reinforced concrete pipe piles, etc.

[0072] Secondary process: A further subdivision of the primary process, such as precast pipe piles and long spiral cast-in-place piles in pile foundation engineering.

[0073] Feature tags: Describe the specific characteristics of the cost item, such as name; drilling method; stratum; pile diameter, length, construction location and flatness; pile type; pile inclination; pile driving method, etc.

[0074] Calculation indicators: Specific numerical indicators of cost data, including quantity of work, unit price and total price.

[0075] Construction of a knowledge graph for construction engineering costs: The graph is constructed in the form of entities (such as engineering projects, materials, and processes) and relationships (such as "belongs to," "consumed," and "calculated"). Its core is the integration of domain expert rules, which mainly include:

[0076] Content calculation rule: Project content = Total project volume ÷ Building area. Used to measure the resource consumption per unit building area.

[0077] Unit cost calculation rule: Unit cost = Total project cost ÷ Total project building area. Used to assess the economic efficiency of project costs.

[0078] These rules are stored in the knowledge graph in the form of logical assertions or executable functions for subsequent data validation.

[0079] Technical effect: By constructing a structured tagging system and knowledge graph, scattered cost data is linked with the system's domain knowledge, giving the data clear semantics and internal logic, and providing a rule-based foundation for intelligent identification and verification.

[0080] Step S3: Based on the five-level cost data labeling system, label the wide table of historical building engineering costs to form a training sample set, and divide the training sample set into a training set and a test set;

[0081] Working principle: Using the established labeling system, the cost wide table (a table containing many fields) that has been compiled in historical projects is manually labeled to form the supervised learning samples required for model training. Detailed implementation method:

[0083] Sample annotation: Domain experts were invited to annotate each field in the historical cost wide table according to the five-level tagging system defined in step S2. The annotation information covers time, project, building, tag (corresponding to the third-level tag, such as "Concrete Engineering - Concrete Pouring - Comprehensive Unit Price"), feature, total amount, and unit.

[0084] Dataset partitioning: The labeled training sample set is randomly divided into a training set and a test set in an 8:2 ratio. The training set is used for model parameter learning, and the test set is used to evaluate the model's generalization performance.

[0085] Technical results: High-quality training data with standard answers were generated, providing the necessary conditions for training deep learning models that can accurately understand the semantics of cost data.

[0086] In this embodiment, historical cost tables of 100 different types of building projects (residential, commercial, and industrial buildings) are selected as sample templates. Each table contains at least 50 cost-related fields. Experts in the field of construction engineering cost annotate the "time-project-building-tag-feature-total-unit" information corresponding to each field to form a training sample set. An example is shown in Table 2 below:

[0087]

[0088] Table 2

[0089] Table 2 is a sample wide table of the training sample set.

[0090] Step S4: Build a deep learning model, train the deep learning model using the training set, and verify the model performance using the test set. When the label matching accuracy reaches the preset threshold, the cost field label matching model is obtained.

[0091] How it works: This step is the core of the method, aiming to train a deep learning model that can automatically map cost data fields to the correct third-level labels. The model achieves accurate classification by fusing field text information, contextual information, and knowledge graph rules. Detailed implementation method:

[0093] Model Architecture Design: The lightweight deep learning model in this embodiment is based on the BERT-ba step Se pre-trained model, fine-tuned for domain adaptation. Through multi-dimensional feature fusion and a phased training strategy, it achieves accurate label matching for construction engineering cost data fields. The model training process strictly follows a standardized workflow of data preprocessing, parameter initialization, iterative optimization, and performance verification. The specific steps are as follows:

[0094] The lightweight deep learning model consists of an input layer, a feature fusion layer, a hidden layer, and an output layer. The functions and parameter configurations of each layer are as follows:

[0095] Input layer: Receives three types of feature vectors and performs dimension alignment; the total input dimension is 768. Specifically, it includes:

[0096] Text vector The semantic vectors of field names (such as "basement steel reinforcement cost" and "concrete foundation cost per cubic meter") are extracted by the BERT encoder, and then segmented into 768-dimensional context semantic vectors.

[0097] Context vector Extract the field name and manually labeled information from the row or three adjacent columns containing the field, and convert them into a 768-dimensional vector through the BERT encoding layer to represent the logical relationship between the fields;

[0098] Knowledge Graph Rule Vector The domain rules related to the current field in the cost knowledge graph (such as "reinforced concrete engineering includes processes such as basement foundation reinforcement and main structure reinforcement") are converted into 768-dimensional vectors through the Tran step SE embedding algorithm, and then injected with domain prior knowledge.

[0099] Feature fusion layer: An attention mechanism is used to adaptively weight and fuse the above three types of vectors. The calculation formula is as follows:

[0100] ;

[0101] Formula Explanation: This is the final fused feature vector used for classification. α, β, and γ (initialized to 1 / 3) are weight parameters automatically learned by the model during training, representing the importance of textual information, contextual information, and knowledge rule information to the current field label judgment, respectively. The purpose of this formula is to dynamically adjust the contribution of different information sources. For example, for the field "unit price," the weight of textual information α may be higher; for the field "content," the weight of knowledge rules may be higher. The weight γ (such as in the content calculation formula) may be higher. This allows the model to flexibly and intelligently integrate multiple features for decision-making.

[0102] Hidden layers: Contain two fully connected network layers with 512 neurons per layer. The activation function is ReLU. Batch normalization is used to suppress overfitting, and the dropout rate is set to 0.1 to enhance the model's generalization ability.

[0103] Output Layer: The hidden layer output is normalized using the step-Softmax function, outputting the joint probability distribution of the three-level labels "Core Project - First-Level Process - Second-Level Process" in the five-level label system (e.g., the probability value of "Reinforcement Engineering - Reinforcement Engineering Above Basement Slab - Cast-in-place Component Reinforcement"). The output dimension is consistent with the total number of categories of the three-level labels in the label system (216 dimensions in this embodiment). The formula for the step-Softmax function is:

[0104] ;

[0105] in, Let Z represent the Softmax function in the step described above, and let Z represent the input vector. This represents the i-th vector in the input vector. Let represent an exponential function. The lower limit of accumulation j=1 indicates that the accumulation starts from the first element and continues to the kth element of the upper limit of accumulation, where k is the dimension of the input vector Z, i.e. the total number of categories.

[0106] Formula Explanation: This function converts the raw scores output by the network. Convert to probability value This ensures that the sum of the probabilities of all labels is 1. This solves the probability output problem in multi-class classification problems, making it easier to select the label with the highest probability as the prediction result.

[0107] Preprocessing of the training dataset for the model:

[0108] The training sample set (including manually annotated three-level labels) constructed in step S3 is preprocessed, including:

[0109] Label encoding: One-hot encoding is used to convert the three-level labels into binary vectors, which are then used as target values ​​for model training;

[0110] Data augmentation: Replace field names in the training set with synonyms (e.g., replace "construction cost" with "expenses" or "cost") and randomly adjust the context order (shuffle the order of non-key columns while maintaining the logical relationship between three adjacent columns) to generate an augmented dataset with 1.5 times the original sample size, thus avoiding overfitting of the model to specific expressions.

[0111] Batch partitioning: The enhanced training set is divided into training batches of 32 samples each, and the test set is divided into batches of the same size to ensure the stability of the training process.

[0112] The specific process of model training is as follows:

[0113] Parameter initialization: Load the weight parameters of the BERT-ba pre-trained model (excluding the output layer). The parameters of the feature fusion layer, hidden layer and output layer are initialized using the Xavier normal distribution. The initial learning rate is set to 0.0001 and the weight decay coefficient is set to 0.01 to suppress excessively large model weights.

[0114] During training, the cross-entropy loss function L is used to measure the difference between the predicted probability distribution and the true label distribution (one-hot encoding). The loss is minimized using the gradient descent algorithm, and the model parameters are iteratively adjusted. The loss function formula is:

[0115] ;

[0116] Formula Explanation: L represents the cross-entropy loss function, and N represents the total number of categories in the three-level label system. This represents the unique hot code value of the manually labeled tag (1 if yes, 0 otherwise). This represents the probability of the i-th class label predicted by the model. When the predicted probability... The closer to reality The smaller the loss L, the better. This function effectively drives the model to learn to classify the correct label.

[0117] Optimizer configuration during model training: The AdamW optimizer is used, with the first-order moment estimation exponential decay rate (β1) set to 0.9, the second-order moment estimation exponential decay rate (β2) set to 0.999, and the numerical stability parameter ε set to le−8, to achieve adaptive updates of model parameters.

[0118] Model performance validation:

[0119] After training is completed, the model is fully validated using a test set. Validation metrics include:

[0120] Label matching accuracy: ≥98% (consistent with the threshold for stopping training);

[0121] Average Precision (Sion): The average precision across all three-level label categories must be ≥97%.

[0122] The two thresholds mentioned above are set based on the high requirements for data accuracy in actual business operations, ensuring the reliability of the model output. Once these two thresholds are met, a usable cost field label matching model is obtained.

[0123] Training stability: The accuracy fluctuation range of three consecutive repeated training sessions should be ≤0.5% to ensure that the model performance is reproducible.

[0124] Through the training methods described above, the model can fully learn the field features, contextual relationships, and domain rules of construction engineering cost data, and achieve efficient matching of the three-level labels of "core project - first-level process - second-level process", providing accurate label support for subsequent standardized processing of cost data.

[0125] Technical results: This step automates and intelligently identifies cost data field labels. The model integrates multi-dimensional information, significantly improving the accuracy and efficiency of labeling, and avoiding the subjectivity and high cost of manual labeling.

[0126] Step S5: Map the initial cost dataset to a standard wide table structure, input the cost field label matching model, and output the labels corresponding to each field to complete the automatic labeling of the wide table fields;

[0127] Working principle: The preprocessed initial cost data is organized according to the standard wide table structure. Then, the trained cost field label matching model is used. The label with the highest probability output by the model is used as the final label for that field. The model automatically predicts and assigns the most likely third-level label to each field. Detailed implementation method:

[0129] The initial cost dataset obtained in step S1 is mapped to a preset standard wide table template, which defines the expected arrangement structure of the fields.

[0130] Input each field name and its context from the wide table into the model trained in step S4.

[0131] The model outputs the corresponding three-level label for each field (such as "reinforcement engineering - fabrication and installation - quantity"), completing the automated labeling of the entire wide table.

[0132] Technical results: It enables rapid, batch, and automated annotation of massive amounts of newly added cost data tables, completely freeing up manpower and improving processing speed by several orders of magnitude compared to manual methods.

[0133] Step S6: Based on the domain rules in the knowledge graph of construction engineering costs, perform field value validation on the labeled wide table and perform data normalization on the field values ​​with the same label;

[0134] Working principle: By utilizing the domain rules integrated in the knowledge graph, the numerical values ​​of the labeled fields are logically verified, and the units and formats of similar data are deeply normalized to ensure the consistency and comparability of the data. Detailed implementation method:

[0136] Field value validation: The system automatically calls rules from the knowledge graph to check the rationality of the data. For example, for a field labeled "Calculation index: Unit cost", the system checks whether its value is equal to the ratio of the total cost of the corresponding project to the building area. If the difference exceeds ±5%, it is marked as an anomaly and the administrator is prompted to check. For a field labeled "Feature tag: Rebar specification", the system checks whether its value conforms to industry standard specifications (such as HPB300, HRB400, etc.). If it does not conform, it is marked as an anomaly.

[0137] Data normalization: For fields with the same label, their values ​​are uniformly converted to a standard unit to facilitate aggregation and comparison. For example:

[0138] All values ​​of the "unit cost" related fields are uniformly converted to "yuan / square meter".

[0139] All values ​​of the "concrete usage" related fields are uniformly converted to "cubic meters / square meter" or "kilograms / square meter", and the project quantity is uniformly based on industry standard units (e.g., steel reinforcement quantity is uniformly "tons", and concrete quantity is uniformly "cubic meters").

[0140] Technical benefits: Automated rule validation promptly detects data errors or anomalies, improving data credibility; deep unit normalization eliminates data analysis obstacles caused by different units of measurement, making cross-project and cross-period data comparison possible.

[0141] Step S7: Output a standardized cost wide table, which includes field names, third-level labels, field values, data sources, and validation results;

[0142] How it works: It integrates the processing results of all previous steps to generate a standardized wide cost table containing complete metadata information.

[0143] Detailed implementation: The output standardized cost wide table is a structured data table, where each row represents a cost item, and each column contains the following information:

[0144] Technical benefits: This standardized cost table is the final data product. It has a clear structure, explicit semantics, standardized values, and controllable quality. It can be directly used for various downstream applications such as cost database construction, BI analysis, and machine learning modeling.

[0145] Step S8: Review the standardized cost wide table. If the review does not meet the requirements, mark it as a negative sample. If the review meets the requirements, mark it as a positive sample and feed it back to step S4 to update the cost field label matching model. At the same time, output the standardized cost wide table that meets the requirements.

[0146] Working principle: The system introduces a manual review process as the final quality control, and feeds the review results back into the model training process, forming a closed loop of "automated processing - manual review - model optimization" to achieve continuous self-evolution of the system. Detailed implementation method:

[0148] Domain experts review the standardized cost table output in step S7.

[0149] If the review is approved, the form will be marked as a positive sample and output for use.

[0150] If the review fails (e.g., labeling errors are found), the table will be marked as a negative sample, and its labels will be corrected.

[0151] These newly generated positive and negative samples are fed back into the training process in step S4, merged with the original training set, and used to incrementally train and update the cost field label matching model.

[0152] Technical benefits: The closed-loop mechanism ensures that the system can learn from errors and continuously adapt to new project characteristics, data formats, and industry terminology, which makes the accuracy and robustness of the model continuously enhanced over time, forming a continuously optimized intelligent data processing system.

[0153] Overall technical effects of Example 1:

[0154] The construction project cost data identification and standardization method proposed in this embodiment, by constructing a unified and complete five-level cost data labeling system and combining it with a construction project cost knowledge graph, realizes the structured expression and semantic association of multi-source heterogeneous cost data throughout the entire life cycle of construction projects. This fundamentally solves the problems of inconsistent field expression, missing labels, semantic ambiguity, and difficulty in cross-project reuse in traditional cost data.

[0155] By constructing a training sample set based on a large number of labeled cost wide tables and utilizing a deep learning model for automatic field label recognition, this embodiment enables automated, batch, and low-human-intervention-required data label matching of cost fields. Compared with traditional methods that rely on manual rules, this embodiment significantly improves the accuracy and processing scale of field recognition, reduces labor costs, and achieves intelligent processing of large-scale engineering cost data.

[0156] This embodiment incorporates multi-dimensional data cleaning, format normalization, field value validation, and aggregation and normalization of fields with the same label. This ensures that the final standardized cost wide table achieves uniformity in field names, third-level labels, field values, and source information, significantly enhancing data consistency, comparability, traceability, and reliability, thus laying the foundation for enterprises to build high-quality cost data assets.

[0157] Furthermore, by feeding back the review results (positive / negative samples) to the model training process, this embodiment innovatively constructs a sustainable, self-evolving closed-loop model optimization mechanism. This mechanism enables the cost field label matching model to continuously iterate, automatically adapting to changes in different construction projects, different cost systems, and different data formats, thereby ensuring the model's stability and generalization ability in long-term operation.

[0158] The final output of the standardized cost spreadsheet can be directly used for various business scenarios such as cross-project cost comparison, budget calculation, cost risk assessment, and cost prediction model training. This greatly improves the informatization and intelligence level of engineering management, significantly enhances the scientific nature of cost management decisions and data support capabilities, and fully taps the business value of construction engineering cost data.

[0159] In a preferred embodiment, step S7 is followed by a manual review step: two or more cost domain experts conduct a cross-review of the standardized cost table. The focus of the manual cross-review includes:

[0160] Focus on verifying the field values ​​marked as "abnormal" in the model;

[0161] Determine whether the third-level labels of the fields conform to the professional understanding of engineering cost;

[0162] Discuss and confirm cost fields that are disputed or have boundary conditions.

[0163] The working principle of manual review includes:

[0164] By cross-judging multiple cost experts and incorporating industry experience and professional knowledge, manual corrections are performed for complex or atypical cases that the model cannot fully cover, thus compensating for the limitations of purely automated methods. Special attention is paid to verifying fields marked as "abnormal" and their tag matching results.

[0165] During the review process, the following processing logic is executed based on the review conclusion:

[0166] If a label error is found: After the correct label is manually confirmed, the field label is corrected, and the corrected data is added to the training sample set as a positive sample;

[0167] If data anomalies are found: further verify the source of the original cost data, correct the field values ​​to ensure the accuracy of the standardized cost wide table, and supplement the training sample set with the corrected data as positive samples;

[0168] If the model output is confirmed to be incorrect and cannot be directly corrected:

[0169] The corresponding samples are added to the training sample set as negative samples, which are then used to focus on constraining this type of error during the subsequent model training phase.

[0170] Through the aforementioned sample feedback mechanism, the training sample set is continuously expanded and gradually covers more real-world business scenarios.

[0171] Example 2, refer to Figure 2 This is the second embodiment of the present invention. Unlike the previous embodiment, this embodiment provides a construction project cost data identification and standardization system, applied to the aforementioned construction project cost data identification and standardization method, including:

[0172] Data acquisition module 1 is used to collect multi-source heterogeneous cost data throughout the entire life cycle of a building project, clean and normalize the data to obtain an initial cost dataset.

[0173] Knowledge construction module 2, connected to data acquisition module 1, is used to construct a five-level cost data labeling system and a construction engineering cost knowledge graph based on the initial cost dataset. The five-level cost data labeling system includes core projects, first-level processes, second-level processes, feature labels, and calculation indicators. The construction engineering cost knowledge graph integrates domain rules based on the labeling system.

[0174] Sample processing module 3, connected to knowledge construction module 2, is used to label the wide table of historical building engineering costs based on the five-level cost data labeling system, form a training sample set, and divide the training sample set into a training set and a test set.

[0175] Model training module 4, connected to sample processing module 3, is used to build a deep learning model, train the deep learning model using the training set, and verify the model performance using the test set. When the label matching accuracy reaches a preset threshold, the cost field label matching model is obtained.

[0176] The label matching module 6 is connected to the data acquisition module 1 and the model training module 4 respectively. It is used to map the initial cost dataset into a standard wide table structure, input the cost field label matching model, and output the labels corresponding to each field to complete the automatic labeling of the wide table fields.

[0177] The verification and normalization module 7 is connected to the tag matching module 6 and the knowledge construction module 2 respectively. It is used to verify the field values ​​of the labeled wide table based on the domain rules in the knowledge graph of construction engineering cost, and to perform data normalization processing on the field values ​​of the same tag.

[0178] The audit and closed-loop update module 8 is connected to the verification and normalization module 7. It is used to output a standardized cost wide table. The standardized cost wide table includes field names, three-level labels, field values, data sources and verification results. The standardized cost wide table is audited. If the audit does not meet the requirements, it is marked as a negative sample. If the audit meets the requirements, it is marked as a positive sample and fed back to step S4 to update the cost field label matching model. At the same time, it outputs a standardized cost wide table that meets the requirements.

[0179] Processor 9 is used to control the operation of the above modules;

[0180] The computing card 10 connects to the processor 9, the model training module 4, and the label matching module 6, and is used to accelerate the computation of model training and label matching.

[0181] The memory 11 is connected to the processor 9 and is used to store the knowledge graph, training sample set, model parameters and standardized cost wide table.

[0182] Specifically, in this embodiment, the system may further include an OCR recognition module and a visualization interaction module. The OCR recognition module is used to extract text information from unstructured data, and its computationally intensive tasks such as high-resolution scan recognition and handwritten cost data parsing can be accelerated by the computing power card 10. The visualization interaction module provides label system configuration, model training status monitoring, standardized wide table display, and manual review interface to facilitate user operation and management.

[0183] In Embodiment 2, the processor 9 can be a central processing unit (CPU), a digital signal processor 9 (D-step SP), an application-specific integrated circuit (A-step SIC), or other programmable logic devices, responsible for overall scheduling, logic verification, and system collaborative management; the computing power card 10 can be a graphics processing unit 9 (GPU), a field-programmable gate array (FPGA), or a dedicated AI acceleration card, focusing on batch recognition and computation of high-concurrency cost data, efficient training of deep learning models (such as cost data classification models, feature extraction model iterations), parallel cleaning and transformation of multi-source heterogeneous data, and rapid generation of large-scale standardized wide tables, improving data processing efficiency in complex scenarios; the memory 11 can be a random access memory 11 (RAM), a read-only memory 11 (ROM), a solid-state drive 11 (step S-step SD), etc., used to store multi-source heterogeneous cost data, tag systems, knowledge graphs, training sample sets, model parameters, and standardized cost wide tables.

[0184] In this embodiment, the specific types of processor 9 and memory 11 are not limited, as long as they can perform the corresponding computation and storage functions. The computer-readable storage medium can be a hard disk, ROM, RAM, etc., used to store the computer program corresponding to the method of this invention.

[0185] Technical effects of Example 2:

[0186] With the above settings, this embodiment can achieve automated identification, accurate labeling and standardized processing of construction project cost data; improve the accuracy of label recognition and clarify threshold constraints; support high-concurrency processing of large-scale engineering data; and improve overall operating efficiency through software and hardware collaboration.

[0187] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for identifying and standardizing construction project cost data, characterized in that, Includes the following steps: Step S1: Collect multi-source heterogeneous cost data throughout the entire life cycle of the building project, clean and normalize the data to obtain the initial cost dataset; Step S2: Construct a five-level cost data labeling system and a construction engineering cost knowledge graph based on the initial cost dataset. The five-level cost data labeling system includes core projects, first-level processes, second-level processes, feature labels, and calculation indicators. The construction engineering cost knowledge graph integrates domain rules based on the labeling system. Step S3: Based on the five-level cost data labeling system, label the historical building project cost wide table to form a training sample set, and divide the training sample set into a training set and a test set; Step S4: Construct a deep learning model, train the deep learning model using the training set, and verify the model performance using the test set. When the label matching accuracy reaches a preset threshold, obtain the cost field label matching model. Step S5: Map the initial cost dataset to a standard wide table structure, input the cost field label matching model, and output the labels corresponding to each field to complete the automatic labeling of the wide table fields; Step S6: Based on the domain rules in the knowledge graph of construction engineering costs, perform field value validation on the labeled wide table and perform data normalization on the field values ​​with the same label; Step S7: Output a standardized cost wide table, which includes field names, third-level labels, field values, data sources, and validation results; Step S8: Review the standardized cost wide table. If the review does not meet the requirements, mark it as a negative sample. If the review meets the requirements, mark it as a positive sample and feed it back to step S4 to update the cost field label matching model. At the same time, output the standardized cost wide table that meets the requirements.

2. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: In step S1, the multi-source heterogeneous cost data includes structured and unstructured data from Excel, PDF, and Word documents; the data cleaning includes removing blank fields and duplicate records; the format normalization process includes converting monetary data into a uniform unit of yuan and retaining two decimal places.

3. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: In step S2, the domain rules of the integrated construction engineering cost knowledge graph include the content calculation rules and the unit cost calculation rules. The calculation rule for the content of the project is as follows: content of the project = total quantity of the project ÷ building area; the calculation rule for the unit cost is as follows: unit cost = total project cost ÷ total building area of ​​the project.

4. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: In step S3, the annotation information of the training sample set includes time, project, building, label, feature, total amount and unit; the ratio of the training set to the test set is 8:

2.

5. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: In step S4, the deep learning model is obtained by fine-tuning the BERT model; The input layer of the deep learning model includes text vectors of field names, field context information, and knowledge graph association rules; The feature fusion layer of the deep learning model uses an attention mechanism to perform weighted fusion of the three types of vectors from the input layer. The formula for weighted fusion is as follows: ; in, Let represent the feature fusion vector, where α, β, and γ represent the preset first, second, and third learnable weight parameters, respectively. , , These represent text vectors, context vectors, and knowledge graph rule vectors, respectively. The output layer of the deep learning model uses a step-Softmax function to normalize the hidden layer output to output the third-level label in the five-level cost data labeling system. The step-Softmax function is expressed as follows: ; in, Let Z represent the Softmax function in the step described above, and let Z represent the input vector. This represents the i-th vector in the input vector. Let represent an exponential function. The lower limit of accumulation j=1 indicates that the accumulation starts from the first element and continues to the kth element of the upper limit of accumulation, where k is the dimension of the input vector Z, i.e. the total number of categories.

6. The method for identifying and standardizing construction project cost data according to claim 5, characterized in that: In step S4, the training of the deep learning model uses the cross-entropy loss function to calculate the error between the predicted label and the manually labeled label, and iteratively adjusts the model parameters through the gradient descent algorithm. The preset threshold is a label matching accuracy of ≥98%; the cross-entropy loss function is expressed as: ; Where L represents the cross-entropy loss function, and N represents the total number of categories in the three-level label. This represents the unique hot code value of the manually labeled tag. This represents the probability of the i-th type of label predicted by the model.

7. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: In step S6, the field value verification includes detecting whether the field value conforms to the domain rules and marking abnormal values; The data normalization process includes converting the unit cost to "yuan / square meter" and the material usage to "cubic meter / square meter".

8. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: In the five-level cost data labeling system, the core projects include earthwork, foundation pit support, steel reinforcement, pile foundation, concrete, and waterproofing; the calculation indicators include quantity, unit price, and total price.

9. The method for identifying and standardizing construction project cost data according to claim 1, characterized in that: Following step S7, an audit and optimization step is also included: the standardized cost wide table is audited by both the model and human review, label errors and data anomalies are corrected, and the corrected data is added to the training sample set as positive samples, while the erroneous data is added to the training sample set as negative samples.

10. A construction project cost data identification and standardization system, applied to the construction project cost data identification and standardization method according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous cost data throughout the entire life cycle of a building project, clean and normalize the data to obtain an initial cost dataset. The knowledge construction module, connected to the data acquisition module, is used to construct a five-level cost data labeling system and a construction engineering cost knowledge graph based on the initial cost dataset. The five-level cost data labeling system includes core projects, first-level processes, second-level processes, feature labels, and calculation indicators. The construction engineering cost knowledge graph integrates domain rules based on the labeling system. The sample processing module, connected to the knowledge construction module, is used to label the historical building engineering cost wide table based on the five-level cost data labeling system, form a training sample set, and divide the training sample set into a training set and a test set. The model training module, connected to the sample processing module, is used to construct a deep learning model, train the deep learning model using the training set, and verify the model performance using the test set. When the label matching accuracy reaches a preset threshold, the cost field label matching model is obtained. The label matching module is connected to the data acquisition module and the model training module respectively. It is used to map the initial cost dataset into a standard wide table structure, input the cost field label matching model, and output the labels corresponding to each field to complete the automatic labeling of the wide table fields. The verification and normalization module is connected to the tag matching module and the knowledge construction module respectively. It is used to verify the field values ​​of the labeled wide table based on the domain rules in the construction engineering cost knowledge graph, and to perform data normalization processing on the field values ​​of the same tag. The audit and closed-loop update module is connected to the verification and normalization module. It is used to output a standardized cost wide table. The standardized cost wide table includes field names, three-level labels, field values, data sources and verification results. The module audits the standardized cost wide table. If the audit does not meet the requirements, it is marked as a negative sample. If the audit meets the requirements, it is marked as a positive sample. The module feeds back to step S4 to update the cost field label matching model. At the same time, it outputs the standardized cost wide table that meets the requirements. The processor is used to control the operation of the above modules; A computing card, connected to the processor, the model training module and the label matching module, is used to accelerate model training and label matching calculations. A memory, connected to the processor, is used to store the knowledge graph, training sample set, model parameters, and standardized cost wide table.

Citation Information

Patent Citations

  • Engineering cost data cleaning method based on artificial intelligence

    CN117290316A

  • Information management method and system for project cost

    CN120746264A

  • Engineering cost big data management and analysis system

    CN120873058A

  • Data label generation and quality control method and system based on pre-training large model

    CN121117627A

  • Power field training set dynamic construction method and system based on BERT and reinforcement learning

    CN121256352A