Sub-item cost prediction method and equipment based on label driving, and storage medium

By using a label-driven method for predicting sub-item costs, feature information is collected and tagged features are generated. The target model is then called to predict the unit price. By combining engineering quantity data and adjusting the error distribution, the problem of low efficiency and insufficient accuracy of traditional manual calculation is solved, and efficient and accurate engineering cost prediction is achieved.

CN121961644APending Publication Date: 2026-05-01深圳市华森建筑工程咨询有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳市华森建筑工程咨询有限公司
Filing Date
2025-12-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the prediction of unit prices for engineering costs relies on manual calculation, which leads to low efficiency. The accuracy of the prediction is limited by personal experience and the timeliness of industry quotas, making it difficult to quantify the range of deviation in cost prediction and failing to provide a reliable reference for decision-making.

Method used

By collecting target project feature information, matching option labels according to preset classification rules and performing feature encoding, generating labeled features, calling the target model of sub-item engineering to predict unit price, combining engineering quantity data to output cost summary data, and adjusting the final result based on model training error distribution information.

Benefits of technology

It improves the efficiency and accuracy of engineering cost forecasting, enables rapid scenario-based configuration and comparison, facilitates audit verification, and provides reliable support for early-stage project decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961644A_ABST
    Figure CN121961644A_ABST
Patent Text Reader

Abstract

The invention discloses a sub-item cost prediction method and device based on label driving and a storage medium, and relates to the technical field of project cost prediction, and the method comprises the steps: collecting target project feature information, matching corresponding option labels for the target project feature information according to a preset classification rule, processing the labels through feature coding, and generating labeled features. And then, disassembling the target project into each sub-item project, calling a target model corresponding to each sub-item, and predicting the unit price of each sub-item in combination with the tagging features to obtain a target unit price. And through a deterministic calculation rule, target unit price and project quantity data are coupled, and sub-item project cost summary data are output. And finally, based on error distribution information obtained by model training, adjusting the cost summary data, and calculating to obtain a prediction result of the total cost of the project. The problem that the efficiency of frequently modifying indexes of an early-stage scheme and the cost prediction precision are difficult to balance is solved through feature tagging, item splitting and unit price pre-processing, cost calculation and error tuning feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Tag-driven itemized cost prediction methods, devices, and storage media Technical Field

[0001] This application relates to the field of engineering cost prediction technology, and in particular to a tag-driven method, device and storage medium for predicting itemized costs. Background Technology

[0002] In engineering cost estimation, cost estimation during the design phase directly impacts the scientific and rational nature of project investment decisions, and accurate prediction of unit prices for each component is the core foundation for total project cost estimation. Currently, related technologies typically rely on cost estimators manually calculating unit prices and total costs by referring to industry standards and combining their personal project experience. This method is not only inefficient, but its prediction accuracy also fluctuates significantly due to limitations imposed by personal experience and the timeliness of industry standards. It is difficult to quantify the range of cost prediction deviations, thus failing to provide risk references for decision-making. Consequently, in scenarios where early-stage design indicators are frequently modified, it becomes difficult to balance efficiency and accuracy in cost prediction.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a tag-driven method, device, and storage medium for itemized cost prediction, aiming to solve the technical problem of balancing efficiency and accuracy in cost prediction.

[0005] To achieve the above objectives, this application proposes a label-driven method for predicting itemized costs. The method includes: collecting feature information related to the target project; matching corresponding option labels to the feature information according to preset classification rules; processing the option labels through feature encoding to generate labeled features; breaking down the target project into various sub-items; calling the target model corresponding to each sub-item; predicting the unit price of each sub-item by combining the labeled features to obtain the target unit price; coupling the target unit price and project quantity data through deterministic calculation rules to output the summary data of sub-item costs; and adjusting the summary data of sub-item costs based on the error distribution information obtained during model training to calculate the predicted result of the total project cost.

[0006] In one embodiment, the data formats of the economic indicator table data of the target project, the information of the questionnaire completion and the data extracted by the model are unified and integrated to obtain an original feature set; the corresponding feature types in the original feature set are identified, and the feature items corresponding to the feature types are filtered and distinguished according to the preset categories; based on the classification rules, the option labels uniquely corresponding to each feature item are matched respectively.

[0007] In one embodiment, the integrity and consistency of the option labels are verified, invalid labels that do not conform to the preset coding specifications are removed, and valid option labels are output. The valid option labels are numerically converted according to the preset discrete feature coding rules to obtain the one-dimensional numerical code corresponding to the feature item. The one-dimensional numerical code, the category division of the original feature, and the item feature dimension sorting are matrix-integrated to generate the labeled feature containing all item feature coding information.

[0008] In one embodiment, the economic indicator table data, engineering technical parameters, and industry classification standards for sub-items of the target project are analyzed to determine the applicable decomposition level and classification boundary of the target project, and decomposition rules are obtained by integration. According to the decomposition rules and combined with the functional attributes of the project, the target project is divided into several major engineering categories, each corresponding to a defined engineering scope and measurement caliber, to obtain the decomposition result. Based on the decomposition rules, the decomposition result is further refined hierarchically to determine the specific work content and measurement object of each sub-item of the project, thus obtaining each sub-item of the project.

[0009] In one embodiment, based on the functional attributes of the sub-items of the project, a corresponding model is matched in a preset model library, and the exclusive feature dimensions of each sub-item are extracted from the labeled features to generate a mapping relationship between each sub-item of the project and the target model, as well as sub-item exclusive features; the inference process of the corresponding target model is triggered sequentially according to the mapping relationship, and the sub-item exclusive features are substituted into the target model one by one to complete the calculation and output the preliminary unit price prediction result; based on the preset reasonable range standard for the unit price of the sub-items of the project and the historical prediction error correction rules, the preliminary unit price prediction result is verified and optimized, outliers exceeding the reasonable range are removed and the prediction deviation is adjusted to obtain the target unit price of each sub-item of the project.

[0010] In one embodiment, the measurement unit of the target unit price and the measurement scope of the corresponding sub-item quantity are verified to establish the correlation between each sub-item and the target unit price and the corresponding quantity; based on the correlation, a preset deterministic calculation rule is applied to multiply the target unit price and the corresponding quantity of each sub-item to obtain the individual cost data of each sub-item; the hierarchical division and project category of the sub-items are summarized level by level, and combined with the individual cost data, the summary cost data of the sub-items is statistically formed.

[0011] In one embodiment, based on the error distribution information obtained during model training, the error characteristics corresponding to each of the sub-items are associated and matched with the total cost data of the sub-items to determine the error fluctuation range and error distribution pattern of each sub-item cost; according to the error distribution pattern, the corresponding deviation correction algorithm is used to adjust the cost of each of the sub-items, while retaining the cost composition relationship corresponding to the project logic, to form correction data; based on the correction data, the correction cost of each sub-item is summed and calculated, and the confidence interval of the total project cost is derived in combination with the error fluctuation range to determine the prediction result.

[0012] In one embodiment, the total engineering cost and the actual cost of each component of the target project are collected. The predicted results are compared with the actual cost data, and the influencing factors are analyzed to output a deviation analysis report. The key features and corresponding option labels that cause deviations in the deviation analysis report are parsed to generate a correlation analysis result between the degree of deviation and the contribution of the features. According to the preset weight adjustment rules, combined with the correlation analysis result, the weight parameters corresponding to each feature are updated to output the optimized feature weight configuration. The optimized feature weight configuration is integrated into the target model training process corresponding to each component of the project. The model is retrained using historical datasets to solidify the weight adjustment effect. The updated target model and weight configuration file are output to complete the optimization loop of the prediction method.

[0013] In addition, to achieve the above objectives, this application also proposes an itemized cost prediction device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the tag-driven itemized cost prediction method described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the tag-driven itemized cost prediction method described above.

[0015] This application provides a label-driven method for predicting itemized costs. The method involves collecting feature information related to the target project, matching corresponding option labels according to preset classification rules, generating labeled features through feature encoding, breaking down the target project into individual sub-items and calling the corresponding target model, predicting the target unit price of each sub-item using the labeled features, coupling the target unit price with the project quantity data through deterministic calculation rules to output a summary of sub-item costs, and finally adjusting the summary data based on the error distribution information obtained during model training to calculate the total project cost prediction result. Simultaneously, it incorporates feedback from subsequent actual cost data to optimize feature weights and the target model. This method solves the technical problems of traditional engineering cost estimation, which relies on manual archives or a single overall model, has a long cycle, low accuracy, poor interpretability, and difficulty in adapting to different construction scenarios. It improves the efficiency, accuracy, and interpretability of engineering cost prediction, enables rapid scenario-based configuration and comparison, facilitates audit verification and continuous method optimization, and provides reliable support for early-stage project decision-making.

[0016] In summary, this application solves the technical problem of balancing efficiency and accuracy in cost prediction by collecting feature information, generating tagged features through label matching and encoding, decomposing the project, calling the target model to predict the target unit price, coupling the engineering quantity to obtain the cost summary data, and combining the error to adjust the total cost and provide feedback optimization. This improves the efficiency and accuracy of cost prediction, enhances interpretability, and supports rapid configuration, comparison, and decision-making. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 is a flowchart of the first embodiment of the tag-driven itemized cost prediction method of this application; Figure 2 is a flowchart of the eighth embodiment of the tag-driven itemized cost prediction method of this application; Figure 3 is a cost unit price prediction result diagram of this application; Figure 4 is a structural schematic diagram of the itemized cost prediction device of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] In related technologies, cost estimators typically rely on industry quota standards and their personal project experience to manually calculate the unit price of each item and the total cost. This method lacks the ability to quantify the uncertainty of the estimation results and cannot provide risk reference for decision-making, thus resulting in poor cost estimation accuracy.

[0023] This application provides a solution: First, collect feature information related to the target project. Then, match corresponding option tags to the feature information according to preset classification rules. Next, process the option tags through feature encoding to generate tagged features. Then, decompose the target project into various sub-items and call the target model corresponding to each sub-item. Combine the tagged features to predict the unit price of each sub-item to obtain the target unit price. Then, through deterministic calculation rules, couple the target unit price and project quantity data to output the summary data of sub-item project costs. Finally, based on the error distribution information obtained during model training, adjust the summary data of sub-item project costs to calculate the predicted result of the total project cost.

[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or itemized cost prediction device capable of performing the above functions. The following description uses an itemized cost prediction device as an example to illustrate this embodiment and the subsequent embodiments.

[0025] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0026] This application provides a tag-driven itemized cost prediction method. Referring to Figure 1, Figure 1 is a flowchart of the first embodiment of the tag-driven itemized cost prediction method of this application.

[0027] In this embodiment, the label-driven itemized cost prediction method includes steps S10 to S40: Step S10: Collect feature information related to the target item, match corresponding option labels to the feature information according to preset classification rules, and process the option labels through feature encoding to generate labeled features.

[0028] In this embodiment, feature information refers to the original related information such as material grade, construction process, number of floors, and structural type associated with the target project. Preset classification rules refer to pre-defined unified standards used to categorize various types of feature information. Option labels refer to unique identifiers corresponding to feature information categories and used to identify feature attributes. Feature encoding processing refers to the operation of converting non-numerical option labels into numerical forms that can be used for subsequent data processing. Tagmed features refer to the numerical set containing all valid feature information of the project, integrated after encoding processing.

[0029] As an optional implementation, this method receives all original feature information of the target project, identifies the attribute category of each feature information according to preset classification rules, and accurately matches a unique corresponding option label for each feature information. Completeness and consistency checks are performed on all features of the matched labels, and invalid labels and their corresponding features are removed. Then, each valid option label is numerically converted according to preset encoding rules, and the converted individual numerical codes are sequentially integrated according to feature category and original dimension order to form a complete tagged feature. This method can comprehensively retain project feature information, and the accuracy of label matching and encoding is high.

[0030] As an alternative implementation, after receiving the original feature information of the target project, the core feature information that plays a crucial role in subsequent prediction is first selected according to a preset priority rule, while secondary features are discarded. Then, the selected core feature information is batch-classified according to a preset classification rule, and corresponding option labels are uniformly matched for core features of the same category. The batch-matched labels are then validated as a whole. After confirming that there is no category confusion, a unified encoding rule is used to batch-convert the labels of the same category into numerical values. All converted numerical codes are then integrated according to the priority order of the core features to form labeled features. This method does not require processing all features; batch operations reduce processing steps, resulting in higher overall efficiency and less computation.

[0031] Step S20: After breaking down the target project into various sub-items, call the target model corresponding to each sub-item and combine it with the labeled features to predict the unit price of each sub-item to obtain the target unit price.

[0032] In this embodiment, a sub-item of work refers to a specific engineering unit formed by breaking down the target project according to its engineering function and work content. The target model refers to a pre-trained, dedicated model specifically adapted for predicting the unit price of each sub-item of work. The unit price of a sub-item refers to the cost price corresponding to a unit of work in each sub-item of work. The target unit price refers to the final unit price of each sub-item of work determined after prediction and verification by the model.

[0033] As an optional implementation method, based on industry-specific engineering classification standards and project technical parameters, the target project is broken down layer by layer from broad engineering categories. The scope of work, measurement objects, and attribute characteristics of each sub-item are determined, forming a complete list of sub-items. The attribute characteristics of each item in the list are precisely matched one-to-one with a pre-set model library to determine a unique target model for each item. Specific feature dimensions matching the attributes of each item are extracted from the tagged features, and these specific features are fully input into the corresponding target model for calculation to obtain the preliminary unit price for each item. The preliminary unit price is then validated against a reasonable price range set based on historical data, and the target unit price is obtained after correcting for deviations. This method involves detailed breakdown and accurate model matching, fully leveraging the impact of the specific characteristics of each item.

[0034] As an alternative implementation method, the target project is divided into several major categories based on its core functions. Each category is then broken down into smaller, more manageable sub-categories, defining the basic scope of each sub-item within that category without excessively refining individual attributes. A general target model is matched to the common characteristics of the major categories, and this model is shared by all sub-items within the same category. Common features of the major categories are extracted in batches from the tagged features, and necessary key individual features for each sub-item are added and integrated. The integrated features are then input into the general model in batches, simultaneously predicting the unit price of all sub-items within the same category. The prediction results are then checked for overall consistency, and significant outliers are removed before determining the unit price of each sub-item. This method simplifies the process through batch operations, offering fast processing speed, low computational load, strong adaptability to large-scale projects with multiple sub-items, and significant efficiency advantages.

[0035] Step S30: Through deterministic calculation rules, the target unit price and project quantity data are coupled and processed to output the summary data of sub-item project costs.

[0036] In this embodiment, deterministic calculation rules refer to pre-defined, fixed, and standardized calculation logic used for engineering cost accounting. Coupling processing refers to the operation of integrating and associating relevant data from different sources according to established rules and performing collaborative calculations. Project quantity data is a data set formed by quantifying and statistically analyzing the physical quantities and workloads of each sub-item of the project based on design documents and measurement specifications. Sub-item project cost summary data is a comprehensive cost data set formed by integrating the individual costs of each sub-item and the hierarchical summary results.

[0037] As an optional implementation method, the target unit price and corresponding project quantity data for each sub-item of the project are verified one by one to confirm that their measurement caliber and units are completely consistent. Then, according to deterministic calculation rules, the target unit price and corresponding project quantity of each sub-item are calculated together to obtain the individual cost of each sub-item of the project. Subsequently, according to the hierarchical division and category of the sub-items of the project, the data is summarized level by level. During the summarization process, the consistency and accuracy of the data at each level are verified in real time, and calculation deviations are corrected. Finally, a summary cost data of the sub-items of the project is formed, including the individual cost, the summary cost of the category, and key measurement indicators. This method has high calculation accuracy, data traceability at each stage, and is easy to audit and verify.

[0038] As an alternative implementation method, the target unit price and project quantity data are first categorized and integrated according to the functional categories of the sub-items, and the consistency of the target unit price and quantity measurement in the same category is verified in batches. Then, according to deterministic calculation rules, batch collaborative calculations are performed on the same category of sub-items to simultaneously obtain the individual cost of all sub-items under the same category. Subsequently, the individual costs of each category are summarized as a whole. After the summary, the rationality of the overall data is checked through preset verification rules, abnormal data is eliminated and quickly corrected, and finally, the summary data of sub-item project costs containing the summary cost of each category and the overall data is output. This method simplifies the processing flow through classified batch operations and has high computational efficiency.

[0039] Step S40: Based on the error distribution information obtained during model training, adjust the total cost summary data of the distributed sub-items of the project, and calculate the predicted result of the total project cost.

[0040] In this embodiment, the error distribution information refers to the distribution pattern and fluctuation range of the prediction errors for each component of the project, recorded during model training. The prediction result of the total project cost is the final prediction data, after error adjustment, which includes the overall project cost and related confidence information.

[0041] As an optional implementation method, error distribution information generated during model training is obtained. This information is then correlated with each sub-item cost in the summary data of sub-item project costs to determine the error fluctuation characteristics and correction coefficients corresponding to each sub-item cost. The corresponding sub-item costs are precisely adjusted according to their specific correction coefficients to ensure that the prediction deviation of each sub-item is specifically corrected. Then, all adjusted sub-item costs are summed, and the fluctuation range of the total project cost is derived by combining the error distribution patterns. Finally, a prediction result containing the specific total cost value and its corresponding confidence interval is obtained. This method can accurately correct the independent errors of each sub-item and effectively avoid error accumulation.

[0042] As an alternative implementation, error distribution information generated during model training is obtained. This error information is then categorized and statistically analyzed according to the functional categories of the sub-projects, determining the average error level and uniform correction ratio for each category. The total cost data for each sub-project is then divided according to its corresponding functional category, and the uniform correction ratio for each category is used to batch adjust the costs of all sub-items within that category. The adjusted cost data for each category is then aggregated to obtain a preliminary total cost. Finally, the confidence interval for the total project cost is calculated by combining the overall fluctuation range of the errors across categories, forming the final prediction result. This method simplifies the process through batch processing, offering fast processing speed, low computational load, and strong adaptability to large-scale projects with multiple sub-items.

[0043] For example, in the context of construction projects, tagged feature engineering and contextualized prediction are used: material grades, construction techniques, number of floors, structural types, etc., are encoded through "option tags," allowing the same model to generate contextualized unit price predictions under different construction conditions. This replaces cumbersome manual archives / rule bases, facilitating rapid configuration and scenario comparison. A computer implementation method couples tag-driven fine-grained unit price prediction for each component with engineering formulas and outputs a probabilistic total cost. The project is subdivided into several (e.g., 25) components, and machine learning predictions are performed on each component based on tagged features (option tags: one-hot), then summarized using engineering formulas (quantity multiplied by unit price, unit selection, unit index, etc.). This combination of "component-specific ML prediction, engineering formula coupling, and probabilistic output" significantly improves accuracy, interpretability, and decision-making value in engineering estimation practice, and is not included in traditional experience-based quotas or single overall models. Hybrid architecture: Coupled with data-driven prediction and engineering deterministic formulas, it couples the output of ML (unit price) with engineering-necessary deterministic calculations (quantity multiplied by unit price, unit selection, unit indicators, etc.) in a composable manner. This preserves interpretable engineering logic while leveraging data to extract nonlinear relationships. It improves industry adoptability and facilitates auditing and verification.

[0044] By combining tagged coding, itemized model prediction, and engineering formula coupling, the problems of long cost estimation cycles, reliance on manual experience, and insufficient accuracy in the early stages of construction projects are solved, thus improving estimation efficiency, accuracy, and interpretability, and providing reliable support for early project decision-making.

[0045] Based on any of the above embodiments, in Embodiment 2 of this application, step S10 includes steps A11 to A13: Step A11, unify the data format of the economic indicator table data of the target project, the information of the questionnaire filling and the data extracted by the model, and integrate them to obtain the original feature set.

[0046] In this embodiment, the economic indicator table data refers to structured data recording core economic parameters such as project land area, building area, and plot ratio. The guided questionnaire information refers to supplementary information collected through questionnaires regarding project construction techniques and location characteristics. The model-extracted data refers to structured data extracted from engineering modeling tools, including the number of components and the area of ​​each zone. The original feature set refers to a comprehensive data set containing all original features of the project, integrated after standardization of formatting.

[0047] As an optional implementation method, this approach analyzes the original formats of the target project's economic indicator table data, guided questionnaire information, and model-extracted data to determine the field types, storage structures, and representation standards for each data category, and establishes a unified data format standard covering all fields. Each field of the three data categories is then converted according to this unified standard, with the validity and consistency of the field data verified during the conversion process, and invalid data removed. The converted data from the three categories are then merged sequentially according to field relationships, missing related fields are added, and finally, a complete, uniformly formatted, and logically coherent set of original features is formed. This method maximizes the retention of all field information for each data category, achieving high completeness and accuracy in data integration, and providing comprehensive support for subsequent feature processing.

[0048] As an alternative implementation, Dynamo scripts incorporate a multi-version Revit model component parameter adaptation library. This library automatically identifies the current Revit version and component attribute definition rules of the model, establishing dynamic mapping relationships for component parameters with inconsistent names but identical functions across different versions. Simultaneously, by analyzing the parameter distribution patterns of similar components, it intelligently completes missing key feature data for components. For example, components without material labels are completed using mainstream materials of the same type. During extraction, the logical correlation between component counts and area / volume is verified in real-time, generating an extraction file containing parameter adaptation logs and completion data, which is then exported in a platform-readable format. This method solves the compatibility issues of data extraction from different Revit models, reduces manual completion workload, significantly improves the completeness and consistency of extracted data, and adapts to multi-version model project scenarios.

[0049] Step A12: Identify the corresponding feature type in the original feature set, and filter and distinguish the feature items corresponding to the feature type according to the preset category.

[0050] In this embodiment, feature type refers to the attribute category to which the feature information belongs. Preset categories are pre-defined standard category systems used for feature filtering. Feature items are specific data units with independent attribute meanings within the original feature set.

[0051] As an optional implementation, this method iterates through each feature item in the original feature set, analyzing its description, data attributes, and associated fields to identify the feature type corresponding to each item. Then, a preset category system is retrieved, and the identified feature types are compared and matched against preset categories one by one. Feature types that meet the preset category requirements are selected, and the successfully matched feature items are categorized according to the preset categories, forming subsets of feature items for each category. During the differentiation process, the compatibility between the feature types and preset categories is verified, and feature items with abnormal compatibility are removed. This method has high accuracy in feature recognition and differentiation, ensuring a high degree of fit between the selected feature items and the preset categories, providing a precise foundation for subsequent implementation.

[0052] Step A13: Based on the classification rules, match the option label that uniquely corresponds to each feature item.

[0053] As an optional implementation, a preset classification rule is retrieved and its core matching logic is parsed. The attribute information, descriptive keywords, and associated features of each feature item are extracted one by one. The core attributes of each feature item are aligned and compared with the label matching conditions in the classification rule. Once a perfect match is confirmed, a corresponding option label is assigned to the feature item, while the uniqueness of the matched labels is verified in real time. If a potential conflict occurs, the matching logic is adjusted backtracking to ensure that each feature item receives a unique and accurate corresponding option label. This method has extremely high label matching accuracy and effectively avoids label conflicts and mismatches.

[0054] For example, in the context of a construction project, the data formats of the economic indicators (building area 30,000 square meters, plot ratio 4.2), the guided questionnaire information (construction process: cast-in-place, material grade: high-strength steel reinforcement), and the data extracted by the random forest model (number of beams: 800, number of columns: 320) for a 20-story reinforced concrete office building project are unified. All data are standardized into a structured format of "feature name, value, unit," resulting in an original feature set containing 12 data items. The feature types in this set are identified and filtered according to preset categories such as structure-related, process-related, material-related, economic-related, and component-related, resulting in 8 valid feature items, including structure type and construction process. Based on the classification rule of "feature attribute, label mapping," unique corresponding option labels are matched for structure type ("reinforced concrete, frame-shear wall"), construction process ("cast-in-place construction"), and material grade ("high-strength steel reinforcement - Grade III"), completing the label matching for all feature items.

[0055] By integrating standardized formats and accurately classifying and matching data, the problems of chaotic multi-source feature data and non-unique label matching in construction projects have been solved, improving the standardization of feature data and the accuracy of label matching, thus laying a solid foundation for subsequent coding processing.

[0056] Based on any of the above embodiments, in Embodiment 3 of this application, step S10 includes steps B11 to B13: Step B11, verify the integrity and consistency of the option tags, remove invalid tags that do not conform to the preset coding specifications, and output valid option tags.

[0057] In this embodiment, completeness refers to the state where all feature items to be matched with the tag have obtained the corresponding tag without any missing items. Consistency refers to the state where the format, expression, and corresponding logic of each tag do not conflict with the preset requirements. The preset coding standard is a pre-set standard for judging the qualification of tag format, attribute correspondence, etc. Invalid tags are tags that do not conform to the preset coding standard, or have missing or conflicting elements. Valid option tags are a set of tags that conform to the standard and are complete and consistent after invalid tags have been verified and removed.

[0058] As an optional implementation, each matched option tag and its corresponding feature item are retrieved one by one. First, completeness is verified to ensure all feature items match the tags and no key information is missing. Next, consistency is verified, checking that the tag's expression and format do not conflict with other tags of the same category or with preset logic. Then, the format requirements and attribute correspondences of the tags are checked against preset coding specifications. Tags with missing, conflicting, or non-compliant elements are directly deemed invalid and removed. Finally, all tags that pass the triple verification of completeness, consistency, and coding specifications are collected to form and output a set of valid option tags. This method has a comprehensive and meticulous verification process, effectively eliminating various invalid tags to the greatest extent possible. The output of valid tags has extremely high accuracy and consistency, providing a reliable guarantee for subsequent feature coding.

[0059] Step B12: The valid option labels are numerically converted according to the preset discrete feature encoding rules to obtain the one-dimensional numerical code corresponding to the feature item.

[0060] In this embodiment, the preset discrete feature encoding rules refer to a pre-defined standardized rule system used to convert non-numerical valid option labels into numerical forms. Numerical conversion is the operation of converting non-numerical labels into numerical forms that can be used for data processing. One-dimensional numerical encoding is the single-dimensional numerical representation of a single feature item after encoding conversion.

[0061] As an optional implementation, a preset discrete feature encoding rule is retrieved, and the one-to-one correspondence and format requirements between labels and values ​​in the rule are analyzed. The core attributes and category information of each valid option label are extracted one by one, and the label information is precisely matched with the encoding rule. A unique corresponding value is assigned to each label according to the matching result, and the uniqueness and rule adaptability of the numerical encoding are simultaneously verified to ensure no encoding conflicts or mismatches. Finally, a one-dimensional numerical code that conforms to the rule requirements and corresponds individually to each feature item is obtained. This method has extremely high encoding accuracy and can strictly follow the rules to achieve a precise mapping between labels and values.

[0062] Step B13: Matrix-integrate the one-dimensional numerical encoding, the category division of the original features, and the item feature dimension sorting to generate the labeled features containing all item feature encoding information.

[0063] In this embodiment, matrix integration refers to the operation of organizing discrete data from multiple dimensions into a structured matrix according to preset rules. Category classification refers to the result of categorizing feature items according to their attributes. The item feature dimension sorting is a preset priority or logical order of feature items.

[0064] As an optional implementation, the original feature classification results are analyzed to determine the feature items and corresponding one-dimensional numerical codes contained in each category. Then, the feature dimension sorting rules are retrieved to determine the column positions of all feature items in the matrix. A matrix framework is constructed hierarchically by category, with each layer corresponding to a feature category. Subsequently, each one-dimensional numerical code is sequentially filled into its corresponding matrix position according to its category and dimension, ensuring complete alignment between the code and the feature item, category, and sorting. After filling, the matching of the number of rows and columns in the matrix with the number of features and categories is verified, misaligned codes are corrected, and finally, hierarchical and accurately mapped labeled features are generated. This method has a clear matrix structure, and the correspondence between features, categories, and sorting is traceable.

[0065] For example, in the scenario of a construction project, the option labels for a 15-story frame structure residential building are verified to include "Structure Type - Frame", "Construction Technology - Prefabricated", "Material Grade - High-Strength Steel Reinforcement", "Location Characteristics - Urban Area", and "Material Grade - Unknown". Completeness is checked to confirm no missing feature items, and consistency is confirmed by the uniform label format. Invalid labels such as "Material Grade - Unknown" that do not conform to the preset coding specifications are removed, resulting in four valid option labels. Following the preset sequential discrete feature coding rules, the valid labels are sequentially converted into numerical values, yielding one-dimensional numerical codes 1, 2, 3, and 4 corresponding to the feature items. Based on the original feature category classification (structure-related, technology-related, material-related, location-related) and the project feature dimension sorting, the above one-dimensional numerical codes are matrix-integrated to generate a labeled feature [1, 2, 3, 4] containing all project feature coding information.

[0066] By employing label verification and purification, standardized encoding conversion, and structured integration, the problems of invalid feature labels, chaotic encoding, and disordered integration in construction projects have been resolved, improving the standardization and encoding accuracy of feature data and providing high-quality input for subsequent model predictions.

[0067] Based on any of the above embodiments, in Embodiment 4 of this application, step S20 includes steps C11 to C13: Step C11, parse the economic indicator table data, engineering technical parameters and industry sub-item engineering classification standards of the target project, determine the applicable dismantling level and classification boundary of the target project, and integrate to obtain dismantling rules.

[0068] In this embodiment, the economic indicator table data consists of structured data recording core economic parameters such as project scale and investment amount. Engineering technical parameters are key indicators describing the project's structural type, construction technology, and other technical attributes. The industry-standard classification of sub-items is a standardized specification commonly used in the industry for project breakdown and classification. The breakdown hierarchy refers to the hierarchical structure of breaking down a project from its whole to its parts. The classification boundary defines the scope of each sub-item. The breakdown rules are standardized operational guidelines formed by integrating the breakdown hierarchy and classification boundaries.

[0069] As an optional implementation method, the economic indicator data of the target project is analyzed to extract core information such as project scale and investment. Then, the engineering technical parameters are broken down to determine the structural type and technical attributes of the construction process. Simultaneously, industry-specific classification standards for sub-items of engineering work are studied to grasp the general breakdown logic. These three types of information are cross-referenced to determine the applicable breakdown level for the project, defining the classification boundaries of each sub-item of engineering work at each level and eliminating overlapping items with ambiguous boundaries. Finally, the breakdown rules are integrated according to the hierarchical order and boundary definitions to form a logically coherent and clearly defined breakdown rule. This method offers high accuracy in breakdown rules, unambiguous classification boundaries, adaptability to the personalized characteristics of projects, and no omissions or overlaps in subsequent breakdowns.

[0070] Step C12: According to the decomposition rules and combined with the functional attributes of the project, the target project is divided into several major engineering categories. Each major category corresponds to a defined engineering scope and measurement caliber, thus obtaining the decomposition result.

[0071] In this embodiment, the functional attributes of an engineering project refer to the core purpose-related attributes of the project. The engineering category is a primary engineering classification based on functional attributes. The engineering scope refers to the specific work content covered by each engineering category. The measurement standard is a unified standard used to quantify the workload of each engineering category. The decomposition result is structured data formed after decomposition, containing each engineering category, its corresponding scope, and the measurement standard.

[0072] As an optional implementation method, this approach comprehensively analyzes the hierarchical requirements and classification boundary standards of the decomposition rules. It then systematically examines the engineering functional attributes of each component of the target project, precisely aligning these attributes with the decomposition rules. The specific engineering scope of each major engineering category is defined according to the rules, determining the boundaries of the work content within that scope. Simultaneously, the corresponding measurement standards for each major category are determined based on industry measurement specifications, ensuring consistency between the standards and the engineering scope. During the decomposition process, real-time verification is performed to ensure no overlap or omissions in the scope of each major category, correcting any ambiguity in the boundaries. The final result is a decomposition that includes all major engineering categories, clearly defined scopes, and unified measurement standards. This method offers high decomposition accuracy, clear scope of each major category, and unified measurement standards.

[0073] Step C13: Based on the decomposition rules, refine the decomposition results hierarchically, determine the specific work content and measurement object of each sub-item project, and obtain each sub-item project.

[0074] In this embodiment, hierarchical refinement refers to a progressive operation that breaks down projects from broad categories to subcategories and individual projects. Specific work content refers to the detailed construction tasks that each sub-item of the project needs to complete. The measurement object is the concrete carrier used to quantify the quantity of that sub-item of the project.

[0075] As an optional implementation method, the decomposition rules and the results of the already decomposed major project categories are retrieved. Following the hierarchical logic set by the rules, the decomposition is refined layer by layer from each major project category downwards. First, it is decomposed into several subcategories, and then further decomposed into individual projects. During the decomposition process, the specific work content of each level of sub-item project is clarified according to the rules, the boundary scope of the construction task is defined, and the corresponding measurement object is determined according to the measurement caliber to ensure accurate matching between work content and measurement object. The decomposition results are verified level by level to confirm that there is no overlap in work content or confusion in measurement objects, and the hierarchical decomposition deviation is corrected, ultimately obtaining complete information for each sub-item project. This method has a clear decomposition hierarchy and accurate definition of work content and measurement object.

[0076] For example, in the context of a construction project, the economic indicator data (building area of ​​40,000 square meters, 25 floors above ground and 2 floors underground), engineering technical parameters (shear wall structure, raft foundation, prefabricated exterior walls), and industry classification standards for sub-items of a 25-story shear wall structure residential project are analyzed to determine the three-level decomposition hierarchy of "major category, sub-category, and individual item" and the functional attribute classification boundaries, and a standardized decomposition rule is obtained. According to this standardized decomposition rule and combined with the engineering functional attributes, the project is divided into three major engineering categories: civil engineering, installation engineering, and decoration engineering. Civil engineering is defined to cover the scope of foundation / main structure, etc., and the measurement caliber is volume / area; installation engineering includes the scope of water supply and drainage / electrical, etc., and the measurement caliber is quantity / length, resulting in the decomposition result. Based on the hierarchical refinement of this result using the decomposition rule, 12 sub-items of engineering are identified, including raft foundation engineering (work content: excavation / rebar binding / concrete pouring, measurement object: volume) and shear wall main structure engineering (work content: formwork erection / rebar installation / pouring, measurement object: area).

[0077] By standardizing hierarchical decomposition and defining boundaries, the problems of chaotic project breakdown and unclear work content and measurement objects have been solved, improving the standardization and accuracy of the breakdown and laying a solid foundation for subsequent unit price forecasting.

[0078] Based on any of the above embodiments, in Embodiment 5 of this application, step S20 includes steps D11 to D13: Step D11, according to the functional attributes of the sub-item project, matching the corresponding model in the preset model library, and extracting the exclusive feature dimension of each sub-item from the labeled features, generating the mapping relationship between each sub-item project and the target model and the sub-item exclusive features.

[0079] In this embodiment, functional attributes refer to the core uses and technical characteristics of each sub-item of the project. The preset model library is a collection of various models pre-stocked to adapt to sub-items with different functional attributes. Specific feature dimensions are a subset of features that are only applicable to a specific sub-item and affect its unit price prediction. Mapping relationships are the one-to-one correspondence between sub-items and their corresponding target models.

[0080] As an optional implementation method, the functional attributes of each sub-item of the project are analyzed one by one to determine the core technical characteristics and prediction requirements of each sub-item. The adaptive functional attributes and input requirements of all models in the preset model library are traversed, and the functional attributes of each sub-item are aligned and compared with the model adaptation conditions to determine the uniquely suitable target model. Simultaneously, from the labeled features, based on the functional attributes and model input specifications of the sub-item, specific feature dimensions related only to that sub-item are precisely selected and extracted, irrelevant features are eliminated, and a unique mapping relationship between each sub-item and its corresponding target model is established, generating highly targeted sub-item-specific features in parallel. This method has high model matching accuracy and no redundant information in the sub-item-specific features.

[0081] Step D12: In accordance with the mapping relationship, the inference process of the corresponding target model is triggered sequentially. The specific features of each item are substituted into the target model one by one to complete the calculation and output the preliminary unit price prediction result.

[0082] In this embodiment, the inference process is the execution flow of the target model performing calculations and predictions based on input features. The preliminary unit price prediction result is the unit price data of the sub-items of the project output by the model after calculation, without verification and correction.

[0083] As an optional implementation method, the reasoning process of the target model corresponding to each sub-item is triggered sequentially according to the order and mapping relationship of the sub-items. The next model is started only after the previous model's reasoning is completed. The unique features of each sub-item are fully extracted and input into the corresponding target model one by one, ensuring complete compatibility between the features and the model input interface. After receiving the features, the model completes the entire process calculation according to preset logic, recording key parameters in real time. After each model's calculation, the preliminary unit price prediction result for that sub-item is output immediately, and the results are temporarily stored for subsequent processing. This method exhibits strong adaptability between model triggering and feature input, traceable calculation process, and high accuracy of preliminary prediction results.

[0084] Step D13: Based on the preset reasonable range standard for unit price of sub-items and the historical prediction error correction rules, verify and optimize the preliminary unit price prediction results, remove outliers that exceed the reasonable range and adjust the prediction deviation to obtain the target unit price of each sub-item.

[0085] In this embodiment, the preset reasonable range standard for the unit price of each sub-item is a criterion for determining the normal fluctuation range of the unit price of each sub-item, set in advance based on industry quotas and market conditions. The historical prediction error correction rule is an adjustment criterion formulated based on the deviation pattern between predicted data and actual data of similar past projects. The prediction deviation is the difference between the preliminary result and the actual reasonable unit price.

[0086] As an optional implementation method, the preliminary unit price prediction results for each sub-item of the project are retrieved one by one. These preliminary unit price prediction results are then precisely compared with the preset reasonable range standard for the corresponding sub-item unit price. It is determined whether the result falls within the range; if it exceeds the range, it is directly identified as an outlier and marked for removal. Next, based on historical prediction error correction rules, combined with the functional attributes, characteristic dimensions, and error patterns of similar sub-items in the past, a correction coefficient is calculated specifically. This correction coefficient is used to adjust the preliminary results that were not removed. After correction, the results are checked again to ensure they conform to the reasonable range. Once it is confirmed that there is no deviation, the target unit price for that sub-item is obtained. This method has extremely high accuracy in verification and optimization, with a small deviation between the target unit price and the actual reasonable value.

[0087] For example, in the scenario of a construction project, based on the functional attributes of the raft foundation (functional attribute: load-bearing foundation), shear wall main structure (functional attribute: structural load-bearing), and water supply and drainage installation (functional attribute: water supply and drainage) of a 25-story shear wall structure residential building, random forest model, gradient boosting model, and support vector machine model are matched respectively in a preset model library. Raft foundation-specific features [1, 3], shear wall main structure-specific features [2, 4], and water supply and drainage installation-specific features [3, 5] are extracted from the labeled features [1, 2, 3, 4, 5]. This generates the mapping relationship between each item and the target model, as well as the item-specific features. The corresponding model inference process is triggered sequentially according to the mapping relationship, and the specific features are substituted into the model one by one to complete the calculation, outputting preliminary unit price prediction results of 230 yuan / square meter, 2800 yuan / square meter, and 320 yuan / square meter. Based on the preset reasonable unit price range for each component of the project (raft foundation 180-220 yuan / square meter, shear wall main body 2600-2900 yuan / square meter, water supply and drainage installation 300-350 yuan / square meter) and the historical prediction error correction rules, the outlier of 230 yuan / square meter for the raft foundation was removed and adjusted to 215 yuan / square meter, resulting in target unit prices of 215 yuan / square meter, 2800 yuan / square meter, and 320 yuan / square meter for each component of the project.

[0088] By employing precise model matching, exclusive feature extraction, and dual-standard verification, the problems of large deviations and outlier interference in the prediction of unit prices for construction projects have been solved, thereby improving the accuracy and reliability of unit price prediction.

[0089] Based on any of the above embodiments, in Embodiment Six of this application, step S30 includes steps E11 to E13: Step E11, verify the measurement unit of the target unit price and the measurement caliber of the corresponding sub-item quantity, and establish the association relationship between each sub-item and the target unit price and the corresponding quantity.

[0090] In this embodiment, the measurement caliber is the standardized basis for defining the statistical scope and calculation rules of engineering quantities.

[0091] As an optional implementation method, the target unit price and its unit of measurement for each sub-item of the project are extracted one by one. Simultaneously, the corresponding quantity data and measurement caliber are retrieved, and the consistency of the expression standards and quantitative logic of the unit of measurement and measurement caliber are verified. It is confirmed that the two are completely matched in statistical dimensions, with no unit conflicts or caliber deviations. After successful verification, the identification information of the sub-item of the project is precisely bound to the corresponding target unit price and quantity data. Key verification nodes in the binding process are recorded, forming an independent association file for each sub-item, ensuring the traceability of the association information. This method has extremely high verification accuracy, no mismatches in the association relationships, and can completely avoid subsequent calculation errors caused by inconsistencies in units or calibers.

[0092] Step E12: Based on the aforementioned relationship, a preset deterministic calculation rule is applied to multiply the target unit price of each sub-item project with its corresponding quantity of work to obtain the individual cost data of each sub-item project.

[0093] In this embodiment, the preset deterministic calculation rules are pre-defined, fixed, and standardized calculation logic used for individual cost accounting. Individual cost data are the independent cost results obtained from the calculation of a single sub-item of the project.

[0094] As an optional implementation method, the relationships between each sub-item of the project are retrieved one by one, and the corresponding target unit price and quantity data are extracted to reconfirm the consistency of their measurement logic. Then, according to preset deterministic calculation rules, the target unit price and quantity of the sub-item are calculated collaboratively. During the calculation process, the calculation steps and key parameters are recorded in real time, and the rationality of the calculation results is verified simultaneously to ensure that the results conform to the normal fluctuation range of the sub-item cost, with no significant calculation deviation. Finally, independent individual cost data for each sub-item of the project is obtained, forming a complete set of individual costs. This method has high calculation accuracy, the calculation process for each sub-item is traceable, facilitating subsequent verification and correction, and effectively avoiding calculation errors.

[0095] Step E13: Summarize the hierarchical division and major category affiliation of the sub-items of the project, and combine the individual cost data to form the summary cost data of the sub-items of the project.

[0096] In this embodiment, hierarchical aggregation refers to the operation of gradually integrating cost data from bottom to top according to the breakdown hierarchy of sub-items. The hierarchical division of sub-items is a multi-level structural system formed after the project is broken down. The project category is the first-level project category to which each sub-item belongs. The total cost data of sub-items is a structured and hierarchical cost set formed after integrating the costs of all levels of sub-items.

[0097] As an optional implementation method, the complete hierarchical division of the sub-projects and their respective major project categories are first determined. Starting from the lowest level of individual projects, the cost data of each individual project is extracted one by one and assigned to the corresponding sub-category according to their affiliation. The completeness and accuracy of the attribution of all individual costs within the sub-category are verified. Then, the cost data of each sub-category is aggregated upwards to the corresponding major project category, and the aggregation logic and data consistency of the sub-category costs within the major category are checked to eliminate duplicates or omissions. Finally, according to the complete hierarchy of individual projects, sub-categories, and major categories, the aggregated cost data of the sub-projects is formed, which includes cost details of each level, totals of the major category, and affiliation relationships. This method has a rigorous aggregation logic, clear hierarchical relationships, and a perfect match between cost data and affiliation.

[0098] For example, in the scenario of a construction project, the target unit price measurement units (raft foundation 215 yuan / square meter, shear wall main body 2800 yuan / square meter, water supply and drainage installation 320 yuan / square meter) of each component of a 25-story shear wall structure residential building are verified to match the measurement scope (area, area, length) of the corresponding component quantities. This confirms that the measurement units and measurement scopes are consistent, establishing the correlation between raft foundation engineering - 215 yuan / square meter - 5000 square meters, shear wall main body engineering - 2800 yuan / square meter - 38000 square meters, and water supply and drainage installation engineering - 320 yuan / m - 45000m. Based on this correlation, the preset deterministic calculation rule "single item cost = target unit price × corresponding quantity" is applied to perform multiplication, yielding single item cost data of 1.075 million yuan, 106.4 million yuan, and 14.4 million yuan. The cost data is divided into "single item, sub-category, and major category" (single items belong to the major categories of civil engineering, civil engineering, and installation engineering), and the cost data of each item is summarized at each level to form the total cost data of each sub-item of the project (107.475 million yuan for the major category of civil engineering, 14.4 million yuan for the major category of installation engineering, and a total of 121.875 million yuan).

[0099] By verifying measurement consistency, standardizing calculations, and hierarchical summarization, the problems of conflicting cost calculation methods and chaotic data associations in construction projects have been resolved, thereby improving the accuracy and structure of cost data.

[0100] Based on any of the above embodiments, in Embodiment 7 of this application, step S40 includes steps F11 to F13: Step F11, based on the error distribution information obtained during model training, associating and matching the error characteristics corresponding to each of the sub-items with the total cost data of the sub-items, and determining the error fluctuation range and error distribution pattern of each sub-item cost.

[0101] In this embodiment, error distribution information is a set of relevant data recorded during model training, including error magnitude, frequency of occurrence, and distribution pattern. Error characteristics are the error-related attributes exhibited by each sub-item in cost prediction. Error fluctuation range is the interval between the maximum and minimum values ​​of the cost error for each sub-item. Error distribution pattern refers to the inherent characteristics of error within the fluctuation range, such as frequency of occurrence and trend of change.

[0102] As an optional implementation method, this approach comprehensively analyzes the error distribution information obtained during model training to determine the overall error distribution pattern and key influencing parameters. Then, it extracts the error characteristics corresponding to each sub-item of the project, including core attributes such as error type and related influencing factors. The error characteristics of each sub-item are precisely matched with the corresponding sub-item cost details in the sub-item cost summary data. The possible range of error values ​​for that sub-item cost is analyzed by comparing it with the error distribution information, and the frequency and trend of error occurrence within that range are identified. The error fluctuation range and error distribution pattern of each sub-item cost are determined one by one, and the consistency between the matching logic and the analysis results is simultaneously verified. This method has high accuracy in matching error characteristics with cost data, and the determined error fluctuation range and distribution pattern closely match the actual situation of individual sub-items, demonstrating strong reliability.

[0103] Step F12: Based on the error distribution pattern, adjust the cost of each of the sub-items of the project using the corresponding deviation correction algorithm, and retain the cost composition relationship corresponding to the project logic to form the correction data.

[0104] In this embodiment, the deviation correction algorithm is a standardized calculation method adapted to the error distribution pattern to adjust cost data deviations. The cost composition relationship corresponding to the engineering logic is a cost association structure formed according to the engineering construction logic and hierarchical affiliation. The corrected data is a precise cost set that retains the cost composition relationship after deviation adjustment.

[0105] As an optional implementation method, this approach analyzes the error distribution patterns of each sub-item of the project, identifies its error change trends and core influencing factors, matches a dedicated deviation correction algorithm adapted to these patterns, and adjusts the cost of each sub-item accordingly based on the algorithm's logic. During the adjustment process, the cost composition relationships corresponding to the project logic are verified in real time to ensure that the inclusion relationships between sub-item costs and their respective subcategories and major categories are not disrupted. After correction, the error is checked again to ensure it meets the expected range, generating correction data for each individual sub-item. Finally, all sub-item correction data are integrated while maintaining the overall cost composition logic consistency. This method boasts extremely high correction accuracy, maximally conforms to the error characteristics of individual sub-items, and ensures unbiased cost composition relationships.

[0106] Step F13: Based on the corrected data, sum and calculate the corrected costs of each item, derive the confidence interval of the total project cost by combining the error fluctuation range, and determine the prediction result.

[0107] In this embodiment, the itemized corrected cost is the adjusted cost data for each sub-item of the project in the corrected data. The confidence interval for the total project cost is a high-probability range of values ​​that includes the true total project cost, derived based on the itemized error.

[0108] As an optional implementation method, the correction cost of each sub-item of the corrected data is extracted one by one, and summed sequentially according to the project level to obtain the total correction cost of each sub-category. Then, the total correction cost of each item is summed to form the total correction cost of the project. At the same time, the error fluctuation range of each item is retrieved one by one, and the superposition logic and influence of each error in the summation are analyzed. Based on the error propagation law, the high-probability fluctuation range of the total project cost is derived, and the upper and lower limits of the confidence interval are defined. Finally, the summation results of the sub-items are integrated with the confidence interval to form a prediction result containing specific values ​​and interval ranges. This method has a rigorous summation logic, the confidence interval derivation closely matches the actual error of each sub-item, and the prediction results have high accuracy and reliability.

[0109] For example, in the scenario of a construction project, a user prepares an Excel template for the engineering economic and technical indicators of an 18-story frame-shear wall structure residential community project, fills in complete data such as land area of ​​8,000 square meters, total building area of ​​32,000 square meters, above-ground saleable area of ​​28,000 square meters, civil defense area of ​​1,200 square meters, plot ratio of 4.0, and greening rate of 35%, and uploads it to the Streamlit platform. After parsing, the platform does not prompt any missing fields. Subsequently, a guided questionnaire was filled out on the Web UI, with the earthwork volume of 12,000 cubic meters and the foundation pit support area of ​​3,000 square meters entered. Features such as "raft foundation" and "cast-in-place construction technology" were selected. After the platform verified the data in real time, an estimation was triggered. The backend merged the Excel and questionnaire data, unified the units, filled in blanks, and trimmed the abnormal range. The corresponding XGBoost model prediction was loaded, and a project estimation table for each sub-item was generated, including an estimated amount of 3.6 million yuan for earthwork (300 yuan / cubic meter) and 5.4 million yuan for foundation pit support (180 yuan / square meter). The table supports exporting to Excel (Project_Estimate_XX Community_v1.2.xlsx), CSV, and PDF formats. Relying on the automated feature encoding matching and lightweight calling mechanism of the target model, the generation and output of the prediction results of the unit price of all sub-items and the total project cost can be completed within seconds.

[0110] By using template-based import, guided questionnaire-based structured input, and XGBoost model prediction, the problems of chaotic input, low efficiency, and insufficient accuracy in construction project cost estimation have been solved, thus improving the standardization of estimation and the reliability of results.

[0111] Based on any of the above embodiments, in Embodiment 8 of this application, referring to Figure 2, which is a flowchart of the eighth embodiment of the tag-driven itemized cost prediction method of this application, after step S40, steps G11~G14 are further included: Step G11, collecting the actual total engineering cost and actual cost data of each item of the target project, comparing the prediction results with the actual cost data, analyzing the influencing factors, and outputting a deviation analysis report.

[0112] In this embodiment, the total actual project cost is the sum of all project expenses actually incurred after project completion. The actual cost data for each sub-item is the actual cost incurred after the completion of each sub-item. Influencing factors are the relevant factors that cause deviations between predicted and actual costs. The deviation analysis report is a structured document integrating the comparison results and the analysis of influencing factors.

[0113] As an optional implementation method, this approach comprehensively collects data on the total actual cost of the target project and the actual cost of each sub-item. It then precisely compares each sub-item with the predicted costs according to its hierarchical classification, calculating the degree of deviation for each sub-item and the total cost. The specific influencing factors for each deviation are analyzed from multiple dimensions, including market prices, construction techniques, and change adjustments, identifying primary and secondary influencing factors and their transmission paths. Finally, the comparison results, deviation data, and influencing factor analysis are integrated to output a detailed deviation analysis report in a standardized structure. This method offers precise comparisons, in-depth analysis of influencing factors, and reports with extremely high reference value.

[0114] Step G12: Analyze the key features and corresponding option labels that cause the deviation in the deviation analysis report, and generate the correlation analysis results between the degree of deviation and the contribution of the features.

[0115] In this embodiment, the key feature causing the deviation is the feature dimension that plays a major role in cost deviation. The degree of deviation is the extent to which the predicted result deviates from the actual cost. Feature contribution is the weight of the influence of a single feature on the deviation.

[0116] As an optional implementation method, this approach extracts the key features and corresponding option labels leading to deviations from the deviation analysis report one by one, and analyzes the attribute type and influence path of each feature. Each key feature is then matched with its corresponding degree of deviation, analyzing the correlation strength between feature frequency, influence range, and degree of deviation. The contribution weight of each feature to the deviation is quantified, and the differences in core contributing features under different degrees of deviation are determined. This integration forms a one-to-one correspondence between each key feature, its corresponding label, the degree of deviation, and its contribution, generating refined correlation analysis results. This method boasts extremely high correlation accuracy, clearly identifying the core contributing features at different degrees of deviation, providing a precise basis for model optimization.

[0117] As an alternative implementation, the platform constructs a feature association model based on the extracted model data, automatically matches questionnaire completion cases from similar historical projects, intelligently recommends highly suitable questionnaire options for unclear features in the current project, and simultaneously verifies the logical correlation between the completed workload and the model-extracted data in real time. If logical conflicts exist, an alert is triggered and the cause of the conflict is analyzed. After the questionnaire is completed, a feature completeness score is generated, indicating the impact of missing key features on the estimation accuracy. This method reduces the randomness of questionnaire completion, lowers the error rate of manual completion, and improves the completeness and rationality of feature input.

[0118] Step G13: According to the preset weight adjustment rules and in combination with the correlation analysis results, update the weight parameters corresponding to each feature and output the optimized feature weight configuration.

[0119] In this embodiment, the preset weight adjustment rule is a standardized criterion set in advance for adjusting weight parameters based on feature contribution. The weight parameter corresponding to each feature is the numerical value of the importance of each feature in the prediction model. The optimized feature weight configuration is a set of feature weights that has been adjusted to better reflect the actual influence patterns.

[0120] As an optional implementation, the contribution data of each feature in the association analysis results is extracted one by one. The adjustment direction and magnitude of each feature's weight are determined by comparing it with preset weight adjustment rules. The corresponding weight parameters are then updated progressively according to the feature's contribution level and the intensity of its bias influence. During the update process, the fit between the weight adjustment and the rules is verified in real time to ensure that each feature weight matches its actual contribution. Finally, all updated weight parameters are integrated to form a refined and optimized feature weight configuration. This method has extremely high weight adjustment accuracy and can maximize the alignment with the actual influence patterns of features.

[0121] Step G14: Integrate the optimized feature weight configuration into the target model training process corresponding to each of the sub-projects, retrain the model using historical datasets to solidify the weight adjustment effect, and output the updated target model and weight configuration file to complete the optimization loop of the prediction method.

[0122] In this embodiment, the historical dataset is a collection of relevant data such as features and costs from similar past projects. Iterative model training is the operation of repeating the training process to enhance the effect of weight adjustment. Solidifying the weight adjustment effect is the process of stabilizing the optimized weights on the model. The updated target model is a prediction model with improved performance after retraining. The weight configuration file is a structured file that records the optimized feature weight data. The optimization loop of the prediction method is a complete optimization cycle from weight adjustment and model training to outputting the updated result.

[0123] As an optional implementation, the optimized feature weight configurations are integrated into the target model training process for each sub-item of the project. The historical dataset corresponding to that sub-item is retrieved, and iterative training is conducted step-by-step according to the model training specifications. After each round of training, the effect of weight adjustments on prediction accuracy is verified, and training parameters are continuously optimized until the weight adjustment effect is stable and solidified. After training is completed, the updated target model and corresponding weight configuration file for that sub-item are output. This iterative optimization of all sub-item models is completed sequentially, ultimately forming a complete set of updated results. This method achieves extremely high model and weight adaptation accuracy, maximizes the effect of weight optimization, and significantly improves prediction accuracy.

[0124] As an alternative implementation, the platform incorporates a model selection decision tree. Based on the functional attributes, feature complexity, and historical prediction accuracy data of the sub-projects, it automatically determines whether to use a single XGBoost model or a multi-model fusion strategy combining XGBoost with random forests and support vector machines. A dynamic weight allocation mechanism is employed for the fused models, adjusting the weight ratio in real time according to the suitability of each model to the current sub-project features. During the computation, the prediction contribution of each model is recorded, generating a report on the model selection criteria and weight allocation. Finally, the fused sub-project-level estimation results are output. This method automates and personalizes model selection and weight allocation, improves the prediction accuracy of complex feature sub-projects, and adapts to the estimation needs of different types of sub-projects.

[0125] As an alternative implementation, an incremental learning framework is introduced, dividing the historical dataset into a base dataset and an incremental dataset. After the optimized feature weight configuration takes effect, incremental training samples are constructed only based on the actual cost data, deviation analysis results, and correlation analysis data of newly added projects, without the need to retrain the entire base dataset. By dynamically adjusting the learning rate to adapt to changes in the feature distribution of incremental samples, the weight parameters of the target model for each sub-item of the project are quickly updated. Simultaneously, the model performance comparison data before and after incremental training is saved, a model iteration optimization report is generated, and rollback to the historical best model version is supported. This method significantly reduces the time cost of model iteration training, lowers computational resource consumption, adapts to high-frequency project data update scenarios, and can quickly respond to market changes and project feature evolution.

[0126] For example, referring to Figure 3, which shows the cost unit price prediction result of this application, the user prepares the "Economic Indicators Table.xlsx" template provided by the platform and fills in the basic data of a residential project: land area of ​​10,000 square meters, total building area of ​​45,000 square meters, above-ground saleable area of ​​40,000 square meters (including 35,000 square meters of residential space and 5,000 square meters of commercial space), building area of ​​each function (underground garage of 5,000 square meters and supporting rooms of 800 square meters), civil defense area of ​​1,200 square meters, plot ratio of 4.5, and greening rate of 30%. After confirming that the key fields are complete, the user uploads the data to the Streamlit platform. The platform automatically parses and displays the preview results, indicating that there are no missing fields or that the types do not match. Subsequently, the user completes the guided questionnaire in the Web UI, filling in the earthwork and foundation engineering section with an earthwork volume of 20,000 cubic meters and a foundation pit support area of ​​3,000 square meters, and selecting features such as "raft foundation," "cast-in-place concrete technology," and "mid-range decoration of residential public areas." The platform verifies the compliance of the fields in real time. Next, the estimation is run. The backend first merges the Excel and questionnaire data, performs data cleaning operations such as unit standardization and blank value filling, and then loads the corresponding XGBoost model based on the selected region to perform prediction. After post-processing and table formatting, the interface displays a breakdown of project estimates (including indicators: earthwork engineering estimated at 160.12 yuan / cubic meter, foundation pit support engineering at 1727.75 yuan / square meter, foundation engineering at 287.70 yuan / square meter, underground civil engineering at 2637.95 yuan / square meter, and above-ground civil engineering at 2268.45 yuan / square meter). The project cost is calculated as follows: square meters, garage and equipment room 570.68 yuan / square meter, residential public area decoration 2280.00 yuan / square meter, interior decoration 1368.23 yuan / square meter, commercial interior decoration 2054.57 yuan / square meter, and ancillary room interior decoration 570.19 yuan / square meter. It also displays "The estimated development cost per square meter based on current project parameters is 2945.99 yuan / square meter." Users can view the file and export it to an Excel file (Project_Estimate_XX Residential Project_v1.2.xlsx).

[0127] By automatically extracting BIM model data, supplementing questionnaires, and combining XGBoost model prediction, the problems of tedious data extraction and incomplete input for construction project cost estimation have been solved, thus improving estimation efficiency and data accuracy.

[0128] This application provides an itemized cost prediction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the tag-driven itemized cost prediction method in Embodiment 1 above.

[0129] Referring to Figure 4 below, a schematic diagram of a structure suitable for implementing the itemized cost prediction device of this application embodiment is shown. The itemized cost prediction device in this application embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, supply chain cost analysis servers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), itemized cost prediction cloud terminals, etc., as well as fixed terminals such as engineering project itemized cost prediction servers, desktop computers, etc. The itemized cost prediction device shown in Figure 4 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0130] As shown in Figure 4, the itemized cost prediction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the itemized cost prediction device. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the cost forecasting device to communicate wirelessly or wiredly with other devices to exchange data. Although the cost forecasting device with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0131] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0132] The itemized cost prediction device provided in this application employs the tag-driven itemized cost prediction method described in the above embodiments, which can solve the technical problem of balancing efficiency and accuracy in cost prediction. Compared with the prior art, the beneficial effects of the itemized cost prediction device provided in this application are the same as those of the tag-driven itemized cost prediction method provided in the above embodiments, and other technical features of this itemized cost prediction device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0133] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0134] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0135] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the tag-driven itemized cost prediction method in the above embodiments.

[0136] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0137] The aforementioned computer-readable storage medium may be included in the itemized cost forecasting device; or it may exist independently and not assembled into the itemized cost forecasting device.

[0138] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the itemized cost prediction device, the itemized cost prediction device performs the following actions: collects feature information related to the target project; matches corresponding option tags to the feature information according to preset classification rules; processes the option tags through feature encoding to generate tagged features; breaks down the target project into various sub-items; calls the target model corresponding to each sub-item; combines the tagged features to predict the unit price of each sub-item to obtain the target unit price; couples the target unit price and project quantity data through deterministic calculation rules to output the summary cost data of the sub-items; and adjusts the summary cost data of the sub-items based on the error distribution information obtained during model training to calculate the predicted result of the total project cost.

[0139] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0141] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0142] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described tag-driven itemized cost prediction method, thereby solving the technical problem of balancing efficiency and accuracy in cost prediction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the tag-driven itemized cost prediction method provided in the above embodiments, and will not be repeated here.

[0143] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A label-driven method for predicting itemized costs, characterized in that, The method includes: collecting feature information related to the target project; matching corresponding option tags to the feature information according to preset classification rules; processing the option tags through feature encoding to generate tagged features; breaking down the target project into various sub-items; calling the target model corresponding to each sub-item; predicting the unit price of each sub-item by combining the tagged features to obtain the target unit price; coupling the target unit price and project quantity data through deterministic calculation rules to output the summary cost data of the sub-items; adjusting the summary cost data of the sub-items based on the error distribution information obtained during model training to calculate the predicted result of the total project cost.

2. The label-driven itemized cost prediction method as described in claim 1, characterized in that, The step of collecting feature information related to the target project and matching corresponding option tags to the feature information according to preset classification rules includes: unifying the data formats of the economic indicator table data, the guidance questionnaire information, and the model extracted data of the target project, and integrating them to obtain an original feature set; identifying the corresponding feature types in the original feature set, filtering and distinguishing the feature items corresponding to the feature types according to preset categories; and matching the option tags uniquely corresponding to each feature item based on the classification rules.

3. The label-driven itemized cost prediction method as described in claim 1, characterized in that, The step of generating tagged features by processing the option labels through feature encoding includes: verifying the completeness and consistency of the option labels, removing invalid labels that do not conform to the preset encoding specifications, and outputting valid option labels; numerically converting the valid option labels according to preset discrete feature encoding rules to obtain one-dimensional numerical codes corresponding to feature items; and matrix-integrating the one-dimensional numerical codes, the category classification of the original features, and the item feature dimension sorting to generate the tagged features containing all item feature encoding information.

4. The label-driven itemized cost prediction method as described in claim 1, characterized in that, The steps of breaking down the target project into its constituent parts include: parsing the economic indicator data, engineering technical parameters, and industry classification standards for constituent parts of the target project; determining the applicable decomposition level and classification boundaries for the target project; and integrating these to obtain decomposition rules. According to these decomposition rules, and in conjunction with the functional attributes of the engineering, the target project is divided into several major engineering categories, each corresponding to a defined engineering scope and measurement caliber, resulting in a decomposition result. Based on these decomposition rules, the decomposition result is further refined hierarchically to determine the specific work content and measurement object of each constituent part of the project, thus obtaining the constituent parts of the project.

5. The label-driven itemized cost prediction method as described in claim 1, characterized in that, The step of calling the target model corresponding to each of the sub-items of the project and predicting the unit price of each sub-item in combination with the labeled features to obtain the target unit price includes: matching the corresponding model in the preset model library according to the functional attributes of the sub-items of the project, and extracting the exclusive feature dimension of each sub-item from the labeled features to generate the mapping relationship between each sub-item of the project and the target model and the sub-item exclusive features; triggering the inference process of the corresponding target model in sequence according to the mapping relationship, substituting the sub-item exclusive features into the target model one by one to complete the calculation, and outputting the preliminary unit price prediction result; verifying and optimizing the preliminary unit price prediction result based on the preset reasonable range standard for the unit price of the sub-items of the project and the historical prediction error correction rules, removing outliers that exceed the reasonable range and adjusting the prediction deviation to obtain the target unit price of each sub-item of the project.

6. The label-driven itemized cost prediction method as described in claim 1, characterized in that, The steps of coupling the target unit price and project quantity data through deterministic calculation rules to output the summary data of sub-item project costs include: verifying the measurement unit of the target unit price and the measurement caliber of the corresponding sub-item project quantity, and establishing the association relationship between each sub-item project and the target unit price and corresponding quantity; based on the association relationship, applying preset deterministic calculation rules to multiply the target unit price and corresponding quantity of each sub-item project to obtain the individual cost data of each sub-item project; and summarizing the hierarchical division and project category of the sub-item projects level by level, and combining the individual cost data to statistically form the summary data of sub-item project costs.

7. The label-driven itemized cost prediction method as described in claim 1, characterized in that, The steps for adjusting the total project cost summary data based on the error distribution information obtained during model training and calculating the predicted project total cost include: based on the error distribution information obtained during model training, associating and matching the error characteristics corresponding to each of the sub-items with the total cost summary data of the sub-items to determine the error fluctuation range and error distribution pattern of each sub-item cost; adjusting the cost of each sub-item using a corresponding deviation correction algorithm according to the error distribution pattern, while retaining the cost composition relationship corresponding to the project logic, to form corrected data; summing and calculating the corrected cost of each sub-item based on the corrected data, and deriving the confidence interval of the total project cost in combination with the error fluctuation range to determine the predicted result.

8. The label-driven itemized cost prediction method as described in claim 1, characterized in that, After the steps of adjusting the aggregated cost data of the distributed sub-items based on the error distribution information obtained during model training and calculating the predicted result of the total project cost, the label-driven sub-item cost prediction method further includes: collecting the actual total project cost and actual cost data of each sub-item of the target project; comparing the predicted result with the actual cost data and analyzing the influencing factors to output a deviation analysis report; parsing the key features and corresponding option labels that cause deviations in the deviation analysis report to generate a correlation analysis result between the degree of deviation and the contribution of features; updating the weight parameters corresponding to each feature according to the preset weight adjustment rules and in combination with the correlation analysis result to output the optimized feature weight configuration; integrating the optimized feature weight configuration into the target model training process corresponding to each sub-item project, iteratively training the model again with historical datasets to solidify the weight adjustment effect, and outputting the updated target model and weight configuration file to complete the optimization closed loop of the prediction method.

9. A device for predicting itemized costs, characterized in that, The itemized cost prediction device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the tag-driven itemized cost prediction method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the tag-driven itemized cost prediction method as described in any one of claims 1 to 8.