Engineering cost prediction method and system based on large model and storage medium
By constructing a dynamic cost feature knowledge base and using a large model with a Transformer architecture for fine-tuning and heterogeneous graph network analysis, the problems of low accuracy and poor adaptability in traditional engineering cost prediction methods are solved, and accurate prediction and dynamic adjustment of engineering costs are achieved.
Patent Information
- Application Number
- CN202510497587.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing engineering cost prediction methods rely on human experience, have low prediction accuracy, are difficult to handle nonlinear relationships and a large number of characteristic variables, cannot fully consider the combined effects of multi-dimensional factors, and are difficult to adapt to the dynamic changes of engineering projects, resulting in a large deviation between the prediction results and the actual cost.
By standardizing historical cost data, a dynamic cost feature knowledge base is constructed. A pre-trained large model based on the Transformer architecture is used for domain-adaptive fine-tuning to generate an adjustment coefficient matrix. Combined with heterogeneous graph network analysis, a hierarchical engineering cost detail table is calculated and corrected for the time dimension.
It significantly improves the accuracy and adaptability of engineering cost forecasting, can accurately capture the complex correlation between engineering cost parameters, solves the problems of low data utilization and rigid knowledge structure in traditional methods, and realizes differentiated processing of price fluctuations and resource type price changes.
Smart Images

Figure CN120410591B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to an engineering cost prediction method and system based on a large model and a storage medium. BACKGROUND
[0002] Engineering cost prediction is an important part of project management. Traditional engineering cost prediction methods mainly include index method, proportion method, analogy estimation method and parameter estimation method. The index method is based on historical engineering cost data, and adjusts the historical cost through the price index to predict the future project cost. The proportion method estimates the cost of a new project by analyzing the proportion of each sub-project in the total cost of historical projects. The analogy estimation method compares and adjusts the cost of similar projects based on cost data, combined with area, volume and other indicators. The parameter estimation method predicts the cost by establishing a mathematical relationship between the cost and the project characteristic parameters. With the development of computer technology, cost prediction methods based on statistics and machine learning have emerged, such as regression analysis, neural networks, support vector machines, etc. These methods describe the relationship between engineering characteristics and cost by constructing mathematical models, which improves the prediction accuracy to some extent.
[0003] However, the existing engineering cost prediction methods have many shortcomings. Traditional methods rely too much on human experience, have limited prediction accuracy and low efficiency. Statistical methods, although introducing mathematical models, are difficult to handle nonlinear relationships and a large number of characteristic variables. Machine learning methods, although improving the complexity of the model, face the problems of high data volume requirement, poor model interpretability, and difficulty in fully utilizing domain knowledge. In addition, existing methods are generally difficult to cope with the dynamic changes of engineering cost, especially in the case of frequent market price fluctuations and diverse engineering types, the predicted results often deviate greatly from the actual cost. More importantly, existing methods lack effective expression and processing ability of complex relationships between engineering elements, and cannot fully consider the comprehensive influence of multi-dimensional factors such as materials, processes and structures on the cost, resulting in limited prediction accuracy. SUMMARY
[0004] The present application provides an engineering cost prediction method and system based on a large model and a storage medium, which realizes accurate prediction and dynamic adjustment of engineering cost, not only solves the problem of insufficient expression of element correlation in traditional cost prediction, but also can be corrected according to time and project characteristics, significantly improving the prediction accuracy and adaptability.
[0005] In a first aspect, the application provides a large model-based engineering cost prediction method, which comprises: standardizing historical cost data to obtain a cost dataset, and constructing an engineering sub-item structure based on the cost dataset to obtain a dynamic cost feature knowledge base; inputting the dynamic cost feature knowledge base into a pre-trained large model based on a Transformer architecture for field adaptability fine-tuning to obtain an engineering cost prediction value; generating a compensation function by historical project similarity clustering and error distribution learning on the engineering cost prediction value to obtain an adjustment coefficient matrix; associating and matching design parameters of a target project with the dynamic cost feature knowledge base, calculating the cost of each sub-item unit through a heterogeneous graph network analysis to obtain a hierarchical engineering cost breakdown; and correcting the hierarchical engineering cost breakdown data in a time dimension based on the adjustment coefficient matrix to obtain a cost prediction result.
[0006] In a second aspect, the application provides a large model-based engineering cost prediction system, which comprises:
[0007] A construction module for standardizing historical cost data to obtain a cost dataset, and constructing an engineering sub-item structure based on the cost dataset to obtain a dynamic cost feature knowledge base;
[0008] An input module for inputting the dynamic cost feature knowledge base into a pre-trained large model based on a Transformer architecture for field adaptability fine-tuning to obtain an engineering cost prediction value;
[0009] A generation module for generating a compensation function by historical project similarity clustering and error distribution learning on the engineering cost prediction value to obtain an adjustment coefficient matrix;
[0010] A matching module for associating and matching design parameters of a target project with the dynamic cost feature knowledge base, calculating the cost of each sub-item unit through a heterogeneous graph network analysis to obtain a hierarchical engineering cost breakdown;
[0011] A correction module for correcting the hierarchical engineering cost breakdown data in a time dimension based on the adjustment coefficient matrix to obtain a cost prediction result.
[0012] In a third aspect, a large model-based engineering cost prediction device is provided, which comprises a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to enable the large model-based engineering cost prediction device to perform the large model-based engineering cost prediction method described above.
[0013] In a fourth aspect, a computer-readable storage medium is provided, which stores instructions that, when executed on a computer, cause the computer to perform the above-mentioned large model-based engineering cost prediction method.
[0014] In the technical solution provided in the present application, by standardizing the historical cost data and constructing a dynamic cost feature knowledge base, the problems of low data utilization and rigid knowledge structure in traditional cost prediction methods are effectively solved. The method inputs the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning, fully utilizes the advantages of the Transformer architecture in processing sequence data and long-distance dependency, and enables the model to accurately capture the complex correlations between engineering cost parameters. By generating a compensation function through historical project similarity clustering and error distribution learning, an adjustment coefficient matrix is formed, overcoming the limitations of traditional prediction methods in adapting to different types of engineering project characteristics, and significantly improving the prediction accuracy. The design parameters of the target project are associated and matched with the dynamic cost feature knowledge base, and the cost of each sub-part unit is calculated through the analysis of the heterogeneous graph network, realizing the fine expression and processing of complex engineering structures, and solving the problem that traditional methods cannot effectively handle the complex correlation between engineering elements. Based on the adjustment coefficient matrix, the data of each level of the hierarchical engineering cost breakdown is corrected in the time dimension, not only solving the cost prediction deviation caused by price fluctuations, but also realizing the differentiated processing of price changes of different resource types, greatly improving the accuracy, timeliness and applicability of engineering cost prediction. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on the drawings without creative labor.
[0016] Figure 1 An embodiment of the large model-based engineering cost prediction method in the present application;
[0017] Figure 2 An embodiment of the large model-based engineering cost prediction system in the present application;
[0018] Figure 3 The structural schematic diagram of the large model-based engineering cost prediction device in the present application. DETAILED DESCRIPTION
[0019] The embodiment of the application provides a large model-based engineering cost prediction method and system and a storage medium. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "includes" or "has" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to the process, method, product or device.
[0020] For ease of understanding, the specific process of the embodiment of the application is described below. Please refer to Figure 1 One embodiment of the large model-based engineering cost prediction method in the embodiment of the application comprises the following steps.
[0021] Step S101, standardizing historical cost data to obtain a cost data set, and constructing an engineering sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base;
[0022] Step S102, inputting the dynamic cost feature knowledge base into a pre-trained large model based on a Transformer architecture for field adaptability fine-tuning to obtain an engineering cost prediction value;
[0023] Step S103, generating a compensation function by historical project similarity clustering and error distribution learning on the engineering cost prediction value to obtain an adjustment coefficient matrix;
[0024] Step S104, associating and matching the design parameters of a target project with the dynamic cost feature knowledge base, calculating the cost of each sub-item unit through a heterogeneous graph network analysis to obtain a hierarchical engineering cost breakdown;
[0025] Step S105, based on the adjustment coefficient matrix, correcting the hierarchical data of the hierarchical engineering cost breakdown in the time dimension to obtain a cost prediction result.
[0026] It can be understood that the execution subject of the application can be a large model-based engineering cost prediction system, and can also be a terminal or a server, which is not limited here. The embodiment of the application takes a server as an execution subject for example.
[0027] Specifically, historical cost data is standardized. A large amount of historical engineering project list data, contract price data and material labor cost data is collected, and then the data is cleaned through outlier detection and missing value filling techniques. Outlier detection identifies outliers by calculating the deviation of data points from the overall distribution, such as a residential project bid that is abnormally low compared to the market average; and missing value filling intelligently completes the data according to similar project data to ensure data integrity. After data cleaning, normalization is performed to convert data of different dimensions into values within a unified interval, so that subsequent analysis is not affected by the differences in original dimensions. Based on the standardized cost data set, engineering projects are classified by type characteristics to obtain an engineering type classification system, covering residential buildings, public buildings, industrial buildings and other types. For each type, sub-item structure characteristics are extracted to generate sub-item structure templates. The templates are associated with regional price indices and time series price fluctuation data to generate price adjustment relationships, build engineering cost knowledge structure, and form a dynamic cost characteristic knowledge base.
[0028] The above dynamic cost characteristic knowledge base is converted into a serialized format and processed into a vector representation by a feature encoder for input into a Transformer model. The pre-trained large model is analyzed for its structure to understand the composition of its multi-head self-attention layers, feedforward neural network layers and layer normalization components, based on which a cost prediction instruction set and reference response pair are constructed, including instructions such as quota inquiry, unit price calculation and engineering quantity conversion, to form a fine-tuning data set. During fine-tuning, most parameters of the Transformer model are frozen, leaving only the last three layers of self-attention modules and the top linear layer parameters trainable. This parameter efficient fine-tuning strategy preserves the model's general knowledge while focusing on specific tasks in the engineering cost field. The engineering cost data is divided into training and validation sets, the gradient is calculated using the cross-entropy loss function, and the parameters are updated using the Adam optimizer to obtain the fine-tuned model. The target engineering parameters are input into the model, which uses the multi-head self-attention mechanism to capture the relationships between the parameters and generate engineering cost prediction values. The generated engineering cost prediction values are compared with the actual cost data of historical projects to calculate the prediction bias values and generate cost prediction error data. The prediction error data is combined with the engineering project characteristics, grouped by size, structure type and functional use to form a project characteristic classification table. The similarity between projects in each category is calculated to build a project association network and generate project clustering groups. The mean and variance of the prediction error within each cluster are analyzed to obtain error statistics, which are used to build a cost prediction adjustment scheme, form prediction adjustment rules and generate adjustment coefficient matrices for different engineering types and sizes.
[0029] The target project design parameters are structured, the quantity of work information, material specifications and construction process information are extracted, and the target project parameter set is formed. The parameters are matched with the dynamic cost characteristic knowledge base to find the corresponding items and obtain the parameter mapping relationship table. Based on the parameter mapping relationship table, a heterogeneous graph network is constructed, the engineering components, materials, processes and prices are taken as different types of nodes, and the engineering cost heterogeneous graph is constructed. The node information transmission calculation is carried out on the heterogeneous graph, the node characteristics and the adjacent node characteristics are weighted and aggregated to obtain the fusion characteristics. According to the fusion characteristics, the unit cost of each sub-item is calculated, the components and materials of the same category are classified and counted, and the unit cost table is generated. Then, according to the quantity list hierarchical structure, the total amount of each level is calculated, and the hierarchical engineering cost detail table is formed.
[0030] Based on the adjustment coefficient matrix, the time dimension of the hierarchical engineering cost detail table is corrected. The material price index, labor cost index and equipment cost index at the current time point are obtained, compared with the benchmark time index, the time span difference value and the price change ratio are calculated, and the time adjustment factor is generated. The time adjustment factor and the adjustment coefficient matrix are combined, the weights are allocated according to the project type and the regional characteristics, and the comprehensive adjustment coefficient is obtained. According to the proportion of material cost, labor cost and mechanical cost, respectively multiply the comprehensive adjustment coefficient, obtain the unit cost after time correction, recalculate the total price and summarize, generate the corrected cost detail table, form the cost prediction result.
[0031] In the embodiments of the present application, by standardizing the historical cost data and constructing a dynamic cost characteristic knowledge base, the problems of low data utilization and rigid knowledge structure of traditional cost prediction methods are effectively solved. The method inputs the dynamic cost characteristic knowledge base into the pre-training large model based on the Transformer architecture for field adaptability fine-tuning, fully utilizes the advantages of the Transformer architecture in processing sequence data and long-distance dependency, and enables the model to accurately capture the complex correlations between engineering cost parameters. Through historical project similarity clustering and error distribution learning to generate a compensation function and form an adjustment coefficient matrix, the limitations of traditional prediction methods in adapting to different types of engineering project characteristics are overcome, and the prediction accuracy is significantly improved. The design parameters of the target project are associated and matched with the dynamic cost characteristic knowledge base, and the unit cost of each sub-item is calculated through the heterogeneous graph network analysis, which realizes the fine expression and processing of complex engineering structure, and solves the problem that traditional methods cannot effectively handle the complex correlation between engineering elements. Based on the adjustment coefficient matrix, the data of each level of the hierarchical engineering cost detail table is corrected in the time dimension, which not only solves the cost prediction deviation caused by price fluctuations, but also realizes the differential processing of price changes of different resource types, greatly improving the accuracy, timeliness and applicability of engineering cost prediction.
[0032] In a specific embodiment, the process of performing step S101 can specifically include the following steps:
[0033] Anomaly detection and missing value imputation are performed on the bill of quantities data, contract price data and material labor cost data of historical engineering projects to obtain a cleaned dataset;
[0034] Normalization processing is performed on the cleaned dataset to obtain a standardized cost dataset;
[0035] Based on the standardized cost dataset, type classification of engineering projects is performed to obtain an engineering type classification system;
[0036] Feature extraction is performed on the sub-item structure of each type of engineering in the engineering type classification system to obtain a sub-item structure template;
[0037] The sub-item structure template is associated with regional price indices and time series price fluctuation data to obtain a price adjustment relationship;
[0038] Based on the sub-item structure template and the price adjustment relationship, an engineering cost knowledge structure is constructed to obtain a dynamic cost feature knowledge base.
[0039] Specifically, anomaly detection and missing value imputation are performed on the bill of quantities data, contract price data and material labor cost data of historical engineering projects. Anomaly detection refers to identifying data points that deviate significantly from the overall data distribution, usually using statistical methods. For example, for the unit cost data of a certain type of engineering project, the mean and standard deviation of all samples are calculated, and if a data point deviates from the mean by more than 3 standard deviations, it is determined to be an anomaly. For missing value imputation, the K-nearest neighbor algorithm is used to find the K most similar data samples to the missing data point, and the weighted average value based on the sample values is used to fill in the missing part. For example, when the labor cost data of the decoration engineering of a residential project is missing, find the 5 most similar projects of the same type and size, calculate the average value of their decoration engineering labor cost to fill in the missing item. Through these two steps, a cleaned dataset is obtained.
[0040] The normalization of the cleaning data set is the process of unifying data of different dimensions and ranges to the same scale. Common normalization methods include maximum and minimum value normalization and Z-score standardization. Maximum and minimum value normalization linearly transforms data to the [0, 1] interval. For example, the total cost data of different engineering projects is mapped from the original range of several million to several hundred million to between 0 and 1, which is convenient for subsequent analysis and processing. Z-score standardization is to subtract the mean and divide by the standard deviation, so that the data conforms to the distribution with a mean of 0 and a standard deviation of 1. Through this process, a standardized cost data set is obtained. Based on the standardized cost data set, the engineering projects are classified by type, and a hierarchical clustering algorithm is used. Hierarchical clustering treats each engineering project as an independent category, then gradually merges the most similar categories until the pre-set number of categories is reached. Similarity calculation uses Euclidean distance or cosine similarity. In this way, engineering projects are classified into different types such as residential buildings, commercial buildings, industrial buildings, and municipal engineering, forming an engineering type classification system.
[0041] The feature extraction of each engineering type in the engineering type classification system is carried out, and the cost composition characteristics of each type of engineering are analyzed. Feature extraction is achieved through frequency statistics and correlation analysis, calculating the frequency, cost proportion and mutual correlation of each sub-item in different engineering types. For example, by analyzing residential building projects, it is found that foundation engineering accounts for about 15% of the total cost, main structure accounts for about 45%, decoration accounts for about 25%, and equipment installation accounts for about 15%. By analyzing a large number of similar project data, the key structural features and cost characteristics of each type of engineering are extracted, and a sub-item structure template is established. The sub-item structure template is associated with regional price indexes and time series price fluctuation data to establish a price change relationship at different times and in different regions. Collect data on material price indexes, labor cost indexes and equipment cost indexes in different regions, as well as data on the changes of these indexes over time. Then calculate the correlation coefficient between each index and the cost of sub-items to identify which price index has the greatest impact on which sub-item. By establishing the mapping relationship between price indexes and sub-item costs, a price adjustment relationship is formed to support cost adjustment in different regions and at different times.
[0042] Based on the sub-item structure template and price adjustment relationship, an engineering cost knowledge structure is constructed to obtain a dynamic cost feature knowledge base. The specific method is to organize engineering types, sub-item structures, price indexes and adjustment relationships into an associated knowledge structure. In this structure, the engineering type node is connected to the corresponding sub-item node, the sub-item node is connected to the related material, labor and equipment resource node, and the resource node is associated with the price index node. Through this network structure, static template data and dynamic price adjustment data are integrated to form a dynamic cost feature knowledge base.
[0043] In a specific embodiment, the process of performing step S102 can specifically include the following steps:
[0044] Convert the dynamic cost feature knowledge base into a serialized input format, process it into a multi-dimensional vector representation through a feature encoder, and obtain a set of cost feature vectors;
[0045] Perform structural analysis on the pre-trained large model based on the Transformer architecture, which includes multi-head self-attention layers, feedforward neural network layers, and layer normalization components, and obtain a model structure mapping;
[0046] Based on the model structure mapping, construct a set of cost prediction instructions and reference response pairs, including sub-item quota queries, material unit price calculations, and engineering quantity conversion instructions, and obtain a fine-tuning data set of instructions;
[0047] According to the set of cost feature vectors and the fine-tuning data set of instructions, freeze the parameters of the Transformer model, retain the last three layers of self-attention modules and the top layer of linear layer parameters trainable, and obtain a model structure to be fine-tuned;
[0048] Divide the engineering cost data into a training set and a validation set in a ratio of 8:2, input it into the model structure to be fine-tuned, calculate the gradient through the cross-entropy loss function, and update the parameters through the Adam optimizer for backpropagation, and obtain a fine-tuned model;
[0049] Encode the target engineering parameters and input them into the fine-tuned model, capture the correlation between the parameters through the multi-head self-attention mechanism in the model, and calculate through the feedforward neural network layer and the output layer to obtain the engineering cost prediction value.
[0050] Specifically, the structured knowledge base data is organized, and information such as engineering type, sub-item, material price, etc. is organized in a fixed format. Serialization refers to converting complex data structures into continuous data sequences for easy input into large models. For example, the information of an engineering project can be arranged in the format of "engineering type + sub-item + material specification + unit price". After conversion, the text and numerical information are converted into vector representations through a feature encoder. The feature encoder is a tool that maps input data to a vector space. For text data, word embedding technology is used, and for numerical features, normalization is performed directly. The obtained set of cost feature vectors is a set of multi-dimensional vectors, each vector representing a piece of information in the engineering cost knowledge base, containing semantic and numerical features.
[0051] The structural analysis of pre-trained large models based on the Transformer architecture is to understand the internal structure of the model, so as to carry out fine-tuning more targetedly. The Transformer architecture is a neural network structure based on attention mechanism, and its core components include multi-head self-attention layer, feed-forward neural network layer and layer normalization component. The multi-head self-attention layer allows the model to focus on different positions of the input sequence simultaneously, capturing the dependencies between elements. Each attention head calculates the weighted sum between queries, keys and values to get the attention score. The feed-forward neural network layer is a fully connected layer composed of two linear transformations, which is used to perform nonlinear transformation on the features output by the attention mechanism. The layer normalization component is used to stabilize the training process and reduce the problem of gradient vanishing. Through the analysis of the model structure, the functions and parameter quantities of each layer are clarified, which prepares for the subsequent selective fine-tuning of parameters.
[0052] Based on the model structure mapping, the construction cost prediction instruction set and the reference response pair are constructed to create training data suitable for fine-tuning. The instruction set contains various queries and operation instructions related to engineering cost, such as sub-item quota query, such as "query the concrete unit price of the main body engineering of frame structure residential building"; material unit price calculation, such as "calculate the comprehensive unit price of C30 concrete including transportation and pumping"; engineering quantity conversion instruction, such as "convert 3000 square meters of building area to concrete cubic meter of foundation engineering". For each instruction, write professional and accurate responses to form instruction-response pairs. The data is organized according to the model input format to form the instruction fine-tuning data set, which is used to teach the model how to correctly understand and respond to queries related to cost prediction.
[0053] According to the cost feature vector set and the instruction fine-tuning data set, the parameters of the Transformer model are frozen and selectively fine-tuned. Parameter freezing refers to setting part of the parameters in the model to a non-trainable state, and the parameters are not updated during back propagation. In this method, the parameters of the first few layers of the model are frozen, and only the last three self-attention modules and the top linear layer parameters are trainable. The reason for this is that the first few layers of the pre-trained model usually capture general language features, while the last few layers focus more on features related to specific tasks. The self-attention module is responsible for capturing the dependencies between elements in the input sequence, and the top linear layer is responsible for mapping features to the output space. Through this selective fine-tuning strategy, the model can retain the general knowledge learned during the pre-training phase while adapting to the specific task of engineering cost, resulting in a fine-tuned model structure with a clear fine-tuning goal. The engineering cost data is divided into training and validation sets in an 8:2 ratio, which is a common data segmentation method in machine learning. The training set is used to learn the model parameters, and the validation set is used to evaluate the model performance and prevent overfitting. During fine-tuning, the cross-entropy loss function is used to calculate the difference between the predicted value and the true value. The cross-entropy loss function is particularly suitable for classification and probability prediction tasks, and calculates the difference between the predicted distribution and the true distribution. The optimizer uses the Adam (Adaptive Moment Estimation) algorithm, which combines the advantages of the momentum method and the RMSProp algorithm, and can adaptively adjust the learning rate to accelerate convergence. In each training batch, the model performs forward propagation calculations on the input data, then calculates the gradient through the back propagation algorithm and updates the trainable parameters. During training, the performance on the validation set is continuously monitored, and training is stopped when the validation loss no longer decreases to prevent overfitting. Through this process, a fine-tuned model optimized specifically for the engineering cost prediction task is obtained.
[0054] Encoding the target engineering parameters and inputting them into the fine-tuned model is the actual application stage. The design parameters of a new engineering project need to be encoded and processed into a vector representation consistent with the model input format. After inputting the model, the multi-head self-attention mechanism works, and each attention head calculates the correlation score between parameters to capture the complex relationship between engineering parameters. For example, the combination of building area and structural form affects the cost, or the interaction between decoration standards and equipment configuration. The associated information is transformed through the feedforward neural network layer to extract higher-level feature representations. Finally, the high-dimensional features are mapped to specific cost prediction values through the output layer.
[0055] In a specific embodiment, the process of performing step S103 can specifically include the following steps:
[0056] Compare the engineering cost prediction value with the actual cost data of the historical engineering project, calculate the prediction deviation value, and obtain the cost prediction error set;
[0057] Based on the cost prediction error set and the engineering project characteristics, a feature-error correspondence relationship is constructed, and the project characteristic classification matrix is obtained by grouping according to the engineering scale, structure type and functional purpose;
[0058] Similarity calculation is performed on each category in the project characteristic classification matrix, a project similarity network is constructed through a cosine similarity function, and a project clustering result is obtained;
[0059] Statistical analysis is performed on the engineering cost prediction error in the project clustering result, the error mean and variance of each cluster center are calculated, and the error distribution parameters are obtained;
[0060] Based on the error distribution parameters, a cost prediction compensation function for each type of engineering project is constructed, the relationship between the error and the project characteristics is fitted through a polynomial regression equation, and a compensation function set is obtained;
[0061] The compensation function set is mapped to the prediction process according to the engineering classification, multi-dimensional adjustment parameters are generated, and an adjustment coefficient matrix is obtained.
[0062] Specifically, the actual settlement data of completed projects are collected as the true construction value. Then the characteristic parameters of these projects are input into the fine-tuned large model to obtain the predicted construction value. By calculating the difference between the predicted value and the actual value of each project, the prediction deviation value is formed. The prediction deviation value refers to the difference between the predicted construction cost and the actual construction cost, which can be represented by absolute error (predicted value-actual value) or relative error ((predicted value-actual value) / actual value). These prediction deviation values are aggregated to form a construction cost prediction error set, providing basic data for subsequent analysis. Building a feature-error correspondence relationship based on the construction cost prediction error set and the characteristics of engineering projects is a key step in understanding the causes of prediction errors. The key features of engineering projects are extracted, including project size, structure type, and functional use, etc. Then a mapping relationship between project features and their corresponding prediction errors is established to form a feature-error correspondence table. According to the feature attributes, projects are grouped, such as small, medium, and large according to area size; frame structure, frame-shear wall structure, steel structure, etc. according to structure type; residential, office, commercial, etc. according to functional use. Through this grouping arrangement, the feature-error correspondence table is converted into a multi-dimensional project feature classification matrix, which is convenient for subsequent similarity analysis. Similarity calculation in each category of the project feature classification matrix is the process of identifying similar project groups. In each feature classification, the similarity between projects needs to be calculated. Cosine similarity is a commonly used similarity calculation method, which measures the similarity between two vectors by calculating the cosine of the angle between them. The smaller the angle between the feature vectors of two projects, the closer the cosine value is to 1, indicating that the projects are more similar. By calculating the cosine similarity between each pair of projects, a similarity network between projects is constructed, which is a graph structure representing the strength of association between projects. Based on the similarity network, a clustering method is used to group similar projects into a group to form the project clustering result. The clustering method can be the K-means algorithm or the hierarchical clustering algorithm, which automatically classifies projects with high similarity into the same cluster.
[0063] Statistical analysis is performed on the engineering cost prediction errors within the project clustering results. In each cluster, the prediction error data of all projects are collected, and the mean error is calculated, which is the average level of prediction deviation for projects in this category, representing the systematic deviation of prediction. The variance of the error is also calculated, reflecting the stability and dispersion of the prediction results of projects in this category. Statistical analysis also includes calculating the distribution form of the error, such as skewness and kurtosis, to determine whether the error distribution conforms to the normal distribution. Through these statistical quantities, the error distribution parameters describing the characteristics of the prediction errors of various projects are obtained, providing a data basis for the subsequent construction of the compensation function.
[0064] For each type of project, analyze the relationship pattern between its error and project characteristics. Through the polynomial regression method, establish the mathematical relationship between project characteristics and prediction error. Polynomial regression is a regression method that can capture nonlinear relationships, by fitting different orders of project characteristic combinations, construct error prediction equation. For each type of project, construct a compensation function respectively, form a compensation function set, which is used to estimate the possible error according to the project characteristics and make correction when predicting.
[0065] According to the characteristics of the project to be predicted, determine its engineering category. Call the compensation function of the corresponding category to calculate the compensation value of the prediction error. Convert the compensation value into an adjustment coefficient to form an adjustment coefficient matrix for different characteristic dimensions. The adjustment coefficient matrix is a multi-dimensional structure that contains adjustment parameters for different dimensions such as engineering scale, structure type, and functional use. In actual prediction, according to the specific characteristics of the project, extract the corresponding adjustment parameters from the adjustment coefficient matrix to correct the original prediction result.
[0066] In a specific embodiment, the process of performing step S104 can specifically include the following steps:
[0067] Structurally process the design parameters of the target project to extract engineering quantity information, material specifications, and construction technology, and obtain a target project parameter set;
[0068] Match the target project parameter set with the dynamic cost feature knowledge base to find the corresponding item in the knowledge base for each design parameter, and obtain a parameter mapping relationship table;
[0069] Based on the parameter mapping relationship table, construct a heterogeneous graph network structure, set engineering components, materials, technology, and prices as nodes of different types, and construct an engineering cost heterogeneous graph according to the set multiple nodes;
[0070] Perform node information transmission calculation on the engineering cost heterogeneous graph, aggregate the characteristics of each node and the characteristics of adjacent nodes by weighting, and obtain fused features;
[0071] Based on the fused features, calculate the unit cost of each sub-item according to the hierarchical structure of the engineering project, classify and count components and materials of the same category, and obtain a unit cost table;
[0072] Organize the unit cost table according to the hierarchical structure of the bill of quantities, and calculate the total amount of each level to obtain a hierarchical engineering cost detail table.
[0073] Specifically, structured processing refers to the conversion of unstructured or semi-structured design data into a standardized structured format. Key information is extracted from design drawings, construction specifications, and bill of quantities, including quantity information (such as concrete volume, steel consumption, wall area, etc.), material specifications (such as concrete grade, steel type, decorative material quality, etc.), and construction technology (such as construction method, process requirements, technical standards, etc.). The extraction process uses text analysis and data normalization techniques to uniformly convert information of different formats into standardized data structures. The converted data is organized according to the hierarchy of engineering sub-items, forming a structured target project parameter set, laying the foundation for subsequent cost calculation.
[0074] Matching the target project parameter set with the dynamic cost feature knowledge base is a key link connecting project features and cost data. The matching process compares each engineering element in the project parameter set with the entries in the knowledge base to find the most matching corresponding item. The matching method includes exact matching and fuzzy matching. Exact matching is used to find completely consistent projects, such as specific specifications of materials or standard components; fuzzy matching is used to handle incomplete consistency, finding the closest entry by calculating similarity. The matching process also considers the particularity of the project and regional differences, and adjusts the standard items appropriately. Finally, a parameter mapping relationship table is formed, with each row containing a parameter of the target project and its corresponding entry in the knowledge base, establishing the mapping relationship between project features and cost knowledge.
[0075] Building a heterogeneous graph network structure based on the parameter mapping relationship table is the innovation point of the engineering cost prediction method. Heterogeneous graph network is a graph structure containing multiple types of nodes and edges, which can represent complex association relationships. In the construction process, different types of elements such as engineering components, materials, processes, and prices are set as nodes in the graph, each type of node has different attribute characteristics. Then according to the parameter mapping relationship table, the connection relationship between nodes is established, for example, the "beam" node and the "C30 concrete" node establish a "use" relationship, and the "C30 concrete" node and its "unit price" node establish a "pricing" relationship. The engineering cost heterogeneous graph constructed in this way contains the complex association network between various elements in the engineering project, providing a structured foundation for subsequent information transmission and feature fusion.
[0076] The node information transmission calculation of the engineering cost heterogeneous graph is the core step of capturing the association relationship by using the graph network. The information transmission calculation is based on the principle of graph attention network (GAT), which enables each node to exchange information with its adjacent nodes and update its own features. In the specific process, each node collects the feature information of the adjacent nodes, and then weights the contributions of different adjacent nodes through the attention mechanism, and the important adjacent nodes obtain higher weights. The weighted aggregated information is combined with the node's own features to form the updated node representation. This information transmission mechanism enables each node to retain its own features and integrate the information of associated nodes, forming a context-aware fusion feature. For example, an architectural component node not only contains its own size information after information transmission, but also integrates the price fluctuation information of the materials used and the complexity information of the construction process. Based on the fusion features, calculating the unit cost of each sub-item according to the hierarchical structure of the engineering project is a key step to convert the graph network learning results into actual cost. The calculation process identifies components and materials with the same features and classifies them for statistics. The unit cost of each component is calculated, which is composed of direct costs (material cost, labor cost, and machinery cost) and indirect costs (management cost, profit, and tax). Direct costs are calculated based on price information and consumption information in the fusion features, and indirect costs are determined based on industry standards and project characteristics. The calculation considers the associated influencing factors contained in the fusion features, such as the price increase of a certain material, which not only affects the material cost but also the labor cost of related construction processes. Finally, a unit cost table is formed, which lists the detailed cost composition and unit price information of each component.
[0077] Organizing the unit cost table according to the hierarchical structure of the bill of quantities and calculating the total amount of each level is the process of forming the final cost prediction result. The hierarchical structure of the bill of quantities usually includes three levels: division, sub-item, and sub-object. Multiply the unit price in the unit cost table with the corresponding quantity to calculate the total price of each sub-object. Then, according to the list structure, the sub-object total price is summarized level by level to obtain the sub-item engineering cost and division engineering cost. In the summary process, consider the associated influence between levels and make necessary adjustments to the cost. Finally, a hierarchical engineering cost breakdown table is formed, which clearly shows the complete cost composition from the total cost to each division, sub-item, and specific sub-object.
[0078] In a specific embodiment, the process of performing step S105 can specifically include the following steps:
[0079] Grouping and organizing the fusion features according to the engineering sub-item coding system to obtain a grouped feature set;
[0080] Based on the feature set of each group, the material, labor, and mechanical consumption information of each sub-item is extracted and matched with the unit price in the market price database to obtain a resource unit price table.
[0081] Each resource in the resource unit price table is multiplied by the corresponding consumption to calculate the direct cost of each sub-item, and the indirect cost is calculated according to the specification requirements to obtain cost composition data.
[0082] The cost composition data is summarized according to the cost calculation rules of the engineering quantity list specification, and the tax, fee, and measure fee are added to obtain the comprehensive unit price of each sub-item.
[0083] The comprehensive unit price is multiplied by the engineering quantity to calculate the total price of each sub-item, and the hierarchical summary is performed according to the hierarchical relationship of the project to obtain a hierarchical cost summary table.
[0084] The data in the hierarchical cost summary table is analyzed for rationality, and the data rationality is confirmed by comparison with historical data to obtain a unit cost table.
[0085] Specifically, the engineering sub-item coding system is a standardized coding system used to identify different engineering parts in engineering cost management, which usually includes main engineering, decoration engineering, installation engineering, etc. Each category is further divided into specific sub-items. The grouping and sorting process matches the feature data with the standard code to identify the category of each feature. For example, the feature data of concrete members is classified into the "main structure engineering" category, and the feature data of wall decoration is classified into the "decoration and decoration engineering" category. Then, according to the coding hierarchical structure, the grouping feature set conforming to the engineering quantity list specification is formed. In the grouping feature set, the items with the same code are grouped into a group, and each group contains all the feature information required for the sub-item engineering, laying a foundation for subsequent refined cost calculation.
[0086] Consumption information refers to the amount of various resources required to complete a unit of engineering, including material consumption, labor days, and mechanical shifts, etc. When extracting this information from the grouping feature set, the influence of engineering characteristics, construction methods, and technical specifications needs to be considered. For example, for reinforced concrete structures, the amount of steel required per cubic meter of concrete, formwork area, concrete consumption, and corresponding labor and mechanical input are extracted according to the structure type and reinforcement ratio. After extraction, these consumption information is matched with the unit price in the market price database. The market price database stores the latest market prices of various materials, labor, and machinery, and through the matching process, appropriate unit prices are determined for each resource to form a resource unit price table. The resource unit price table is a detailed data set listing all resource types, specifications, consumption, and unit prices in the project, providing a data basis for direct cost calculation.
[0087] The direct cost includes material cost, labor cost and mechanical cost, which are calculated by multiplying the consumption of each type of resource by the corresponding unit price. For example, the material cost is equal to the total of the usage of various materials multiplied by the unit price, the labor cost is equal to the total of the number of working days of each type of work multiplied by the working day unit price, and the mechanical cost is equal to the total of the number of shifts of each type of machinery multiplied by the shift unit price. After calculating the direct cost, the indirect cost is calculated according to the engineering construction industry specifications and contract requirements. The indirect cost usually includes enterprise management cost, site management cost, etc., which are generally calculated according to a certain percentage of the direct cost. The direct cost and the indirect cost are added together to obtain the cost composition data, which is the basic component of the engineering cost. The engineering quantity list specification specifies the calculation method and procedure from the direct cost, the indirect cost to the final comprehensive unit price. During the aggregation process, tax, fee and measure fee need to be added. The tax refers to various taxes paid during the engineering construction process, such as value-added tax; the fee refers to the fees that must be paid according to the national regulations, such as social insurance fees; the measure fee refers to the non-direct production cost occurred during the construction process, such as temporary facility fee, safe and civilized construction fee, etc. These costs are calculated according to the calculation method and proportion specified in the specification, and together with the direct cost and the indirect cost, form a complete engineering cost. Through this process, the comprehensive unit price of each sub-item is obtained, i.e. the complete cost per unit of engineering quantity. The engineering quantity refers to the actual quantity of each sub-item engineering, such as the quantity of concrete, the weight of steel bars, the area of wall surface, etc. The comprehensive unit price calculated in the previous step is multiplied by the corresponding engineering quantity to obtain the total cost of each sub-item engineering, i.e. the combined price. After the combined price is calculated, the cost is aggregated according to the hierarchical structure of the engineering project. The engineering project is usually divided into unit engineering, sub-division engineering, sub-item engineering, etc. The aggregation process starts from the bottom layer of the sub-item engineering and accumulates upward level by level to form the cost amount of each level. This hierarchical aggregation method meets the actual needs of engineering management and is convenient for project management personnel to understand the cost composition of each part of the engineering and obtain the hierarchical cost summary table.
[0088] The rationality analysis adopts the comparison method to compare the calculated cost data with the cost data of the same type of historical project to check whether there is a significant deviation. The comparison indexes include the unit cost (the cost per square meter of building area), the proportion of sub-division engineering (such as the proportion of structure engineering in the total cost), the main material usage index (such as the steel bar usage per square meter of building area), etc. If a certain index is found to have a large deviation from the historical data, the calculation process needs to be traced back to find the reason and make adjustments. Through this rationality analysis, it is ensured that the final cost prediction result meets the engineering actuality and obtains the verified unit cost table.
[0089] In a specific embodiment, the process of performing step S106 can specifically include the following steps:
[0090] The material price index, labor cost index and equipment cost index at the current time point are obtained and compared with the benchmark time index used when calculating the hierarchical engineering cost breakdown table, to obtain the time span difference value;
[0091] Based on the time span difference value, the price change ratio is calculated, the price changes of various resources are classified and counted, and the time adjustment factor is obtained;
[0092] The time adjustment factor and the adjustment coefficient matrix are weighted and combined, the weight distribution is allocated according to the project type and the project area characteristics, and the comprehensive adjustment coefficient is obtained;
[0093] The unit cost of each sub-item in the hierarchical engineering cost breakdown table is multiplied by the comprehensive adjustment coefficient according to the composition proportion of material cost, labor cost and mechanical cost, to obtain the time-corrected unit cost;
[0094] The time-corrected unit cost is multiplied by the engineering quantity, the unit price of each sub-item is recalculated, and the data is summarized according to the hierarchical relationship to obtain the corrected cost breakdown table;
[0095] The corrected cost breakdown table is comprehensively analyzed, the data at each level is sorted according to importance, a cost prediction report is generated, and the cost prediction result is obtained.
[0096] Specifically, the price index is a relative value reflecting the change of price level, which is used to measure the change of price in a specific period. In the field of engineering cost, commonly used price indexes include material price index, labor cost index and equipment cost index. The sources of current index data include official indexes released by industry associations and market research data. The benchmark time index is the price index used when preparing the hierarchical engineering cost breakdown table, which may be several months ago or even earlier. The comparison process calculates the difference between the current index and the benchmark index to obtain the time span difference value, which directly reflects the degree of change of the price of various resources in the time dimension.
[0097] Based on the time span difference value, the price change ratio is calculated, the price changes of various resources are classified and counted, and the time adjustment factor is obtained;
[0098] The adjustment coefficient matrix is generated through historical project similarity clustering and error distribution learning in the preceding steps, reflecting the influence of project type, scale, and other characteristics on cost prediction. The weighted combination process assigns weights based on the project type and project area characteristics, for example, for steel structure projects that are more affected by material prices, the time adjustment factor of the material price index has a higher weight; while for labor-intensive decoration projects, the time adjustment factor of the labor cost index has a higher weight. Through this weighted combination, the comprehensive adjustment coefficient is obtained, which is a comprehensive correction parameter considering the time dimension and project characteristics, used for comprehensive adjustment of the cost data. Multiplying the unit cost of each sub-item in the hierarchical engineering cost breakdown table by the comprehensive adjustment coefficient according to the composition ratio of material cost, labor cost, and mechanical cost is the key step to achieve fine adjustment. The composition of engineering cost is usually divided into material cost, labor cost, and mechanical cost, and the proportion of each category in the total cost varies due to the nature of the project. For each sub-item project, the cost composition needs to be analyzed to determine the proportion of material cost, labor cost, and mechanical cost. Then, according to different cost categories, the corresponding adjustment coefficient is extracted from the comprehensive adjustment coefficient for multiplication operation. For example, for a sub-item with a material cost proportion of 60%, the adjustment coefficient related to material price is used; for a sub-item with a labor cost proportion of 30%, the adjustment coefficient related to labor cost is used. Through this classification adjustment, the time-corrected unit cost is obtained, which is more accurate cost data considering the price time change factor.
[0099] Multiplying the time-corrected unit cost by the engineering quantity to recalculate the unit price of each sub-item is the process of converting unit price adjustment to total price adjustment. Engineering quantity data comes from project design documents and bill of quantities, which is a specific quantity index reflecting the engineering scale. Multiplying the corrected unit cost by the corresponding engineering quantity gives the updated unit price of each sub-item project. After the unit price calculation is completed, the hierarchical relationship of the project is used for hierarchical aggregation, starting from the bottom layer of sub-item projects and gradually accumulating upwards to form the total cost of sub-projects, unit projects, and the entire project. This bottom-up aggregation method ensures the integrity and consistency of cost correction, and the corrected cost breakdown table is obtained, which is a complete cost data set reflecting the current price level.
[0100] The comprehensive analysis includes the reasonableness check of total cost and sub-item cost, the influence evaluation of key materials and engineering changes, and the identification of risk factors. The importance ranking is based on the three factors of cost amount, change amplitude, and sensitivity, and the projects that have a greater impact on the total cost are placed in the front row to facilitate decision-makers to focus on key content. Through this analysis and sorting, a structured cost prediction report is formed, including total cost data, sub-item cost breakdown, time adjustment factor analysis, and key attention items.
[0101] The method for predicting engineering cost based on a large model in the embodiments of the present application is described above. The system for predicting engineering cost based on a large model in the embodiments of the present application is described below. Please refer to Figure 2 One embodiment of the system for predicting engineering cost based on a large model in the embodiments of the present application includes:
[0102] The construction module 201 is configured to perform standardization processing on historical cost data to obtain a cost data set, and construct an engineering sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base.
[0103] The input module 202 is configured to input the dynamic cost feature knowledge base into a pre-trained large model based on a Transformer architecture for domain adaptability fine-tuning to obtain an engineering cost prediction value.
[0104] The generation module 203 is configured to generate a compensation function by historical project similarity clustering and error distribution learning on the engineering cost prediction value to obtain an adjustment coefficient matrix.
[0105] The matching module 204 is configured to associate and match design parameters of a target project with the dynamic cost feature knowledge base, calculate the cost of each sub-item unit through heterogeneous graph network analysis, and obtain a hierarchical engineering cost breakdown.
[0106] The correction module 205 is configured to correct each level of data of the hierarchical engineering cost breakdown in the time dimension based on the adjustment coefficient matrix to obtain a cost prediction result.
[0107] Through the synergistic cooperation of the above-mentioned various components, by standardizing the historical cost data and constructing a dynamic cost characteristic knowledge base, the problems of low data utilization and rigid knowledge structure of traditional cost prediction methods are effectively solved. The method inputs the dynamic cost characteristic knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning, fully utilizes the advantages of the Transformer architecture in processing sequence data and long-distance dependencies, and enables the model to accurately capture the complex correlations between engineering cost parameters. Through historical project similarity clustering and error distribution learning to generate a compensation function and form an adjustment coefficient matrix, the limitations of traditional prediction methods in adapting to different types of engineering project characteristics are overcome, and the prediction accuracy is significantly improved. The design parameters of the target project are associated and matched with the dynamic cost characteristic knowledge base, and the unit cost of each sub-item is calculated through the analysis of the heterogeneous graph network, realizing the fine expression and processing of complex engineering structures, and solving the problem that traditional methods cannot effectively handle the complex correlation between engineering elements. Based on the adjustment coefficient matrix, the time dimension of the hierarchical engineering cost breakdown table is corrected, not only solving the cost prediction deviation caused by price fluctuations, but also realizing the differentiated processing of price changes of different resource types, greatly improving the accuracy, timeliness and applicability of engineering cost prediction.
[0108] The above Figure 2 From the perspective of modular functional entities, the large model-based engineering cost prediction system in the embodiments of the present application is described in detail below. From the perspective of hardware processing, the large model-based engineering cost prediction device in the embodiments of the present application is described in detail.
[0109] Figure 3 is a structural schematic diagram of a large model-based engineering cost prediction device provided by the embodiments of the present application. The large model-based engineering cost prediction device 300 can have relatively large differences due to different configurations or performances, and can include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device ends) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be temporary storage or persistent storage. The programs stored in the storage medium 330 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the large model-based engineering cost prediction device 300. Further, the processor 310 can be configured to communicate with the storage medium 330, execute a series of instruction operations in the storage medium 330 on the large model-based engineering cost prediction device 300, to realize the steps of the above-mentioned large model-based engineering cost prediction method.
[0110] The large model-based engineering cost prediction device 300 can also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, and the like. Those skilled in the art can understand that, Figure 3 The large model-based engineering cost prediction device structure shown does not constitute a limitation on the large model-based engineering cost prediction device provided by the present application, and can include more or fewer components than shown, or combine certain components, or different component arrangements.
[0111] The present application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium, and can also be a volatile computer readable storage medium, and the computer readable storage medium has instructions stored therein, and when the instructions are run on a computer, the computer executes the steps of the large model-based engineering cost prediction method.
[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, system and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0113] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a large model-based engineering cost prediction device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0114] The above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features can be replaced by equivalent features; the modification or replacement does not make the essence of the corresponding technical solution deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A large model-based engineering cost prediction method, characterized in that, The method comprises: standardizing historical cost data to obtain a cost dataset, and constructing an engineering sub-item structure based on the cost dataset to obtain a dynamic cost feature knowledge base; inputting the dynamic cost feature knowledge base into a pre-training large model based on a Transformer architecture for field adaptability fine-tuning to obtain an engineering cost prediction value, including: converting the dynamic cost feature knowledge base into a serialized input format, processing it into a multi-dimensional vector representation through a feature encoder to obtain a cost feature vector set; performing structural analysis on the pre-training large model based on the Transformer architecture, which includes a multi-head self-attention layer, a feedforward neural network layer, and a layer normalization component to obtain a model structure mapping; based on the model structure mapping, constructing a cost prediction instruction set and a reference response pair, including sub-item quota inquiry, material unit price calculation, and engineering quantity conversion instructions to obtain an instruction fine-tuning dataset; according to the cost feature vector set and the instruction fine-tuning dataset, freezing the parameters of the Transformer model, keeping the last three layers of self-attention modules and the top layer of linear layer parameters trainable to obtain a model structure to be fine-tuned; dividing the engineering cost data into a training set and a validation set in a ratio of 8:2, inputting it into the model structure to be fine-tuned, calculating the gradient through a cross-entropy loss function, and updating the parameters through backpropagation using an Adam optimizer to obtain a fine-tuned model; encoding the target engineering parameters and inputting them into the fine-tuned model, capturing the parameter correlation through the multi-head self-attention mechanism in the model, and calculating through the feedforward neural network layer and the output layer to obtain the engineering cost prediction value; generating a compensation function by historical project similarity clustering and error distribution learning based on the engineering cost prediction value to obtain an adjustment coefficient matrix; associating and matching the design parameters of the target project with the dynamic cost feature knowledge base to calculate the cost of each sub-item unit through heterogeneous graph network analysis to obtain a hierarchical engineering cost breakdown; based on the adjustment coefficient matrix, correcting the data of each level of the hierarchical engineering cost breakdown in the time dimension to obtain the cost prediction result.
2. The large model-based engineering cost prediction method of claim 1, wherein, The standardizing historical cost data to obtain a cost dataset, and constructing an engineering sub-item structure based on the cost dataset to obtain a dynamic cost feature knowledge base comprises: performing outlier detection and missing value filling on the list data, contract price data, and material labor cost data of historical engineering projects to obtain a cleaned dataset; performing normalization processing on the cleaned dataset to obtain a standardized cost dataset; based on the standardized cost dataset, classifying engineering projects by type to obtain an engineering type classification system; extracting features from the sub-item structure of each engineering type in the engineering type classification system to obtain a sub-item structure template; associating and mapping the sub-item structure template with regional price indexes and time series price fluctuation data to obtain a price adjustment relationship; Based on the sub-item structure template and the price adjustment relationship, a project cost knowledge structure is constructed, and a dynamic cost characteristic knowledge base is obtained. 3.The large model-based engineering cost prediction method according to claim 1, characterized in that, The project cost prediction value is compared with the actual cost data of the historical engineering projects, a prediction deviation value is calculated, and a cost prediction error set is obtained. The project cost prediction value is compared with the actual cost data of the historical engineering projects, a prediction deviation value is calculated, and a cost prediction error set is obtained. Based on the cost prediction error set and the engineering project characteristics, a feature-error correspondence relationship is constructed, and the project feature classification matrix is obtained by grouping according to the engineering scale, structure type and functional purpose. The project cost prediction error in the project clustering result is statistically analyzed, the error mean and variance of each cluster center are calculated, and the error distribution parameters are obtained. The project cost prediction value is compared with the actual cost data of the historical engineering projects, a prediction deviation value is calculated, and a cost prediction error set is obtained. Based on the error distribution parameters, a cost prediction compensation function for each type of engineering project is constructed, and the relationship between the error and the project characteristics is fitted by a polynomial regression equation to obtain a compensation function set. The compensation function set is mapped to the prediction process according to the engineering classification, and a multi-dimensional adjustment parameter is generated to obtain an adjustment coefficient matrix.
4. The large model-based engineering cost prediction method of claim 1, wherein, The design parameters of the target project are associated and matched with the dynamic cost characteristic knowledge base, the unit cost of each sub-item unit is calculated through heterogeneous graph network analysis, and a hierarchical project cost breakdown is obtained, including: The design parameters of the target project are structured and processed to extract the quantity of work, material specifications and construction technology, and a target project parameter set is obtained. The target project parameter set is matched with the dynamic cost characteristic knowledge base, and the corresponding item in the knowledge base is found for each design parameter to obtain a parameter mapping relationship table. Based on the parameter mapping relationship table, a heterogeneous graph network structure is constructed, engineering components, materials, processes and prices are set as nodes of different types, and an engineering cost heterogeneous graph is constructed according to the set multiple nodes. The node information transmission calculation is performed on the engineering cost heterogeneous graph, the features of each node and the features of adjacent nodes are weighted and aggregated, and the fusion features are obtained. Based on the fusion features, the unit cost of each sub-item is calculated according to the hierarchical structure of the engineering project, and the components and materials of the same category are classified and counted to obtain a unit cost table. The unit cost table is organized according to the hierarchical structure of the bill of quantities, and the total amount of each level is calculated to obtain a hierarchical project cost breakdown.
5. The large model-based engineering cost prediction method of claim 4, wherein, The fusion features are grouped and arranged according to the engineering sub-item coding system to obtain a grouped feature set. Based on the grouped feature set, the material, labor and mechanical consumption information of each sub-item is extracted and matched with the unit price in the market price database to obtain a resource unit price table. The cost composition data is summarized according to the cost calculation rules of the bill of quantities specification, and tax, fee and measure fee are added to obtain the comprehensive unit price of each sub-item; The comprehensive unit price is multiplied by the engineering quantity to calculate the total price of each sub-item, and the cost summary table of each level is obtained by hierarchical summary according to the hierarchical relationship of the project; The data in the cost summary table of each level is analyzed for rationality, and the data rationality is confirmed by comparison with historical data to obtain a unit cost table. The cost prediction result is obtained by correcting the data of each level of the hierarchical engineering cost breakdown table in the time dimension based on the adjustment coefficient matrix, including: 6.The large model-based engineering cost prediction method according to claim 1, characterized in that, Obtain the material price index, labor cost index and equipment cost index at the current time point, and compare them with the benchmark time index used when calculating the hierarchical engineering cost breakdown table to obtain the time span difference value; Based on the time span difference value, calculate the price change ratio to classify and count the price changes of various resources, and obtain the time adjustment factor; The time adjustment factor and the adjustment coefficient matrix are combined by weighting, and the weight is distributed according to the type of engineering project and the characteristics of the project area to obtain a comprehensive adjustment coefficient; The unit cost of each sub-item in the hierarchical engineering cost breakdown table is multiplied by the comprehensive adjustment coefficient according to the composition proportion of material cost, labor cost and mechanical cost to obtain the time-corrected unit cost; The time-corrected unit cost is multiplied by the engineering quantity to recalculate the total price of each sub-item, and the cost breakdown table after correction is obtained by summarizing according to the hierarchical relationship. The cost prediction result is obtained by generating a cost prediction report by sorting the data of each level according to importance and performing comprehensive analysis on the cost breakdown table after correction. The engineering cost prediction system based on a large model for implementing the method according to any one of claims 1 to 6 comprises: 7.A large model-based engineering cost prediction system, characterized in that, A construction module for standardizing historical cost data to obtain a cost dataset, and constructing an engineering sub-item structure based on the cost dataset to obtain a dynamic cost feature knowledge base; An input module for inputting the dynamic cost feature knowledge base into a pre-trained large model based on a Transformer architecture for domain adaptability fine-tuning to obtain an engineering cost prediction value; A generation module for generating a compensation function by historical project similarity clustering and error distribution learning based on the engineering cost prediction value to obtain an adjustment coefficient matrix; A matching module for associating and matching the design parameters of a target project with the dynamic cost feature knowledge base to calculate the unit cost of each sub-item through heterogeneous graph network analysis and obtain a hierarchical engineering cost breakdown table; A correction module for correcting the data of each level of the hierarchical engineering cost breakdown table in the time dimension based on the adjustment coefficient matrix to obtain a cost prediction result. 8. A large model-based engineering cost prediction device, characterized by, The computer program is stored in the memory and includes a computer program capable of running on the processor, and the processor executes the computer program to realize the large model-based engineering cost prediction method in any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is stored in the memory and includes a computer program capable of running on the processor, and the processor executes the computer program to realize the large model-based engineering cost prediction method in any one of claims 1 to 6.
Citation Information
Patent Citations
House building project cost estimation method and device and storage medium
CN117934035A
Cost evaluation model generation method and device and cost evaluation method and device
CN119379324A