Project cost prediction method and system based on large model, and storage medium

By constructing a dynamic cost feature knowledge base and fine-tuning of Transformer model, combined with heterogeneous graph network analysis, the accuracy and adaptability of engineering cost prediction in traditional methods are solved, and the complex relationship between engineering elements is refined is achieved, and the accuracy and applicability of prediction are improved.

CN120410591AActive Publication Date: 2025-08-01广东中建普联科技股份有限公司

Patent Information

Application Number
CN202510497587.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing engineering cost prediction methods rely on manual experience, have low prediction accuracy, are difficult to deal with nonlinear relationships and a large number of feature variables, cannot fully consider the comprehensive impact of multi-dimensional factors, and are difficult to adapt to the complex correlation between engineering elements, resulting in large deviations in prediction results.

Method used

By standardizing the historical cost data, a knowledge base for dynamic cost characteristics is constructed, and a pre-trained large model based on Transformer architecture is used for field adaptability fine-tuning, combining heterogeneous graph network analysis and adjustment coefficient matrix to achieve accurate prediction and dynamic adjustment of engineering cost.

Benefits of technology

It significantly improves the accuracy and adaptability of project cost prediction, can handle complex engineering structures, adapt to the characteristics of different types of engineering projects, reduces prediction deviations caused by price fluctuations, and improves the timeliness and applicability of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410591A_ABST
    Figure CN120410591A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a project cost prediction method and system based on a large model, and a storage medium. The method comprises the following steps: standardizing historical cost data to construct a dynamic cost feature knowledge base; inputting the knowledge base into a Transform pre-training large model, and performing fine tuning to obtain a predicted value; obtaining an adjustment coefficient matrix through item similarity clustering and error learning; the design parameters are matched with a knowledge base, and a hierarchical cost table is calculated by using a heterogeneous graph network; and performing time correction on the cost table based on the adjustment coefficient to obtain a prediction result. According to the method, accurate prediction and dynamic adjustment of the project cost are realized, the problem of insufficient element association expression in traditional cost prediction is solved, targeted correction can be carried out according to time and project features, and the prediction accuracy and adaptability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and particularly to a project cost prediction method, system, and storage medium based on large models. Background Art

[0002] Project cost prediction is an important link in project management. Traditional project cost prediction methods mainly include the index method, proportion method, analogy estimation method, and parameter estimation method, etc. The index method is based on historical project cost data and adjusts historical costs through price indices to predict future project costs; the proportion method estimates project costs for new projects by analyzing the proportion relationship of each sub - project in the total cost of historical projects; the analogy estimation method is based on the cost data of similar projects and makes comparison and adjustment in combination with indicators such as area and volume; the parameter estimation method predicts costs by establishing a mathematical relationship between costs and project characteristic parameters. With the development of computer technology, cost prediction methods based on statistics and machine learning have emerged, such as regression analysis, neural networks, support vector machines, etc. These methods describe the relationship between project characteristics and costs by constructing mathematical models, and to a certain extent, improve the prediction accuracy.

[0003] However, the existing project cost prediction methods have many deficiencies. Traditional methods rely too much on manual experience, with limited prediction accuracy and low efficiency; although statistical methods introduce mathematical models, they are difficult to handle non - linear relationships and a large number of characteristic variables; although machine learning methods improve the model complexity, they face problems such as high requirements for data volume, poor model interpretability, and difficulty in making full use of domain knowledge. In addition, existing methods generally have difficulty in coping with the dynamic changes of project costs, especially in the case of frequent market price fluctuations and diverse project types, there are often large deviations between the prediction results and the actual costs. More critically, existing methods lack the ability to effectively express and process the complex correlation relationships between project elements, and cannot comprehensively consider the comprehensive impact of multi - dimensional factors such as materials, processes, and structures on costs, resulting in limited prediction accuracy. Summary of the Invention

[0004] This application provides a project cost prediction method, system, and storage medium based on large models, which are used to achieve accurate prediction and dynamic adjustment of project costs. It not only solves the problem of insufficient expression of element correlations in traditional cost prediction, but also can make targeted corrections according to time and project characteristics, significantly improving the prediction accuracy and adaptability.

[0005] In the first aspect, the present application provides a method for predicting engineering cost based on a large model, which includes: standardizing historical cost data to obtain a cost data set, and constructing an engineering sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base; inputting the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning to obtain an engineering cost prediction value; generating a compensation function for the engineering cost prediction value through historical project similarity clustering and error distribution learning to obtain an adjustment coefficient matrix; associating and matching the design parameters of the target project with the dynamic cost feature knowledge base, calculating the unit cost of each sub-item through heterogeneous graph network analysis, and obtaining a hierarchical engineering cost detailed list; based on the adjustment coefficient matrix, correcting the time dimension of each level of data in the hierarchical engineering cost detailed list to obtain a cost prediction result.

[0006] In a second aspect, the present application provides a large-scale model-based engineering cost prediction system, comprising:

[0007] A construction module is used to standardize historical cost data to obtain a cost data set, and to construct a project sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base;

[0008] An input module is used to input the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning to obtain a predicted value of the project cost;

[0009] A generation module is used to generate a compensation function for the project cost forecast value by historical project similarity clustering and error distribution learning to obtain an adjustment coefficient matrix;

[0010] A matching module is used to associate and match the design parameters of the target project with the dynamic cost feature knowledge base, calculate the unit cost of each sub-item through heterogeneous graph network analysis, and obtain a hierarchical engineering cost detailed list;

[0011] The correction module is used to correct the time dimension of each level data of the hierarchical engineering cost detailed list based on the adjustment coefficient matrix to obtain the cost prediction result.

[0012] In a third aspect, a large-model-based engineering cost prediction device is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the large-model-based engineering cost prediction device executes the above-mentioned large-model-based engineering cost prediction method.

[0013] Fourthly, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is enabled to execute the above-mentioned project cost prediction method based on a large model.

[0014] In the technical solution provided by this application, by standardizing historical cost data and constructing a dynamic cost feature knowledge base, the problems of low data utilization rate and rigid knowledge structure in traditional cost prediction methods are effectively solved. This method inputs the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptation fine-tuning, making full use of the advantages of the Transformer architecture in processing sequence data and long-distance dependency relationships, enabling the model to accurately capture the complex correlations between project cost parameters. By clustering historical project similarities and learning error distributions to generate a compensation function and form an adjustment coefficient matrix, the limitation that traditional prediction methods are difficult to adapt to the characteristics of different types of engineering projects is overcome, and the prediction accuracy is significantly improved. Associating and matching the design parameters of the target project with the dynamic cost feature knowledge base, and calculating the cost of each sub-item unit through heterogeneous graph network analysis, realizes the refined expression and processing of complex engineering structures, and solves the problem that traditional methods cannot effectively handle the complex correlation relationships between engineering elements. Based on the adjustment coefficient matrix, time dimension correction is performed on the data at each level of the hierarchical project cost breakdown table, which not only solves the cost prediction deviation caused by price fluctuations, but also realizes the differential processing of price changes of different resource types, greatly improving the accuracy, timeliness and applicability of project cost prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the accompanying drawings without creative efforts.

[0016] Figure 1 It is a schematic diagram of an embodiment of the project cost prediction method based on a large model in the embodiments of this application;

[0017] Figure 2 It is a schematic diagram of an embodiment of the project cost prediction system based on a large model in the embodiments of this application;

[0018] Figure 3 It is a schematic block diagram of the structure of the project cost prediction device based on a large model in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The present application embodiment provides a kind of engineering cost prediction method, system and storage medium based on large model. The term "first", "second", "third", "fourth" etc. (if any) in the specification and claims of the present application and the above-mentioned drawings is used to distinguish similar objects, and is not necessarily used to describe a specific order or precedence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiment described here can be implemented in an order other than the content illustrated or described here. In addition, the term "comprise" or "have" and any variation thereof are intended to cover non-exclusive inclusion, for example, the process, method, system, product or equipment comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to the process, method, product or equipment.

[0020] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of the method for predicting construction cost based on a large model includes:

[0021] Step S101: Standardize historical cost data to obtain a cost data set, and construct a project sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base;

[0022] Step S102: Input the dynamic cost feature knowledge base into the pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning to obtain the project cost prediction value;

[0023] Step S103: Generate a compensation function for the project cost forecast value by clustering historical project similarity and learning error distribution to obtain an adjustment coefficient matrix;

[0024] Step S104: Associating and matching the design parameters of the target project with the dynamic cost feature knowledge base, calculating the unit cost of each sub-item through heterogeneous graph network analysis, and obtaining a hierarchical project cost detailed list;

[0025] Step S105: Based on the adjustment coefficient matrix, the time dimension of each level of the hierarchical engineering cost detailed list is corrected to obtain the cost prediction result.

[0026] It is understandable that the execution subject of this application can be a large-scale model-based engineering cost prediction system, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.

[0027] Specifically, standardize the historical cost data. Collect the list data, contract price data, and material and labor cost data of a large number of historical engineering projects, and then clean the data through outlier detection and missing value imputation techniques. Outlier detection identifies outliers by calculating the deviation of data points from the overall distribution. For example, the quotation of a residential project is abnormally lower than the market average price. Missing value imputation intelligently fills in the missing values based on the data of similar projects to ensure data integrity. After data cleaning, perform normalization to convert data with different dimensions into values within a unified range, so that subsequent analysis is not affected by the differences in the original dimensions. Based on the standardized cost data set, classify the engineering projects according to type characteristics to obtain an engineering type classification system, covering various types such as residential buildings, public buildings, and industrial buildings. Extract the structural characteristics of each part and item for each type to generate a part and item structural template. Establish an associated mapping between the template and the regional price index and time series price fluctuation data to generate a price adjustment relationship, construct a project cost knowledge structure, and form a dynamic cost feature knowledge base.

[0028] Convert the above dynamic cost feature knowledge base into a serialized format, and process it into a vector representation through a feature encoder for input into the Transformer model. Analyze the structure of the pre-trained large model to understand the composition of its multi-head self-attention layer, feed-forward neural network layer, and layer normalization components. Based on this, construct a cost prediction instruction set and reference response pairs, including instructions such as quota query, unit price calculation, and engineering quantity conversion, to form a fine-tuning data set. During the fine-tuning process, freeze most of the parameters of the Transformer model, and only keep the parameters of the last three self-attention modules and the top linear layer trainable. This parameter-efficient fine-tuning strategy retains the general knowledge of the model and at the same time enables it to focus on specific tasks in the field of project cost. Divide the project cost data into a training set and a validation set, calculate the gradient through the cross-entropy loss function, and use the Adam optimizer to update the parameters to obtain the fine-tuned model. Input the target project parameters into this model, use the multi-head self-attention mechanism to capture the correlation between parameters, and generate the predicted project cost value. Compare the generated predicted project cost value with the actual cost data of historical projects, calculate the prediction deviation value, and generate cost prediction error data. Combine the prediction error data with the project characteristics, group them according to scale, structural type, and functional use, and form a project characteristic classification table. Calculate the similarity between projects in each category, construct a project association network, and generate project clustering groups. Analyze the mean and variance of the prediction errors within each cluster to obtain the error statistical values, and based on this, construct a cost prediction adjustment plan, form a prediction adjustment rule, and generate an adjustment coefficient matrix for different engineering types and scales.

[0029] The target project design parameters are structured, and engineering quantity information, material specifications, and construction process information are extracted to form the target project parameter set. The parameters are matched with the dynamic cost feature knowledge base to identify corresponding items and obtain a parameter mapping relationship table. A heterogeneous graph network is constructed based on the parameter mapping relationship table, and engineering components, materials, processes, and prices are used as different types of nodes to construct a heterogeneous engineering cost graph. Node information transmission calculations are performed on the heterogeneous graph, and node features are weighted and aggregated with adjacent node features to obtain fused features. Based on the fused features, the unit cost of each sub-item is calculated. Components and materials of the same category are classified and counted to generate a unit cost table. Then, according to the hierarchical structure of the bill of quantities, the total amount of each level is calculated to form a hierarchical engineering cost detailed table.

[0030] Based on the adjustment coefficient matrix, the hierarchical project cost schedule is corrected for time dimensions. The material price index, labor cost index, and equipment cost index at the current time point are obtained and compared with the benchmark time index. The time span difference and price change ratio are calculated to generate the time adjustment factor. The time adjustment factor is combined with the adjustment coefficient matrix, and weights are assigned based on project type and regional characteristics to obtain a comprehensive adjustment coefficient. Based on the proportion of material costs, labor costs, and machinery costs, each is multiplied by the comprehensive adjustment coefficient to obtain the time-corrected unit cost. The total price is recalculated and summarized to generate a revised cost schedule, forming the cost forecast result.

[0031] In an embodiment of the present application, by standardizing historical cost data and constructing a dynamic cost feature knowledge base, the problems of low data utilization and rigid knowledge structure in traditional cost forecasting methods are effectively solved. This method inputs the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning, making full use of the advantages of the Transformer architecture in processing sequence data and long-distance dependencies, so that the model can accurately capture the complex correlations between engineering cost parameters. By clustering the similarity of historical projects and learning the error distribution to generate a compensation function and form an adjustment coefficient matrix, the limitation of traditional forecasting methods that are difficult to adapt to the characteristics of different types of engineering projects is overcome, and the forecast accuracy is significantly improved. The design parameters of the target project are associated with the dynamic cost feature knowledge base, and the unit cost of each sub-item is calculated through heterogeneous graph network analysis, which realizes the refined expression and processing of complex engineering structures and solves the problem that traditional methods cannot effectively handle the complex correlations between engineering elements. Based on the adjustment coefficient matrix, the time dimension of the data at each level of the hierarchical engineering cost details table is corrected. This not only solves the cost forecast deviation caused by price fluctuations, but also realizes differentiated treatment of price changes of different resource types, greatly improving the accuracy, timeliness and applicability of engineering cost forecasts.

[0032] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0033] Perform outlier detection and missing value imputation on the list data, contract price data, and material and labor cost data of historical engineering projects to obtain a cleaned dataset;

[0034] Normalize the cleaned dataset to obtain a standardized cost dataset;

[0035] Classify the engineering projects based on the standardized cost dataset to obtain an engineering type classification system;

[0036] Extract features from the sub - item structure of each engineering type in the engineering type classification system to obtain a sub - item structure template;

[0037] Establish an association mapping between the sub - item structure template, regional price index, and time - series price fluctuation data to obtain a price adjustment relationship;

[0038] Construct a project cost knowledge structure based on the sub - item structure template and price adjustment relationship to obtain a dynamic cost feature knowledge base.

[0039] Specifically, perform outlier detection and missing value imputation on the list data, contract price data, and material and labor cost data of historical engineering projects. Outlier detection refers to identifying data points that deviate significantly from the overall data distribution, usually using statistical methods. For example, for the unit cost data of a certain type of engineering project, calculate the mean and standard deviation of all samples. If a data point deviates from the mean by more than 3 standard deviations, it is determined as an outlier. For missing value imputation, the K - Nearest Neighbor algorithm is used. By finding the K most similar data samples to the data point to be imputed, calculate the weighted average value based on the values of the samples to fill in the missing part. For example, when the labor cost data of the decoration project of a certain residential project is missing, find 5 most similar projects of the same type and scale, and calculate the average value of their decoration project labor costs to fill in the missing item. Through these two steps of processing, a cleaned dataset is obtained.

[0040] Normalizing the cleaned dataset is a process of unifying data with different dimensions and ranges to the same scale. Common normalization methods include min-max normalization and Z-score standardization. Min-max normalization linearly transforms the data to the interval [0,1]. For example, the total construction cost data of different engineering projects, which originally ranges from several million to hundreds of millions, is uniformly mapped to the range between 0 and 1 for subsequent analysis and processing. Z-score standardization is to divide by the standard deviation after subtracting the mean, making the data conform to a distribution with a mean of 0 and a standard deviation of 1. Through this processing, a standardized construction cost dataset is obtained. Based on the standardized construction cost dataset, a hierarchical clustering algorithm is used to classify the engineering projects by type. Hierarchical clustering treats each engineering project as an independent category and then gradually merges the most similar categories until the preset number of categories is reached. The similarity calculation uses Euclidean distance or cosine similarity. In this way, the engineering projects are classified into different types such as residential buildings, commercial buildings, industrial buildings, and municipal engineering, forming an engineering type classification system.

[0041] Feature extraction is performed on the sub - item structure of each engineering type in the engineering type classification system to analyze the cost composition characteristics of various types of projects. Feature extraction is achieved through statistical frequency and correlation analysis, calculating the occurrence frequency, cost proportion, and mutual correlation of each sub - item in different engineering types. For example, by analyzing residential building projects, it is found that the foundation project accounts for about 15% of the total cost, the main structure accounts for about 45%, the decoration accounts for about 25%, and the equipment installation accounts for about 15%. By analyzing a large amount of data of similar projects, the key structural features and cost characteristics of each type of project are extracted to establish a sub - item structure template. The sub - item structure template is associated with the regional price index and time - series price fluctuation data to establish the price change relationship at different regions and time points. The material price index, labor cost index, and equipment cost index of different regions, as well as the data of the index change over time, are collected. Then, the correlation coefficient between each index and the sub - item cost is calculated to identify which price index has the greatest impact on which sub - item. By establishing the mapping relationship between the price index and the sub - item cost, a price adjustment relationship is formed to support cost adjustment in different regions and times.

[0042] Based on the sub - item structure template and the price adjustment relationship, a construction cost knowledge structure is constructed to obtain a dynamic cost feature knowledge base. The specific approach is to organize the engineering type, sub - item structure, price index, and adjustment relationship into a relational knowledge structure. In this structure, the engineering type node is connected to the corresponding sub - item node, the sub - item node is connected to the relevant material, labor, and equipment resource nodes, and the resource node is associated with the price index node. Through this network structure, the static template data and the dynamic price adjustment data are integrated to form a dynamic cost feature knowledge base.

[0043] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0044] Convert the dynamic cost feature knowledge base into a serialized input format, and process it into a multi-dimensional vector representation through a feature encoder to obtain a set of cost feature vectors;

[0045] Conduct a structural analysis on the pre-trained large model based on the Transformer architecture. The Transformer architecture includes a multi-head self-attention layer, a feed-forward neural network layer, and a layer normalization component to obtain a model structure mapping;

[0046] Based on the model structure mapping, construct a cost prediction instruction set and reference response pairs, including itemized quota queries, material unit price calculations, and engineering quantity conversion instructions, to obtain an instruction fine-tuning data set;

[0047] According to the set of cost feature vectors and the instruction fine-tuning data set, freeze the parameters of the Transformer model, and keep the parameters of the last three self-attention modules and the top linear layer trainable to obtain a model structure to be fine-tuned;

[0048] Divide the project cost data into a training set and a validation set according to an 8:2 ratio, input them into the model structure to be fine-tuned, calculate the gradient through the cross-entropy loss function, and use the Adam optimizer to perform backpropagation to update the parameters to obtain a fine-tuned model;

[0049] Encode the target project parameters and input them into the fine-tuned model. Through the multi-head self-attention mechanism in the model, capture the correlation relationships between the parameters, and after the calculations of the feed-forward neural network layer and the output layer, obtain the predicted project cost value.

[0050] Specifically, organize the structured knowledge base data, and organize information such as project types, itemized parts, and material prices in a fixed format. Serialization refers to converting a complex data structure into a data sequence that can be continuously processed, which is convenient for input into a large model. For example, the information of a project can be arranged in the format of "project type + itemized parts + material specifications + unit price". After the conversion, use a feature encoder to convert the text and numerical information into vector representations. A feature encoder is a tool that maps input data to a vector space. For text data, word embedding technology is used, and for numerical features, they are directly represented after normalization processing. The obtained set of cost feature vectors is a group of multi-dimensional vectors, and each vector represents a piece of information in the project cost knowledge base, including semantic and numerical features.

[0051] The structural analysis of the pre-trained large model based on the Transformer architecture is to understand the internal structure of the model, so as to perform fine-tuning more pertinently. The Transformer architecture is a neural network structure based on the attention mechanism, and its core components include the multi-head self-attention layer, the feed-forward neural network layer, and the layer normalization component. The multi-head self-attention layer allows the model to simultaneously focus on different positions of the input sequence and capture the dependencies between elements. Each attention head calculates the weighted sum between the query, key, and value to obtain the attention score. The feed-forward neural network layer is a fully connected layer composed of two linear transformations, which is used to perform non-linear transformations on the features output by the attention mechanism. The layer normalization component is used to stabilize the training process and reduce the problem of gradient disappearance. Through the analysis of the model structure, the functions and the number of parameters of each layer are clarified, preparing for the subsequent selective parameter fine-tuning.

[0052] Based on the model structure mapping, constructing a cost prediction instruction set and reference response pairs is to create training data suitable for fine-tuning. The instruction set contains various queries and operation instructions related to project cost, such as the query of sub-item quotas, such as "query the concrete unit price of the main project of a framed residential building"; the calculation of material unit prices, such as "calculate the comprehensive unit price of C30 concrete including transportation and pumping"; the engineering quantity conversion instruction, such as "convert 3000 square meters of building area into the cubic meters of concrete for the foundation project". For each instruction, a professional and accurate response is written to form an instruction-response pair. The data is organized in the model input format, constituting an instruction fine-tuning data set, which is used to teach the model how to correctly understand and respond to queries related to cost prediction.

[0053] Based on the cost feature vector set and the instruction fine-tuning data set, parameter freezing and selective fine-tuning are performed on the Transformer model. Parameter freezing means setting some parameters in the model to an untrainable state and not updating the parameters during the backpropagation process. In this method, the parameters of the first few layers of the model are frozen, and only the parameters of the last three self-attention modules and the top linear layer are kept trainable. The reason for this is that the first few layers of the pre-trained model usually capture general language features, while the last few layers are more focused on feature processing related to specific tasks. The self-attention module is responsible for capturing the dependencies between elements in the input sequence, and the top linear layer is responsible for mapping the features to the output space. Through this selective fine-tuning strategy, the model can retain the general knowledge learned in the pre-training stage and at the same time adapt to specific tasks in the field of project cost, resulting in a fine-tuning model structure with clear fine-tuning objectives. The project cost data is divided into a training set and a validation set according to the ratio of 8:2, which is a commonly used data splitting method in machine learning. The training set is used for the model to learn parameters, and the validation set is used to evaluate the model performance and prevent overfitting. During the fine-tuning process, the cross-entropy loss function is used to calculate the difference between the predicted value and the true value. The cross-entropy loss function is particularly suitable for classification and probability prediction tasks and calculates the degree of difference between the predicted distribution and the true distribution. The optimizer uses the Adam (Adaptive Moment Estimation) algorithm, which is an optimization algorithm that combines the advantages of the momentum method and RMSProp and can adaptively adjust the learning rate to accelerate convergence. In each training batch, the model performs forward propagation calculation on the input data, and then calculates the gradient and updates the trainable parameters through the backpropagation algorithm. The performance on the validation set is continuously monitored during the training process, and the training is stopped when the validation loss no longer decreases to prevent overfitting. Through this process, a fine-tuned model optimized specifically for the project cost prediction task is obtained.

[0054] Encoding the target project parameters and inputting them into the fine-tuned model is the actual application stage. It is necessary to encode the design parameters of the new engineering project and convert them into a vector representation consistent with the model input format. After inputting into the model, the multi-head self-attention mechanism comes into play, and each attention head calculates the correlation scores between the parameters respectively to capture the complex correlation relationships between the engineering parameters. For example, the combined impact of the building area and the structural form on the cost, or the interaction between the decoration standard and the equipment configuration. The associated information undergoes non-linear transformation through the feed-forward neural network layer to extract higher-level feature representations. Finally, the high-dimensional features are mapped into specific cost prediction values through the output layer.

[0055] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0056] Compare the project cost prediction value with the actual cost data of historical engineering projects, calculate the prediction deviation value, and obtain the cost prediction error set;

[0057] Based on the cost prediction error set and the characteristics of engineering projects, establish the correspondence between characteristics and errors, and group them according to the project scale, structure type and functional use to obtain the project characteristic classification matrix;

[0058] Calculate the similarity for each category in the project characteristic classification matrix, construct the similarity network between projects through the cosine similarity function, and obtain the project clustering results;

[0059] Conduct statistical analysis on the cost prediction errors within the project clustering results, calculate the error mean and variance of each clustering center, and obtain the error distribution parameters;

[0060] Based on the error distribution parameters, construct the cost prediction compensation function for various types of engineering projects, fit the relationship between errors and project characteristics through the polynomial regression equation, and obtain the compensation function set;

[0061] Map the compensation function set into the prediction process according to the project classification, generate multi-dimensional adjustment parameters, and obtain the adjustment coefficient matrix.

[0062] Specifically, collect the actual settlement data of completed projects as the true construction value. Subsequently, input the characteristic parameters of these projects into the fine-tuned large model to obtain the predicted construction value. By calculating the difference between the predicted value and the actual value for each project, a prediction deviation value is formed. The prediction deviation value refers to the difference between the predicted construction cost and the actual construction cost, which can be expressed as the absolute error (predicted value - actual value) or the relative error ((predicted value - actual value) / actual value). These prediction deviation values are aggregated to form a construction cost prediction error set, providing basic data for subsequent analysis. Constructing the feature-error correspondence relationship based on the construction cost prediction error set and the engineering project characteristics is a key step in deeply understanding the causes of prediction errors. Extract the key characteristics of the engineering project, including project scale, structural type, and functional use, etc. Then establish a mapping relationship between the project characteristics and their corresponding prediction errors to form a feature-error correspondence table. Group the projects according to the feature attributes, such as dividing them into small, medium, and large according to the area scale; dividing them into frame structure, frame-shear wall structure, steel structure, etc. according to the structural type; dividing them into residential, office, commercial, etc. according to the functional use. Through this grouping and sorting, transform the feature-error correspondence table into a multi-dimensional project feature classification matrix for subsequent similarity analysis. Calculating the similarity for each category in the project feature classification matrix is the process of identifying similar project groups. In each feature classification, it is necessary to calculate the similarity degree between projects. Cosine similarity is a commonly used similarity calculation method, which measures their similarity by calculating the cosine value of the angle between two vectors. The smaller the angle between the feature vectors of two projects, the closer the cosine value is to 1, indicating that the projects are more similar. By calculating the cosine similarity between each pair of projects, construct a project similarity network, which is a graph structure representing the association strength between projects. Based on the similarity network, use a clustering method to group similar projects into one group to form a project clustering result. The clustering method can be the K-means algorithm or the hierarchical clustering algorithm, which automatically classifies projects with high similarity into the same cluster.

[0063] Conduct statistical analysis on the construction cost prediction errors within the project clustering results. Within each cluster, collect the prediction error data of all projects and calculate the mean of the errors, which represents the average level of prediction deviation for projects in this category and indicates the systematic deviation of the prediction. At the same time, calculate the variance of the errors, which reflects the stability and dispersion degree of the prediction results for projects in this category. Statistical analysis also includes calculating the distribution form of the errors, such as skewness and kurtosis, to judge whether the error distribution conforms to the normal distribution. Through these statistics, obtain the error distribution parameters describing the prediction error characteristics of various types of projects, providing a data basis for the subsequent construction of the compensation function.

[0064] For each project category, the relationship between error and project characteristics is analyzed. Polynomial regression is used to establish a mathematical relationship between project characteristics and forecast error. Polynomial regression is a regression method that can capture nonlinear relationships. By fitting combinations of project characteristics of varying orders, an error prediction equation is constructed. A compensation function is constructed for each project category, forming a set of compensation functions that are used to estimate and correct potential errors based on project characteristics during forecasting.

[0065] Based on the characteristics of the project to be forecasted, the project category to which it belongs is determined. The corresponding compensation function is called to calculate the compensation value for the forecast error. The compensation value is converted into an adjustment coefficient, forming an adjustment coefficient matrix tailored to different feature dimensions. The adjustment coefficient matrix is a multidimensional structure that includes adjustment parameters for different dimensions, such as project scale, structural type, and functional purpose. In actual forecasting, the corresponding adjustment parameters are extracted from the adjustment coefficient matrix based on the specific characteristics of the project to modify the original forecast results.

[0066] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0067] Structural processing is performed on the design parameters of the target project to extract engineering quantity information, material specifications, and construction processes to obtain the target project parameter set;

[0068] Match the target project parameter set with the dynamic cost feature knowledge base, find the corresponding item in the knowledge base for each design parameter, and obtain a parameter mapping relationship table;

[0069] Construct a heterogeneous graph network structure based on the parameter mapping relationship table, set engineering components, materials, processes and prices as different types of nodes, and construct a heterogeneous engineering cost graph based on the set multiple nodes;

[0070] Perform node information transmission calculation on the heterogeneous graph of engineering cost, perform weighted aggregation on the features of each node and the features of adjacent nodes to obtain fusion features;

[0071] Based on the fusion features, the unit cost of each sub-item is calculated according to the hierarchical structure of the project, and components and materials of the same category are classified and counted to obtain a unit cost table;

[0072] Organize the unit cost table according to the hierarchical structure of the bill of quantities, and calculate the total amount of each level to obtain a hierarchical project cost detailed table.

[0073] Specifically, the structuring process refers to converting unstructured or semi-structured design data into a standardized structured format. Key information is extracted from design drawings, construction specifications, and bill of quantities, including engineering quantity information (such as concrete volume, steel usage, wall area, etc.), material specifications (such as concrete grade, steel bar type, quality of decorative materials, etc.), and construction techniques (such as construction methods, process requirements, technical standards, etc.). The extraction process uses text analysis and data normalization techniques to uniformly convert information in different formats into a standardized data structure. The converted data is organized according to the hierarchical levels of the project's sub - items, forming a structured target project parameter set, which lays the foundation for subsequent cost calculation.

[0074] Data matching between the target project parameter set and the dynamic cost feature knowledge base is a key link connecting project features and cost data. In the matching process, each engineering element in the project parameter set is compared with the entries in the knowledge base to find the most matching corresponding item. The matching methods include exact matching and fuzzy matching. Exact matching is used to find exactly the same items, such as materials with specific specifications or standard components; fuzzy matching is used to handle cases where they are not exactly the same, and the most similar entry is found by calculating the similarity. The particularity of the project and regional differences are also considered during the matching process to appropriately adjust the standard items. Finally, a parameter mapping relationship table is formed. Each row in the table contains a parameter of the target project and its corresponding entry in the knowledge base, establishing the mapping relationship between project features and cost knowledge.

[0075] Constructing a heterogeneous graph network structure based on the parameter mapping relationship table is an innovation point of the engineering cost prediction method. A heterogeneous graph network is a graph structure that contains multiple types of nodes and edges and can represent complex association relationships. During the construction process, different types of elements such as engineering components, materials, techniques, and prices are set as nodes in the graph, and each type of node has different attribute characteristics. Then, according to the parameter mapping relationship table, connection relationships between nodes are established. For example, a "beam" node and a "C30 concrete" node establish a "usage" relationship, and a "C30 concrete" node and its "unit price" node establish a "pricing" relationship. The constructed heterogeneous graph of engineering cost contains a complex association network between various elements in the engineering project, providing a structured basis for subsequent information transmission and feature fusion.

[0076] Performing node information passing calculations on the engineering cost heterogeneous graph is the core step of using the graph network to capture correlation relationships. The information passing calculation is based on the principle of the Graph Attention Network (GAT), enabling each node to exchange information with its adjacent nodes and update its own features. In the specific process, each node collects the feature information of its adjacent nodes, and then weights the contributions of different adjacent nodes through the attention mechanism, with more important adjacent nodes obtaining higher weights. The information after weighted aggregation is combined with the features of the node itself to form an updated node representation. This information passing mechanism enables each node to retain its own features while integrating the information of associated nodes, forming context-aware integrated features. For example, after information passing, a building component node not only contains its own dimension information but also integrates the price change information of the materials used and the complexity information of the construction process. Based on the integrated features, calculating the unit cost of each sub-item according to the hierarchical structure of the engineering project is the key step in transforming the graph network learning results into the actual cost. The calculation process identifies components and materials with the same features and classifies and counts them. The unit cost of each type of component is calculated, and the unit cost consists of direct costs (material costs, labor costs, machinery costs) and indirect costs (management fees, profits, taxes, etc.). The direct costs are calculated based on the price information and consumption information in the integrated features, and the indirect costs are determined according to industry norms and project characteristics. The associated influencing factors included in the integrated features are considered in the calculation. For example, an increase in the price of a certain material not only affects the material cost but may also affect the labor cost of the relevant construction process. Finally, a unit cost table is formed, which lists the detailed cost composition and unit price information of each type of component.

[0077] Organizing the unit cost table according to the hierarchical structure of the bill of quantities and calculating the total amount of each level is the process of forming the final cost prediction result. The hierarchical structure of the bill of quantities usually includes three levels: sub-projects, sub-items, and sub-titles. Multiply the unit prices in the unit cost table by the corresponding quantities to calculate the total price of each sub-title. Then, according to the list structure, sum up the total prices of the sub-titles level by level to obtain the sub-project cost and sub-project cost. During the summarization process, consider the associated influences between each level and make necessary adjustments to the cost. Finally, a hierarchical engineering cost breakdown table is formed, which clearly shows the complete cost composition from the total cost to each sub-project, sub-item, and specific sub-title.

[0078] In a specific embodiment, the process of performing step S105 may specifically include the following steps:

[0079] Group and organize the integrated features according to the engineering sub-item coding system to obtain a grouped feature set;

[0080] Extract the material, labor, and machinery consumption information for each sub - item based on the grouped feature set, match it with the unit prices in the market price database, and obtain the resource unit price table;

[0081] Multiply each resource in the resource unit price table by the corresponding consumption quantity, calculate the direct cost of each sub - item, and calculate the indirect cost according to the specification requirements to obtain the cost composition data;

[0082] Summarize the cost composition data according to the cost calculation rules of the bill of quantities specification, add taxes, fees, and measure fees to obtain the comprehensive unit price of each sub - item;

[0083] Multiply the comprehensive unit price by the quantity of work, calculate the total price of each sub - item, and summarize them hierarchically according to the hierarchical relationship of the project to obtain the hierarchical cost summary table;

[0084] Analyze the rationality of the data in the hierarchical cost summary table, confirm the data rationality by comparing with historical data, and obtain the unit cost table.

[0085] Specifically, the project sub - item coding system is a standardized coding system used in project cost management to identify different project parts. It usually includes major categories such as the main project, decoration project, and installation project. Each major category is further divided into specific sub - items. The grouping and sorting process matches the feature data with the standard coding to identify the sub - item category to which each feature belongs. For example, the feature data of concrete components is classified into the "main structure project" category, and the feature data of wall decoration is classified into the "decoration and fit - out project" category. Then it is organized according to the coding hierarchical structure to form a grouped feature set that conforms to the bill of quantities specification. In the grouped feature set, items with the same coding are grouped together, and each group contains all the feature information required for this sub - item project, laying a foundation for subsequent refined cost calculation.

[0086] Consumption information refers to the quantity of various resources required to complete a unit of work quantity, including material usage, labor days, and machinery shifts, etc. When extracting this information from the grouped feature set, the influence of project characteristics, construction methods, and technical specifications needs to be considered. For example, for a reinforced concrete structure, according to the structure type and reinforcement ratio, extract the steel usage, formwork area, concrete and other material consumptions per cubic meter of concrete, as well as the corresponding labor and machinery inputs. After extraction, match this consumption information with the unit prices in the market price database. The market price database stores the latest market prices of various materials, labor, and machinery. Through the matching process, an appropriate unit price is determined for each resource to form the resource unit price table. The resource unit price table is a data set that details all resource types, specifications, consumption quantities, and unit prices in the project, providing a data basis for direct cost calculation.

[0087] Direct costs include three parts: material costs, labor costs, and machinery costs, which are calculated by multiplying the consumption of various resources by the corresponding unit prices. For example, material costs equal the sum of the usage of various materials multiplied by their unit prices, labor costs equal the sum of the man-days of each type of work multiplied by the daily wage rate, and machinery costs equal the sum of the number of machine working shifts of various types multiplied by the shift unit price. After calculating the direct costs, indirect costs are calculated according to the industry norms of engineering construction and the requirements of the contract. Indirect costs usually include enterprise management fees, on-site management fees, etc., and are generally calculated as a certain percentage of the direct costs. Adding the direct costs and indirect costs together gives the cost composition data, which is the basic component of the project cost. The Bill of Quantities Specification stipulates the calculation methods and procedures from direct costs, indirect costs to the final comprehensive unit price. During the summarization process, taxes, fees, and measure costs need to be added. Taxes refer to various taxes payable during the engineering construction process, such as value-added tax; fees refer to the fees that must be paid according to national regulations, such as social insurance fees; measure costs refer to the non-direct production costs incurred during the construction process, such as temporary facility costs, safety and civilized construction costs, etc. These costs are calculated according to the calculation methods and ratios stipulated in the specification and, together with the direct costs and indirect costs, constitute the complete project cost. Through this processing, the comprehensive unit price of each sub-item is obtained, that is, the complete cost of the unit engineering quantity. The engineering quantity refers to the actual quantity of each sub-item project, such as the volume of concrete, the weight of steel bars, the area of the wall surface, etc. Multiplying the comprehensive unit price calculated in the previous step by the corresponding engineering quantity gives the total cost of each sub-item project, that is, the combined price. After the combined price calculation is completed, hierarchical summarization is carried out according to the hierarchical structure of the engineering project. Engineering projects are usually divided into levels such as unit projects, sub-projects, and sub-items. The summarization process starts from the lowest-level sub-items and accumulates step by step upward to form the cost amounts at each level. This hierarchical summarization method meets the actual needs of project management, facilitates project managers to understand the cost composition of each part of the project, and obtains the cost summary tables at each level.

[0088] The rationality analysis uses the comparison method to compare the calculated cost data with the cost data of historical similar projects to check whether there are obvious deviations. The comparison indicators include the unit cost (the cost per square meter of building area), the proportion of sub-projects (such as the proportion of the structural project in the total cost), the main material consumption index (such as the amount of steel bars per square meter of building area), etc. If a large deviation is found in a certain indicator compared with the historical data, it is necessary to trace back the calculation process, find the reasons and make adjustments. Through this rationality analysis, it is ensured that the final cost prediction result conforms to the actual situation of the project, and a verified unit cost table is obtained.

[0089] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0090] Obtain the material price index, labor cost index, and equipment cost index at the current time point, compare them with the benchmark time index used in the hierarchical project cost breakdown schedule calculation, and obtain the time span difference value;

[0091] Calculate the price change ratio based on the time span difference value, classify and statistically analyze the price changes of various resources, and obtain the time adjustment factor;

[0092] Perform a weighted combination of the time adjustment factor and the adjustment coefficient matrix, and allocate weights according to the project type and project area characteristics to obtain the comprehensive adjustment coefficient;

[0093] Multiply the unit cost of each sub - item unit in the hierarchical project cost breakdown schedule by the comprehensive adjustment coefficient according to the composition ratio of material cost, labor cost, and machinery cost to obtain the unit cost after time correction;

[0094] Multiply the unit cost after time correction by the engineering quantity, recalculate the total price of each sub - item, and summarize according to the hierarchical relationship to obtain the corrected cost breakdown schedule;

[0095] Conduct a comprehensive analysis of the corrected cost breakdown schedule, sort the data of each level according to importance, generate a cost prediction report, and obtain the cost prediction result.

[0096] Specifically, the price index is a relative value reflecting the change in price level and is used to measure the price change index during a specific period. In the field of project cost, the commonly used price indices include material price index, labor cost index, and equipment cost index. The sources for obtaining the current index data include the official indices released by industry associations and market research data. The benchmark time index refers to the price index used when compiling the hierarchical project cost breakdown schedule, which may be data from several months ago or earlier. The comparison process obtains the time span difference value by calculating the difference between the current index and the benchmark index, and this difference value directly reflects the degree of change in the prices of various resources in the time dimension.

[0097] Calculating the price change ratio based on the time span difference value is a key step in quantifying the impact of price fluctuations. The price change ratio is obtained by dividing the current index by the benchmark index, indicating the proportion of price increase or decrease. For different types of resources, there may be significant differences in price changes. For example, the price of steel may increase significantly, while the price of some decorative materials may be relatively stable. By classifying and statistically analyzing the price changes of various resources, grouping the resources according to the direction and amplitude of price changes, a time adjustment factor containing different resource categories is obtained. The time adjustment factor is a set of coefficients reflecting the time changes in the prices of different resource categories and is used for subsequent time - dimension correction of cost data.

[0098] The adjustment coefficient matrix is generated through historical project similarity clustering and error distribution learning in the previous steps, reflecting the impact of characteristics such as project type and scale on cost prediction. The weighted combination process assigns weights according to the engineering project type and project area characteristics. For example, for steel structure projects that are greatly affected by material prices, the time adjustment factor weight of the material price index is relatively high; while for labor-intensive decoration projects, the time adjustment factor weight of the labor cost index is relatively high. Through this weighted combination, a comprehensive adjustment coefficient is obtained, which is a comprehensive correction parameter considering the time dimension and project characteristics and is used to comprehensively adjust the cost data. Multiplying the cost of each sub-item unit in the hierarchical project cost breakdown table by the comprehensive adjustment coefficient according to the composition ratios of material costs, labor costs, and machinery costs respectively is the key step to achieve refined adjustment. The composition of project costs is usually divided into three categories: material costs, labor costs, and machinery costs, and the proportion of each category in the total cost varies depending on the nature of the project. For each sub-item project, it is necessary to analyze its cost composition and determine the proportions of material costs, labor costs, and machinery costs. Then, according to different cost categories, the corresponding adjustment coefficients are extracted from the comprehensive adjustment coefficient for multiplication operations. For example, for a sub-item with a material cost proportion of 60%, the adjustment coefficient related to material prices is used; for a sub-item with a labor cost proportion of 30%, the adjustment coefficient related to labor costs is used. Through this classified adjustment, the unit cost after time correction is obtained, which is more accurate cost data considering the time change factor of prices.

[0099] Multiplying the unit cost after time correction by the engineering quantity and recalculating the total price of each sub-item is the process of converting unit price adjustment into total price adjustment. The engineering quantity data comes from the project design documents and the bill of quantities and is a specific quantity indicator reflecting the project scale. Multiplying the corrected unit cost by the corresponding engineering quantity gives the updated total price of each sub-item project. After the total price calculation is completed, hierarchical summarization is carried out according to the hierarchical relationship of the project, starting from the lowest-level sub-item projects and accumulating step by step upwards to form the total cost of the sub-project, unit project, and finally the entire project. This bottom-up summarization method ensures the integrity and consistency of cost correction, resulting in a corrected cost breakdown table, which is a complete cost data set reflecting the current price level.

[0100] Comprehensive analysis includes checking the reasonableness of the total cost and sub-item costs, evaluating the impact of key materials and project changes, and identifying risk factors. Importance ranking is based on three factors: cost amount, change range, and sensitivity, placing the projects with a greater impact on the total cost at the forefront to facilitate decision-makers to focus on key contents. Through this analysis and ranking, a structured cost prediction report is formed, including total cost data, sub-item cost details, time adjustment factor analysis, and key attention items.

[0101] The above describes the method for predicting project cost based on a large model in the embodiments of the present application. Next, the system for predicting project cost based on a large model in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the system for predicting project cost based on a large model in the embodiments of the present application includes:

[0102] A construction module 201, configured to perform standardization processing on historical cost data to obtain a cost data set, and construct an engineering sub - item structure based on the cost data set to obtain a dynamic cost feature knowledge base;

[0103] An input module 202, configured to input the dynamic cost feature knowledge base into a pre - trained large model based on the Transformer architecture for domain - adaptation fine - tuning to obtain a project cost prediction value;

[0104] A generation module 203, configured to generate a compensation function for the project cost prediction value through historical project similarity clustering and error distribution learning to obtain an adjustment coefficient matrix;

[0105] A matching module 204, configured to associate and match the design parameters of a target project with the dynamic cost feature knowledge base, and calculate the cost of each sub - item unit of the project through heterogeneous graph network analysis to obtain a hierarchical project cost breakdown table;

[0106] A correction module 205, configured to perform time - dimension correction on the data of each level of the hierarchical project cost breakdown table based on the adjustment coefficient matrix to obtain a cost prediction result.

[0107] Through the collaborative cooperation of the above-mentioned various components, by standardizing historical cost data, a dynamic cost feature knowledge base is constructed, effectively solving the problems of low data utilization rate and rigid knowledge structure in traditional cost prediction methods. This method inputs the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptation fine-tuning, making full use of the advantages of the Transformer architecture in processing sequence data and long-distance dependency relationships, enabling the model to accurately capture the complex correlations between project cost parameters. By clustering historical project similarities and learning the error distribution to generate a compensation function, an adjustment coefficient matrix is formed, overcoming the limitations of traditional prediction methods that are difficult to adapt to the characteristics of different types of engineering projects and significantly improving the prediction accuracy. Associating and matching the design parameters of the target project with the dynamic cost feature knowledge base, and calculating the cost of each sub-item unit through heterogeneous graph network analysis, realizing the refined expression and processing of complex engineering structures, and solving the problem that traditional methods cannot effectively handle the complex correlation relationships between engineering elements. Based on the adjustment coefficient matrix, the data at each level of the hierarchical project cost breakdown schedule is corrected in the time dimension, not only solving the cost prediction deviation caused by price fluctuations, but also realizing the differential processing of price changes of different resource types, greatly improving the accuracy, timeliness, and applicability of project cost prediction.

[0108] Above Figure 2 The engineering cost prediction system based on a large model in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the engineering cost prediction device based on a large model in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0109] Figure 3 FIG. is a schematic structural diagram of an engineering cost prediction device based on a large model provided by an embodiment of the present invention. The engineering cost prediction device 300 based on a large model may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) for storing application programs 333 or data 332. Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the engineering cost prediction device 300 based on a large model. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the engineering cost prediction device 300 to implement the steps of the above-mentioned engineering cost prediction method based on a large model.

[0110] The large model-based project cost prediction device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 The structure of the large model-based project cost prediction device shown does not constitute a limitation on the large model-based project cost prediction device provided by the present invention. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0111] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the large model-based project cost prediction method.

[0112] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a large model-based project cost prediction device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0114] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and the modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for predicting engineering cost based on large models, characterized in that, The method comprises: Standardize historical cost data to obtain a cost data set, and construct a project sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base; Inputting the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning to obtain a predicted value of the project cost; Generate a compensation function for the project cost forecast value by historical project similarity clustering and error distribution learning to obtain an adjustment coefficient matrix; Associating and matching the design parameters of the target project with the dynamic cost feature knowledge base, calculating the unit cost of each sub-item through heterogeneous graph network analysis, and obtaining a hierarchical engineering cost detailed list; Based on the adjustment coefficient matrix, the time dimension of each level data of the hierarchical engineering cost detailed list is corrected to obtain the cost prediction result.

2. The method for predicting project cost based on a large model according to claim 1, wherein The historical cost data is standardized to obtain a cost data set, and a project sub-item structure is constructed based on the cost data set to obtain a dynamic cost feature knowledge base, including: Perform outlier detection and missing value filling on the list data, contract price data, and material and labor cost data of historical engineering projects to obtain a clean data set; Normalizing the cleaned data set to obtain a standardized cost data set; Classifying engineering projects based on the standardized cost data set to obtain an engineering type classification system; Extracting features of the sub-item structure of each engineering type in the engineering type classification system to obtain a sub-item structure template; Establishing an association mapping between the sub-item structure template and the regional price index and time series price fluctuation data to obtain a price adjustment relationship; Based on the sub-item structure template and the price adjustment relationship, a project cost knowledge structure is constructed to obtain a dynamic cost feature knowledge base.

3. The method for predicting project cost based on a large model according to claim 1, wherein The dynamic cost feature knowledge base is input into a pre-trained large model based on the Transformer architecture for domain adaptability fine-tuning to obtain a project cost prediction value, including: Converting the dynamic cost feature knowledge base into a serialized input format and processing it into a multi-dimensional vector representation through a feature encoder to obtain a cost feature vector set; Structural analysis of a large pre-trained model based on the Transformer architecture, which includes multi-head self-attention layers, feedforward neural network layers, and layer normalization components, to obtain a model structure map; Based on the model structure mapping, a cost prediction instruction set and a reference response pair are constructed, including sub-item quota query, material unit price calculation and engineering quantity conversion instructions, to obtain an instruction fine-tuning data set; Freeze the parameters of the Transformer model based on the cost feature vector set and the instruction fine-tuning dataset, retain the parameters of the last three layers of self-attention modules and the top linear layer for training, and obtain the model structure to be fine-tuned; The construction cost data is divided into a training set and a validation set in a ratio of 8:2, and input into the model structure to be fine-tuned. The gradient is calculated by the cross entropy loss function, and the Adam optimizer is used to perform back propagation to update the parameters to obtain the fine-tuned model; Encode the target engineering parameters and input them into the fine-tuned model. Through the multi-head self-attention mechanism in the model, capture the correlation relationships between the parameters. After the calculations of the feed-forward neural network layer and the output layer, obtain the predicted value of the project cost.

4. The method for predicting project cost based on a large model according to claim 1, wherein Generate a compensation function for the predicted value of the project cost through historical project similarity clustering and error distribution learning to obtain an adjustment coefficient matrix, including: Compare the predicted value of the project cost with the actual cost data of historical engineering projects, calculate the prediction deviation value, and obtain a cost prediction error set. Based on the cost prediction error set and the engineering project characteristics, construct a feature-error correspondence relationship, and group them according to the project scale, structural type, and functional use to obtain a project feature classification matrix. Calculate the similarity for each category in the project feature classification matrix, construct a project similarity network through the cosine similarity function, and obtain the project clustering result. Conduct statistical analysis on the project cost prediction errors within the project clustering result, calculate the error mean and variance of each clustering center, and obtain the error distribution parameters. Based on the error distribution parameters, construct a cost prediction compensation function for various types of engineering projects, fit the relationship between the error and the project characteristics through a polynomial regression equation, and obtain a set of compensation functions. Map the set of compensation functions into the prediction process according to the engineering classification to generate multi-dimensional adjustment parameters and obtain an adjustment coefficient matrix.

5. The method for predicting engineering cost based on large models according to claim 1, wherein Associate and match the design parameters of the target project with the dynamic cost feature knowledge base, and calculate the cost of each sub-item unit through heterogeneous graph network analysis to obtain a hierarchical project cost breakdown table, including: Conduct structured processing on the design parameters of the target project, extract the engineering quantity information, material specifications, and construction technology to obtain a target project parameter set. Match the target project parameter set with the dynamic cost feature knowledge base, find the corresponding items in the knowledge base for each design parameter, and obtain a parameter mapping relationship table. Based on the parameter mapping relationship table, construct a heterogeneous graph network structure, set engineering components, materials, processes, and prices as different types of nodes, and construct a project cost heterogeneous graph according to the set multiple nodes. Conduct node information transfer calculations on the project cost heterogeneous graph, perform weighted aggregation on the features of each node and the features of adjacent nodes to obtain fused features. Based on the fused features, calculate the cost of each sub-item unit according to the hierarchical structure of the engineering project, classify and count the components and materials of the same category, and obtain a unit cost table. Organize the unit cost table according to the hierarchical structure of the bill of quantities, and calculate the total amount of each level to obtain a hierarchical project cost breakdown table.

6. The method for predicting engineering cost based on large models according to claim 5, wherein, Based on the fused features, calculate the cost of each sub-item unit according to the hierarchical structure of the engineering project, classify and count the components and materials of the same category, and obtain a unit cost table, including: Group and sort the fused features according to the engineering sub-item coding system to obtain a grouped feature set. Extract the material, labor, and machinery consumption information of each sub-item based on the grouped feature set, match it with the unit prices in the market price database, and obtain a resource unit price table. Multiply each resource in the resource unit price table by the corresponding consumption volume, calculate the direct cost of each sub-item, and calculate the indirect cost according to the specification requirements to obtain the cost composition data; Summarize the cost composition data according to the cost calculation rules of the bill of quantities specification, add taxes, fees and measures costs to obtain the comprehensive unit price of each sub-item; Multiply the comprehensive unit price by the quantity of work, calculate the total price of each sub-item, and summarize them hierarchically according to the hierarchical relationship of the project to obtain the hierarchical cost summary table; Analyze the rationality of the data in the hierarchical cost summary table, and confirm the data rationality by comparing with historical data to obtain the unit cost table.

7. The method for predicting project cost based on a large model according to claim 1, wherein Based on the adjustment coefficient matrix, correct the data of each level of the hierarchical project cost breakdown table in the time dimension to obtain the cost prediction result, including: Obtain the material price index, labor cost index and equipment cost index at the current time point, compare them with the benchmark time index used in the calculation of the hierarchical project cost breakdown table, and obtain the time span difference value; Calculate the price change ratio based on the time span difference value, classify and count the price changes of various resources to obtain the time adjustment factor; Perform a weighted combination of the time adjustment factor and the adjustment coefficient matrix, and allocate weights according to the project type and project area characteristics to obtain the comprehensive adjustment coefficient; Multiply the unit cost of each sub-item in the hierarchical project cost breakdown table by the comprehensive adjustment coefficient respectively according to the composition ratio of material cost, labor cost and machinery cost to obtain the unit cost after time correction; Multiply the unit cost after time correction by the quantity of work, recalculate the total price of each sub-item, and summarize them according to the hierarchical relationship to obtain the corrected cost breakdown table; Conduct a comprehensive analysis of the corrected cost breakdown table, sort the data of each level according to the importance, generate a cost prediction report, and obtain the cost prediction result.

8. A project cost prediction system based on a large model, characterized in that, For implementing the engineering cost prediction method based on the large model as described in any one of claims 1-7, the engineering cost prediction system based on the large model includes: A construction module for standardizing historical cost data to obtain a cost data set, and constructing an engineering sub-item structure based on the cost data set to obtain a dynamic cost feature knowledge base; An input module for inputting the dynamic cost feature knowledge base into a pre-trained large model based on the Transformer architecture for domain adaptation fine-tuning to obtain an engineering cost prediction value; A generation module for generating a compensation function for the engineering cost prediction value through historical project similarity clustering and error distribution learning to obtain an adjustment coefficient matrix; A matching module for associatively matching the design parameters of the target project with the dynamic cost feature knowledge base, and calculating the unit cost of each sub-item through heterogeneous graph network analysis to obtain a hierarchical project cost breakdown table; A correction module for correcting the data of each level of the hierarchical project cost breakdown table in the time dimension based on the adjustment coefficient matrix to obtain the cost prediction result.

9. An engineering cost prediction device based on a large model, characterized in that, It includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the large model-based project cost prediction method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the large model-based project cost prediction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dynamic management control method for cost of power transmission and transformation project

    CN105956788A

  • Estimation method for compiling new substation project

    CN113159841A

  • Power transmission and transformation project cost model construction method, system and device based on big data and medium

    CN116611785A

  • Engineering cost management method and system based on big data

    CN117808499A

  • House building project cost estimation method and device and storage medium

    CN117934035A

Cited By

  • Road material cost measuring and calculating method based on generative adversarial network and conditional regression

    CN120912249A

  • A Road Material Cost Calculation Method Based on Generative Adversarial Networks and Conditional Regression

    CN120912249B

  • Project construction cost data management method, system, equipment and medium

    CN121235494A

  • Engineering cost intelligent evaluation method based on graph neural network

    CN121258187A

  • Classification pricing method and system for 10 kV distribution network non-power-cut operation cost

    CN121504561A