A multi-source data-based engineering cost dynamic prediction system and method
By constructing a dynamic engineering cost prediction system based on multi-source data, the problems of insufficient data utilization and poor model generalization ability in existing technologies are solved. It achieves high-precision and dynamic engineering cost prediction and risk assessment, has self-learning and closed-loop optimization capabilities, and can adapt to the high-frequency changes of complex engineering projects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN XIAO KELP DATA TECH CO LTD
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing engineering cost prediction systems suffer from insufficient data utilization, weak multi-source data fusion capabilities, poor model generalization ability, and a lack of effective feedback loop mechanisms and risk perception capabilities, resulting in low prediction accuracy and an inability to meet the intelligent needs of modern engineering management.
A dynamic engineering cost prediction system based on multi-source data is constructed, including a data acquisition and preprocessing module, a mechanism decoupling and correction module, a model training and optimization module, a credibility feedback and prediction module, and a model closed-loop optimization module. Through mechanism decoupling and correction of multi-source data, model training and optimization, and credibility assessment and feedback, dynamic prediction and risk assessment of engineering costs are achieved, and closed-loop optimization is performed.
It improves the accuracy and stability of engineering cost prediction, achieves high-precision cost prediction and risk assessment, has self-learning and dynamic optimization capabilities, adapts to the high frequency of changes and complex environments of engineering projects, and enhances risk prevention capabilities.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of engineering cost prediction, and more specifically, to a dynamic engineering cost prediction system and method based on multi-source data. Background Technology
[0002] Construction cost is a core component of construction project management, and its accuracy directly impacts investment control, risk management, and decision-making efficiency. Traditional construction cost forecasting methods rely heavily on expert experience, historical cost indicators, and linear regression models. These methods have limited responsiveness to changes in the external environment, construction dynamics, and policy shifts, making them unsuitable for the high-frequency changes and high uncertainty inherent in today's complex construction projects.
[0003] With the rapid development of IoT, AI, and big data technologies, engineering projects have gradually generated massive amounts of multi-source heterogeneous data during construction, such as economic indicator data, construction behavior data, and environmental disturbance data. This data contains rich information on the evolution patterns of engineering costs and risk factors, providing new insights for achieving more accurate, dynamic, and intelligent cost prediction.
[0004] However, existing cost prediction systems suffer from insufficient data utilization, weak multi-source data fusion capabilities, and failure to fully explore the interactions between economic, construction, and environmental factors; poor model generalization ability, ignoring the heterogeneity and complexity of data, resulting in weak robustness and limited accuracy of prediction models; lack of feedback loops, with a lack of effective prediction credibility assessment and historical feedback mechanisms, preventing models from achieving dynamic optimization; and weak risk perception capabilities, relying solely on qualitative judgments to identify risk levels, lacking data-driven quantitative evaluation methods.
[0005] Therefore, there is an urgent need to build an engineering cost prediction system that integrates multi-source data, has high-precision prediction and risk perception capabilities, and possesses self-learning and closed-loop optimization capabilities, in order to adapt to the intelligent development needs of modern engineering management. Summary of the Invention
[0006] The purpose of this invention is to provide a dynamic engineering cost prediction system and method based on multi-source data. This system addresses the shortcomings of existing cost prediction systems, such as insufficient data utilization, weak multi-source data fusion capabilities, failure to fully explore the interactions between economic, construction, and environmental factors, poor model generalization ability (ignoring data heterogeneity and complexity leading to weak robustness and limited accuracy), lack of feedback loops (lacking effective prediction reliability assessment and historical feedback mechanisms, preventing dynamic model optimization), and weak risk perception capabilities (relying solely on qualitative judgments for risk level identification, lacking data-driven quantitative evaluation methods, and failing to meet user needs).
[0007] This invention achieves the above objective through the following technical solution: a dynamic prediction system for engineering costs based on multi-source data, the system comprising: The module includes: data acquisition and preprocessing, mechanism decoupling and correction, model training and optimization, credibility feedback and prediction, and model closed-loop optimization. The data acquisition and preprocessing module is used to acquire multi-source heterogeneous raw data required for engineering cost prediction and to preprocess the multi-source heterogeneous raw data. The mechanism decoupling and correction module is used to perform mechanism decoupling and interference factor inversion correction on the preprocessed multi-source data to generate corrected feature data. The model training and optimization module is used to build a multi-source cost risk dynamic prediction model and complete the training and iterative optimization of the model. The credibility feedback and prediction module is used to input the corrected feature data into the trained model to realize dynamic prediction and risk assessment of engineering costs, and output the prediction results and the corresponding prediction credibility and risk level. The model closed-loop optimization module is used to dynamically update model parameters and data source weights based on prediction results and credibility assessment results, thereby achieving closed-loop optimization of engineering cost prediction.
[0008] Furthermore, the data acquisition and preprocessing module includes the following steps: Multi-source heterogeneous raw data is obtained through engineering databases, industry monitoring platforms, IoT sensing terminals, and policy release channels. The multi-source heterogeneous raw data includes three major categories: economic data, construction data, and environmental data. The data acquisition and preprocessing module sequentially performs data cleaning, missing value completion, outlier removal, and normalization on the multi-source heterogeneous raw data, and divides the preprocessed standardized multi-source data into training set, validation set, and test set according to a preset ratio.
[0009] Furthermore, the data acquisition and preprocessing module also includes the following steps: K-nearest neighbor interpolation is used to complete missing time series data, and mean interpolation is used to complete missing numerical data. Outliers are removed in principle, and normalization is performed using the Z-score standardization method. The preprocessed standardized multi-source data is divided into training set, validation set and test set in a ratio of 7:2:1.
[0010] Furthermore, the mechanism decoupling and correction module includes the following steps: Based on the mechanism and transmission path of factors affecting engineering cost, the multi-source data is divided into the economic driving mechanism layer, the construction behavior mechanism layer, and the environmental disturbance mechanism layer, thus completing the mechanism decoupling and hierarchical mapping of the multi-source data. An inversion model for interference factors is constructed. Through this model, the contribution of data from different mechanism layers to the cost prediction results is obtained, and the implicit influence weights of non-target disturbance variables within each mechanism layer are inferred.
[0011] Furthermore, the mechanism decoupling and correction module also includes the following steps: Based on the deviation contribution of each mechanism layer, the implicit influence weight of non-target perturbation variables in each mechanism layer is determined by combining the analytic hierarchy process, expert scoring method and coefficient of variation method for weighting. The initial parameters of the prediction model are dynamically adjusted based on the implicit influence weights and the actual perturbation errors of each perturbation variable. The preprocessed multi-source hierarchical data is input into the prediction model after correction parameters to generate corrected feature data after eliminating coupling interference.
[0012] Furthermore, the model training and optimization module includes the following steps: The constructed multi-source cost risk dynamic prediction model consists of a feature extraction layer, a fusion layer, a prediction layer, and a credibility assessment layer. The model training and optimization module trains the model in stages, and uses early stopping and L2 regularization to prevent overfitting. The test set is used to evaluate the performance of the trained model. The model is considered to have passed training when the coefficient of determination R² is greater than or equal to 0.85.
[0013] Furthermore, the feature extraction layer uses a deep neural network (DNN), the fusion layer uses an attention mechanism, the prediction layer uses a fully connected network, and the credibility evaluation layer is built based on a fuzzy comprehensive evaluation model. The model training and optimization module establishes a credibility self-learning feedback mechanism based on long short-term memory networks, through which the weights of each data source are automatically and dynamically adjusted.
[0014] Furthermore, the credibility feedback and prediction module includes the following steps: A reliability assessment model for prediction output is constructed based on fuzzy comprehensive evaluation, with the completeness, timeliness, and accuracy of the corrected feature data and the feature indicators of each data source as evaluation factors. The objective weights of each evaluation factor are determined using the entropy weight method, and the overall reliability of the prediction results is calculated. The credibility feedback and prediction module combines the deviation range between the prediction results and the actual cost, as well as the risk coefficient of the environmental disturbance mechanism layer, to classify the risk level of the engineering cost prediction into three levels: low risk, medium risk, and high risk.
[0015] Furthermore, the model closed-loop optimization module includes the following steps: Establish a historical database for engineering cost forecasting, and perform structured storage and classification management of all forecast information; The actual deviation value of the prediction is used as the loss function, and the correction parameters of the prediction model are optimized by gradient descent using the stochastic gradient descent method. The initial weights of each data source are updated based on the credibility self-learning feedback mechanism. The optimized model parameters and the updated data source weights are then applied to the next engineering cost prediction, forming a closed-loop optimization system. The system can also iteratively correct existing prediction results in real time by collecting new multi-source data, and trigger the full-process parameter and weight update and incremental training of the model.
[0016] A method for dynamic prediction of engineering costs based on multi-source data is applied to the aforementioned dynamic prediction system for engineering costs based on multi-source data. The method includes: S1. Obtain the multi-source heterogeneous raw data required for engineering cost prediction, and perform data cleaning, missing value completion, outlier removal and normalization on the data in sequence to complete the data preprocessing. Then, divide the preprocessed standardized multi-source data into training set, validation set and test set. S2. The preprocessed multi-source data is decoupled and mapped hierarchically according to the mechanism and transmission path of the factors affecting engineering cost. An interference factor inversion model is constructed to inversely deduce the implicit influence weight of non-target disturbance variables in each mechanism layer. The initial parameters of the prediction model are corrected according to the implicit influence weight to generate corrected feature data after eliminating coupling interference. S3. Construct a multi-source cost risk dynamic prediction model consisting of a feature extraction layer, a fusion layer, a prediction layer, and a credibility assessment layer. Train the model in stages, use early stopping and regularization to prevent overfitting, evaluate the model performance through a test set, and establish a credibility self-learning feedback mechanism to dynamically adjust the data source weights. S4. Input the corrected feature data into the trained model with dynamically adjusted weights. Calculate the overall credibility of the prediction results through fuzzy comprehensive evaluation. Combine the deviation range between the prediction results and the actual cost, and the risk coefficient of the environmental disturbance mechanism layer, determine the risk level of the engineering cost prediction, and output the engineering cost prediction results and the corresponding prediction credibility and risk level. S5. Based on the prediction results and the credibility assessment results, establish a prediction history database, use the gradient descent method to optimize the model correction parameters, update the data source weights, and apply the optimized parameters and weights to the next prediction to form a closed-loop optimization system; at the same time, collect new multi-source data in real time, iteratively correct the original prediction results, trigger the full-process update and incremental training of model parameters and weights, and realize real-time dynamic prediction of the entire life cycle of engineering cost.
[0017] The beneficial effects of this invention are as follows: 1. The system integrates multi-source data from engineering databases, industry monitoring platforms, IoT terminals, and policy channels, covering three dimensions: economic, construction, and environmental. It employs K-nearest neighbor interpolation, mean interpolation, and other methods. Outlier removal and Z-score normalization are among the data processing techniques used to standardize data and provide high-quality input, laying a solid foundation for subsequent predictions.
[0018] 2. Based on the mechanism of engineering cost impact, a three-layer mechanism mapping is constructed, consisting of an economic driving layer, a construction behavior layer, and an environmental disturbance layer. An interference factor inversion and weight correction model is introduced to effectively identify and weaken the interference of non-target disturbance factors, thereby improving the accuracy and stability of the prediction results.
[0019] 3. Construct a multi-layer prediction structure that integrates DNN, attention mechanism, fully connected network and fuzzy comprehensive evaluation model to achieve high-precision cost prediction and credibility assessment. Introduce a weight self-learning feedback mechanism based on LSTM to enable the model to dynamically adjust.
[0020] 4. By establishing a historical prediction database and combining prediction error feedback with gradient descent optimization algorithms, model parameters and data source weights are continuously adjusted to ensure that the model maintains high prediction capability under different stages and project conditions.
[0021] 5. By combining the prediction error range with the environmental disturbance risk coefficient, the system can intelligently classify the risk level of the project cost prediction results into low, medium and high levels, thereby improving the risk prevention capability of cost control.
[0022] 6. Real-time acquisition of new multi-source data and iterative model training enable dynamic correction of existing prediction results, ensuring the timeliness and foresight of the engineering cost prediction system in practical applications. Attached Figure Description
[0023] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a system block diagram of the present invention; Figure 2 This is a flowchart of the mechanism decoupling and correction module of the present invention; Figure 3 This is a flowchart of the model training and optimization module of the present invention; Figure 4 This is a flowchart of the overall method of the present invention. Detailed Implementation
[0024] The present application will now be described in further detail with reference to the accompanying drawings. It should be noted that the following specific embodiments are only used to further illustrate the present application and should not be construed as limiting the scope of protection of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application based on the above application content. Example
[0025] Please see Figure 1-3 This invention provides a technical solution: a dynamic prediction system for engineering costs based on multi-source data, the system comprising: The module includes: data acquisition and preprocessing, mechanism decoupling and correction, model training and optimization, credibility feedback and prediction, and model closed-loop optimization. The data acquisition and preprocessing module is used to acquire multi-source heterogeneous raw data required for engineering cost prediction and to preprocess the multi-source heterogeneous raw data. Among them, multi-source heterogeneous raw data refers to data from a wide range of sources, possibly from different channels, systems, or departments. For example, in engineering cost prediction, data may come from multiple sources such as project planning documents, construction records, market price databases, and policy and regulatory documents. Heterogeneous means that the structure and type of data are different; these data may be structured, semi-structured, or unstructured. Raw data is unprocessed initial data that may contain noise, errors, missing values, etc., and needs to be preprocessed before it can be used for subsequent analysis and modeling. Data preprocessing involves cleaning, transforming, and integrating the acquired multi-source heterogeneous raw data to improve data quality and make it suitable for subsequent analysis and modeling. Common data preprocessing operations include removing duplicate data, handling missing values, data standardization or normalization, and data encoding conversion. The mechanism decoupling and correction module is used to construct a dynamic prediction framework for engineering cost based on heterogeneous data mechanism decoupling and interference factor inversion correction. It performs mechanism decoupling and interference factor inversion correction on preprocessed multi-source data to generate corrected feature data. In the heterogeneous data mechanism decoupling-interference factor inversion correction, the heterogeneous data mechanism decoupling is necessary because there may be complex correlations and mutual influences between multi-source heterogeneous data. The purpose of mechanism decoupling is to decompose and analyze these complex relationships to find the independent mechanisms and inherent laws of the impact of each data source on engineering cost. For example, it analyzes how factors such as material prices, construction technology, and policies and regulations affect engineering cost, and whether there are interactions between them. Interference factor inversion correction is necessary because there are factors in engineering cost prediction that may interfere with the accuracy of prediction. Inversion correction refers to the identification of these interference factors from the preprocessed data through certain methods and algorithms, and their correction or elimination to improve the accuracy and reliability of the data. For example, market price fluctuations may be affected by multiple factors. Inversion correction can remove the interference of some irrelevant factors and more accurately reflect the impact of material prices on engineering cost. The corrected feature data is the data obtained after mechanism decoupling and interference factor inversion correction. This data more accurately reflects the key features and factors related to engineering cost and can provide more valuable information for subsequent model training. The model training and optimization module is used to build a multi-source cost risk dynamic prediction model based on prediction credibility self-learning feedback, and to complete the training and iterative optimization of the model. In the self-learning feedback of prediction credibility, prediction credibility is a measure of the reliability of engineering cost prediction results. It can be evaluated in various ways, such as verification based on historical data and uncertainty analysis of the model. The higher the prediction credibility, the more reliable the prediction result; conversely, the lower the prediction credibility, the greater the potential error in the prediction result. Self-learning feedback means that the model can automatically adjust and optimize its parameters and structure based on the credibility assessment of the prediction results to improve the accuracy and reliability of the prediction. Through continuous learning and feedback, the model can gradually adapt to different data and scenarios, improving its generalization ability. The multi-source cost risk dynamic prediction model is a model specifically designed to predict engineering costs and their risks. It can comprehensively consider information from multiple sources of data, capture the dynamic characteristics of engineering costs changing over time and with other factors, and assess and warn of potential risks. Model training and iterative optimization involves training the multi-source cost risk dynamic prediction model using corrected feature data. By adjusting the model's parameters, the model's prediction results are made as close as possible to the actual values. Iterative optimization refers to continuously repeating this process during training, improving and adjusting the model based on its performance to enhance its prediction accuracy and generalization ability. The credibility feedback and prediction module is used to input the corrected feature data into the trained model to complete the dynamic prediction and risk assessment of the project cost, and output the project cost prediction results and the corresponding prediction credibility and risk level. In the dynamic forecasting and risk assessment of project costs, dynamic forecasting takes into account that project costs will change with time, market changes, project progress, and other factors. Dynamic forecasting can predict project costs in real time or periodically, reflecting their latest trends. Risk assessment identifies, analyzes, and evaluates potential risks during the project cost forecasting process, determines the likelihood and impact of risks, and takes corresponding measures to address and manage them based on the risk level. The project cost forecasting results, along with the corresponding forecasting credibility and risk level, are the outputs of this module. The forecasting results provide specific numerical values or ranges for project costs; forecasting credibility reflects the reliability of the forecasting results; and the risk level indicates the degree of risk that the project cost may face, typically categorized as low, medium, and high, providing a reference for decision-makers. The model closed-loop optimization module is used to dynamically update the prediction model parameters and data source weights based on the prediction results and the credibility assessment results, so as to realize the closed-loop optimization of engineering cost prediction. Among these components, the prediction results and the credibility assessment results are as follows: the prediction results are the model's predicted values or ranges for project costs, while the credibility assessment results are a measure of the reliability of the prediction results. These two results together reflect the model's performance and prediction quality. Dynamically updating the prediction model parameters and data source weights involves adjusting and optimizing the model parameters based on the prediction results and credibility assessment results, enabling the model to better adapt to new data and situations and improve prediction accuracy. Dynamically updating the data source weights is crucial because different data sources may have varying importance and influence on project cost prediction. By dynamically updating the data source weights, the contribution of each data source in the model can be adjusted according to actual conditions, allowing the model to utilize multi-source data more rationally. The closed-loop optimization of project cost prediction involves continuously updating the model parameters and data source weights dynamically based on the prediction results and credibility assessment results, forming a closed-loop optimization process. This continuously improves the model's performance, making the prediction results more accurate and reliable, thereby achieving continuous optimization and improvement of project cost prediction.
[0026] It should be noted that during use, the data acquisition and preprocessing module can collect and process a wide range of heterogeneous raw data from multiple sources, ensuring data integrity and accuracy, laying the foundation for subsequent analysis. The mechanism decoupling and correction module processes the data by constructing a prediction framework, eliminating interference factors, generating more accurate correction feature data, and improving prediction accuracy. The model training and optimization module builds and optimizes the model based on self-learning feedback, enabling the model to continuously adapt to new data and improve generalization ability. The credibility feedback and prediction module not only provides prediction results but also credibility and risk levels, providing a comprehensive reference for decision-making. The model closed-loop optimization module dynamically updates parameters and weights based on the results, forming a closed loop, continuously optimizing prediction performance, allowing the system to keep up with changes in actual conditions, and providing scientific, dynamic, accurate, and comprehensive support for engineering cost prediction.
[0027] In one embodiment, multi-source heterogeneous raw data required for engineering cost prediction is acquired, and the multi-source heterogeneous raw data is preprocessed, including: By utilizing engineering databases, industry monitoring platforms, IoT sensing terminals, and policy dissemination channels, we acquire multi-source heterogeneous raw data related to engineering costs. This multi-source heterogeneous raw data covers three major categories: economic data, construction data, and environmental data. The economic data includes building material price indices, construction industry labor wage indices, and benchmark lending rates in the financial market. The construction data includes actual project construction progress rates, construction resource input density, and construction machinery utilization rates. The environmental data includes climate monitoring data for the project construction area, industry policy fluctuation documents, and building material supply chain disruption risk assessment data. For multi-source heterogeneous raw data, data cleaning, missing value imputation, outlier removal, and normalization are performed sequentially. Data cleaning removes redundant and duplicate data; missing value imputation uses K-nearest neighbor interpolation to imput missing time-series data and mean interpolation to imput missing numerical data; outlier removal is based on… The principle is to identify and remove outlier samples that deviate from the data distribution. Normalization is performed using the Z-score standardization method, and the normalization expression is:
[0028] in, For the first The first class of data sources The normalized values of the data. For the first The first class of data sources One set of raw data, For the first The mean of all samples in the data source class. For the first The standard deviation of all samples in the data source class For the first The number of valid samples from the data source class; The preprocessed standardized multi-source data is divided into training set, validation set and test set in a ratio of 7:2:1. The training set is used for basic training of the model, the validation set is used for hyperparameter tuning and overfitting monitoring, and the test set is used for final performance evaluation of the model.
[0029] This design illustrates the sources of the multi-source heterogeneous raw data, covering three major categories: economic, construction, and environmental data. The acquired data undergoes preprocessing such as cleaning, completion, outlier removal, and normalization. Furthermore, it is divided into training, validation, and test sets. This multi-source data comprehensively covers factors influencing project costs, ensuring data integrity. Preprocessing improves data quality, eliminates noise and errors, and makes the data more standardized and consistent, facilitating subsequent analysis. The reasonable division of the dataset—using the training set for model learning, the validation set for parameter tuning, and the test set for performance evaluation—allows the model to be fully trained and accurately evaluated at different stages, improving its generalization ability and laying the foundation for accurate project cost prediction.
[0030] In one embodiment, a dynamic prediction framework for engineering costs is constructed based on heterogeneous data mechanism decoupling and interference factor inversion correction. Mechanism decoupling and interference factor inversion correction are performed on preprocessed multi-source data to generate corrected feature data, including: Based on the mechanisms and transmission paths of factors influencing engineering cost, these factors are divided into three layers: the economic driving mechanism layer, the construction behavior mechanism layer, and the environmental disturbance mechanism layer. The economic driving mechanism layer corresponds to economic data and is the core driving layer of engineering cost, reflecting the fundamental impact of macroeconomic changes on cost. The construction behavior mechanism layer corresponds to construction data and is the direct impact layer of engineering cost, reflecting the real-time impact of actual operational behaviors on cost during project construction. The environmental disturbance mechanism layer corresponds to environmental data and is the external disturbance layer of engineering cost, reflecting the random impact of uncontrollable external factors on cost. This completes the decoupling and hierarchical mapping of multi-source data. A multi-source data interference factor inversion model based on gradient boosting regression was constructed and trained. The training parameters were set as follows: learning rate 0.01, number of decision tree base learners 300, maximum depth of decision tree 8, minimum number of leaf node samples 10, subsampling ratio 0.8, feature sampling ratio 0.9, and mean squared error (MSE) as the loss function. The model was trained iteratively on the training set until it converged. Overfitting was monitored on the validation set. Training was stopped when the validation set loss increased for 10 consecutive iterations. Hierarchical multi-source data from the training and validation sets are input into the trained interference factor inversion model. Data from individual mechanistic layers are sequentially removed using the controlled variable method. The deviation between the predicted results after removal and the baseline predicted results is calculated, thus obtaining the contribution of different mechanistic layers of data to the cost prediction results. The calculation formula is:
[0031] in, These represent the economic driving mechanism layer, the construction behavior mechanism layer, and the environmental disturbance mechanism layer, respectively. This represents the deviation between the predicted results and the baseline predicted results after removing data from the first-order mechanism layer. The baseline predicted results are the model predictions when all layered data are input. The deviation contribution is... The range of values is The larger the value, the greater the impact of the corresponding mechanism layer on the cost prediction deviation; Based on the deviation contribution of each mechanism layer The weights of the criterion layer are determined by combining the analytic hierarchy process (AHP). Then, by combining expert scoring and the coefficient of variation method to assign weights, the judgment matrix of the indicator layer is determined, and the implicit influence weights of non-target disturbance variables in each mechanism layer are inferred. In the combined weighting, the expert scoring method accounts for 0.4% of the weight, and the coefficient of variation method accounts for 0.6%. For the first The perturbation variable number in the mechanism layer implies the influence weight. Satisfying nonnegativity and normalization constraints, i.e. , For the first The number of perturbation variables in the mechanism layer The larger the value, the stronger the implicit interference of the corresponding non-target disturbance variable on cost prediction; Based on the implicit influence weights obtained from the inversion Combining the actual disturbance errors of each disturbance variable The initial parameters of the prediction model Dynamic correction is performed to obtain the corrected model parameters:
[0032] in, For the first In the mechanism layer, the first The perturbation error between the actual observed value and the theoretical predicted value of each perturbation variable is used to input the preprocessed multi-source hierarchical data into the prediction model after correction parameters. Through feature extraction and fusion, corrected feature data after eliminating coupling interference is generated.
[0033] This design stratifies the factors influencing engineering costs according to their mechanisms of action, completing the decoupling and hierarchical mapping of multi-source data. Next, an interference factor inversion model is constructed to calculate the contribution of deviations at different mechanism levels, determine the weights of hidden influences, and finally correct the model parameters to generate corrected feature data. The hierarchical mapping clearly presents the levels of influence of each factor on costs, facilitating targeted analysis. The inversion model accurately identifies interference factors, and by calculating the contribution of deviations and the weights of hidden influences, it delves into the intrinsic relationships within the data. Correcting the model parameters eliminates coupling interference, making the generated feature data more accurate and effectively improving the accuracy of subsequent engineering cost predictions.
[0034] In one embodiment, a multi-source cost risk dynamic prediction model is constructed based on prediction confidence self-learning feedback. After training and iterative optimization of the model, the corrected feature data is input into the trained model to complete the dynamic prediction and risk assessment of the project cost, and output the project cost prediction result and the corresponding prediction confidence and risk level, including: A multi-source cost risk dynamic prediction model is constructed, consisting of a feature extraction layer, a fusion layer, a prediction layer, and a credibility assessment layer. The feature extraction layer adopts a deep neural network (DNN), the fusion layer adopts an attention mechanism, the prediction layer adopts a fully connected network, and the credibility assessment layer is built based on a fuzzy comprehensive evaluation model. The multi-source cost risk dynamic prediction model is trained in stages, and the steps are as follows: The corrected feature data of the training set is input into the feature extraction layer. The DNN layer parameters are set as follows: the number of neurons in the input layer is consistent with the dimension of the corrected feature, the hidden layer is set to 3 layers, and the number of neurons is 512, 256 and 128 respectively. The activation function is ReLU, the dropout rate is set to 0.2, and the root mean square error (RMSE) is used as the loss function. The feature extraction layer is trained until convergence. The output of the feature extraction layer is input into the fusion layer, and the fusion weights of each feature are learned using an attention mechanism. The number of training iterations of the fusion layer is set to 200 rounds, the optimizer is Adam, and the learning rate is 0.001. The fusion features of the fusion layer are input into the prediction layer. The number of neurons in the output layer of the fully connected network is 1, which corresponds to the predicted value of the engineering cost. The loss of the prediction layer is stabilized after training. The prediction layer output and feature indicators from each data source are input into the credibility evaluation layer to complete the training of the credibility evaluation layer; Early stopping is used during model training to prevent overfitting. The early stopping parameters are set as follows: the number of validation set loss tolerance rounds is 15, and the model weights are saved as the weights when the validation set loss is minimized. At the same time, L2 regularization is used to suppress overfitting, and the regularization coefficient is set to 0.0005. The trained model is evaluated using a test set. Evaluation metrics include mean absolute error (MAE), mean squared error (MSE), and coefficient of determination. ,when If the model is deemed to have passed training, then... Then readjust the model hyperparameters and train again; A reliability assessment model for prediction output based on fuzzy comprehensive evaluation is constructed. The completeness, timeliness, and accuracy of the corrected feature data and various data sources are used as evaluation factors. The entropy weight method is employed to determine the objective weights of each evaluation factor based on its information entropy value; the smaller the information entropy value, the higher the weight of the corresponding evaluation factor. Weight coefficients for each evaluation factor are set, and the contribution of various data sources to the reliability of the prediction results is quantified using fuzzy membership functions. Reliability contribution Satisfying the normalization constraint , Given the total number of data source categories, the overall reliability of the prediction results is calculated based on the reliability contribution of each data source. The calculation formula is:
[0035] in, For the first The relative error rate of the data source. , To use only the first Predicted values obtained from similar data sources The actual statistical value of the project cost, overall reliability The range of values is The closer the value is to 1, the higher the reliability of the prediction result; A self-learning feedback mechanism for reliability based on a Long Short-Term Memory (LSTM) network was established. The LSTM network was trained with the following parameters: hidden layer dimension 128, time step 10, batch size 32, and 300 iterations. The AdamW optimizer was used with a learning rate of 0.0005, and the mean squared error of historical prediction bias was used as the loss function. Training continued until the network converged. Historical prediction bias values throughout the project's lifecycle, historical weights of each data source, and prediction reliability results were used as training samples. This enabled the model to automatically and dynamically adjust the weights of each data source based on their prediction performance at different stages of the project cost, including early-stage estimation, mid-stage accounting, and late-stage settlement. The weight update formula is:
[0036] in, For the first The first prediction The initial weights of the data sources are preset based on the industry importance of the data sources. The initial weights for economic, construction, and environmental data sources are set to 0.45, 0.4, and 0.15 respectively. For the first The first prediction Local trustworthiness corresponding to the data source class For the first The local confidence mean of all data sources is used for the next prediction. When the local confidence of a certain data source is higher than the mean, its weight is positively increased, and vice versa. After the weight is updated, it needs to be normalized to ensure that the sum of the weights of all data sources is 1. The corrected feature data is input into the multi-source cost risk dynamic prediction model with dynamically adjusted weights. The model then calculates and outputs the numerical prediction result of the project cost through inference. Simultaneously, considering the deviation range between the predicted results and the actual cost, as well as the risk coefficient of the environmental disturbance mechanism layer, the risk level of engineering cost prediction is divided into three levels: low risk, medium risk, and high risk. The risk coefficient is based on the implicit influence weight of each disturbance variable in the environmental disturbance mechanism layer. The weighted summation is obtained, and the calculation formula is:
[0037] in, This is the risk quantification value for the environmental disturbance variable, with a value of 0. 1; Deviation range Internal risk coefficient below 0.2 is considered low risk; Deviation range ~ And the risk coefficient is Medium risk; Deviation range exceeds Or the risk factor is higher than High risk; Determine the prediction confidence threshold based on the project type and prediction stage. The thresholds for municipal engineering, building construction, and road and bridge engineering are set at 0.75, 0.8, and 0.85 respectively; the thresholds for the preliminary cost estimation, mid-term accounting, and final settlement stages are set at 0.7, 0.8, and 0.9 respectively. When the overall reliability of the prediction results... When the system is in a critical state, it automatically triggers a fault tolerance mechanism, introducing alternative data sources through a preset alternative data source library, or switching to a pre-trained backup prediction model structure to achieve prediction fault tolerance and robust optimization.
[0038] This design constructs a multi-source dynamic cost risk prediction model, trains each layer in stages, employs multiple methods to prevent overfitting, evaluates model performance, constructs a credibility assessment model, establishes a self-learning feedback mechanism, classifies risk levels and sets up a fault tolerance mechanism. Staged training allows each layer of the model to focus on learning specific tasks, improving overall performance. Multiple overfitting prevention methods ensure the model's generalization ability. Credibility assessment and self-learning feedback mechanisms allow the model to dynamically adjust according to actual conditions, improving prediction reliability. Risk level classification and fault tolerance mechanisms provide more comprehensive information for decision-making, enhance the system's ability to cope with abnormal situations, and ensure the robustness of engineering cost prediction.
[0039] In one embodiment, based on the prediction results and the credibility assessment results, the prediction model parameters and data source weights are dynamically updated to achieve closed-loop optimization of engineering cost prediction, including: Establish a historical database for engineering cost prediction, and record in real time the predicted engineering cost values, overall credibility, weights and local credibility of each data source, prediction deviation, risk level and model training parameters, etc. Complete the structured storage and classification management of historical data, and provide data support for model iteration and optimization. Using the actual deviation from the prediction as the loss function, the stochastic gradient descent method is used to correct the parameters of the prediction model. Gradient descent optimization is performed, and the optimization formula is:
[0040] Among them, learning rate The model's convergence speed is dynamically adjusted, with an initial value set to 0.005. When the model's loss function decreases by less than [a certain value] for five consecutive iterations... When the learning rate decreases, the learning rate is halved; when the rate of decline is higher than 1%, the learning rate increases. The gradient descent iteration count is set to 50 rounds, and the batch size is set to 16. The gradient of the loss function. For absolute value loss function, For the first The predicted project cost results For the first The actual value of the project cost corresponding to the next prediction; Based on the output of the credibility self-learning feedback mechanism, the initial weights of each data source are iteratively updated, while the alternative data source library and the backup prediction model library are dynamically maintained and updated. The optimized model parameters and the updated data source weights are applied to the next engineering cost prediction process, forming a closed-loop optimization system of prediction-evaluation-optimization-re-prediction. After each model parameter update, the performance of the updated model is retested using a test set. If the model performance deteriorates after the retest compared to before the update, the updated parameters are discarded and the original model parameters are used instead, ensuring the effectiveness of model optimization and continuously improving the model's prediction accuracy.
[0041] This design establishes a historical database to record prediction-related information, uses stochastic gradient descent to optimize model parameters, iteratively updates data source weights, maintains a candidate database, forms a closed-loop optimization system, and retests model performance. The historical database provides rich data support for model optimization, facilitating the analysis of model performance and problems. Stochastic gradient descent can quickly and effectively optimize model parameters and improve prediction accuracy. Iteratively updating data source weights allows the model to adapt to data changes. The closed-loop optimization system allows the model to continuously improve, and retesting model performance ensures the effectiveness of optimization, avoids ineffective updates, and ensures that the model is always in an optimal state, thereby improving the accuracy of engineering cost prediction.
[0042] In one embodiment, the method further includes dynamic updating of the project cost prediction results. New, multi-source data on project costs is collected in real time via IoT sensor terminals and the project management system. After preprocessing the new data, the steps of mechanism decoupling, interference correction, confidence feedback, and model optimization are repeated to iteratively correct the original prediction results in real time. A data update threshold is set, and when the amount of newly added valid data accounts for a certain percentage of the current dataset... And above, or the amount of new data from a single data source accounts for a significant portion of the total amount of data in that type of dataset. When the time reaches 100 or above, a full-process model parameter and weight update is triggered, and incremental training is performed on the model. The incremental training parameters are set as follows: the learning rate is 1 / 5 of the original model training learning rate, the number of iterations is 50 rounds, and the batch size is the same as the original model, so as to realize real-time dynamic prediction of the entire life cycle of engineering cost.
[0043] This design, by collecting new data in real time and executing a series of processing steps to iteratively correct the original prediction results, sets a data update threshold to trigger model parameter and weight updates and incremental training. Real-time collection of new data can reflect the latest changes in engineering costs in a timely manner, making the prediction results closer to the actual situation. Setting an update threshold avoids frequent updates that waste computing resources, while ensuring that the model is updated in a timely manner when there are sufficient changes in the data. Incremental training quickly adapts to new data based on the original model, improving training efficiency and realizing real-time dynamic prediction of engineering costs throughout the entire life cycle, providing timely and accurate information for engineering decisions. Example
[0044] Please see Figure 4 A method for dynamic prediction of engineering costs based on multi-source data is proposed, applied to the aforementioned dynamic prediction system for engineering costs based on multi-source data. The method includes: S1. Obtain the multi-source heterogeneous raw data required for engineering cost prediction, and perform data cleaning, missing value completion, outlier removal and normalization on the data in sequence to complete the data preprocessing. Then, divide the preprocessed standardized multi-source data into training set, validation set and test set. S2. The preprocessed multi-source data is decoupled and mapped hierarchically according to the mechanism and transmission path of the factors affecting engineering cost. An interference factor inversion model is constructed to inversely deduce the implicit influence weight of non-target disturbance variables in each mechanism layer. The initial parameters of the prediction model are corrected according to the implicit influence weight to generate corrected feature data after eliminating coupling interference. S3. Construct a multi-source cost risk dynamic prediction model consisting of a feature extraction layer, a fusion layer, a prediction layer, and a credibility assessment layer. Train the model in stages, use early stopping and regularization to prevent overfitting, evaluate the model performance through a test set, and establish a credibility self-learning feedback mechanism to dynamically adjust the data source weights. S4. Input the corrected feature data into the trained model with dynamically adjusted weights. Calculate the overall credibility of the prediction results through fuzzy comprehensive evaluation. Combine the deviation range between the prediction results and the actual cost, and the risk coefficient of the environmental disturbance mechanism layer, determine the risk level of the engineering cost prediction, and output the engineering cost prediction results and the corresponding prediction credibility and risk level. S5. Based on the prediction results and the credibility assessment results, establish a prediction history database, use the gradient descent method to optimize the model correction parameters, update the data source weights, and apply the optimized parameters and weights to the next prediction to form a closed-loop optimization system; at the same time, collect new multi-source data in real time, iteratively correct the original prediction results, trigger the full-process update and incremental training of model parameters and weights, and realize real-time dynamic prediction of the entire life cycle of engineering cost.
[0045] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0046] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A dynamic cost prediction system for engineering projects based on multi-source data, characterized in that, The system includes: The module includes: data acquisition and preprocessing, mechanism decoupling and correction, model training and optimization, credibility feedback and prediction, and model closed-loop optimization. The data acquisition and preprocessing module is used to acquire multi-source heterogeneous raw data required for engineering cost prediction and to preprocess the multi-source heterogeneous raw data. The mechanism decoupling and correction module is used to perform mechanism decoupling and interference factor inversion correction on the preprocessed multi-source data to generate corrected feature data. The model training and optimization module is used to build a multi-source cost risk dynamic prediction model and complete the training and iterative optimization of the model. The credibility feedback and prediction module is used to input the corrected feature data into the trained model to realize dynamic prediction and risk assessment of engineering costs, and output the prediction results and the corresponding prediction credibility and risk level. The model closed-loop optimization module is used to dynamically update model parameters and data source weights based on prediction results and credibility assessment results, thereby achieving closed-loop optimization of engineering cost prediction.
2. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 1, characterized in that, The data acquisition and preprocessing module includes the following steps: Multi-source heterogeneous raw data is obtained through engineering databases, industry monitoring platforms, IoT sensing terminals, and policy release channels. The multi-source heterogeneous raw data includes three major categories: economic data, construction data, and environmental data. The data acquisition and preprocessing module sequentially performs data cleaning, missing value completion, outlier removal, and normalization on the multi-source heterogeneous raw data, and divides the preprocessed standardized multi-source data into training set, validation set, and test set according to a preset ratio.
3. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 2, characterized in that, The data acquisition and preprocessing module further includes the following steps: K-nearest neighbor interpolation is used to complete missing time series data, and mean interpolation is used to complete missing numerical data. Outliers are removed in principle, and normalization is performed using the Z-score standardization method. The preprocessed standardized multi-source data is divided into training set, validation set and test set in a ratio of 7:2:
1.
4. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 1, characterized in that, The mechanism decoupling and correction module includes the following steps: Based on the mechanism and transmission path of factors affecting engineering cost, the multi-source data is divided into the economic driving mechanism layer, the construction behavior mechanism layer, and the environmental disturbance mechanism layer, thus completing the mechanism decoupling and hierarchical mapping of the multi-source data. An inversion model for interference factors is constructed. Through this model, the contribution of data from different mechanism layers to the cost prediction results is obtained, and the implicit influence weights of non-target disturbance variables within each mechanism layer are inferred.
5. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 4, characterized in that, The mechanical decoupling and correction module further includes the following steps: Based on the deviation contribution of each mechanism layer, the implicit influence weight of non-target perturbation variables in each mechanism layer is determined by combining the analytic hierarchy process, expert scoring method and coefficient of variation method for weighting. The initial parameters of the prediction model are dynamically adjusted based on the implicit influence weights and the actual perturbation errors of each perturbation variable. The preprocessed multi-source hierarchical data is input into the prediction model after correction parameters to generate corrected feature data after eliminating coupling interference.
6. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 1, characterized in that, The model training and optimization module includes the following steps: The constructed multi-source cost risk dynamic prediction model consists of a feature extraction layer, a fusion layer, a prediction layer, and a credibility assessment layer. The model training and optimization module trains the model in stages, and uses early stopping and L2 regularization to prevent overfitting. The test set is used to evaluate the performance of the trained model. The model is considered to have passed training when the coefficient of determination R² is greater than or equal to 0.
85.
7. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 6, characterized in that: The feature extraction layer uses a deep neural network (DNN), the fusion layer uses an attention mechanism, the prediction layer uses a fully connected network, and the credibility evaluation layer is built based on a fuzzy comprehensive evaluation model. The model training and optimization module establishes a credibility self-learning feedback mechanism based on long short-term memory networks, through which the weights of each data source are automatically and dynamically adjusted.
8. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 1, characterized in that, The credibility feedback and prediction module includes the following steps: A reliability assessment model for prediction output is constructed based on fuzzy comprehensive evaluation, with the completeness, timeliness, and accuracy of the corrected feature data and the feature indicators of each data source as evaluation factors. The objective weights of each evaluation factor are determined using the entropy weight method, and the overall reliability of the prediction results is calculated. The credibility feedback and prediction module combines the deviation range between the prediction results and the actual cost, as well as the risk coefficient of the environmental disturbance mechanism layer, to classify the risk level of the engineering cost prediction into three levels: low risk, medium risk, and high risk.
9. The dynamic cost prediction system for engineering projects based on multi-source data according to claim 1, characterized in that, The model closed-loop optimization module includes the following steps: Establish a historical database for engineering cost forecasting, and perform structured storage and classification management of all forecast information; The actual deviation value of the prediction is used as the loss function, and the correction parameters of the prediction model are optimized by gradient descent using the stochastic gradient descent method. The initial weights of each data source are updated based on the credibility self-learning feedback mechanism. The optimized model parameters and the updated data source weights are then applied to the next engineering cost prediction, forming a closed-loop optimization system. The system can also iteratively correct existing prediction results in real time by collecting new multi-source data, and trigger the full-process parameter and weight update and incremental training of the model.
10. A method for dynamic prediction of engineering costs based on multi-source data, characterized in that, The method, applied to the engineering cost dynamic prediction system based on multi-source data as described in any one of claims 1-9, comprises: S1. Obtain the multi-source heterogeneous raw data required for engineering cost prediction, and perform data cleaning, missing value completion, outlier removal and normalization on the data in sequence to complete the data preprocessing. Then, divide the preprocessed standardized multi-source data into training set, validation set and test set. S2. The preprocessed multi-source data is decoupled and mapped hierarchically according to the mechanism and transmission path of the factors affecting engineering cost. An interference factor inversion model is constructed to inversely deduce the implicit influence weight of non-target disturbance variables in each mechanism layer. The initial parameters of the prediction model are corrected according to the implicit influence weight to generate corrected feature data after eliminating coupling interference. S3. Construct a multi-source cost risk dynamic prediction model consisting of a feature extraction layer, a fusion layer, a prediction layer, and a credibility assessment layer. Train the model in stages, use early stopping and regularization to prevent overfitting, evaluate the model performance through a test set, and establish a credibility self-learning feedback mechanism to dynamically adjust the data source weights. S4. Input the corrected feature data into the trained model with dynamically adjusted weights. Calculate the overall credibility of the prediction results through fuzzy comprehensive evaluation. Combine the deviation range between the prediction results and the actual cost, and the risk coefficient of the environmental disturbance mechanism layer, determine the risk level of the engineering cost prediction, and output the engineering cost prediction results and the corresponding prediction credibility and risk level. S5. Based on the prediction results and the credibility assessment results, establish a prediction history database, use the gradient descent method to optimize the model correction parameters, update the data source weights, and apply the optimized parameters and weights to the next prediction to form a closed-loop optimization system; at the same time, collect new multi-source data in real time, iteratively correct the original prediction results, trigger the full-process update and incremental training of model parameters and weights, and realize real-time dynamic prediction of the entire life cycle of engineering cost.