Project cost automatic review method based on big data, medium and equipment
The method leverages big data and advanced models to automate engineering cost estimation and review processes, addressing inefficiencies and improving accuracy and reliability in complex project management.
Patent Information
- Application Number
- CN202510361165.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
There are low efficiency, poor dynamic adaptability in the process of traditional engineering cost evaluation, limited data collection and single model, which leads to low review efficiency and incomplete risk coverage, making it difficult to meet the dynamic management needs of complex engineering projects.
The automatic evaluation method of engineering cost based on big data is generated by obtaining project information to generate dynamically adjusted data acquisition strategies, heterogeneous cost data is collected and normalized. The cost prediction model, XGBoost risk prediction model and Monte Carlo simulated expected return model are used to generate an evaluation report.
It realizes dynamic optimization and multi-dimensional analysis of data acquisition, improves the automation level of engineering cost evaluation and the reliability of results, reduces manual intervention errors, and is suitable for rapid decision-making of large-scale or high-complex projects.
Smart Images

Figure CN120317809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to an automatic project cost review method, medium, and device based on big data. Background Art
[0002] In the traditional project cost review process, data collection and analysis usually rely on manual experience, resulting in problems such as low efficiency and poor dynamic adaptability. Specifically, it is manifested in the limited data collection of the cost plan, the long cycle of making the project plan, the need to separately organize materials during the review, and the complex process. At the same time, with high confidentiality requirements, the whole process is cumbersome and has a long time span, resulting in some materials not being updated in a timely manner and the management being not standardized enough. That is, there are problems such as poor data adaptability, low review efficiency caused by a single model, and incomplete risk coverage, making it difficult to meet the dynamic management needs of complex engineering projects. Summary of the Invention
[0003] In view of the above problems, the present invention provides an automatic project cost review method based on big data, which solves the problem that the existing project cost process from data collection and integration to the review stage lacks complete and standardized management.
[0004] To achieve the above object, in the first aspect, the present application provides an automatic project cost review method based on big data, including:
[0005] Obtain project information, and generate a data collection strategy according to the project information. The data collection strategy includes collection frequency, collection items, collection methods, and data preprocessing logic. The project information includes project type, design requirements, contract terms, and industry standards. When the data collection strategy is configured to meet the preset conditions, a dynamic adjustment mechanism is triggered and updated;
[0006] Collect according to the data collection strategy to obtain heterogeneous project cost data;
[0007] Perform normalization processing on the heterogeneous project cost data to obtain a first data set;
[0008] Input the first data set into a project cost prediction model to obtain a project cost prediction coefficient. The project cost prediction model is configured to be generated based on the fusion of an ARIMA model and a LightGBM model;
[0009] And input the first data set into a risk prediction model to obtain the divided risk levels and avoidance suggestions. The risk prediction model is configured to be an XGBoost model constructed with audit items as adjustment parameters;
[0010] And input the first data set into an expected return model to obtain an expected return range;
[0011] Generate a review report based on the cost prediction coefficient, risk level, avoidance suggestions, and expected return range.
[0012] In some embodiments, obtaining project information and generating a data collection strategy based on the project information includes:
[0013] Determine the collection frequency and collection items according to the project type. The collection frequency is generated based on the complexity and scale of the project type. The collection items include material prices, labor costs, equipment rental fees, and construction progress data;
[0014] Determine the collection method and data preprocessing logic according to the design requirements. The collection methods include real-time collection and batch collection. The data preprocessing logic includes data cleaning, data deduplication, and data format conversion;
[0015] Determine the trigger conditions for the dynamic adjustment mechanism according to the contract terms. The trigger conditions include contract amount thresholds, project duration change events, and warning events;
[0016] Supplement and correct the collection frequency and collection items according to industry standards. The industry standards include material price fluctuation ranges, labor cost benchmark values, and average equipment rental market prices;
[0017] The collection frequency being generated based on the complexity and scale of the project type includes:
[0018] Define the complexity level, determine the complexity level to which the current project belongs according to the project type, and generate the first configuration weight;
[0019] And determine the scale level to which the current project belongs according to the project type, and generate the second configuration weight;
[0020] Obtain the preset collection frequency, and perform calculations based on the first configuration weight, the second configuration weight, and the preset collection frequency to obtain the initial collection frequency, which is the collection frequency.
[0021] In some embodiments, the preset conditions include a first preset condition and a second preset condition, and the dynamic adjustment mechanism includes a first dynamic adjustment mechanism and a second dynamic adjustment mechanism;
[0022] When the data collection strategy is configured to meet the preset conditions, trigger the dynamic adjustment mechanism and update, including:
[0023] When the data collection strategy meets the first preset condition, trigger the first dynamic adjustment mechanism to update the data collection strategy;
[0024] When the data collection strategy meets the second preset condition, trigger the second dynamic adjustment mechanism to update the data collection strategy.
[0025] In some embodiments, when the data acquisition strategy meets the first preset condition, triggering the first dynamic adjustment mechanism to update the data acquisition strategy includes:
[0026] Collect data in real time and record the collected data as the first raw data;
[0027] Obtain multiple first raw data within the first preset time period up to now and plot the first data fluctuation amplitude;
[0028] Judge whether the first data fluctuation amplitude is within the first preset fluctuation threshold. If not, it means that the data acquisition strategy meets the first preset condition, triggering the first dynamic adjustment mechanism, and the first dynamic adjustment mechanism is configured to increase or decrease the acquisition frequency of the data acquisition strategy;
[0029] When the data acquisition strategy meets the second preset condition, triggering the second dynamic adjustment mechanism to update the data acquisition strategy includes:
[0030] Judging one by one whether the first raw data is within the preset numerical range. If not, it means that the data acquisition strategy meets the second preset condition, triggering the second dynamic adjustment mechanism, and the second dynamic adjustment mechanism is configured as:
[0031] Obtain multiple first raw data within the second preset time period up to now and plot the second data fluctuation amplitude;
[0032] Judge whether the second data fluctuation amplitude is within the second preset fluctuation threshold;
[0033] If so, record the first raw data that is not within the preset numerical range as an abnormal error value, and use the interpolation method to correct the abnormal error value;
[0034] If not, record the first raw data that is not within the preset numerical range as an abnormal value, and generate an abnormal warning message based on the abnormal value and the second data fluctuation amplitude.
[0035] In some embodiments, inputting the first data set into the cost prediction model to obtain the cost prediction coefficient includes:
[0036] Obtain the first data set and perform time series analysis on the first data set to extract time series features, where the time series features include trend features, periodic features, and random fluctuation features;
[0037] Perform a stationarity test on the time series features. If the time series features do not meet the stationarity threshold, perform a differencing operation on the time series features until the time series features meet the stationarity threshold;
[0038] Input the time series features into the trained ARIMA model to generate the first prediction result;
[0039] Fuse the first prediction result with the first data set to obtain a second data set. The feature fusion includes adding the first prediction result as a new feature to the first data set and performing standardization processing;
[0040] Evaluate the feature importance of the second data set, and screen out the features that have a significant impact on the prediction result, denoted as the second data features;
[0041] Input the second data features into the trained LightGBM model to generate a second prediction result;
[0042] Calculate the first prediction error of the first prediction result on the historical data, and calculate the second prediction error of the second prediction result on the historical data;
[0043] Judge the magnitudes of the first prediction error and the second prediction error. If the first prediction error is smaller, increase the weight value corresponding to the first prediction error. If the second prediction error is smaller, increase the weight value corresponding to the second prediction error;
[0044] Add the weighted first prediction result and the second prediction result to generate a cost prediction coefficient.
[0045] In some embodiments, inputting the time series features into the trained ARIMA model to generate the first prediction result includes:
[0046] Generate the first model parameters of the ARIMA model according to the autocorrelation function and the partial autocorrelation function. The first model parameters include the autoregressive order, the differencing order, and the moving average order;
[0047] Train the ARIMA model according to the first model parameters and the first sample data, and predict the time series feature data after training is completed to generate the first prediction result;
[0048] Inputting the second data features into the trained LightGBM model to generate the second prediction result includes:
[0049] Generate the second model parameters based on the gradient boosting decision tree algorithm. The second model parameters include the first learning rate, the first maximum depth of the tree, and the number of first leaf nodes;
[0050] Iteratively train the LightGBM model according to the second model parameters and the second sample data, and predict the second data features after training is completed to generate the second prediction result.
[0051] In some embodiments, inputting the first data set into the risk prediction model to obtain the divided risk levels and avoidance suggestions. The risk prediction model is configured as an XGBoost model constructed with audit items as adjustment parameters, including:
[0052] Generate audit entries based on project information, and extract features from the first data set according to the audit entries to generate audit feature data. The audit entries include contract compliance, cost overrun rate, project duration delay rate, and quality standard compliance;
[0053] Clean the audit feature data to generate a third data set. Data cleaning includes missing value filling, outlier handling, and feature standardization;
[0054] And, label the third sample data according to the audit entries to generate training samples. The labels include low risk, medium risk, and high risk;
[0055] Construct an XGBoost model based on the gradient boosting algorithm, and initialize the third model parameters of the XGBoost model through cross-validation. The third model parameters include the second learning rate, the second maximum depth of the tree, the second number of leaf nodes, and the regularization parameter;
[0056] Use the labeled third sample data to train the XGBoost model, and predict the third data set after training is completed to generate a risk level classification result;
[0057] According to the risk level classification result, combine it with a preset risk avoidance rule library to generate avoidance suggestions.
[0058] In some embodiments, input the first data set into an expected return model, and the obtained expected return range includes:
[0059] Extract feature data related to returns from the first data set and perform data cleaning to obtain a fourth data set. The feature data includes project budget, cost distribution, market volatility, and historical return data;
[0060] Perform random sampling on the fourth data set based on the Monte Carlo simulation method to generate multiple return prediction scenarios;
[0061] Input each return prediction scenario into a random forest model to obtain a corresponding return prediction value. The random forest model is configured to integrate prediction results through multiple decision trees;
[0062] Perform statistical analysis on multiple return prediction values to generate an expected return range. The expected return range includes the lowest return value, the highest return value, and the median expected return.
[0063] In a second aspect, the present invention also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described in the first aspect is implemented.
[0064] In a third aspect, the present invention further provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.
[0065] Different from the prior art, the above technical solution has the following beneficial effects:
[0066] The above technical solution provides an automatic project cost review method, medium and device based on big data. The method includes: generating a dynamically adjustable data collection strategy through project information, and the data collection strategy triggers an update mechanism according to preset conditions; collecting heterogeneous project cost data based on the data collection strategy and performing normalization processing to form a first data set; inputting the first data set into a project cost prediction model, a risk prediction model, and an expected return model respectively, where the project cost prediction model is constructed by fusing ARIMA and LightGBM, and the risk prediction model adopts an XGBoost model with audit items as adjustment parameters; finally, generating a review report by comprehensively outputting the project cost prediction coefficient, risk level, avoidance suggestions, and expected return range. This technical solution realizes the dynamic optimization of data collection and the collaborative decision-making of multi-dimensional analysis models, improving the automation level and result reliability of project cost review.
[0067] The above relevant records of the invention content are only an overview of the technical solution of this application. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of this application, and then can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features, and advantages of this application more easily understood, the following is described in conjunction with the specific embodiments and drawings of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The drawings are only used to illustrate the principles, implementation methods, applications, features, and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation to this application.
[0069] In the drawings of the specification:
[0070] Figure 1 It is a step schematic diagram of steps S101 to S105 of the automatic project cost review method described in the specific embodiment;
[0071] Figure 2 It is a step schematic diagram of steps S201 to S204 of the automatic project cost review method described in the specific embodiment;
[0072] Figure 3 It is a step schematic diagram of steps S301 to S309 of the automatic project cost review method described in the specific embodiment;
[0073] Figure 4 It is a schematic diagram of steps S401 to S404 of the automatic project cost review method described in the specific implementation manner;
[0074] Figure 5 It is a schematic diagram of the structure of the electronic device described in the specific implementation manner.
[0075] The descriptions of the reference numerals involved in the above-mentioned respective drawings are as follows:
[0076] 1. Electronic device;
[0077] 11. Memory;
[0078] 12. Processor. Specific implementation manner
[0079] To illustrate in detail the possible application scenarios, technical principles, specific implementable solutions, achievable purposes and effects, etc. of the present application, the following will be described in detail with reference to the listed specific examples and in conjunction with the drawings. The examples described herein are only used to more clearly illustrate the technical solutions of the present application, and therefore are only used as examples and cannot be used to limit the protection scope of the present application.
[0080] Referring to "embodiment" in this text means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The term "embodiment" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present application, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0081] Unless otherwise defined, the meanings of the technical terms used in this text are the same as those generally understood by those skilled in the technical field to which the present application belongs; the use of the relevant terms in this text is only for describing specific embodiments and is not intended to limit the present application.
[0082] In the description of the present application, the phrase "and / or" is an expression used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " in this text generally represents an "or" logical relationship between the associated objects before and after.
[0083] In the present application, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, primary-secondary or order relationship, etc. between these entities or operations.
[0084] Without further limitations, in this application, the use of "including", "comprising", "having" or other similar open-ended expressions in a statement is intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method or product that includes the said elements. Thus, in a process, method or product that includes a series of elements, it can include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such a process, method or product.
[0085] Similar to the understanding in the "Examination Guidelines", in this application, expressions such as "greater than", "less than", "exceeding" are understood not to include the base number; expressions such as "above", "below", "within" are understood to include the base number. In addition, in the description of the embodiments of this application, the meaning of "multiple" is two or more (including two). Similar expressions related to "many", such as "multiple groups", "multiple times", etc., are understood in the same way, unless otherwise specifically defined.
[0086] In the description of the embodiments of this application, the spatially related expressions used, such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the specific embodiment or the drawing. It is only for the convenience of describing the specific embodiments of this application or for the reader's understanding, rather than indicating or implying that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, it should not be construed as a limitation to the embodiments of this application.
[0087] Please refer to Figure 1 , in the first aspect, this embodiment provides a big data-based automatic project cost review method, including:
[0088] S101. Obtain project information and generate a data collection strategy according to the project information. The data collection strategy includes collection frequency, collection items, collection methods, and data preprocessing logic. The project information includes project type, design requirements, contract terms, and industry standards. When the data collection strategy is configured to meet the preset conditions, it triggers a dynamic adjustment mechanism and updates;
[0089] S102. Collect according to the data collection strategy to obtain heterogeneous project cost data;
[0090] S103. Perform normalization processing on the heterogeneous project cost data to obtain the first data set;
[0091] S104. Input the first data set into the cost prediction model to obtain the cost prediction coefficient. The cost prediction model is configured to be generated based on the fusion of the ARIMA model and the LightGBM model;
[0092] Also, input the first data set into the risk prediction model to obtain the divided risk level and the avoidance suggestions. The risk prediction model is configured as an XGBoost model constructed with the audit items as adjustment parameters;
[0093] Also, input the first data set into the expected return model to obtain the expected return range;
[0094] S105. Generate a review report according to the cost prediction coefficient, the risk level, the avoidance suggestions, and the expected return range.
[0095] In step S101, the project information is the basic attributes and constraints of the project, including project type (such as building, municipal), design requirements (such as structural specifications), contract terms (such as payment methods), and industry standards (such as material specifications). The data collection strategy is a dynamic rule formulated to adapt to the project characteristics. Among them, the collection frequency is the time interval for data collection. Preferably, it is jointly determined by the complexity and scale; the collection items are preferably core cost elements such as material prices and labor costs, and need to conform to the benchmark values in the industry standards (such as the material price fluctuation range); the triggering conditions of the dynamic adjustment mechanism include the contract amount threshold (such as 10% overspending) or the construction period change event. After being triggered, the strategy is updated through interpolation methods (such as linear interpolation) or rule engines.
[0096] In step S102, the heterogeneous cost data refers to multi-source heterogeneous data, such as text contracts and sensor real-time data, which are obtained through different collection methods (such as API interfaces, batch imports).
[0097] In step S103, the goal of the normalization process is to convert the heterogeneous cost data (such as numerical values with different dimensions or ranges like material prices and labor costs) into standardized data with a unified dimension to ensure the stability of subsequent model training. The normalization process adopts the Min-Max normalization or the Z-score method to eliminate the dimension difference. The specific process is as follows:
[0098] For fields with missing values (such as the lack of labor cost in a certain period), if the data distribution is uniform, the mean value is used for filling; if there are outliers, the median value is used for filling;
[0099] For cases where the data distribution has a clear boundary (such as the contract amount), the Min-Max normalization is adopted, and the formula is as follows:
[0100]
[0101] where x is the original value, xmin is the minimum value of this feature in the dataset, x max is the maximum value of this feature in the dataset, x norm is the value after normalization, with a range of (0, 1).
[0102] For cases where the data distribution has no obvious boundary or there are outliers (such as market volatility), Z-score standardization is used, and the formula is as follows:
[0103]
[0104] Among them, μ is the mean of this feature, σ is the standard deviation, and z is the value after standardization (with a mean of 0 and a standard deviation of 1).
[0105] For non-numerical data (such as construction progress status "normal / delayed"), one-hot encoding is used to convert it into a binary vector.
[0106] In step S104, the cost prediction model is generated by fusing the ARIMA model and the LightGBM model. Preferably, the ARIMA model extracts time series features, the LightGBM model filters variables based on feature importance (such as the Gini index), and finally outputs the prediction coefficient through error-weighted fusion.
[0107] In the risk prediction model, preferably, audit items are used as features, and the risk level (such as low risk, medium risk, high risk) is generated through the gradient boosting tree of XGBoost, and the avoidance suggestions are based on the rule library (such as a 5% overspending triggering a cost review).
[0108] The expected return model can generate random scenarios through Monte Carlo simulation (such as the market volatility following a normal distribution), and the random forest model integrates multiple decision trees to predict the return range.
[0109] In step S105, the review report is a structured document integrating the outputs of multiple models, including cost prediction analysis, risk level and avoidance suggestions, expected return range, dynamic strategy adjustment records, visualization charts, etc.
[0110] The cost prediction coefficient and key influencing factors are recorded in the cost prediction analysis. The cost prediction coefficient represents the prediction deviation degree of the total project cost. For example, a coefficient of 1.05 means that the predicted cost is overspent by 5%; the key influencing factors are the top 5 features ranked by importance determined by the LightGBM model (such as material price fluctuations, project duration delay rate).
[0111] The risk level can be classified by different colors. For example, green indicates low risk (e.g., cost deviation < 3%), yellow indicates medium risk (e.g., cost deviation is 3% - 10%), and red indicates high risk (e.g., cost deviation > 10%); the avoidance suggestions are obtained by matching the rule base. For example, "When the risk of over - expenditure on labor costs is detected, it is recommended to initiate Article 5 of the contract terms to renegotiate the price."
[0112] The expected return range includes the return interval and sensitivity analysis. Examples of the return interval are "The median of the expected return is 12 million yuan, and the 90% confidence interval is (10 million, 15 million) yuan"; sensitivity analysis refers to marking the variable that has the greatest impact on the return. For example, when the market volatility increases by 1%, the return decreases by 2%.
[0113] The dynamic strategy adjustment record refers to the events that trigger the dynamic adjustment mechanism (such as the contract amount exceeding the threshold), and the updated acquisition strategy parameters (such as the acquisition frequency being adjusted from once a day to once an hour).
[0114] The visualization charts include time - series prediction curves (ARIMA output), risk distribution histograms (XGBoost classification probabilities), and Monte Carlo simulation return distribution density plots.
[0115] The method provided in this embodiment ensures that the data acquisition strategy flexibly adapts to the project characteristics through a dynamic adjustment mechanism, guarantees the comprehensive and accurate acquisition of heterogeneous construction cost data, and avoids problems of incomplete data coverage or redundancy caused by traditional fixed strategies. The dimensional difference problem of heterogeneous construction cost data is solved through normalization processing, providing high - quality input for the model and ensuring the stability of the analysis. At the same time, the construction cost prediction model is used for construction cost deviation prediction to accurately calculate the construction cost prediction coefficient; the risk prediction model is used for audit risk classification, quantifying the risk level and generating avoidance suggestions to strengthen the pertinence of risk control; the expected return model is used for return volatility assessment, predicting the expected return range by evaluating random scenarios such as market volatility to support the robustness of return decisions. Combining the above three models realizes complementary analysis of multiple models, collaboratively outputs multi - dimensional analysis results, covers the entire chain of construction cost prediction, risk identification, and return assessment, and significantly reduces the prediction deviation of a single model. The finally generated review report includes the construction cost prediction coefficient, risk level, avoidance suggestions, and expected return range, which not only intuitively displays the analysis results but also provides an operable optimization path for project management. The overall process from data acquisition, pre - processing to model prediction and report generation is fully automated, reducing manual intervention links, greatly improving the review efficiency, especially suitable for the rapid decision - making needs of large - scale or highly complex projects, systematically solving the pain points of poor data adaptability, incomplete risk coverage, and high manual dependence in traditional methods, and providing efficient and reliable technical support for project cost control and risk management.
[0116] Please refer to Figure 2, in some embodiments, obtaining project information and generating a data collection strategy based on the project information includes:
[0117] S201. Determine the collection frequency and collection items according to the project type. The collection frequency is generated based on the complexity and scale of the project type. The collection items include material prices, labor costs, equipment rental fees, and construction progress data;
[0118] S202. Determine the collection method and data preprocessing logic according to the design requirements. The collection methods include real-time collection and batch collection. The data preprocessing logic includes data cleaning, data deduplication, and data format conversion;
[0119] S203. Determine the trigger conditions for the dynamic adjustment mechanism according to the contract terms. The trigger conditions include contract amount thresholds, project duration change events, and warning events;
[0120] S204. Supplement and correct the collection frequency and collection items according to industry standards. The industry standards include material price fluctuation ranges, labor cost benchmark values, and average market prices for equipment rental;
[0121] The collection frequency generated based on the complexity and scale of the project type includes:
[0122] Define complexity levels, determine the complexity level to which the current project belongs according to the project type, and generate a first configuration weight;
[0123] And determine the scale level to which the current project belongs according to the project type, and generate a second configuration weight;
[0124] Obtain a preset collection frequency, and perform calculations based on the first configuration weight, the second configuration weight, and the preset collection frequency to obtain an initial collection frequency, which is the collection frequency.
[0125] In step S201, the project type is a professional classification attribute of the project (such as construction, municipal), and its complexity and scale are the core parameters for generating the data collection strategy. The complexity levels include low complexity, medium complexity, and high complexity. Among them, low-complexity projects are standardized projects, medium-complexity projects are partially customized projects, and high-complexity projects are fully customized projects; the scale levels are divided by floor area, total investment amount, etc.
[0126] The initial collection frequency (i.e., the collection frequency) is jointly calculated by the preset collection frequency, the first configuration weight, and the second configuration weight. The calculation formula is as follows:
[0127] F init =F base ×(W1 + W2);
[0128] Where F initTo initialize the acquisition frequency, F base is the preset acquisition frequency, W1 is the first configuration weight, and W2 is the second configuration weight. The preset acquisition frequency is set based on industry general specifications or historical project experience, and is determined by statistically updating the data requirement frequency of similar projects to ensure coverage of the fluctuation cycle of key cost elements. The first configuration weight reflects the impact intensity of complexity on the acquisition frequency, which can be quantified by expert scoring or the analytic hierarchy process. For example, map the complexity level to weight values (0.3, 0.5, 0.7). High-complexity projects require a higher acquisition density due to the variability of technical parameters, so a larger weight is assigned. The second configuration weight reflects the impact intensity of scale on the acquisition strategy, which can be set by grading according to the project investment amount or the amount of work. The larger the scale, the higher the data volume and fluctuation risk, and more frequent acquisition is required to ensure real-time performance. The calculation of the initial acquisition frequency of the dynamic adjustment mechanism needs to satisfy the weight normalization constraint (W1 + W2 ≤ 1). When the sum of the complexity and scale weights exceeds 1, it is scaled proportionally to a total of 1.
[0129] The acquisition items need to cover material prices (such as steel bars, concrete), labor costs (classified by types of work), equipment rental fees (such as crane shift fees), and construction progress data (such as completion percentage). Among them, the material prices need to match the benchmark values in the industry standards (such as the steel price fluctuation range of ±5%).
[0130] In step S202, the design requirements are the set of engineering technical specifications (such as seismic grade, waterproof standard), which determine the data acquisition method and preprocessing rules. Real-time acquisition is applicable to dynamically changing parameters (such as the concrete temperature monitored by sensors) and is transmitted through the API interface; batch acquisition is used for static data (such as contract texts) and is imported using files.
[0131] Data cleaning targets outliers (such as negative values in labor cost records) and uses threshold filtering; duplicate records (such as multiple device status reports at the same time point) are eliminated through data deduplication. Data deduplication can use hash value comparison, and the latest record is retained in case of conflicts; through data format conversion, unstructured text (such as "progress delay" in construction logs) is converted into structured fields (such as the progress status is marked as 0 / 1). Data format conversion can map discrete states to binary vectors through one-hot encoding to avoid ambiguity introduced by numerical magnitudes.
[0132] In step S203, when the difference between the actual contract amount and the contract-agreed amount reaches the contract amount threshold, the data acquisition strategy is triggered to be updated. The construction period change event matches the response actions in the event coding rule library (such as increasing the acquisition frequency of progress data). Warning events include material price fluctuations exceeding the benchmark or safety accident records, triggering real-time warnings and activating the emergency acquisition mode.
[0133] In step S204, the industry standard is a regulatory document issued by an authoritative agency and is used to modify the acquisition strategy parameters. The fluctuation range of material prices restricts the rationality of the real-time acquired data. If the market price exceeds this range, it is marked as abnormal and triggers manual verification; the labor cost benchmark value is set according to regions and types of work and is used for cost deviation analysis; the average market price of equipment leasing is calculated by rolling historical data and serves as a reference benchmark for the rationality of equipment costs. When making supplementary modifications, if the industry standard is updated (such as a newly issued concrete strength test specification), the detection parameters in the acquisition items are adjusted synchronously (such as adding sampling points for compressive strength).
[0134] This embodiment divides the complexity level and scale level based on the project type, combines the preset acquisition frequency with the first configuration weight and the second configuration weight to generate the initial acquisition frequency, and realizes the precise matching of the acquisition density and project characteristics; designs a demand-driven data preprocessing logic to improve the data quality and the degree of structuring, and supports the reliability of subsequent model analysis; triggers a dynamic adjustment mechanism through the contract amount threshold, construction period change events, and warning events to respond in real time to cost overruns, schedule delays, or abnormal market fluctuations. For example, when the threshold is exceeded, the strategy parameters are automatically updated to enhance the risk response ability; introduces the material price fluctuation range, labor cost benchmark value, and average market price of equipment leasing in the industry standard to ensure the compliance and rationality of the acquired data. This embodiment realizes the full-link closed-loop management from data acquisition to strategy adjustment through parametric configuration, a dynamic rule engine, and a standard linkage mechanism, and significantly improves the timeliness, accuracy, and risk control ability of engineering cost review.
[0135] In some embodiments, the preset conditions include a first preset condition and a second preset condition, and the dynamic adjustment mechanism includes a first dynamic adjustment mechanism and a second dynamic adjustment mechanism;
[0136] When the data acquisition strategy is configured to meet the preset conditions, the dynamic adjustment mechanism is triggered and updated to include:
[0137] When the data acquisition strategy meets the first preset condition, the first dynamic adjustment mechanism is triggered to update the data acquisition strategy;
[0138] When the data acquisition strategy meets the second preset condition, the second dynamic adjustment mechanism is triggered to update the data acquisition strategy.
[0139] In this embodiment, the update of the data acquisition strategy is finely regulated by hierarchical conditions, significantly enhancing the flexibility and adaptability of the data acquisition strategy. When the data acquisition strategy meets the first preset condition, the first dynamic adjustment mechanism is triggered. Preferably, the first preset condition means that within the first preset time period, the fluctuation range of the first raw data collected in real time exceeds the preset first fluctuation threshold. When the fluctuation range exceeds the first preset fluctuation threshold, the first dynamic adjustment mechanism is triggered, and the real-time response to the dynamic change trend is achieved by increasing or decreasing the data acquisition frequency. It can not only increase the acquisition density to capture key change details when the data fluctuates violently, but also reduce the acquisition frequency when the data tends to be stable to reduce resource redundancy, thereby optimizing the overall acquisition efficiency and cost.
[0140] When the data acquisition strategy meets the second preset condition, the second dynamic adjustment mechanism is triggered. Preferably, the second preset condition means that the first raw data exceeds the preset numerical range. By analyzing the fluctuation range within the second preset time period, abnormal error values and systematic abnormal values are distinguished, and interpolation correction or generation of abnormal warning information is respectively used for processing, effectively avoiding the interference of abnormal data on subsequent analysis and enhancing data reliability and decision support capabilities.
[0141] In this embodiment, through the coordinated operation of the first dynamic adjustment mechanism and the second dynamic adjustment mechanism, it not only ensures the adaptive adjustment ability of the data acquisition strategy to the fluctuation scenario, but also strengthens the recognition and processing logic of abnormal data, forming a closed-loop management from dynamic frequency optimization to abnormal risk prevention and control, and finally realizing the comprehensive improvement of data acquisition quality, efficiency and stability.
[0142] In some embodiments, when the data acquisition strategy meets the first preset condition and triggers the first dynamic adjustment mechanism to update the data acquisition strategy, it includes:
[0143] Collect data in real time and record the collected data as the first raw data;
[0144] Obtain multiple first raw data within the first preset time period up to now and draw the first data fluctuation range;
[0145] Judge whether the first data fluctuation range is within the first preset fluctuation threshold. If not, it means that the data acquisition strategy meets the first preset condition and triggers the first dynamic adjustment mechanism. The first dynamic adjustment mechanism is configured to increase or decrease the acquisition frequency of the data acquisition strategy;
[0146] When the data acquisition strategy meets the second preset condition and triggers the second dynamic adjustment mechanism to update the data acquisition strategy, it includes:
[0147] Judging one by one whether the first raw data is within a preset numerical range. If not, it means that the data acquisition strategy meets the second preset condition, and the second dynamic adjustment mechanism is triggered. The second dynamic adjustment mechanism is configured as follows:
[0148] Obtain multiple first raw data within the second preset time period up to now and plot the second data fluctuation amplitude;
[0149] Judge whether the second data fluctuation amplitude is within the second preset fluctuation threshold;
[0150] If so, mark the first raw data that is not within the preset numerical range as an abnormal error value, and use the interpolation method to correct the abnormal error value;
[0151] If not, mark the first raw data that is not within the preset numerical range as an abnormal value, and generate an abnormal warning message based on the abnormal value and the second data fluctuation amplitude.
[0152] The automatic engineering cost review method provided in this embodiment is aimed at the pre - data acquisition in the field of engineering cost. The data acquisition strategy described in this embodiment is for multiple historical project data sets in the database that match the target project characteristics, such as engineering cost data with similar building scale, structure type or regional attributes.
[0153] In this embodiment, the first raw data refers to the historical engineering cost records extracted from the database that have similar characteristics (such as building area, structural complexity, regional attributes) to the target project. For example, the unit price of concrete consumption or the consumption of man - hours of similar projects is used to construct a benchmark data sequence.
[0154] The first preset time period is a continuous time period (such as 24 hours or the typical cycle during the project preparation stage) for evaluating the fluctuation characteristics of the historical project data set. Its length is set according to the update frequency of the engineering cost data and is used to determine the data window range for calculating the standard deviation or range.
[0155] The first data fluctuation amplitude quantifies the dispersion degree of the historical data set through the standard deviation. The calculation formula of the standard deviation is as follows:
[0156]
[0157] Where, S represents the standard deviation, y k represents the k - th historical project data point, is the arithmetic mean of the data within the first preset time period, and n is the total amount of data included within the first preset time period.
[0158] The first data fluctuation amplitude characterizes the data span through the range. The calculation formula of the range is as follows:
[0159] R = max(yk ) - min(y k );
[0160] where max(y k ) is the maximum value of the data within the first preset time period, and min(y k ) is the minimum value of the data within the first preset time period.
[0161] The first preset fluctuation threshold refers to the allowable fluctuation range preset according to industry experience or historical data distribution. When the standard deviation S or the range R exceeds the first preset fluctuation threshold, it is determined that the data fluctuates abnormally and the acquisition frequency needs to be adjusted.
[0162] The preset numerical range refers to the reasonable interval defined based on the project cost specification and is used to screen abnormal data.
[0163] The second preset time period is the time window (such as 72 hours or the complete preparation period across project phases) used to analyze the persistence of abnormal data in the second dynamic adjustment mechanism. Its length is usually longer than the first preset time period and is used to evaluate whether the abnormal data has the characteristics of time continuity or systematic deviation.
[0164] The second data fluctuation amplitude is characterized by variance or the moving average deviation D. Variance reflects the overall dispersion, and its calculation formula is as follows:
[0165]
[0166] where h q is the q-th data point within the second preset time period, is the arithmetic mean of the data within the second preset time period, and m is the total amount of data included within the second preset time period.
[0167] The moving average deviation measures the time series stability, and its calculation formula is as follows:
[0168]
[0169] where h r is the actual value, is the moving average predicted value, l is the calculation window length (i.e., the second preset time period),
[0170] The second preset fluctuation threshold refers to the judgment boundary set for systematic anomalies, preset according to industry experience or historical data distribution. When the variance or the moving average deviation D does not exceed the second preset fluctuation threshold, the first raw data is an abnormal error value; when the variance or the moving average deviation D exceeds the second preset fluctuation threshold, it is determined that the first raw data is an abnormal value.
[0171] An abnormal error value refers to a data deviation caused by local interference (such as a single input error or a temporary price fluctuation), whose fluctuation range is within the second preset fluctuation threshold and can be corrected by the interpolation method. The linear interpolation formula is as follows:
[0172]
[0173] Where v corrected is the first original data after correction, t′ is the abnormal time, t′1 is the nearest time before the abnormal time, t′2 is the nearest time after the abnormal time, v1 is the data point corresponding to the nearest time before the abnormal time, and v2 is the data point corresponding to the nearest time after the abnormal time.
[0174] For example, the unit price of steel in a certain project was misfilled as 8,500 yuan / ton (the reasonable range is [6,000, 7,500]), but the data fluctuation in adjacent periods was gentle ( not exceeding the limit), then it was marked as an abnormal error value and corrected.
[0175] An outlier refers to a significant deviation that exceeds the second preset fluctuation threshold (such as the sand and gravel price in a certain area continuously being higher than the historical extreme value due to policy adjustments), indicating the existence of structural changes or systematic errors. It is necessary to generate an abnormal warning message and trigger an artificial verification process. For example, the hydropower installation cost of multiple similar projects in the database suddenly increased to 200 yuan / square meter (the benchmark range is [80, 120]), and the variance increased significantly, then it was determined as an outlier.
[0176] The automatic engineering cost review method provided in this embodiment realizes the efficient acquisition and accurate abnormal processing of historical engineering cost data by constructing a hierarchical dynamic adjustment mechanism. Based on the first dynamic adjustment mechanism, the historical data set matching the characteristics of the target project is obtained in real time, and the standard deviation or range is calculated within the first preset time period. When the first data fluctuation range exceeds the first preset fluctuation threshold, the acquisition frequency is automatically increased or decreased to adapt to the dynamic change of the data (such as high-frequency acquisition to cope with violent price fluctuations). Through the second dynamic adjustment mechanism, the persistence of abnormal data is analyzed in a longer period. For abnormal error values that exceed the preset numerical range but the second data fluctuation range does not exceed the second preset fluctuation threshold, linear interpolation is used to correct based on adjacent valid data points; for outliers with excessive fluctuations, warning messages are generated to trigger artificial verification. This embodiment adopts a two-level threshold determination and differential processing strategy, taking into account both data smoothing correction and systematic risk warning, effectively solving the problem of identifying local interference and structural deviation in the early-stage data collection of engineering cost. It not only avoids the masking of systematic anomalies by single interpolation correction but also reduces the low efficiency of manual review, significantly improving the reliability and dynamic adaptability of the review data, and providing high-quality benchmark data support for subsequent cost estimation and risk analysis.
[0177] Please refer to Figure 3 , in some embodiments, inputting the first data set into the cost prediction model to obtain the cost prediction coefficient includes:
[0178] S301. Obtain the first data set, and perform time series analysis on the first data set to extract time series features, where the time series features include trend features, periodic features, and random fluctuation features;
[0179] S302. Conduct a stationarity test on the time series features. If the time series features do not meet the stationarity threshold, perform a differencing operation on the time series features until the time series features meet the stationarity threshold;
[0180] S303. Input the time series features into the trained ARIMA model to generate the first prediction result;
[0181] S304. Perform feature fusion on the first prediction result and the first data set to obtain the second data set. Feature fusion includes adding the first prediction result as a new feature to the first data set and performing standardization processing;
[0182] S305. Evaluate the feature importance of the second data set, and screen out the features that have a significant impact on the prediction result, denoted as the second data features;
[0183] S306. Input the second data features into the trained LightGBM model to generate the second prediction result;
[0184] S307. Calculate the first prediction error of the first prediction result on the historical data, and calculate the second prediction error of the second prediction result on the historical data;
[0185] S308. Judge the magnitudes of the first prediction error and the second prediction error. If the first prediction error is smaller, increase the weight value corresponding to the first prediction error. If the second prediction error is smaller, increase the weight value corresponding to the second prediction error;
[0186] S309. Add the weighted first prediction result and the second prediction result to generate the cost prediction coefficient.
[0187] In step S301, the trend feature represents the long-term upward or downward trend of the data over time (such as the labor cost increasing year by year due to inflation); the periodic feature refers to the repeated fluctuations at fixed intervals (such as the seasonal material price fluctuations), and its cycle length is determined according to the data spectrum analysis; the random fluctuation feature is the residual term after removing the trend and cycle, reflecting accidental interference or noise.
[0188] In step S302, the stationarity test uses the ADF test to determine whether the time series characteristics meet the stationarity threshold through the statistic. The stationarity threshold is usually < 0.05, and the statistic calculation formula is as follows:
[0189]
[0190] where, is the estimated value of the autoregressive coefficient, and SE is the standard error.
[0191] If the test fails, the sequence is differenced until the ADF statistic of the differenced sequence meets the stationarity threshold. The differencing operation formula is as follows:
[0192]
[0193] where, d represents the differencing order, y t represents the original observation value of the time series at time t (such as the concrete price in the t-th month), and y t-d represents the observation value lagged by d time units (such as the price of the previous month when d = 1). This operation makes the data meet the stationarity assumption of models such as ARIMA by eliminating the non-stationary trend (such as linear growth) or periodic fluctuations (such as seasonal fluctuations) of the time series.
[0194] In step S303, the input of the ARIMA model is the stationary time series characteristics, and its structural parameters are (p, d, q), where p is the order of the autoregressive term, q is the order of the moving average term, and d is the number of differencing times determined in step S302. The ARIMA model outputs the first prediction result The specific content will be described in detail later.
[0195] In step S304, feature fusion adds as a new feature to the first dataset to form the second dataset, and performs standardization processing. The calculation formula for standardizing each feature x m (including the original feature and the predicted value) is as follows:
[0196]
[0197] where, μ m is the mean of this feature, σ m is the standard deviation, and z m is the value after standardization.
[0198] In step S305, the feature importance assessment uses the permutation importance method to calculate the importance score I m for each feature z in the second dataset m, the permutation importance method usually calculates importance by shuffling the values of a certain feature and observing the change in model performance. Its calculation formula is as follows:
[0199]
[0200] Where, L is the loss function (such as mean squared error), N is the number of permutations, is the predicted value of the model on the second dataset when no feature is shuffled, only shuffle feature z m After that, the predicted value of the model on the same dataset. Select features with I m ≥θ (the threshold θ takes the top 20% quantile) as the second data features.
[0201] In step S306, the LightGBM model takes the second data features as input and generates a second prediction result through gradient boosting decision trees The specific content will be described in detail later.
[0202] In step S307, the first prediction error E1 is expressed by the following formula:
[0203]
[0204] The second prediction error E2 is expressed by the following formula:
[0205]
[0206] In step S309, the cost prediction coefficient is expressed by the following formula:
[0207]
[0208] Where, ω1 is the weight value corresponding to the first prediction error, and ω2 is the weight value corresponding to the second prediction error.
[0209] In this embodiment, by integrating time series analysis and an ensemble learning model, the accuracy and interpretability of cost prediction are improved. Specifically, first, the original data is decomposed by time series. After stationarity test and differencing operation, it is input into the ARIMA model to capture linear time series dependencies and generate the first prediction result. Subsequently, the predicted values are used as new features and fused with the original data. The dimensional differences are eliminated through standardization processing, and the permutation importance method is used to screen key features and remove redundant noise to ensure that the LightGBM model focuses on features with high contribution degrees. Further, based on the historical prediction errors of the two models, the weights are dynamically adjusted so that the final prediction coefficient adaptively combines the linear advantages of ARIMA and the non-linear fitting ability of LightGBM. This embodiment not only enhances the input information density through feature fusion but also optimizes the model collaboration efficiency using the error feedback mechanism, effectively balancing long-term trends, short-term fluctuations, and complex non-linear relationships. While improving the prediction accuracy, it can analyze the influence weights of each feature on the cost, providing a quantifiable decision basis for cost control.
[0210] Please refer to Figure 4 , in some embodiments, inputting the time series features into the trained ARIMA model to generate the first prediction result includes:
[0211] S401. Generate the first model parameters of the ARIMA model according to the autocorrelation function and partial autocorrelation function. The first model parameters include the autoregressive order, differencing order, and moving average order.
[0212] S402. Train the ARIMA model according to the first model parameters and the first sample data, and predict the time series feature data after training is completed to generate the first prediction result.
[0213] Inputting the second data features into the trained LightGBM model to generate the second prediction result includes:
[0214] S403. Generate the second model parameters based on the gradient boosting decision tree algorithm. The second model parameters include the first learning rate, the first maximum depth of the tree, and the first number of leaf nodes.
[0215] S404. Iteratively train the LightGBM model according to the second model parameters and the second sample data, and predict the second data features after training is completed to generate the second prediction result.
[0216] In step S401, the input of the ARIMA model is the stationary time series features, and its first model parameters are (p, d, q), where p is the autoregressive order, q is the moving average order, and d is the differencing order determined in step S302. The first sample data refers to the stationary time series training set that has completed the stationary processing, which can be the stationary data after the differential transformation of the historical steel price series.
[0217] In step S402, training the ARIMA model includes the following steps:
[0218] Based on the differencing order d determined in step S302, the original non-stationary sequence is converted into a stationary sequence through d-order differencing as the model input. At this time, p and q in the first model parameters need to be determined by the truncation characteristics of the autocorrelation function (ACF) and the partial autocorrelation function (PACF): If the ACF decays sharply within the confidence interval after lag q steps, then q takes this critical value; if the PACF decays sharply within the confidence interval after lag p steps, then p takes this critical value.
[0219] For the autoregressive coefficient α i and the moving average coefficient β j Solve them using conditional sum of squares minimization or maximum likelihood estimation.
[0220] Evaluate the rationality of the first model parameters (p, d, q) through the Akaike information criterion or the Bayesian information criterion, and select the model with the minimum criterion value as the final training result.
[0221] The calculation formula for the ARIMA model to output the first prediction result is expressed as:
[0222]
[0223] where γ is the constant term, that is, the intercept, which adjusts the prediction benchmark value (such as the long-term equilibrium price), α i is the autoregressive coefficient (order p), representing the linear weight of the historical observation value y t-i (such as the steel price in the past i periods) for the current prediction; β j is the moving average coefficient (order q), representing the correction weight of the historical prediction error ∈ t-j (such as the prediction deviation j periods ago) for the current prediction.
[0224] In step S403, the LightGBM model takes the second data feature as the input and generates the second prediction result through gradient boosting decision trees Its formula is expressed as:
[0225]
[0226] where B is the number of trees, f b (z t ) is the prediction output of the b-th tree, η is the first learning rate, which is used to shrink the weight of each tree to prevent overfitting. The first maximum depth of the tree is used to limit the number of vertical splitting layers of a single decision tree to prevent the model from being overly complex; the first number of leaf nodes of the tree is used to constrain the number of terminal nodes (leaves) of a single decision tree, which directly affects the partitioning granularity of the feature space.
[0227] In step S404, the iterative training of the LightGBM model includes the following steps:
[0228] Input the second sample data, including feature z t and the corresponding true value y t ;
[0229] Minimize the loss function L′ through gradient descent, and its formula is expressed as:
[0230]
[0231] where T is the number of samples, λ is the regularization coefficient, and Ω(f b ) is the tree complexity penalty term;
[0232] Based on the training data, split the nodes through the greedy algorithm, and dynamically determine the structure (split feature, threshold) of each tree f b (z t ) and the weights of the leaf nodes.
[0233] Use the trained LightGBM model to predict the second data feature z t and accumulate the weighted outputs of all the trees to obtain the final second prediction result
[0234] In this embodiment, by integrating the ARIMA and LightGBM dual - model architectures, comprehensive modeling of time - series features and non - linear features is achieved, improving the robustness and accuracy of the prediction system. Specifically, the ARIMA model accurately determines the autoregressive order, moving average order, and differencing order based on the truncation characteristics of the autocorrelation function and partial autocorrelation function, solves the autoregressive coefficients and moving average coefficients by combining conditional sum of squares minimization or maximum likelihood estimation, and screens the optimal parameter combination through the Akaike information criterion or Bayesian information criterion, ensuring that the model has a strong ability to capture the linear dynamic relationship of the stationary time - series features (such as the trend term of steel prices), and can effectively analyze the linear contribution of historical observations and prediction errors to the current prediction. At the same time, the LightGBM model uses the gradient - boosting decision - tree algorithm to dynamically adjust the weight of each tree with the first learning rate, restricts the complexity of a single tree by combining the first maximum depth of the tree and the first number of leaf nodes, and introduces a regularization coefficient and a complexity penalty term to suppress the overfitting risk, enabling it to efficiently mine the non - linear associations hidden in the second data features (such as market supply - demand fluctuations, unexpected event interferences, etc.). ARIMA focuses on the explicit modeling of linear time - series dependencies, and LightGBM strengthens the high - dimensional mapping of non - linear features. Finally, the prediction results are synergistically optimized through a weighted or ensemble strategy. This embodiment significantly improves the prediction stability and generalization performance for complex scenarios (such as multi - factor coupled fluctuations of steel prices) while ensuring the interpretability of the model.
[0235] In some embodiments, the first data set is input into a risk prediction model to obtain the divided risk levels and avoidance suggestions. The risk prediction model configured as an XGBoost model constructed with audit items as adjustment parameters includes:
[0236] Generate audit items according to the project information, and extract features from the first data set according to the audit items to generate audit feature data. The audit items include contract compliance, cost over - run rate, construction period delay rate, and quality standard compliance;
[0237] Clean the audit feature data to generate a third data set. The data cleaning includes missing value filling, outlier handling, and feature standardization;
[0238] And, label the third sample data according to the audit items to generate training samples. The labels include low - risk, medium - risk, and high - risk;
[0239] Construct an XGBoost model based on the gradient - boosting algorithm, and initialize the third model parameters of the XGBoost model through cross - validation. The third model parameters include the second learning rate, the second maximum depth of the tree, the second number of leaf nodes, and the regularization parameter;
[0240] Train the XGBoost model using the third sample data after label annotation, and predict the third data set after training is completed to generate the risk level classification result;
[0241] According to the risk level classification result, combined with the preset risk avoidance rule library, generate avoidance suggestions.
[0242] In this embodiment, the audit items can be contract compliance (checking the integrity of clause performance), cost overrun rate (the deviation ratio of actual cost to budget), construction period delay rate (the proportion of the lag time in the schedule), and quality standard compliance (the passing rate of acceptance indicators). The audit items are mapped to the first data set through preset rules. For example, contract compliance is calculated through clause matching degree, and the cost overrun rate is quantified by (actual expenditure - budget) / budget. The generated audit feature data needs to include the numerical values or classification indicators corresponding to each item.
[0243] The third data set refers to the audit feature data after data cleaning, and its generation process includes:
[0244] Missing value filling: For the fields that cannot be directly obtained in the audit items, fill them with the mean value of similar projects or interpolation method;
[0245] Outlier processing: Based on the or quantile threshold, correct or remove the observed values that deviate from the normal range (such as extreme values where the cost overrun rate exceeds 200%);
[0246] Feature standardization: Perform Z-score normalization on features with large dimensional differences (such as construction period delay rate and quality standard score) to make them follow a distribution with a mean of 0 and a standard deviation of 1.
[0247] The third sample data is the training set of the third data set after label annotation, and its label is generated based on the comprehensive evaluation of the audit items. Preferably, low risk is that contract compliance ≥ 90%, cost overrun rate ≤ 5%, construction period delay rate ≤ 10% and quality standard compliance ≥ 95%; medium risk is that any audit item exceeds the low risk threshold but does not reach the high risk critical value; high risk is that contract compliance < 70%, cost overrun rate > 15%, construction period delay rate > 20% or quality standard compliance < 80%.
[0248] The second learning rate of the third model parameter controls the contribution weight of each decision tree to the final prediction, reducing the risk of overfitting by shrinking the step size; the second maximum depth of the tree limits the splitting layer of a single tree to prevent the model from memorizing noise due to excessive complexity; the second number of leaf nodes restricts the number of terminal nodes of a single tree to balance the feature space division granularity and calculation efficiency; the regularization parameter adds a penalty term to the leaf node weights in the loss function to suppress the model complexity.
[0249] When training the model, the objective function is defined as:
[0250]
[0251] Among them, O is the total number of samples of the third sample data, y o The true risk label of the o-th sample, is the model's predicted output for the oth sample, B is the total number of boosted trees, and g b is the function representation of the b-th tree.
[0252] The trained XGBoost model is used to predict the third data set, generate risk level classification results, and generate avoidance suggestions in combination with the preset risk avoidance rule base. The risk avoidance rule base is a predefined set of strategies that maps to specific operational suggestions based on risk levels, and is associated with risk level classification results through logical conditions to ensure that the generation process of suggestions is traceable and in compliance with business specifications. Avoidance suggestions include: for low risk levels, general monitoring suggestions are generated, including regular data review and progress tracking; for medium risk levels, active intervention suggestions are generated, including cost control measures and construction period optimization plans; for high risk levels, emergency handling suggestions are generated, including contract terms revision and resource reallocation.
[0253] This embodiment achieves multi-dimensional accurate quantification and dynamic response of project risks by constructing a risk prediction framework based on audit items and XGBoost model. Feature extraction is performed on the first data set according to the audit items to generate audit feature data to ensure that the feature space covers the core risk dimensions of project execution. In the data cleaning stage, missing value filling, outlier correction and Z-score standardization are used to generate a high-quality third data set, eliminate noise interference and unify the dimension, and provide reliable input for model training. Label annotation is based on the threshold rule of audit items to ensure the objectivity of risk level division and consistency of business logic. The XGBoost model integrates the sum of squares of prediction errors and tree complexity penalty terms in the objective function through the coordinated optimization of the second learning rate, the second maximum depth of the tree, the second number of leaf nodes and the regularization parameter, effectively improving the modeling ability and generalization performance of nonlinear risk associations. Finally, the risk avoidance rule library is mapped to hierarchical recommendations based on the level division results: low risk triggers regular monitoring, medium risk starts active intervention, and high risk drives emergency treatment, forming a closed-loop risk management mechanism. This embodiment takes into account both prediction accuracy and business operability through structured definition of audit items, standardized processing of data cleaning, interpretable configuration of model parameters, and logical output of avoidance suggestions, thereby significantly improving risk identification efficiency and the pertinence of response strategies.
[0254] In some embodiments, the first data set is input into the expected return model to obtain the expected return range including:
[0255] Extract revenue-related feature data from the first dataset and perform data cleaning to obtain the fourth dataset. The feature data includes project budget, cost distribution, market volatility, and historical revenue data;
[0256] Perform random sampling on the fourth dataset based on the Monte Carlo simulation method to generate multiple revenue prediction scenarios;
[0257] Input each revenue prediction scenario into the random forest model to obtain the corresponding revenue prediction value. The random forest model is configured to integrate the prediction results through multiple decision trees;
[0258] Perform statistical analysis on multiple revenue prediction values to generate an expected revenue range, which includes the minimum revenue value, the maximum revenue value, and the median expected revenue.
[0259] In this embodiment, the fourth dataset includes project budget (total funds allocation amount), cost distribution (probability density of each link's cost), market volatility (standard deviation of price or demand change range), and historical revenue data (past return records of similar projects), which are used to quantify the dynamic characteristics of revenue influencing factors. Data cleaning needs to handle missing features (e.g., filling blank items of historical revenue with industry benchmark values), outliers (e.g., truncating when market volatility exceeds 3 times the industry peak), and standardization (converting cost distribution into a probability density function with a uniform dimension).
[0260] Performing random sampling on the fourth dataset based on the Monte Carlo simulation method includes the following steps:
[0261] Define the probability distribution for each feature in the fourth dataset: The project budget follows a uniform distribution (assuming the budget adjustment range is known); the cost distribution is fitted to a lognormal distribution (right-skewed characteristic); the market volatility is based on historical data to establish an autoregressive conditional heteroskedasticity (ARCH) model; the historical revenue data uses kernel density estimation to generate a continuous probability density function;
[0262] Execute random sampling: Independently draw groups of feature values to form revenue prediction scenarios, where is the number of simulation times (e.g., ), and each group of scenarios includes budget, cost distribution parameters, volatility, and revenue benchmark values.
[0263] The training process of the random forest model is as follows:
[0264] Each decision tree h t″ (x) is trained from a subset of the fourth dataset (Bootstrap sampling) and randomly selected partial features (such as cost distribution and market volatility);
[0265] The final predicted value is the mean of the outputs of all the trees, and its formula is as follows:
[0266]
[0267] where \(T'\) is the total number of trees;
[0268] The model parameters include the third maximum depth of the tree (limiting the complexity of a single tree), the number of randomly selected features (controlling the size of the feature subset), and the minimum number of samples in a leaf node (preventing overfitting).
[0269] The method for generating the expected return range is as follows: Input scenarios into the random forest model to obtain the predicted return set Determine the 5% quantile corresponding to the lowest return value through quantile calculation The 95% quantile corresponding to the highest return value The 50% quantile corresponding to the median of the expected return is the predicted return value for the th scenario.
[0270] In this embodiment, through the synergistic effect of Monte Carlo simulation and the random forest model, the robustness enhancement and uncertainty quantification of return prediction are achieved. The fourth dataset comprehensively depicts the dynamic association and random characteristics of the return influencing factors. The data cleaning process standardizes missing values, outliers, and dimensional differences to ensure the input quality. The Monte Carlo simulation method constructs a multi-dimensional joint probability distribution by defining the uniform distribution of the budget, the log-normal distribution of the cost, the ARCH model of the volatility, and the kernel density estimation of the return, and performs random sampling to generate a set of return prediction scenarios, fully covering extreme cases and normal distributions. The random forest model uses Bootstrap sampling and random feature selection to train multiple decision trees, reduces the risk of overfitting of a single tree through the integrated output, and improves the generalization ability for non-linear relationships. Finally, based on the predicted return set, the lowest return value, the highest return value, and the median of the expected return are determined through quantile calculation to form an expected return range that takes into account both conservative estimates and optimistic expectations. This embodiment captures the uncertainties of market fluctuations and cost distributions through probability sampling, combines the robust prediction and integrated noise reduction capabilities of the random forest, provides a statistically significant return range reference for decision-makers, and supports the adaptive adjustment of risk preferences and the optimization of resource allocation.
[0271] In a second aspect, this embodiment also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described in the first aspect is implemented.
[0272] The computer program involved in this embodiment can be stored in a computer-readable storage medium. The computer-readable storage medium includes, but is not limited to, magnetic disks, magnetic tapes, magnetic cards, floppy disks, flash memories, optical discs, optical cards, read-only memories (ROMs), random access memories (RAMs), erasable programmable ROMs (EPROMs), and electrically erasable programmable ROMs (EEPROMs), etc. It also includes other biological, physical, or chemical structures that can achieve functions similar to or equivalent to the above-listed storage media, such as units with information storage capabilities like DNA, RNA, proteins, etc. In a specific embodiment, the storage medium involved can be one of the above medium types or a combination of the above medium types. In different embodiments, the computer program involved in the embodiment can be stored centrally in a single medium or distributedly stored in multiple media. The memory containing the computer-readable storage medium can be a non-volatile memory or a random access memory. These computer-readable storage media can be built into the device or can be connected to the device involved in the embodiment as an external device or a part of an external device. In some embodiments, the memory with the computer-readable storage medium is deployed locally; in other embodiments, a scheme of deploying the memory away from the processor can also be adopted, such as a network-attached memory accessed via an RF circuit or an external port and a communication network, where the communication network can be the Internet, one or more intranets, local area networks (LANs), wide area wireless networks (WLANs), storage area networks (SANs), etc., or a suitable combination thereof, as long as the computer device can access the memory. In addition, the computer program involved in the embodiment can be stored in plaintext / ciphertext form or can be designed as training data and be integrally reorganized and implicitly stored in the parameter states of a deep neural network or other machine learning models through model training.
[0273] Please refer to Figure 5 , in a third aspect, this embodiment also provides an electronic device 1, including a memory 11 and a processor 12. The memory 11 is used to store one or more computer program instructions, where the one or more computer program instructions are executed by the processor 12 to implement the method described in the first aspect.
[0274] The processor described in this embodiment can be implemented by hardware, firmware, software, or a combination thereof. It can use circuits, one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, microprocessors, or at least one of them. It also includes other physical, biological, or chemical structures that can achieve the same or equivalent functions as the above-listed processors, such as biological neurons, quantum computing units, DNA computing units, etc., so that the processor can execute some steps, all steps, or any combination of the steps mentioned in the computer programs or methods involved in the various embodiments of this application.
[0275] Different from the prior art, the above technical solution has the following beneficial effects:
[0276] The present invention realizes the intelligent upgrade of project cost review by integrating a multi-dimensional prediction model and an adaptive data acquisition mechanism. A dynamic data acquisition strategy is constructed based on project characteristics, and the acquisition frequency and acquisition method are dynamically adjusted by real-time sensing of project complexity and scale. Trigger conditions are set in combination with contract terms and industry standards to ensure the timeliness and integrity of data acquisition. A collaborative architecture of a cost prediction model, a risk prediction model, and an expected return model is adopted to improve the review accuracy: the cost prediction model combines time series analysis and gradient boosting decision trees to comprehensively capture cost trends and feature correlations; the risk prediction model uses the XGBoost algorithm with parameterized adjustment of audit items to achieve accurate risk level classification and generation of avoidance strategies; the expected return model constructs multi-dimensional return prediction scenarios with the help of Monte Carlo simulation and ensemble learning. The three work together to form a closed-loop review system from cost control, risk warning to return evaluation, effectively solving the problems of data lag and single prediction dimension in traditional reviews. Through automated data preprocessing and model calculation, the review efficiency is significantly improved, the manual intervention error is reduced, and multi-dimensional quantitative basis is provided for engineering decisions.
[0277] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of this application, the patent protection scope of this application cannot be limited thereby. Any technical solutions obtained by equivalent structure or equivalent process substitution or modification based on the essential concept of this application and using the content recorded in the text and drawings of the specification of this application, as well as those directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, are all included within the patent protection scope of this application.
Claims
1. An automatic evaluation method for project cost based on big data, characterized in that, Including: Obtain project information and generate a data collection strategy based on the project information. The data collection strategy includes collection frequency, collection items, collection methods, and data preprocessing logic. The project information includes project type, design requirements, contract terms, and industry standards. The data collection strategy is configured to trigger a dynamic adjustment mechanism and be updated when meeting preset conditions; Collect according to the data collection strategy to obtain heterogeneous cost data; Perform normalization processing on the heterogeneous cost data to obtain a first data set; Input the first data set into a cost prediction model to obtain a cost prediction coefficient. The cost prediction model is configured to be generated based on the fusion of an ARIMA model and a LightGBM model; And input the first data set into a risk prediction model to obtain the divided risk level and avoidance suggestions. The risk prediction model is configured to be an XGBoost model constructed with audit items as adjustment parameters; And input the first data set into an expected return model to obtain an expected return range; Generate a review report based on the cost prediction coefficient, risk level, avoidance suggestions, and expected return range.
2. The automatic project cost review method based on big data according to claim 1, wherein Obtaining project information and generating a data collection strategy based on the project information includes: Determine the collection frequency and collection items according to the project type. The collection frequency is generated according to the complexity and scale of the project type. The collection items include material prices, labor costs, equipment rental fees, and construction progress data; Determine the collection method and data preprocessing logic according to the design requirements. The collection method includes real-time collection and batch collection. The data preprocessing logic includes data cleaning, data deduplication, and data format conversion; Determine the trigger conditions for the dynamic adjustment mechanism according to the contract terms. The trigger conditions include contract amount threshold, project duration change event, and warning event; Supplement and correct the collection frequency and collection items according to the industry standards. The industry standards include material price fluctuation range, labor cost benchmark value, and equipment rental market average price; The collection frequency being generated according to the complexity and scale of the project type includes: Define complexity levels, determine the complexity level to which the current project belongs according to the project type, and generate a first configuration weight; And determine the scale level to which the current project belongs according to the project type, and generate a second configuration weight; Obtain a preset collection frequency, and perform calculations based on the first configuration weight, second configuration weight, and preset collection frequency to obtain an initial collection frequency, which is the collection frequency.
3. The automatic engineering cost review method based on big data according to claim 1 or 2, characterized in that The preset conditions include a first preset condition and a second preset condition. The dynamic adjustment mechanism includes a first dynamic adjustment mechanism and a second dynamic adjustment mechanism; The data collection strategy being configured to trigger a dynamic adjustment mechanism and be updated when meeting preset conditions includes: The data collection strategy triggers the first dynamic adjustment mechanism to update the data collection strategy when meeting the first preset condition; The data collection strategy triggers the second dynamic adjustment mechanism to update the data collection strategy when meeting the second preset condition.
4. The automatic engineering cost review method based on big data according to claim 3, characterized in that When the data acquisition strategy meets the first preset condition, it triggers the first dynamic adjustment mechanism to update the data acquisition strategy, including: Collect data in real time and record the collected data as the first raw data; Obtain multiple first raw data within the first preset time period up to now and plot the first data fluctuation amplitude; Judge whether the first data fluctuation amplitude is within the first preset fluctuation threshold. If not, it means that the data acquisition strategy meets the first preset condition and triggers the first dynamic adjustment mechanism, which is configured to increase or decrease the acquisition frequency of the data acquisition strategy; When the data acquisition strategy meets the second preset condition, it triggers the second dynamic adjustment mechanism to update the data acquisition strategy, including: Judging one by one whether the first raw data is within the preset numerical range. If not, it means that the data acquisition strategy meets the second preset condition and triggers the second dynamic adjustment mechanism, which is configured to: Obtain multiple first raw data within the second preset time period up to now and plot the second data fluctuation amplitude; Judge whether the second data fluctuation amplitude is within the second preset fluctuation threshold; If so, record the first raw data that is not within the preset numerical range as an abnormal error value, and use the interpolation method to correct the abnormal error value; If not, record the first raw data that is not within the preset numerical range as an abnormal value, and generate an abnormal warning message according to the abnormal value and the second data fluctuation amplitude.
5. The automatic engineering cost review method based on big data according to claim 1, characterized in that Inputting the first data set into the cost prediction model to obtain the cost prediction coefficient, including: Obtain the first data set and perform time series analysis on the first data set to extract time series features, where the time series features include trend features, periodic features, and random fluctuation features; Conduct a stationarity test on the time series features. If the time series features do not meet the stationarity threshold, perform a differencing operation on the time series features until the time series features meet the stationarity threshold; Input the time series features into the trained ARIMA model to generate the first prediction result; Fuse the features of the first prediction result and the first data set to obtain the second data set. The feature fusion includes adding the first prediction result as a new feature to the first data set and performing standardization processing; Evaluate the feature importance of the second data set and screen out the features that have a significant impact on the prediction result, denoted as the second data features; Input the second data features into the trained LightGBM model to generate the second prediction result; Calculate the first prediction error of the first prediction result on the historical data, and calculate the second prediction error of the second prediction result on the historical data; Judge the magnitudes of the first prediction error and the second prediction error. If the first prediction error is smaller, increase the weight value corresponding to the first prediction error. If the second prediction error is smaller, increase the weight value corresponding to the second prediction error; Add the weighted first prediction result and the second prediction result to generate the cost prediction coefficient.
6. The automatic engineering cost review method based on big data according to claim 5, wherein, Inputting the time series features into the trained ARIMA model to generate the first prediction result, including: Generate the first model parameters of the ARIMA model according to the autocorrelation function and the partial autocorrelation function, where the first model parameters include the autoregressive order, the differencing order, and the moving average order; Train the ARIMA model according to the first model parameters and the first sample data, and predict the time series feature data after training is completed to generate a first prediction result; Input the second data feature into the trained LightGBM model to generate a second prediction result, including: Generate second model parameters based on the gradient boosting decision tree algorithm, where the second model parameters include a first learning rate, a first maximum depth of the tree, and a first number of leaf nodes; Iteratively train the LightGBM model according to the second model parameters and the second sample data, and predict the second data feature after training is completed to generate a second prediction result.
7. The automatic project cost review method based on big data according to claim 1, characterized in that Input the first data set into the risk prediction model to obtain the divided risk levels and avoidance suggestions. The risk prediction model is configured as an XGBoost model constructed with audit items as adjustment parameters, including: Generate audit items according to the project information, and extract features from the first data set according to the audit items to generate audit feature data. The audit items include contract compliance, cost overrun rate, project duration delay rate, and quality standard compliance; Perform data cleaning on the audit feature data to generate a third data set. The data cleaning includes missing value filling, outlier handling, and feature standardization; And, label the third sample data according to the audit items to generate training samples, where the labels include low risk, medium risk, and high risk; Construct an XGBoost model based on the gradient boosting algorithm, and initialize the third model parameters of the XGBoost model through cross-validation. The third model parameters include a second learning rate, a second maximum depth of the tree, a second number of leaf nodes, and a regularization parameter; Train the XGBoost model using the labeled third sample data, and predict the third data set after training is completed to generate a risk level classification result; Generate avoidance suggestions according to the risk level classification result in combination with a preset risk avoidance rule library.
8. The automatic engineering cost review method based on big data according to claim 1, wherein Input the first data set into the expected return model to obtain the expected return range, including: Extract feature data related to returns from the first data set and perform data cleaning to obtain a fourth data set. The feature data includes project budget, cost distribution, market volatility, and historical return data; Perform random sampling on the fourth data set based on the Monte Carlo simulation method to generate multiple return prediction scenarios; Input each return prediction scenario into the random forest model to obtain the corresponding return prediction value. The random forest model is configured to integrate prediction results through multiple decision trees; Perform statistical analysis on multiple return prediction values to generate an expected return range, where the expected return range includes the lowest return value, the highest return value, and the median expected return.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.
10. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data acquisition frequency control method and device and air conditioner system
CN108981069A
Combined model sales prediction method based on sales prediction system
CN112348571A
Prediction method and device for grassland land state and storage device
CN117131402A
Photovoltaic module income prediction model system based on big data model and method thereof
CN117634695A
Engineering cost management method and system based on big data
CN117808499A
Cited By
Engineering cost intelligent evaluation method based on graph neural network
CN121258187A