Water conservancy project cost prediction method and device, readable storage medium and equipment
Through the cost prediction model of water conservancy engineering with multi-source data sets and feature fusion layers, the problem of inability to capture nonlinear relationships and insufficient data fusion in the prior art is solved, and more accurate cost prediction and resource optimization management are achieved.
Patent Information
- Application Number
- CN202510490166.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-04
AI Technical Summary
The cost prediction method of water conservancy engineering in the prior art cannot effectively capture the complex nonlinear relationship between engineering parameters and cost, and fails to fully integrate multiple types of data, resulting in unreliable prediction results.
The multi-source data set is used to obtain engineering correlation data, and the engineering feature vectors are extracted and fused through the trained water conservancy project cost prediction model. The feature fusion layer is used to combine feature vectors from different sources, and combined with the multi-head self-attention mechanism to perform feature fusion to predict the engineering cost of the specified time interval.
It improves data processing speed, mines hidden information, improves prediction accuracy, helps to reasonably plan resource allocation, reduces unnecessary expenses, and provides scientific basis to make informed investment decisions and adapt to market changes.
Smart Images

Figure CN120258239A_ABST
Abstract
Description
Background Art
[0002] As a core area of national infrastructure construction, the cost prediction of water conservancy projects needs to consider multiple complex factors.
[0003] The cost analysis methods in the related technologies are generally based on linear regression models or empirical formulas. However, the above solutions cannot capture the complex non-linear relationship between engineering parameters and costs, and cannot fully integrate various types of data, resulting in unreliable prediction results.
[0004] In view of this, it is urgent to develop a new water conservancy project cost prediction method and device in this field.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a water conservancy project cost prediction method, a water conservancy project cost prediction device, a computer-readable storage medium and an electronic device, so as to at least overcome to some extent the technical problem of unreliable prediction results caused by the limitations of the related technologies.
[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be learned in part through the practice of the present disclosure.
[0008] According to a first aspect of the present disclosure, there is provided a water conservancy project cost prediction method, including:
[0009] Obtain a multi-source data set corresponding to the target water conservancy project; the multi-source data set contains engineering-related data from multiple sources;
[0010] Extract the engineering feature vectors corresponding to the engineering-related data of each source through a trained water conservancy project cost prediction model, and fuse the multiple engineering feature vectors corresponding to the multiple sources to obtain a fused feature vector;
[0011] Predict the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector through the trained water conservancy project cost prediction model.
[0012] In an exemplary embodiment of the present disclosure, after predicting the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector through the trained water conservancy project cost prediction model, the method further includes:
[0013] Obtain the deviation value between the project cost prediction data corresponding to the specified time interval and the actual project cost data;
[0014] In response to the deviation value being greater than a preset deviation threshold, optimize the trained water conservancy project cost prediction model.
[0015] In an exemplary embodiment of the present disclosure, the water conservancy project cost prediction model is obtained through training in the following manner:
[0016] Obtain a basic sample set; the basic sample set contains a plurality of basic multi-source data sets corresponding to a plurality of sample water conservancy projects; each basic multi-source data set contains basic sample data from multiple sources;
[0017] Divide the basic sample set into a training set, a validation set, and a test set according to preset data division conditions;
[0018] Use the training set to iteratively train the water conservancy project cost prediction model to be trained, so as to update the model parameters of the water conservancy project cost prediction model to be trained;
[0019] Use the validation set to validate the water conservancy project cost prediction model to be trained, so as to update the hyperparameters of the water conservancy project cost prediction model to be trained;
[0020] Use the test set to evaluate the model performance of the water conservancy project cost prediction model to be trained. When the model performance meets the preset performance requirements, obtain the trained water conservancy project cost prediction model.
[0021] In an exemplary embodiment of the present disclosure, the obtaining of the basic sample set includes:
[0022] Obtain an original sample set; the original sample set contains a plurality of original multi-source data sets corresponding to the plurality of sample water conservancy projects; each original multi-source data set contains original sample data from multiple sources;
[0023] Perform preprocessing on each original multi-source data set corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set.
[0024] In an exemplary embodiment of the present disclosure, the performing of preprocessing on each original multi-source data set corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set includes:
[0025] Perform preprocessing on each original multi-source data set to obtain a first data set;
[0026] Perform data cleaning on the first data set to obtain a second data set;
[0027] Perform data integration on the second data set to obtain a third data set;
[0028] Normalize the third data set to obtain the basic sample set.
[0029] In an exemplary embodiment of the present disclosure, the preprocessing of each original multi-source data set to obtain the first data set includes:
[0030] Perform missing value processing on the original sample data of each source in the original multi-source data set, and perform spatio-temporal alignment processing on the original sample data in the original multi-source data set to obtain the first data set.
[0031] In an exemplary embodiment of the present disclosure, the data cleaning of the first data set to obtain the second data set includes:
[0032] Obtain the outliers in the first data set;
[0033] Remove the outliers to obtain the second data set.
[0034] In an exemplary embodiment of the present disclosure, the data integration of the second data set to obtain the third data set includes:
[0035] Calculate composite engineering indicators based on the second data set, and perform vectorization processing on the text data in the second data set to obtain the third data set.
[0036] In an exemplary embodiment of the present disclosure, the data normalization of the third data set to obtain the basic sample set includes:
[0037] Perform normalization processing on the continuous variables in the third data set;
[0038] Perform interval discretization on the target sample data in the third data set; the target sample data includes non-linear sensitive sample data;
[0039] And perform encoding processing on the discrete variables in the third data set to obtain the basic sample set.
[0040] In an exemplary embodiment of the present disclosure, the encoding processing of the discrete variables in the third data set includes:
[0041] Perform sparse encoding on the first type of discrete variables, and splice a preset semantic vector on the basis of the obtained sparse vector;
[0042] Perform word embedding encoding processing on the second type of discrete variables;
[0043] Wherein, the number of categories included in the first type of discrete variables is less than a preset category threshold, and the number of categories included in the second type of discrete variables is greater than the preset category threshold.
[0044] In an exemplary embodiment of the present disclosure, the water conservancy project cost prediction model to be trained includes a feature extraction layer, a feature fusion layer, and an output layer;
[0045] Iteratively training the water conservancy project cost prediction model to be trained using the training set to update the model parameters of the water conservancy project cost prediction model to be trained includes:
[0046] Extracting the sample project feature vectors corresponding to the basic sample data of each source through the feature extraction layer corresponding to each source;
[0047] Fusing the multiple sample project feature vectors corresponding to the multiple sources through the feature fusion layer to obtain a sample fusion feature vector;
[0048] Predicting the project cost prediction sample data corresponding to each sample water conservancy project in a preset time interval based on the sample fusion feature vector through the output layer;
[0049] Updating the model parameters of the water conservancy project cost prediction model to be trained according to the difference value between the project cost prediction sample data and the actual project cost data label.
[0050] In an exemplary embodiment of the present disclosure, the basic sample data of the multiple sources includes engineering design sample data, geospatial environment sample data, building material price time series sample data, and construction log sample data;
[0051] The extracting the sample project feature vectors corresponding to the basic sample data of each source through the feature extraction layer corresponding to each source includes:
[0052] Extracting features from the engineering design sample data through a fully connected layer to obtain an engineering design sample vector;
[0053] Extracting features from the geospatial environment sample data through a convolutional neural network to obtain an environmental sample vector;
[0054] Extracting features from the building material price time series sample data through a recurrent neural network to obtain a price sample vector;
[0055] Extracting features from the construction log sample data through a natural language processing network to obtain a construction sample vector.
[0056] In an exemplary embodiment of the present disclosure, the fusing the multiple sample project feature vectors through the feature fusion layer to obtain a sample fusion feature vector includes:
[0057] Fuse the multiple sample engineering feature vectors based on the multi-head self-attention mechanism to obtain the fused feature vector.
[0058] In an exemplary embodiment of the present disclosure, after obtaining the trained water conservancy project cost prediction model, the method further includes:
[0059] When a preset trigger condition is satisfied, update the training set;
[0060] Based on the updated training set, perform model update on the trained water conservancy project cost prediction model.
[0061] In an exemplary embodiment of the present disclosure, the preset trigger condition includes any one of the following:
[0062] The difference degree between the first probability distribution of the newly added building material price time series data and the second probability distribution of the building material price time series data in the training set is greater than the preset difference degree range;
[0063] The fluctuation degree of the building material price within a preset time period exceeds the preset fluctuation range.
[0064] In an exemplary embodiment of the present disclosure, the performing model update on the trained water conservancy project cost prediction model based on the updated training set includes:
[0065] Fix the network parameters of some networks in the trained water conservancy project cost prediction model; the some networks include the fully connected layer, the convolutional neural network, and the natural language processing network;
[0066] Update the network parameters of the remaining networks other than the some networks.
[0067] According to a second aspect of the present disclosure, there is provided a water conservancy project cost prediction device, including:
[0068] A data acquisition module, configured to acquire a multi-source data set corresponding to a target water conservancy project; the multi-source data set contains engineering-related data from multiple sources;
[0069] A data processing module, configured to extract engineering feature vectors corresponding to the engineering-related data of each source through the trained water conservancy project cost prediction model, and fuse the multiple engineering feature vectors corresponding to the multiple sources to obtain a fused feature vector;
[0070] A prediction module, configured to predict the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector through the trained water conservancy project cost prediction model.
[0071] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for predicting the cost of a water conservancy project described in the first aspect is implemented.
[0072] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the water conservancy project cost prediction method described in the first aspect above by executing the executable instructions.
[0073] It can be seen from the above technical solutions that the water conservancy project cost prediction method, water conservancy project cost prediction device, computer-readable storage medium and electronic device in the exemplary embodiments of the present disclosure have at least the following advantages and positive effects:
[0074] In the technical solutions provided in some embodiments of the present disclosure, on the one hand, by integrating data from multiple sources, it is possible to more comprehensively capture various factors affecting the cost of water conservancy projects, thereby providing more accurate input data; further, by automatically extracting features from the data through a trained water conservancy project cost prediction model, and combining feature vectors from different sources together through a feature fusion layer, not only the speed of data processing is improved, but also it helps to dig out valuable information hidden in complex data, laying the foundation for subsequent accurate predictions; further, by using the fused feature vectors to predict the project cost prediction data for a specific time period, more dimensional influencing factors can be taken into account. Compared with traditional methods that only rely on a single data source or limited variables, this method can better reflect the actual situation and improve the accuracy of predictions; on the other hand, based on accurate cost predictions, it helps to rationally plan resource allocation, reduce unnecessary expenses, and improve the efficiency of capital use, providing a scientific basis for project managers, helping them make more informed investment decisions, select the best design scheme, and adjust construction plans in time to cope with possible changes, providing strong support for the sustainable and healthy development of the industry.
[0075] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0077] Figure 1Shows the schematic flowchart of the water conservancy project cost prediction method in the embodiments of the present disclosure;
[0078] Figure 2 Shows the schematic diagram of the training process of the engineering data processing model in the embodiments of the present disclosure;
[0079] Figure 3 Shows the schematic flowchart of how to obtain the basic sample set in the embodiments of the present disclosure;
[0080] Figure 4 Shows the schematic flowchart of how to preprocess each original multi-source data set corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set in the embodiments of the present disclosure;
[0081] Figure 5 Shows the schematic flowchart of how to iteratively train the water conservancy project cost prediction model to be trained with the training set to update the model parameters of the water conservancy project cost prediction model to be trained in the embodiments of the present disclosure;
[0082] Figure 6 Shows the overall architecture diagram of the water conservancy project cost prediction method in the embodiments of the present disclosure;
[0083] Figure 7 Shows the overall schematic flowchart of the water conservancy project cost prediction method in the embodiments of the present disclosure;
[0084] Figure 8 Shows the schematic structural diagram of the water conservancy project cost prediction device in the exemplary embodiments of the present disclosure;
[0085] Figure 9 Shows the schematic structural diagram of the electronic device in the exemplary embodiments of the present disclosure. Detailed implementation manners
[0086] Now, the exemplary embodiments will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0087] In this specification, the terms "a", "an", "the", and "said" are used to indicate the presence of one or more elements / components / etc.; the terms "comprising" and "having" are used to mean an open inclusion and refer to the presence of additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first", "second", etc. are only used as labels and do not limit the quantity of their objects.
[0088] In addition, the drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0089] In the related art, generally, linear regression models or artificial empirical formulas are used to predict water conservancy project cost data. However, the above solutions have at least the following problems:
[0090] First, the complex non-linear relationship between engineering parameters and cost is difficult to be captured by a linear model;
[0091] Second, the existing methods do not fully integrate multi-source heterogeneous data such as geographic information systems, meteorological data, and market dynamics, resulting in the omission of key influencing factors;
[0092] Third, the static model has a long update cycle and is difficult to respond to sudden market changes in a timely manner.
[0093] In an embodiment of the present disclosure, first, a water conservancy project cost prediction method is provided, which at least overcomes the defect of low reliability of prediction results in the related art to a certain extent.
[0094] Figure 1 The flowchart of the water conservancy project cost prediction method in the embodiment of the present disclosure is shown. The execution subject of the water conservancy project cost prediction method can be a server for predicting water conservancy project cost.
[0095] Refer to Figure 1 , according to an embodiment of the present disclosure, the water conservancy project cost prediction method includes the following steps:
[0096] Step S110, obtaining a multi-source data set corresponding to the target water conservancy project; the multi-source data set contains engineering-related data from multiple sources;
[0097] Step S120, extracting the engineering feature vectors corresponding to the engineering-related data of each source through a trained water conservancy project cost prediction model, and fusing the multiple engineering feature vectors corresponding to the multiple sources to obtain a fused feature vector;
[0098] Step S130: Based on the trained water conservancy project cost prediction model, predict the project cost prediction data corresponding to the target water conservancy project for a specified time interval using the fused feature vector.
[0099] In Figure 1 In the technical solution provided by the illustrated embodiment, on the one hand, by integrating data from multiple sources, various factors affecting the water conservancy project cost can be captured more comprehensively, thus providing more accurate input data. Further, the trained water conservancy project cost prediction model automatically extracts features from the data, and combines feature vectors from different sources through the feature fusion layer, which not only improves the speed of data processing, but also helps to discover valuable information hidden in complex data, laying a foundation for accurate subsequent prediction. Further, by using the fused feature vector to predict the project cost prediction data for a specific time interval, more dimensional influencing factors can be considered. Compared with traditional methods that only rely on a single data source or limited variables, this method can better reflect the actual situation and improve the prediction accuracy. On the other hand, accurate cost prediction helps to reasonably plan resource allocation, reduce unnecessary expenses, improve the efficiency of fund use, provides a scientific basis for project managers, helps them make more informed investment decisions, select the optimal design scheme, and timely adjust the construction plan to cope with possible changes, providing strong support for the sustainable and healthy development of the industry.
[0100] The following Figure 1 elaborates in detail the specific implementation process of each step in
[0101] First of all, it should be noted that the water conservancy projects in this disclosure can be reservoirs, levee projects, etc., which can be set according to the actual situation, and this disclosure does not make special limitations on this.
[0102] Next, before step S110, the training process of the water conservancy project cost prediction model in this disclosure is first described. Refer to Figure 2 , Figure 2 which shows the schematic diagram of the training process of the engineering data processing model in the embodiment of this disclosure, including steps S201 - S205:
[0103] In step S201, obtain the basic sample set.
[0104] In this step, the basic sample set can be obtained. The basic sample set can include multiple basic multi-source data sets corresponding to multiple sample water conservancy projects, that is, each sample water conservancy project can be associated with a basic multi-source data set, and each basic multi-source data set contains basic sample data from multiple sources.
[0105] The above basic sample data can be water conservancy project data related to each sample water conservancy project and sourced from multiple dimensions. Exemplarily, it can be sourced from the following four data sources: engineering design data, geospatial environment data, building material price data, and construction log data, which can be set according to actual circumstances, and the present disclosure does not make special limitations thereon.
[0106] Reference Figure 3 , Figure 3 FIG. shows a flowchart of how to obtain a basic sample set in an embodiment of the present disclosure, including steps S301 - S302:
[0107] In step S301, an original sample set is obtained.
[0108] In this step, an original sample set can be obtained. The original sample set may include multiple original multi-source data sets corresponding to multiple sample water conservancy projects, and each original multi-source data set may include original sample data from multiple sources.
[0109] Exemplarily, each original multi-source data set may include the following types of data:
[0110] ① Engineering design data: Structured engineering design data can be extracted through a BIM model (Building Information Modeling), engineering design documents, and engineering design drawings. Exemplarily, the engineering design data may include dam height, reservoir capacity, concrete strength grade, volume of earthwork and stonework (in 10,000 m 3 ), dam type (gravity dam / arch dam), reservoir capacity, and construction technology, etc., which can be set according to actual circumstances, and the present disclosure does not make special limitations thereon; after extracting the engineering design data, it can be stored in a relational database such as MySQL.
[0111] ② Geospatial environment data: Geospatial environment data can be obtained through a GIS platform (Geographic Information System), meteorological satellites, and drone aerial photography data, etc. Exemplarily, the geospatial environment data may include altitude, terrain slope, depth of the groundwater level, and historical rainfall data provided by the meteorological bureau; after extracting the geospatial environment data, it can be stored in a GeoJSON file in raster form (a format for encoding various geographical data structures. It is based on JSON (JavaScript Object Notation) and thus has the characteristics of being lightweight and easy to read).
[0112] ③ Building material price data: The building material price data can be extracted or captured in real time from building material trading platforms and publicly available data of statistical bureaus. Exemplarily, the building material price data may include steel / cement prices, mechanical rental fee indices, CPI fluctuations, etc.; after the building material price data is extracted, it can be stored in a time series database such as InfluxDB according to the time stamp.
[0113] ④ Construction log data: Key texts can be extracted from the construction management system or geological exploration reports to obtain construction log data. Exemplarily, the construction log data may include formation lithology, karst development, suspension of work during the flood season, etc., which can be set according to the actual situation, and the present disclosure does not make special limitations thereon.
[0114] In step S302, preprocessing is performed on each original multi-source data set corresponding to each sample water conservancy project in the original sample set to obtain a basic sample set.
[0115] In this step, preprocessing can be performed on each original multi-source data set corresponding to each sample water conservancy project in the above-mentioned original sample set to obtain a basic sample set. Among them, preprocessing is to ensure the quality, consistency, and applicability of the data, so as to provide a reliable basis for subsequent data analysis and model training.
[0116] Exemplarily, refer to Figure 4 , Figure 4 which shows a schematic flow chart of how to perform preprocessing on each original multi-source data set corresponding to each sample water conservancy project in the original sample set of the embodiment of the present disclosure to obtain a basic sample set, including step S401-step S404:
[0117] In step S401, preprocessing is performed on each original multi-source data set to obtain a first data set.
[0118] In this step, missing value processing can be performed on the original sample data of each source in the above-mentioned original multi-source data set, and spatio-temporal alignment processing can be performed on the original sample data in the above-mentioned original multi-source data set to obtain a first data set.
[0119] Optionally, for missing numerical values, the spatio-temporal KNN (Spatio-Temporal K-Nearest Neighbors) interpolation method can be used to fill in the data. For example, assuming that a certain sample water conservancy project is missing the "groundwater depth", the groundwater depth data of 5 similar sample water conservancy projects in the same basin can be weighted for filling; for missing texts, the BERT model can be used to predict the missing keywords, or, alternatively, the missing keywords can be extracted and supplemented from the relevant data of the same type of project based on keyword matching, which can be set according to the actual situation, and the present disclosure does not make special limitations thereon.
[0120] Optionally, the original sample data in the above original multi-source dataset can be subjected to spatio-temporal alignment processing. For example, the GIS coordinates can be matched with the engineering pile number system to establish a spatial grid (the resolution can be 10m×10m), and the rainfall, geological data, and engineering design data within the same grid can be associated. The timestamp can also be unified with a specified date (e.g., the project start date) as the initial date, and the building material time series data can be interpolated and aligned at a frequency of weeks or months. All of these can be set according to the actual situation, and the present disclosure does not make special limitations on this.
[0121] In step S402, the first dataset is subjected to data cleaning to obtain a second dataset.
[0122] In this step, after the first dataset is obtained through preprocessing, the first dataset can be subjected to data cleaning to obtain a second dataset. Specifically, outliers in the first dataset can be removed (e.g., removing outlier samples with a cost deviation exceeding ±30% based on the 3σ principle) to obtain the second dataset.
[0123] In step S403, the second dataset is subjected to data integration to obtain a third dataset.
[0124] In this step, the second dataset can be subjected to data integration to obtain a third dataset. Specifically, composite engineering indicators can be calculated based on the second dataset. For example, for each sample water conservancy project, the concrete consumption per unit reservoir capacity, the transportation difficulty coefficient (distance × slope), the construction efficiency index (engineering quantity / number of mechanical shifts labor input), etc. can be calculated; at the same time, the text data in the second dataset can also be vectorized. For example, the BERT model can be used to encode the risk descriptions in the construction logs of each sample water conservancy project into 128-dimensional vectors to obtain the third dataset.
[0125] In step S404, the third dataset is subjected to data standardization to obtain a basic sample set.
[0126] In this step, the third dataset can be subjected to data standardization to obtain a basic sample set.
[0127] Specifically, first, the continuous variables in the third dataset can be normalized. For example, for variables such as price and engineering quantity, the Z-Score algorithm can be used for normalization.
[0128] Secondly, interval discretization can be performed on the target sample data in the above-mentioned third dataset. The target sample data may include non-linear sensitive sample data. Non-linear sensitive sample data refers to data whose variation law is not a simple linear relationship under different conditions. Such data usually exhibits complex patterns or depends on the interaction of multiple factors. Exemplarily, the non-linear sensitive sample data can be seismic intensity, so that the seismic intensity can be discretized into ordered categories according to intervals;
[0129] Thirdly, encoding processing can be performed on the discrete variables in the third dataset to obtain a basic sample set. Specifically, the discrete variables may include first-type discrete variables and second-type discrete variables. The first-type discrete variables can also be called discrete variables with a low number of categories, that is, the specific categories it contains are less than a preset category threshold (for example: 10, which can also be set according to the actual situation, and the present disclosure does not make special limitations in this regard). The second-type discrete variables can also be called discrete variables with a high number of categories, that is, the specific categories it contains are greater than the preset category threshold.
[0130] Thus, for the above-mentioned first-type discrete variables, such as: dam type, encoding processing can be performed on it based on a hybrid encoding method. Specifically, sparse encoding can be performed on it first, such as: One-Hot encoding. After that, a preset semantic vector can be concatenated after the obtained sparse vector. For example, the preset semantic vector corresponding to the dam type "concrete gravity dam" can be set to (1, 0, 0), and the preset semantic vector corresponding to "concrete arch dam" can be set to (0, 1, 0). The preset semantic vector can be configured according to engineering experience, and the present disclosure does not make special limitations in this regard.
[0131] For the above-mentioned second-type discrete variables, a word embedding encoding network can be pre-trained, and then, the word embedding encoding network is used to perform word embedding encoding processing on it. The embedding dimension can be expressed as
[0132] After obtaining the basic sample set through the above-mentioned preliminary processing, reference can be made to Figure 2 In step S202, the basic sample set is divided into a training set, a validation set, and a test set according to preset data division conditions.
[0133] In this step, the basic sample set can be divided into a training set, a validation set, and a test set according to preset data division conditions. Exemplarily, the sample set can be divided according to the following data ratios:
[0134] Training set: 70% of historical water conservancy project data (uniform sampling can be performed according to the regions where the projects are located to avoid regional biases);
[0135] Validation set: 15% of recent water conservancy project data (which can include cases of price mutations to test the dynamic response ability of the model).
[0136] Test set: 15% of cross-regional water conservancy project data (to verify the generalization ability of the model).
[0137] In step S203, the training set is used to iteratively train the water conservancy project cost prediction model to be trained, so as to update the model parameters of the water conservancy project cost prediction model to be trained.
[0138] In this step, the training set can be used to iteratively train the water conservancy project cost prediction model to be trained, so as to update the model parameters of the water conservancy project cost prediction model to be trained.
[0139] Specifically, the water conservancy project cost prediction model to be trained can be built first. This model can include an input layer, a feature extraction layer, a feature fusion layer, and an output layer. The input layer is used to receive input data from multiple sources. The feature extraction layer is used to extract the respective features corresponding to the input data of each source. The feature fusion layer is used to fuse the multiple features extracted. The output layer (i.e., the fully connected layer, 1 node, linear activation) is used to output the prediction result based on the fused features.
[0140] Exemplarily, the above feature extraction layer can include four feature extraction layers corresponding to the above four sources of data: a fully connected layer (256 nodes, ReLU activation) for feature extraction of engineering design sample data, a lightweight convolutional neural network CNN (MobileNetV3, outputting 64-dimensional features) for feature extraction of geospatial environment sample data, a recurrent neural network GRU (128 hidden units, tanh activation) for feature extraction of building material price time series sample data, and a natural language processing network BERT (outputting 128-dimensional vectors) for feature extraction of construction log sample data.
[0141] After building the water conservancy project cost prediction model to be trained, reference can be made to Figure 5 , Figure 5 which shows a schematic flowchart of how to use the training set to iteratively train the water conservancy project cost prediction model to be trained to update the model parameters of the water conservancy project cost prediction model to be trained in the embodiments of the present disclosure, including step S501 - step S504:
[0142] In step S501, the sample project feature vectors corresponding to the basic sample data of each source are extracted through the feature extraction layer corresponding to each source.
[0143] In this step, the sample engineering feature vectors corresponding to the basic sample data of each source can be extracted through the feature extraction layer corresponding to each source. Specifically, the basic sample data of multiple sources includes engineering design sample data, geospatial environment sample data, building material price time series sample data, and construction log sample data. Therefore, the engineering design sample vectors can be obtained by extracting features from the engineering design sample data through a fully connected layer; the environmental sample vectors can be obtained by extracting features from the geospatial environment sample data through a convolutional neural network CNN; the price sample vectors can be obtained by extracting features from the building material price time series sample data through a recurrent neural network GRU; and the construction sample vectors can be obtained by extracting features from the construction log sample data through a natural language processing network BERT.
[0144] Among them, GRU controls the flow of information through a gating mechanism, and its core gating mechanism includes a reset gate, an update gate, a candidate hidden state, and a final hidden state, where:
[0145] Reset Gate: Referring to the following formula 1, the reset gate is used to control the degree of forgetting of historical information. It calculates a value through the Sigmod function, and this value ranges from 0 to 1, indicating the influence degree of the combination of the previous hidden state h j-1 and the current input t j on the candidate hidden state:
[0146] r j = σ(W r [h j-1 , t j + b j ) Formula 1
[0147] Among them, r j represents the above reset gate, W r is the weight matrix, and b j is the bias term.
[0148] Update Gate: Referring to the following formula 2, the update gate is used to control the proportion of new information incorporated. It also calculates a value through the Sigmod function, and this value ranges from 0 to 1, indicating the degree of fusion between new information and old information h j-1 :
[0149] z j = σ(W z [h j-1 , t j + b z ) Formula 2
[0150] Among them, z j represents the above update gate, W zis the weight matrix, and b z is the bias term.
[0151] Candidate hidden state: Referring to Equation 3 below, the candidate hidden state is based on the reset gate r j to adjust the previous hidden state h j-1 and the current input t j , and the hyperbolic tangent function tanh is used to ensure that the output is between -1 and 1:
[0152]
[0153] where represents new information, W h is the weight matrix, ⊙ is element-wise multiplication, and b h is the bias term.
[0154] Final hidden state: Referring to Equation 4 below, the fusion of new and old information is controlled by the update gate z j If z j is close to 1, more reliance is placed on new information If it is close to 0, more reliance is placed on the old information h j-1 :
[0155]
[0156] where h j is the final hidden state and z j is the output of the update gate.
[0157] The above equations work together to enable the GRU to effectively handle long-term dependencies in time series data while avoiding the vanishing gradient problem in traditional RNNs.
[0158] In step S502, the feature fusion layer fuses multiple sample engineering feature vectors corresponding to multiple sources to obtain a sample fusion feature vector.
[0159] In this step, the feature fusion layer can fuse the multiple sample engineering feature vectors corresponding to the above multiple sources to obtain a sample fusion feature vector. Specifically, the feature fusion layer can fuse the above multiple sample engineering feature vectors based on the multi-head attention mechanism (for example: 8 heads, which can be set according to the actual situation, and the present disclosure does not make special limitations thereon) to obtain a 512-dimensional sample fusion feature vector.
[0160] Specifically, first, the feature fusion layer can concatenate the above multiple sample engineering feature vectors to form a concatenated vector X = [X1, X2,......X n , where X iRepresents the embedding vector of the i-th feature.
[0161] The multi-head self-attention mechanism can include multiple attention heads. Thus, for each attention head (corresponding to a set of query weight matrices W Q , key weight matrices W k and value weight matrices W V ), for each feature X i , its query vector Q i , key vector K i and value vector V i can be calculated based on the following Formulas 5-7:
[0162] Q i = X i W Q Formula 5
[0163] K i = X i W k Formula 6
[0164] V i = X i W V Formula 7
[0165] where is a learnable weight matrix, d in is the dimension of the input vector, and d k is the dimension scaling factor.
[0166] After that, the attention weights Attention(Q i , K i , V i ) corresponding to each attention head can be calculated based on the following Formula 8:
[0167]
[0168] Furthermore, the attention weights corresponding to each attention head can be multiplied by the value vector V i to obtain the output matrix head i corresponding to each attention head. Next, the output matrices of multiple attention heads can be concatenated based on the following Formula 9 to obtain the global feature matrix H trans corresponding to each type of source data of each sample's water conservancy project:
[0169] H trans = Concat(head1,…,head h )W o Formula 9
[0170] where H transis the above-mentioned global feature matrix, is the output projection matrix, dmodel is a hyperparameter, representing the feature dimension inside the model or the dimension of the hidden layer. Specifically, dmodel is the dimension of the input and output vectors, and also the dimension of the main operation space of each layer in the model (including the self-attention mechanism and the feed-forward neural network).
[0171] After that, the global feature matrices from the above multiple sources can be concatenated based on the following formula 10 to obtain the sample feature fusion vector F fusion :
[0172] F fusion = Concat(Mean Pooling(H trans )) Formula 10
[0173] where F fusion is the above-mentioned sample feature fusion vector, Mean Pooling represents mean pooling, and Concat represents feature concatenation.
[0174] In step S503, based on the sample fusion feature vector, the output layer predicts the project cost prediction sample data corresponding to each sample water conservancy project for a preset time interval.
[0175] In this step, the output layer can predict the project cost prediction sample data corresponding to each sample water conservancy project for a preset time interval (for example: a certain day or a certain week, which can be set according to the actual situation, and this disclosure does not make special limitations). Exemplarily, the Monte Carlo Dropout sampling method can be used to predict the project cost prediction sample data corresponding to the preset time interval.
[0176] In step S504, according to the difference value between the project cost prediction sample data and the actual project cost data label, the model parameters of the water conservancy project cost prediction model to be trained are updated.
[0177] In this step, the model parameters of the water conservancy project cost prediction model to be trained can be updated according to the difference value between the project cost prediction sample data and the actual project cost data label. Specifically, the early stopping method can be used to iteratively train the above-mentioned water conservancy project cost prediction model to be trained, and monitor the change of the above difference value until the above difference value converges (for example: no improvement in 15 consecutive rounds of training is considered convergence).
[0178] Then refer to Figure 2 In step S204, the validation set is used to validate the water conservancy project cost prediction model to be trained, so as to update the hyperparameters of the water conservancy project cost prediction model to be trained.
[0179] In this step, the validation set can be used to validate the above-mentioned water conservancy project cost prediction model to be trained, so as to update the hyperparameters of the water conservancy project cost prediction model to be trained. Among them, the hyperparameters can be the learning rate, batch size, regularization coefficient, etc., which can be set according to the actual situation, and the present disclosure does not make special limitations thereto. Exemplarily, during the model validation process, 20% of the input features can be randomly masked to verify the fault tolerance ability of the model. If the error increase is less than 8%, it can be determined that the validation is passed.
[0180] In step S205, the test set is used to evaluate the model performance of the water conservancy project cost prediction model to be trained. When the model performance meets the preset performance requirements, the trained water conservancy project cost prediction model is obtained.
[0181] In this step, the test set can be used to evaluate the model performance of the water conservancy project cost prediction model to be trained. When the model performance meets the preset performance requirements, the trained water conservancy project cost prediction model is obtained. Exemplarily, the evaluation indexes of the model performance can adopt MAE (Mean Absolute Error), RMSE (Root Mean Squared Error), R 2 (Coefficient of Determination), etc., which can be set according to the actual situation, and the present disclosure does not make special limitations thereto.
[0182] After obtaining the trained water conservancy project cost prediction model, in view of the fact that the building material prices may change in real time, thus, the present disclosure also provides a scheme for dynamically incremental learning of the trained water conservancy project cost prediction model. Dynamically incremental learning means that after the model training is completed, as time goes by and new data arrives, the model parameters are gradually updated to maintain its accuracy and adaptability, rather than retraining the entire model from scratch. This method enables the model to effectively utilize historical data and can quickly adapt to new data or changes. Specifically, a preset trigger condition can be set. When the preset trigger condition is met, the training set can be updated. After that, the trained water conservancy project cost prediction model can be updated based on the updated training set.
[0183] Exemplarily, the above-mentioned preset trigger conditions can include: the difference degree between the first probability distribution of the newly added building material price time series data and the second probability distribution of the building material price time series data in the training set is greater than the preset difference degree range (KL divergence is greater than 0.1); the fluctuation degree of the building material price within the preset time period exceeds the preset fluctuation range (for example: the single-day fluctuation of the building material price is greater than 10%); changes in engineering-related policies or regulations, etc., which can be set according to the actual situation, and the present disclosure does not make special limitations thereto.
[0184] When the above preset trigger conditions are met, the training set can be updated. Specifically, the time series sample data of building material prices in the training set can be updated based on the sliding window mechanism. For example, the 50 earliest building material price data corresponding to each sample of the water conservancy project in the above training set can be eliminated, and the training set of building material price time series samples can be updated with the newly added 200 latest building material price sample data.
[0185] After updating the training set, the network parameters of some networks in the trained water conservancy project cost prediction model can be fixed. Some networks can include the above-mentioned fully connected layer, convolutional neural network, and natural language processing network. The network parameters of the remaining networks except the above-mentioned part of the networks are updated, that is, only the network parameters of the recurrent neural network GRU, feature fusion layer, and output layer are updated. A verification mechanism can be set during the model update training process, and the retained samples selected from the newly added samples are used to verify the model stability. When the stability error fluctuation is less than 3%, the update training can be stopped to obtain the updated water conservancy project cost prediction model.
[0186] After obtaining the trained water conservancy project cost prediction model, the model application stage can be entered, and then refer to Figure 1 , in step S110, a multi-source data set corresponding to the target water conservancy project is obtained.
[0187] In this step, the target water conservancy project can be a water conservancy project for which the project cost needs to be predicted. Therefore, the multi-source data corresponding to the target water conservancy project can be obtained. The multi-source data set can contain project-related data from multiple sources. Specifically, it can be engineering design data, geospatial environment data, building material price time series data, and construction log data.
[0188] Exemplarily, the engineering design parameters and construction log data can be obtained through a Web form uploaded by an API (Application Programming Interface, an interactive element set embedded in a web page that allows users to input data and submit it to the server for processing); the geospatial environment data can be obtained by parsing the input GIS file. It should be noted that if the target water conservancy project is located in an earthquake zone, the geospatial environment data can also include recent earthquake monitoring data; the building material price time series data can be captured in real time from a building material trading platform, or the building material price time series data can also be input through an API, which can be set according to the actual situation, and the present disclosure does not make special limitations on this.
[0189] In step S120, the trained water conservancy project cost prediction model is used to extract the project feature vectors corresponding to the project-related data of each source, and multiple project feature vectors corresponding to multiple sources are fused to obtain a fused feature vector.
[0190] In this step, after obtaining the multi-source dataset corresponding to the target water conservancy project, the trained water conservancy project cost prediction model can be used to extract the project feature vectors corresponding to the project-related data of each source, and multiple project feature vectors corresponding to the above-mentioned multiple sources are fused to obtain a fused feature vector.
[0191] In step S130, the trained water conservancy project cost prediction model is used to predict the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector.
[0192] In this step, the above-mentioned trained water conservancy project cost prediction model can, according to the Monte Carlo Dropout sampling method, predict the project cost prediction data corresponding to the above-mentioned target water conservancy project for a specified time interval and the confidence interval corresponding to the project cost prediction data. At the same time, it can also output a feature importance analysis report, which can visually display the attention weight matrix corresponding to each project feature vector calculated in the intermediate processing process of the model, so as to improve the interpretability of the model.
[0193] It should be noted that after obtaining the project cost prediction data corresponding to the specified time interval, the present disclosure can also obtain the deviation value between the project cost prediction data corresponding to the specified time interval and the actual project cost data. If the deviation value is greater than the preset deviation threshold, the above-mentioned trained water conservancy project cost prediction model can be triggered to optimize, so as to improve the prediction accuracy of the model.
[0194] Reference Figure 6 , Figure 6 shows the overall architecture diagram of the water conservancy project cost prediction method in the embodiments of the present disclosure, as Figure 6 shown:
[0195] First, project-related data from four sources can be collected, including engineering design sample data, geospatial environment sample data, building material price time series sample data, and construction log sample data;
[0196] After that, the collected data is preprocessed and aligned;
[0197] Next, the fully connected layer can be used to extract features from the engineering design sample data to obtain an engineering design sample vector; the convolutional neural network CNN (specifically, it can be CNNMoblileNet V3) can be used to extract features from the geospatial environment sample data to obtain an environmental sample vector; the recurrent neural network GRU can be used to extract features from the building material price time series sample data to obtain a price sample vector; the natural language processing network BERT can be used to extract features from the construction log sample data to obtain a construction sample vector;
[0198] After that, the feature fusion layer can calculate the attention weight matrix corresponding to each of the above four types of vectors based on the multi-head self-attention mechanism, and perform feature splicing to obtain a fused feature vector;
[0199] Next, the output layer can perform cost prediction based on the fused feature vector;
[0200] In addition, the present disclosure also provides a dynamic update and feedback mechanism for the model, that is, when there is a large fluctuation in building material prices, the model can be triggered to be forcibly updated;
[0201] At the same time, the present disclosure can also display a feature importance bar chart according to the attention weight matrix calculated during the model processing process (each element in the matrix represents the importance parameter corresponding to each input feature of the model), so as to improve the interpretability of the model output result.
[0202] Reference Figure 7 , Figure 7 shows the overall flowchart of the water conservancy project cost prediction method in the embodiments of the present disclosure, including step S701-step S706:
[0203] In step S701, data collection;
[0204] In step S702, data cleaning and alignment;
[0205] In step S703, data feature extraction;
[0206] In step S704, use the trained water conservancy project cost prediction model for dynamic prediction;
[0207] In step S705, abnormal detection of building material price changes can be performed or abnormal feedback information for the model can be received;
[0208] In step S706, incremental update of the model is performed.
[0209] Based on the above technical solutions, the present disclosure can at least achieve the following technical effects:
[0210] First, by integrating data from multiple sources (such as meteorological data, geographic information, historical engineering data, etc.), it is possible to more comprehensively capture various factors affecting water conservancy projects, thereby providing more accurate input data; using machine learning or deep learning models to automatically extract features from raw data, and combining feature vectors from different sources through a feature fusion layer, not only improves the speed of data processing, but also helps to dig out valuable information hidden in complex data, laying the foundation for subsequent accurate predictions;
[0211] Second, by using the fused feature vector to predict the project cost in a specified time interval, more influencing factors can be taken into account, thereby improving the accuracy of the prediction results. Compared with the traditional solution that only relies on a single data source or limited variables, the method in the present disclosure can better reflect the actual situation and improve the prediction accuracy;
[0212] Third, accurate cost forecasts help to rationally plan resource allocation, reduce unnecessary expenses, and improve the efficiency of capital use. They provide a scientific basis for project managers, helping them make more informed investment decisions, choose the best design solutions, and adjust construction plans in a timely manner to cope with possible changes, providing strong support for the sustained and healthy development of the industry.
[0213] The present disclosure also provides a water conservancy project cost prediction device, Figure 8 FIG. 2 shows a schematic diagram of the structure of a water conservancy project cost prediction device in an exemplary embodiment of the present disclosure; Figure 8 As shown, the water conservancy project cost prediction device 800 may include a data acquisition module 810, a data processing module 820 and a prediction module 830. Among them:
[0214] The data acquisition module 810 is used to acquire a multi-source data set corresponding to a target water conservancy project; the multi-source data set includes project-related data from multiple sources;
[0215] The data processing module 820 is used to extract the engineering feature vector corresponding to the engineering related data from each source through the trained water conservancy project cost prediction model, and to fuse the multiple engineering feature vectors corresponding to the multiple sources to obtain a fused feature vector;
[0216] The prediction module 830 is used to predict the engineering cost prediction data of the target water conservancy project corresponding to the specified time interval based on the fused feature vector by using the trained water conservancy project cost prediction model.
[0217] In an exemplary embodiment of the present disclosure, after predicting the engineering cost prediction data of the target water conservancy project corresponding to the specified time interval based on the fused feature vector by using the trained water conservancy project cost prediction model, the prediction module 830 is configured as follows:
[0218] Obtain the deviation value between the predicted project cost data and the actual project cost data corresponding to the specified time interval;
[0219] In response to the deviation value being greater than a preset deviation threshold, optimize the trained water conservancy project cost prediction model.
[0220] In an exemplary embodiment of the present disclosure, the above device further includes a model training module 840, which is configured to obtain a trained water conservancy project cost prediction model in the following manner:
[0221] Obtain a basic sample set; the basic sample set contains multiple basic multi-source data sets corresponding to multiple sample water conservancy projects; each basic multi-source data set contains basic sample data from multiple sources;
[0222] Divide the basic sample set into a training set, a validation set, and a test set according to preset data division conditions;
[0223] Use the training set to iteratively train the water conservancy project cost prediction model to be trained, so as to update the model parameters of the water conservancy project cost prediction model to be trained;
[0224] Use the validation set to validate the water conservancy project cost prediction model to be trained, so as to update the hyperparameters of the water conservancy project cost prediction model to be trained;
[0225] Use the test set to evaluate the model performance of the water conservancy project cost prediction model to be trained, and obtain a trained water conservancy project cost prediction model when the model performance meets the preset performance requirements.
[0226] In an exemplary embodiment of the present disclosure, the model training module 840 obtains a basic sample set, including:
[0227] Obtain an original sample set; the original sample set contains multiple original multi-source data sets corresponding to the multiple sample water conservancy projects; each original multi-source data set contains original sample data from multiple sources;
[0228] Perform preprocessing on each original multi-source data set corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set.
[0229] In an exemplary embodiment of the present disclosure, the model training module 840 performs preprocessing on each original multi-source data set corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set, including:
[0230] Perform preprocessing on each original multi-source data set to obtain a first data set;
[0231] Clean the first data set to obtain a second data set;
[0232] Integrate the second data set to obtain a third data set;
[0233] Normalize the third data set to obtain the basic sample set.
[0234] In an exemplary embodiment of the present disclosure, the model training module 840 preprocesses each of the original multi-source data sets to obtain a first data set, including:
[0235] Process missing values for the original sample data of each source in the original multi-source data set, and perform spatio-temporal alignment processing on the original sample data in the original multi-source data set to obtain the first data set.
[0236] In an exemplary embodiment of the present disclosure, the cleaning of the first data set to obtain a second data set includes:
[0237] Obtain the outliers in the first data set;
[0238] Remove the outliers to obtain the second data set.
[0239] In an exemplary embodiment of the present disclosure, the model training module 840 integrates the second data set to obtain a third data set, including:
[0240] Calculate composite engineering metrics based on the second data set, and perform vectorization processing on the text data in the second data set to obtain the third data set.
[0241] In an exemplary embodiment of the present disclosure, the model training module 840 normalizes the third data set to obtain the basic sample set, including:
[0242] Normalize the continuous variables in the third data set;
[0243] Discretize the target sample data in the third data set; the target sample data includes non-linear sensitive sample data;
[0244] And encode the discrete variables in the third data set to obtain the basic sample set.
[0245] In an exemplary embodiment of the present disclosure, the model training module 840 encodes the discrete variables in the third data set, including:
[0246] Sparsely encode the first type of discrete variable, and splice a preset semantic vector based on the obtained sparse vector;
[0247] Perform word embedding encoding on the second type of discrete variable;
[0248] Wherein, the number of categories included in the first type of discrete variable is less than a preset category threshold, and the number of categories included in the second type of discrete variable is greater than the preset category threshold.
[0249] In an exemplary embodiment of the present disclosure, the water conservancy project cost prediction model to be trained includes a feature extraction layer, a feature fusion layer, and an output layer;
[0250] The model training module 840 uses the training set to iteratively train the water conservancy project cost prediction model to be trained, so as to update the model parameters of the water conservancy project cost prediction model to be trained, including:
[0251] Extract the sample project feature vectors corresponding to the basic sample data of each source through the feature extraction layer corresponding to each source;
[0252] Fuse the multiple sample project feature vectors corresponding to the multiple sources through the feature fusion layer to obtain a sample fusion feature vector;
[0253] Predict the project cost prediction sample data corresponding to each sample water conservancy project for a preset time interval based on the sample fusion feature vector through the output layer;
[0254] Update the model parameters of the water conservancy project cost prediction model to be trained according to the difference value between the project cost prediction sample data and the actual project cost data label.
[0255] In an exemplary embodiment of the present disclosure, the basic sample data of the multiple sources includes engineering design sample data, geospatial environment sample data, building material price time series sample data, and construction log sample data;
[0256] The model training module 840 extracts the sample project feature vectors corresponding to the basic sample data of each source through the feature extraction layer corresponding to each source, including:
[0257] Extract features from the engineering design sample data through a fully connected layer to obtain an engineering design sample vector;
[0258] Extract features from the geospatial environment sample data through a convolutional neural network to obtain an environmental sample vector;
[0259] Extract features from the building material price time series sample data through a recurrent neural network to obtain a price sample vector;
[0260] Feature extraction is performed on the construction log sample data through a natural language processing network to obtain a construction sample vector.
[0261] In an exemplary embodiment of the present disclosure, the model training module 840 fuses the multiple sample project feature vectors through the feature fusion layer to obtain a sample fusion feature vector, including:
[0262] Fusing the multiple sample project feature vectors based on the multi-head self-attention mechanism to obtain the fusion feature vector.
[0263] In an exemplary embodiment of the present disclosure, after obtaining the trained water conservancy project cost prediction model, the model training module 840 is configured to:
[0264] Update the training set when a preset trigger condition is met;
[0265] Based on the updated training set, perform model update on the trained water conservancy project cost prediction model.
[0266] In an exemplary embodiment of the present disclosure, the preset trigger condition includes any one of the following:
[0267] The difference degree between the first probability distribution of the newly added building material price time series data and the second probability distribution of the building material price time series data in the training set is greater than the preset difference degree range;
[0268] The fluctuation degree of the building material price within a preset time period exceeds the preset fluctuation range.
[0269] In an exemplary embodiment of the present disclosure, the model training module 840 performs model update on the trained water conservancy project cost prediction model based on the updated training set, including:
[0270] Fix the network parameters of some networks in the trained water conservancy project cost prediction model; the some networks include the fully connected layer, the convolutional neural network, and the natural language processing network;
[0271] Update the network parameters of the remaining networks except the some networks.
[0272] The specific details of each module in the above water conservancy project cost prediction device have been described in detail in the corresponding water conservancy project cost prediction method, so they will not be elaborated here.
[0273] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0274] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0275] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.
[0276] The present disclosure also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or can exist separately without being assembled into the electronic device.
[0277] A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0278] A computer-readable storage medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0279] The computer-readable storage medium carries one or more programs that, when executed by the electronic device, cause the electronic device to implement the method described in the above embodiments.
[0280] In addition, in the embodiments of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0281] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0282] The following refers to Figure 9 to describe the electronic device 900 according to this embodiment of the present disclosure. Figure 9 The illustrated electronic device 900 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0283] As Figure 9 shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include but are not limited to: at least one processor 910, at least one memory 920, a bus 930 connecting different system components (including the memory 920 and the processor 910), and a display 940.
[0284] Among them, the memory stores program code that can be executed by the processor 910, causing the processor 910 to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processor 910 can execute as Figure 1As shown in: Step S110, obtaining a multi-source dataset corresponding to the target water conservancy project; the multi-source dataset contains project-related data from multiple sources; Step S120, extracting, through the trained water conservancy project cost prediction model, the project feature vectors corresponding to the project-related data of each source, and fusing the multiple project feature vectors corresponding to multiple sources to obtain a fused feature vector; Step S130, predicting, through the trained water conservancy project cost prediction model, the project cost prediction data corresponding to the target water conservancy project for a specified time interval.
[0285] The memory 920 may include a readable medium in the form of volatile storage, such as a random access memory (RAM) 9201 and / or a cache memory 9202, and may further include a read-only memory (ROM) 9203.
[0286] The memory 920 may also include a program / utilities 9204 having a set (at least one) of program modules 9205. Such program modules 9205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0287] The bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures.
[0288] The electronic device 900 may also communicate with one or more external devices 1000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or may communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 950. And, the electronic device 900 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 960. As shown in the figure, the network adapter 960 communicates with other modules of the electronic device 900 through the bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0289] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A method for predicting the cost of water conservancy projects, characterized in that, Including: Obtain a multi-source dataset corresponding to a target water conservancy project; the multi-source dataset contains project-related data from multiple sources; Extract the project feature vectors corresponding to the project-related data of each source through a trained water conservancy project cost prediction model, and fuse the multiple project feature vectors corresponding to the multiple sources to obtain a fused feature vector; Predict the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector through the trained water conservancy project cost prediction model.
2. The method according to claim 1, wherein After predicting the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector through the trained water conservancy project cost prediction model, the method further includes: Obtain the deviation value between the project cost prediction data corresponding to the specified time interval and the actual project cost data; In response to the deviation value being greater than a preset deviation threshold, optimize the trained water conservancy project cost prediction model.
3. The method according to claim 1, characterized in that, The water conservancy project cost prediction model is obtained through the following training method: Obtain a basic sample set; the basic sample set contains multiple basic multi-source datasets corresponding to multiple sample water conservancy projects; each basic multi-source dataset contains basic sample data from multiple sources; Divide the basic sample set into a training set, a validation set, and a test set according to preset data division conditions; Use the training set to perform iterative training on the water conservancy project cost prediction model to be trained to update the model parameters of the water conservancy project cost prediction model to be trained; Use the validation set to validate the water conservancy project cost prediction model to be trained to update the hyperparameters of the water conservancy project cost prediction model to be trained; Use the test set to evaluate the model performance of the water conservancy project cost prediction model to be trained, and when the model performance meets the preset performance requirements, obtain the trained water conservancy project cost prediction model.
4. The method according to claim 3, wherein The obtaining of the basic sample set includes: Obtain an original sample set; the original sample set contains multiple original multi-source datasets corresponding to the multiple sample water conservancy projects; each original multi-source dataset contains original sample data from multiple sources; Perform preprocessing on each original multi-source dataset corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set.
5. The method according to claim 4, wherein The performing of preprocessing on each original multi-source dataset corresponding to each sample water conservancy project in the original sample set to obtain the basic sample set includes: Perform preprocessing on each original multi-source dataset to obtain a first dataset; Perform data cleaning on the first dataset to obtain a second dataset; Perform data integration on the second dataset to obtain a third dataset; Perform data standardization on the third dataset to obtain the basic sample set.
6. The method according to claim 5, wherein The performing of preprocessing on each original multi-source dataset to obtain a first dataset includes: Perform missing value processing on the original sample data of each source in the original multi-source dataset, and perform spatio-temporal alignment processing on the original sample data in the original multi-source dataset to obtain the first dataset.
7. The method according to claim 5, characterized in that, Performing data cleaning on the first data set to obtain a second data set, including: Obtaining the outliers in the first data set; Removing the outliers to obtain the second data set.
8. The method according to claim 5, wherein Performing data integration on the second data set to obtain a third data set, including: Calculating composite engineering metrics based on the second data set, and performing vectorization processing on the text data in the second data set to obtain the third data set.
9. The method according to claim 5, characterized in that, Performing data standardization on the third data set to obtain the basic sample set, including: Performing normalization processing on the continuous variables in the third data set; Performing interval discretization on the target sample data in the third data set; the target sample data includes non-linear sensitive sample data; And performing encoding processing on the discrete variables in the third data set to obtain the basic sample set.
10. The method according to claim 9, characterized in that, Performing encoding processing on the discrete variables in the third data set includes: Performing sparse encoding on the first type of discrete variables, and splicing a preset semantic vector on the basis of the obtained sparse vector; Performing word embedding encoding processing on the second type of discrete variables; Wherein, the number of categories included in the first type of discrete variables is less than a preset category threshold, and the number of categories included in the second type of discrete variables is greater than the preset category threshold.
11. The method according to claim 3, wherein The water conservancy project cost prediction model to be trained includes a feature extraction layer, a feature fusion layer and an output layer; Using the training set to perform iterative training on the water conservancy project cost prediction model to be trained to update the model parameters of the water conservancy project cost prediction model to be trained, including: Extracting the sample project feature vectors corresponding to the basic sample data of each source through the feature extraction layer corresponding to each source; Fusing the multiple sample project feature vectors corresponding to the multiple sources through the feature fusion layer to obtain a sample fusion feature vector; Predicting the project cost prediction sample data corresponding to each sample water conservancy project in a preset time interval based on the sample fusion feature vector through the output layer; Updating the model parameters of the water conservancy project cost prediction model to be trained according to the difference value between the project cost prediction sample data and the actual project cost data label.
12. The method according to claim 11, wherein The basic sample data of the multiple sources includes engineering design sample data, geospatial environment sample data, building material price time series sample data and construction log sample data; The extracting the sample project feature vectors corresponding to the basic sample data of each source through the feature extraction layer corresponding to each source includes: Performing feature extraction on the engineering design sample data through a fully connected layer to obtain an engineering design sample vector; Performing feature extraction on the geospatial environment sample data through a convolutional neural network to obtain an environment sample vector; Performing feature extraction on the building material price time series sample data through a recurrent neural network to obtain a price sample vector; Performing feature extraction on the construction log sample data through a natural language processing network to obtain a construction sample vector.
13. The method according to claim 11, characterized in that, The fusing the multiple sample project feature vectors through the feature fusion layer to obtain a sample fusion feature vector includes: Fusing the multiple sample engineering feature vectors based on the multi-head self-attention mechanism to obtain the fused feature vector.
14. The method according to claim 12, wherein After obtaining the trained water conservancy project cost prediction model, the method further includes: Updating the training set when a preset trigger condition is satisfied; Updating the trained water conservancy project cost prediction model based on the updated training set.
15. The method according to claim 14, wherein The preset trigger condition includes any one of the following: The difference degree between the first probability distribution of the newly added building material price time series data and the second probability distribution of the building material price time series data in the training set is greater than the preset difference degree range; The fluctuation degree of the building material price within a preset time period exceeds the preset fluctuation range.
16. The method according to claim 14, wherein The updating the trained water conservancy project cost prediction model based on the updated training set includes: Fixing the network parameters of some networks in the trained water conservancy project cost prediction model; the some networks include the fully connected layer, the convolutional neural network, and the natural language processing network; Updating the network parameters of the remaining networks except the some networks.
17. A device for predicting the cost of water conservancy projects, characterized in that, Including: A data acquisition module, configured to acquire a multi-source data set corresponding to a target water conservancy project; the multi-source data set contains engineering-related data from multiple sources; A data processing module, configured to extract the engineering feature vector corresponding to the engineering-related data of each source through the trained water conservancy project cost prediction model, and fuse the multiple engineering feature vectors corresponding to the multiple sources to obtain a fused feature vector; A prediction module, configured to predict the project cost prediction data corresponding to the target water conservancy project for a specified time interval based on the fused feature vector through the trained water conservancy project cost prediction model.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the water conservancy project cost prediction method according to any one of claims 1 to 16.
19. An electronic device, characterized in that, Including: A processor; And A memory, configured to store the executable instructions of the processor; Wherein, the processor is configured to execute the water conservancy project cost prediction method according to any one of claims 1 to 16 by executing the executable instructions.
Citation Information
Cited By
Engineering cost big data management and analysis system
CN120873058A
An engineering cost big data management and analysis system
CN120873058B