Water supply pipe network water demand prediction method based on multi-source data fusion and hybrid model

Through multi-source data fusion and hybrid model, combined with clustering and classification models, the problems of data feature extraction and insufficient universality in water demand prediction are solved, and high-precision water demand prediction for different independent measurement areas are achieved, which improves the management efficiency and resource utilization of the water supply system.

CN120278331AActive Publication Date: 2025-07-08TIANJIN UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510392708.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

When existing water volume prediction methods face nonlinear and non-uniformly distributed long-term series data, it is difficult to extract effective features. The prediction effect of a single model on independent metrology areas of different categories is unstable, with poor universality, and obvious interference from external factors, resulting in insufficient prediction accuracy and robustness.

Method used

A mixed model of integrated clustering model, classification model and prediction model is adopted. Through the fusion of multi-source data, including historical water consumption, meteorological data and holiday information, daily type clustering is performed using the K-means algorithm, and the daily type classification and water demand prediction model are trained in combination with the XGBoost algorithm to optimize the feature matrix to improve prediction accuracy and robustness.

Benefits of technology

It improves the universality and accuracy of the water demand prediction model, and can maintain good performance when user water use occurs temporarily irregularly. It is suitable for different types of independent metering areas, improving the operating efficiency and resource utilization of the water supply system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278331A_ABST
    Figure CN120278331A_ABST
Patent Text Reader

Abstract

The invention discloses a water supply pipe network water demand prediction method based on multi-source data fusion and a hybrid model. The hybrid model comprises a clustering model, a daily type classification model and a prediction model. The method mainly comprises the following steps: collecting multi-source data including historical water consumption data, meteorological data and holiday and festival information data, and preprocessing the historical water consumption data and the meteorological data; performing feature extraction by using a clustering method, and forming a clustering result vector; respectively making a feature matrix and an output label for training a classification model and a prediction model in combination with the multi-source data and the clustering result, and training a daily type classification and water demand prediction model by using an XGBoost algorithm; using the trained day type classification and water demand prediction model to predict the water demand at a certain moment in a certain day in the future, and finally obtaining a water demand prediction value at a corresponding moment of a to-be-predicted day. According to the method, the capability of the XGBoost model can be better played, and the method has good robustness and has a good prediction effect on different types of DMA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a water demand prediction method based on multi-source data fusion, which is used to solve the problem of water demand prediction for users in the independent metering area of the water supply network. Background Art

[0002] Water demand prediction is one of the core issues in the optimal management of the water supply system and the rational allocation of water resources. The purpose of water demand prediction is to predict the water demand in a specific future period based on historical water consumption data and influencing factors of water consumption behavior. The factors affecting water demand are diverse, including seasonality, holidays, meteorological conditions (such as temperature, rainfall), user attributes in the prediction area, social and economic activities, population growth, water-saving policies, etc. Common prediction time scales include short-term (such as hours or days), medium-term (such as weeks or months), and long-term (such as years) predictions. The focus of predictions at different time scales is different. This article mainly discusses short-term water demand prediction, and the prediction results mainly serve water supply scheduling and operation optimization. After achieving accurate water demand prediction for different independent metering areas (DMAs), the scheduling strategies of the pipe network pumps, valves, and secondary water supply systems can be modified according to the prediction results, thereby improving the operating efficiency of the water supply system, reducing energy consumption, reducing resource waste, and providing support for formulating long-term water resource plans. In recent years, with the development of data collection and analysis technologies, water demand prediction methods have experienced a transformation from traditional statistical models to data-driven intelligent models.

[0003] Currently, the relatively popular water demand prediction methods can be roughly divided into three categories. The first is the early statistical methods. Researchers use traditional statistical time series regression analysis methods such as linear regression, ARMA, ARIMA, and SVR. The second is the machine learning and deep learning methods that have emerged in recent years. The third is the method of hybrid models, which improves the prediction accuracy of the model by combining the advantages of different methods.

[0004] However, historical water demand data is usually non-linear and non-uniformly distributed time series data with a large time span. Traditional models often still have difficulty in fully extracting the characteristics of long sequence data sets with a large time span. In addition, the prediction effect of a single model on the water consumption of different categories of DMAs is not stable, and the universality is poor. For the current hybrid models, the feature input is mostly simple lagged water consumption data. In addition to lagged water consumption, the water consumption of users is also easily interfered by various external factors. Due to their special user attributes, the water demand of some users is prone to significant changes due to the changes of multiple external factors such as season, temperature, rainfall, and social activities. Summary of the Invention

[0005] In view of the above-mentioned prior art, the present invention provides a hybrid model method integrating a clustering model, a classification model and a prediction model, which can combine the advantages of the clustering model, the classification model and the prediction model, fuse multi-source data information such as holidays, meteorology, historical water consumption, etc., improve the universality and accuracy of the water demand prediction model, and enable the prediction model to still take into account the characteristics of recent water consumption and water consumption on similar days even when there are short-term irregular fluctuations in user water use, and make the model have good performance when predicting the water consumption of different types of DMAs.

[0006] To solve the above technical problems, a water demand prediction method for a water supply network based on multi-source data fusion and a hybrid model is proposed in the present invention. The hybrid model is a combination of multiple models, and the hybrid model includes a clustering model, a daily type classification model and a prediction model; the water demand prediction method for the water supply network includes the following steps:

[0007] Step 1: Collect and preprocess multi-source data:

[0008] The multi-source data includes historical water consumption data, meteorological data and holiday information data, which are relatively easy for water companies to obtain; since the present invention is a water demand prediction method for long time series, it has certain requirements for the quality of data. Among them, the granularity of the historical water consumption data is one number per hour, and the time span is more than 1 year; the meteorological data includes air temperature, rainfall, humidity and wind speed, and the granularity is one number per hour, and the time span is more than 1 year; the holiday information includes working days, weekends, legal holidays and daylight saving time information, and the granularity is one number per day, and the time span is more than one year; the collected historical water consumption data and meteorological data are preprocessed, and the z-score method is used to detect and filter outliers, and the K-nearest neighbor method is used to implement missing value imputation.

[0009] Step 2: Use the clustering model to perform daily type clustering based on the daily historical water consumption data:

[0010] Use the clustering method to extract features from the preprocessed historical water consumption data, divide the historical water consumption data into several daily water consumption samples on a daily basis, and use the k-means algorithm to cluster the daily water consumption samples to obtain the clustering result of the daily historical water consumption data, thereby forming a clustering result vector y clusterSince the water demand data fluctuates greatly due to external factors such as seasonality, holidays, and meteorological conditions, it has the characteristics of time-series data with non-uniform distribution, and the time span of the water demand data is relatively long. The traditional time-series prediction method that only uses the lagged water demand has a poor prediction effect on this type of long-sequence data. Therefore, before training the prediction model of the present invention, a clustering method is used to extract the features of the preprocessed historical water consumption data. The historical water consumption data is segmented into several daily water consumption samples on a daily basis, and a clustering algorithm is used to cluster the daily water consumption samples to obtain the clustering results of the daily historical water consumption data, thereby forming a clustering result vector y cluster ; In the selection of the clustering method, after trying the partitioning clustering model K-means and the time-series clustering model K-shape, the present invention preferably uses the K-means algorithm to cluster the daily water consumption samples.

[0011] Step 3: Produce a feature matrix for training the classification model and train the daily type classification model:

[0012] A feature matrix for training the classification model is produced according to the holiday information, the preprocessed historical water consumption data, and the meteorological data, denoted as feature matrix A, and the label Y A is the clustering result vector y cluster ; The feature matrix A is input into the XGBoost algorithm for training to form a daily type classification model; the feature information of the feature matrix A corresponding to the day to be predicted is input into the daily type classification model, and the output of the daily type classification model is the daily type classification result of the day to be predicted; the XGBoost algorithm is an efficient gradient boosting algorithm, which is widely used in classification, regression, and ranking tasks, and is used in the present invention to solve the problems of daily type classification and water demand prediction. The principle of XGBoost is based on the Gradient Boosting Trees (GBT), and several important optimizations have been made on this basis, especially suitable for large-scale data sets. This algorithm combines multiple weak learners through the additive model. In the process of training to form the daily type classification model, the weak learner is a decision tree.

[0013] The feature matrix A and the output label Y A are structured as follows:

[0014]

[0015] Y A = y cluster

[0016] where x se is the seasonal factor vector, and the daylight saving time and winter time to which the day belongs are represented by 0 and 1 respectively; xmon is the monthly factor vector. The months are encoded sequentially, and the encoding result of the month to which the day belongs is used as the value of the feature; x day is the date factor vector. The dates are encoded sequentially, and the encoding corresponding to the day is used as the value of the feature. The meteorological factors include vectors in four dimensions, namely rainfall x rain , temperature x temp , humidity x hum , wind speed x wind , and the average value of the daily meteorological factors is used as the value of the feature; x holi is the holiday factor vector. According to the local calendar information, weekdays and rest days are encoded as 0 and 1 respectively, and the encoding result is used as the value of the feature, where legal holidays and weekends are both recognized as rest days; The lagged daily water consumption information includes the average water consumption of the lagged day and the standard deviation of the water consumption Among them, the lagged daily water consumption refers to the water consumption a certain number of days before the current day. For example, the water consumption on the 1st lag day is the water consumption 1 day before the current day; respectively include the average hourly water consumption feature vectors from the 1st lag day to the nth lag day; respectively include the standard deviation of hourly water consumption feature vectors from the 1st lag day to the nth lag day; is the average hourly water consumption feature vector on the kth lag day; is the standard deviation of the hourly water consumption feature vector on the kth lag day; x ik is the water consumption at the ith hour on the kth lag day.

[0017] Since the dimensionality of the features in each dimension is different and there are scale differences between the features, it is necessary to uniformly standardize the features in each dimension to align the feature scales. The present invention uses the Z-score standardization method to standardize the features in each dimension of the feature matrix A, and its calculation formula is as follows:

[0018]

[0019] Among them, x is the feature value, μ is the mean of the feature vector, and σ is the standard deviation of the feature vector.

[0020] Step 4: Produce a multi-source data fusion feature matrix for training the prediction model, and train the water demand prediction model:

[0021] The clustering result vector of the daily historical water consumption data obtained by the clustering model is combined with multi-source data to form a multi-source data fusion feature matrix for training the prediction model, denoted as feature matrix B; the feature matrix B is input into the XGBoost algorithm for training to form a water demand prediction model; the XGBoost algorithm combines multiple weak learners through the additive model therein to gradually reduce the prediction error. During the process of training to form the water demand prediction model, the weak learner is a regression tree; the feature matrix B and the output label Y B has the following structure:

[0022]

[0023] Among them, the steps for obtaining the long-term lag water consumption of the same type of day include: First, according to the clustering result y of the daily historical water consumption data obtained by the clustering model cluster , the daily historical water consumption is grouped and rearranged in chronological order to obtain the historical water consumption sequence of each type of day; then, according to the historical water consumption sequence of the same type of day to which the current moment water consumption belongs, after performing a lag process on the current moment water consumption, the long-term lag water consumption of the same type of day is obtained.

[0024] After obtaining the day type classification model and the water demand prediction model, the water demand at a certain moment on a future day can be predicted. Specifically, it includes the following steps: According to the structure of the feature matrix A, the features of the day to be predicted are made into a feature sample of the day to be predicted and input into the day type classification model to obtain the day type classification result of the day to be predicted; according to the structure of the feature matrix B, combined with the day type classification result of the day to be predicted, a feature sample of the moment to be predicted on the day to be predicted is made and input into the water demand prediction model, and the output of the water demand prediction model is the water demand prediction value at that moment on that day.

[0025] In the present invention, according to the feature importance of the short-term lag water consumption feature and the feature importance of the long-term lag water consumption feature in the water demand prediction model, the water use behavior of users is characterized. Among them, the feature importance of the short-term lag water consumption feature represents the inertia of the user's water use behavior, and the feature importance of the long-term lag water consumption feature represents the periodicity of the user's water use behavior. The larger the feature importance value, the more obvious the corresponding water use behavior characteristics. The feature importance is calculated after using the XGBoost algorithm to complete the training of the water demand prediction model, and the contribution degree of each dimension data in the feature matrix to the prediction result is calculated. The result is represented by the Gain value, and its calculation method is the same as the method for calculating the information gain, that is, calculating the entropy increment when a certain feature is split. The larger the increment, the stronger the importance of the feature.

[0026] Compared with the prior art, the beneficial effects of the present invention are:

[0027] (1) The optimized hybrid feature matrix obtained by fusing multi-source data during the training process of the present invention can better exert the prediction ability of the XGBoost regression model for future water demand, and the features after multi-source data fusion have a higher proportion of importance in the prediction model compared to simple time-lag features.

[0028] (2) The prediction model optimized by multi-source data fusion in the present invention has better robustness compared to existing prediction methods and has good prediction effects for DMAs of different user types.

[0029] (3) The XGBoost algorithm used in the present invention has a feature importance evaluation function, and can characterize to a certain extent the water use behavior of the predicted DMA users according to the proportion of importance of each feature in the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a schematic diagram of the water demand prediction method for the water supply network of the present invention;

[0031] Figure 2 Schematic diagram of the XGBoost algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following further describes the water demand prediction method based on multi-source data fusion and hybrid model of the present invention in conjunction with the drawings and through specific embodiments:

[0033] As Figure 1 shown, a water demand prediction method for a water supply network based on multi-source data fusion and a hybrid model proposed by the present invention, the hybrid model is a combination of multiple models, and the hybrid model includes a clustering model, a daily type classification model, and a prediction model; the water demand prediction method for the water supply network mainly includes: classifying and preprocessing multi-source data; implementing daily type clustering based on daily historical water consumption data using a clustering algorithm; making a feature matrix based on multi-source data and daily type clustering labels to train the daily type classification model; making a multi-source data fusion feature matrix for training the prediction model based on the classification results of the classification model and training the prediction model; predicting the water demand at a certain moment on a certain future day after obtaining the daily type classification model and the water demand prediction model. The specific content of each step is as follows:

[0034] The first step: Classify and preprocess multi-source data.

[0035] In the present invention, the multi-source data includes historical water consumption data, meteorological data, and holiday information data, which are relatively easy to obtain for water companies. Since the present invention is a long-term series water demand prediction method, it has certain requirements for the quality of data. Among them, the historical water consumption data has a granularity of one number per hour and a time span of more than 1 year; the meteorological data includes temperature, rainfall, humidity, and wind speed with a granularity of one number per hour and a time span of more than 1 year; the holiday information includes working days, weekends, legal holidays, and daylight saving time information with a granularity of one number per day and a time span of more than one year. The collected historical water consumption data and meteorological data are preprocessed, and the z-score method is used to detect and filter outliers, and the K-nearest neighbor method is used to implement missing value imputation.

[0036] Step 2: Use a clustering algorithm to achieve daily type clustering based on the daily historical water consumption data.

[0037] Since the water demand data is greatly affected by external factors such as seasonality, holidays, and meteorological conditions, showing large fluctuations, having the characteristics of non-uniformly distributed time series data, and the time span of the water demand data is relatively long, the traditional time series prediction method that only uses lagged water demand has poor prediction effects on this type of long series data. Therefore, in the present invention, before training the prediction model, the clustering method is first used to extract features from the preprocessed historical water consumption data. The historical water consumption data is sliced into several daily water consumption samples in units of days, and a clustering algorithm is used to cluster the daily water consumption samples to obtain the clustering results of the daily historical water consumption data, thus forming a clustering result vector y cluster , and then corresponding prediction model feature matrices are formulated according to the water consumption data of different daily types. The long series data is divided into individual samples in units of days, and the features in similar daily type samples are taken for model training, thereby improving the fitting accuracy of the prediction model for different daily types.

[0038] Since the historical water consumption data is typical time series data and has obvious time series shape characteristics of peaks and valleys, this problem can be transformed into a time series clustering problem. The historical water consumption data sample set in the present invention can be recorded as a set D = {D1, D2,..., D n} containing n time series samples, where each sample is a time series composed of the water consumption values at 24 moments of a day. The clustering process of the historical water consumption data is the process of mapping D to the clustering result C through the similarity measure F, and this process can be represented by the formula C = F(D), where C = {C1, C2,..., C m}, and

[0039] In the selection of the clustering method, after trying the partitioning clustering model K-means and the clustering model K-Shape for time series, the present invention decides to use the K-means algorithm to cluster the daily time-sharing historical water consumption data.

[0040] The core of the k-means algorithm mainly consists of three steps. The first is to calculate the centroid of each cluster, the second is to determine the category to which each data belongs according to the centroid and the similarity measurement method, and the third is to perform iteration according to the objective function. In the present invention, the Euclidean distance formula is used to calculate the distance Dist(D i ,D j ) between data sample points, and its calculation formula is as follows:

[0041]

[0042] Among them, D i ,D j represent the i-th and j-th historical water consumption data samples respectively; D ik ,D jk represent the values at the k-th moment of the i-th and j-th samples respectively; T represents the total number of variables of each sample. In the present invention, each historical water consumption data sample is a time series composed of the water consumption values at 24 moments of a day, that is, D. Therefore, T takes 24 here.

[0043] The centroid D i of each cluster C io is calculated as follows:

[0044]

[0045] Among them, r i is the total number of samples in cluster C i ; D tk represents the value of the t-th variable in the k-th sample, and in this article, it represents the water consumption value at the t-th moment in each historical water consumption data sample.

[0046] The goal of the clustering in the present invention is to minimize the sum of the squares of the distances between the sample points within the cluster and the centroid, and its calculation formula is as follows:

[0047]

[0048] Among them, m is the total number of clusters in the set C; C i is the i-th cluster in the set C; D is the historical water consumption data sample in C i , and D io is the centroid of the i-th cluster.

[0049] After completing the daily type clustering task of historical water consumption data, it is necessary to import the historical water consumption similarity features extracted in the daily type clustering into the feature matrix of the prediction model, and this process is implemented through the daily type classification module.

[0050] Step 3: Based on multi-source data and daily type clustering labels, make a feature matrix to train the daily type classification model.

[0051] Since the water consumption information of the day to be predicted is unknown, the only available information for judging the daily type of the day to be predicted is the influencing factors of water use behavior other than water consumption. In order to add the clustering information of the day to be predicted to the feature matrix of the prediction model, it is necessary to classify the daily type of the day to be predicted based on multi-source data and clustering results. The present invention makes a feature matrix for training the classification model according to holiday information, preprocessed historical water consumption data and meteorological data, denoted as feature matrix A, and label Y A is the clustering result vector y cluster ; Input the feature matrix A into the XGBoost algorithm for training to form a daily type classification model; input the feature information of the feature matrix A corresponding to the day to be predicted into the daily type classification model, and the output of the daily type classification model is the daily type classification result of the day to be predicted.

[0052] The XGBoost algorithm combines multiple weak learners through the additive model therein. In the process of training to form the daily type classification model, the weak learner is a decision tree.

[0053] The feature matrix A and the output label Y A are structured as follows:

[0054]

[0055] Y A = y cluster

[0056] where, x se is the seasonal factor vector, and the daylight saving time and winter time to which the day belongs are represented by 0 and 1 respectively; x mon is the month factor vector, the months are encoded in sequence, and the encoding result of the month to which the day belongs is used as the value of the feature; x day is the date factor vector, the dates are encoded in sequence, and the encoding corresponding to the day is used as the value of the feature. The meteorological factors include four-dimensional vectors, namely rainfall x rain , temperature x temp , humidity x hum , wind speed x wind , and the average value of the daily meteorological factors is used as the value of the feature; x holiLet \(H\) be the holiday factor vector. According to the local calendar information, weekdays and rest days are encoded as 0 and 1 respectively, and the encoded results are used as the values of the features. Among them, legal holidays and weekends are both recognized as rest days.

[0057] The lagged daily water consumption information includes the average water consumption of the lagged day and the standard deviation of the water consumption Among them, the lagged daily water consumption refers to the water consumption a certain number of days before the current day. For example, the water consumption of the 1st lag day is the water consumption 1 day before the current day; respectively include the average hourly water consumption feature vectors from the 1st lag day to the \(n\)th lag day; respectively include the standard deviation of hourly water consumption feature vectors from the 1st lag day to the \(n\)th lag day; is the average hourly water consumption feature vector of the \(k\)th lag day; is the standard deviation of hourly water consumption feature vector of the \(k\)th lag day; \(x\) ik is the water consumption at the \(i\)th hour of the \(k\)th lag day.

[0058] Since the dimensionality of each dimension feature is different and there is a scale difference between features, it is necessary to uniformly standardize each dimension feature to align the feature scales. The present invention uses the Z - score standardization method to standardize each dimension feature of the feature matrix \(A\), and its calculation formula is as follows:

[0059]

[0060] Among them, \(x\) is the feature value, \(\mu\) is the mean of the feature vector, and \(\sigma\) is the standard deviation of the feature vector.

[0061] The present invention uses the XGBoost algorithm as the algorithm core of the classification model. The XGBoost algorithm is an efficient gradient boosting algorithm, which is widely used in classification, regression and ranking tasks, and is used to solve the daily type classification and water demand prediction problems in the present invention. The principle of XGBoost is based on the gradient boosting tree (GBT), and a number of important optimizations have been carried out on this basis, which is especially suitable for large - scale data sets. The algorithm structure of XGBoost is as Figure 2 shown. Its core idea is to combine multiple weak learners through an additive model to gradually reduce the prediction error. The prediction form of this model is shown as the following formula:

[0062]

[0063] Among them, \(\varphi(x\) i ) represents the total model of water consumption prediction, \(f\) k(x) represents the k-th sub-tree model, where K is the total number of sub-tree models in the total water consumption prediction model. The goal of each sub-tree model is to fit the residual of water demand prediction by minimizing the loss function. In the present invention, the goal of the XGBoost model is to minimize the loss function with a regularization term, and this loss function can be expressed by the following formula:

[0064]

[0065] where, represents the loss function, and in the present invention, the mean squared residual between the water consumption prediction value and the true value is taken, that is Ω(f k ) represents the regularization term, which is used to control the complexity of each tree. XGBoost uses L2 regularization to limit the number of leaf nodes and the weights of leaf nodes to prevent overfitting.

[0066] The specific process of the algorithm is as follows:

[0067] 1) Initialize the model. The goal of XGBoost is to minimize the objective function by gradually adding multiple weak learners to the model. To achieve this effect, in the present invention, the mean value of the prediction values is first used as the initial value of the total model.

[0068] 2) Construct the weak learner sub-model. XGBoost fits the residual of the previous generation model by constructing a new weak learner sub-model in each round of iteration, and optimizes the internal structure of the weak learner sub-model by calculating the first-order and second-order derivatives of the loss function, thereby reducing the value of the objective function. This optimization process mainly includes two parts:

[0069] 2-1) Select the best split point. The best split point is determined by calculating the gain of all possible split points of each feature, and the split point with the largest gain is selected as the optimal split point of the current node. XGBoost adopts a split gain calculation method based on the first-order and second-order derivatives, and the gain of each split point can be expressed by the following formula:

[0070]

[0071] where, G L and G R are the sums of the gradients of the left and right child nodes respectively, H L and H R are the corresponding sums of the second-order gradients, λ is the regularization parameter, and γ is the regularization term;

[0072] 2-2) Calculate the leaf node weights. The weights are optimized by calculating the first-order and second-order derivative information of the loss function, and the optimal value of the weight can be expressed by the following formula:

[0073]

[0074] where \(I\) j is the sample set included in leaf node \(j\); \(g\) i is the gradient of the predicted value of the current sample water volume; \(h\) i is the second-order derivative of the predicted value of the current sample water demand, and \(\lambda\) is the regularization parameter;

[0075] 3) Repeat step 2) for iteration until the termination condition is met. The termination condition is to reach the preset number of weak learners, reach the maximum number of iterations, or the residual change is small.

[0076] In addition to adding regularization and penalty terms to the objective function, the XGBoost algorithm also performs column subsampling, parallel computing, shrinkage, and missing value handling optimizations, making it have a certain improvement in the prediction accuracy and the spatio-temporal complexity of calculation compared with the traditional gradient boosting algorithm.

[0077] In the present invention, the daily water consumption type classification model uses a decision tree as the weak learner of the XGBoost algorithm, and its model training parameters can be set with reference to the following table:

[0078] Hyperparameter Name Value or Range Step Size objective 'multi:softmax' - eval_metric 'logloss' - num_class 2-9 1 max_depth 3-9 3 learning_rate 0.01-0.3 0.02 n_estimators 50-100 10 gamma 0.5 - colsample_bytree 0.5-0.9 0.2

[0079] Among them, the values of the hyperparameters objective and eval_metric of the classification model indicate that the classification model uses the softmax function for solving multi-classification tasks as the training set loss function and the logarithmic function as the validation set evaluation function. num_class represents the number of classifications in the multi-classification task; max_depth represents the maximum depth of each subtree model; learning_rate represents the learning rate; n_estimator represents the number of subtree models. Since the number of classification samples is relatively small, this parameter in the classification model is set to 50 - 100; gamma represents the second-order derivative regularization term coefficient, which is used to control the model complexity and takes 0.5 in the present invention; colsample_bytree represents the number of column samplings, that is, the subtree model will only extract a specific proportion of columns for training during training. This operation can reduce overfitting, accelerate the training process, and improve the model generalization ability, usually taking 0.5 - 0.9.

[0080] Step 4: Based on the classification results of the classification model, make a multi-source data fusion feature matrix for training the prediction model and train the prediction model.

[0081] The clustering result vector of the daily historical water consumption data obtained by the clustering model is combined with multi-source data to form a multi-source data fusion feature matrix for training the prediction model, denoted as feature matrix B; this feature matrix B is input into the XGBoost algorithm for training to form a water demand prediction model.

[0082] In the present invention, the XGBoost algorithm is also used in the training process of the water demand prediction model. Different from the classification model, in the process of training to form the water demand prediction model, the weak learner is a regression tree.

[0083] In the process of training to form the water demand prediction model, the training parameters of the XGBoost algorithm can be set with reference to the following table:

[0084] Hyperparameter Name Value or Range Step Size objective 'reg:squarederror' - max_depth 3-9 3 learning_rate 0.01-0.3 0.02 n_estimators 100-1000 100 gamma 0.5 - colsample_bytree 0.5-0.9 0.2

[0085] Among them, the value of the hyperparameter objective of the prediction model indicates that the mean squared error is used as the model loss function; n_estimator represents the number of sub-tree models. Since the number of prediction samples is relatively large, this parameter in the prediction model is set to 100 - 1000; the parameter settings of max_depth, learning_rate, gamma, and colsample_bytree are the same as those of the classification model.

[0086] The present invention uses the feature importance method of XGBoost to analyze the contribution degree of each dimension data to the prediction result. Feature importance refers to the method of measuring the contribution degree of each dimension data to the prediction result after using the XGBoost algorithm to complete the training of the water demand prediction model. In the present invention, the "Gain value" method is used to measure the contribution degree of each dimension feature to the prediction result, and its calculation method is the same as the method of calculating information gain, that is, calculating the entropy increment when a certain feature is split. The larger the increment, the stronger the importance of the feature. Usually, the existing feature matrix can be optimized according to the performance of feature importance, deleting the features with smaller feature importance, so as to make the model more focused on the features with high contribution degree while improving the model training rate.

[0087] To improve the prediction accuracy and robustness of the prediction model for long time series data, the present invention improves the input matrix structure of the prediction model by combining the clustering results of the clustering model and the feature importance of the lagged water consumption. The traditional water demand prediction model usually only uses the lagged water consumption input feature matrix, and contains a lot of useless lagged water consumption with low importance, resulting in high training cost, low training accuracy, and poor robustness of the model.

[0088] Based on the above experience, a new feature engineering is designed in the present invention, and the specific operation is as follows:

[0089] 1) The present invention adds a time feature to the periodic information of water consumption retention to align the time dimension of the prediction result.

[0090] 2) Based on the feature importance of lagged water consumption, the feature dimension of lagged water consumption in the feature matrix is simplified, reducing the original 24 dimensions to 4 dimensions, where there are 2 dimensions for short-term lagged water consumption features and 2 dimensions for long-term lagged water consumption features respectively.

[0091] 3) To retain the statistical law of long-term lagged water consumption, the present invention adds statistical features of water consumption at 24 lagged moments.

[0092] 4) To ensure that the model has good performance for different types of DMAs, the present invention adds daily type features and long-term lagged water consumption features of the same type of day in combination with the clustering results.

[0093] 5) The present invention simultaneously introduces meteorological factor features, aiming to achieve the fusion of multi-source data and absorb the hidden user water consumption behavior information in the multi-source data.

[0094] The multi-source data fusion feature matrix B and output label Y for training the prediction model in the present invention B are structured as follows:

[0095]

[0096] Among them, the steps for obtaining the long-term lagged water consumption of the same type of day include: First, the clustering result y of the daily historical water consumption data is obtained according to the clustering model. cluster, group the daily historical water consumption and rearrange it in chronological order to obtain the historical water consumption sequence of each type of day; then, based on the historical water consumption sequence of the same type of day to which the current water consumption belongs, perform lag processing on the current water consumption to obtain the long-term lag water consumption of the same type of day; this feature can more accurately capture the periodic characteristics of the water use pattern by comparing the lag water consumption of the same type of day, and then more accurately reflect the periodic water use behavior of users. Re-input the finally determined feature matrix into the XGBoost algorithm for training to obtain the final water demand prediction model. After obtaining the daily type classification model and the water demand prediction model, the water demand at a certain moment on a future day can be predicted, which specifically includes the following steps: 1) According to the feature matrix structure of the daily type classification model in step 3, make the features of the day to be predicted into a feature sample of the day to be predicted and input it into the daily type classification model to obtain the daily type classification result of the day to be predicted; 2) According to the multi-source data fusion feature matrix structure of the prediction model, combine the daily type classification result of this day to make a feature sample of the moment to be predicted on the day to be predicted and input it into the water demand prediction model, and the final model will output the water demand prediction value at this moment on this day. At the same time, the present invention can also characterize the water use behavior of users according to the feature importance of the short-term lag water consumption feature and the long-term lag water consumption feature of the prediction model. Among them, the short-term lag water consumption feature can characterize the inertia of the user's water use behavior, and the long-term lag water consumption feature can characterize the periodicity of the user's water use behavior. The larger the feature importance value, the more obvious the corresponding water use behavior characteristics are.

[0097] Taking the public dataset (https: / / wdsa-ccwi2024.it) provided by the Battle of Water Demand Forecasting (BWDF) competition organized under the background of the 3rd International WDSA-CCWI Joint Conference as an example for water demand forecasting, and comparing the forecasting results with those of Perelman et al. [1] The prediction methods using the combined regional cross-correlation coefficient and lag water consumption, the prediction method only using the lag water consumption feature, and the present invention, that is, the prediction method based on multi-source data fusion and hybrid model, are compared. The comparison results are quantified by the Mean Absolute Percentage Error (MAPE) of the prediction results of each model on the test set, and the comparison results are as follows:

[0098]

[0099] The results show that the prediction method based on multi-source data fusion and hybrid model used in the present invention performs significantly better than the lag water consumption feature method combined with regional cross-correlation and the method using lag water consumption features in most DMAs, indicating that the feature matrix optimized by multi-source data fusion can better exert the prediction ability of the XGBoost regression algorithm for future water demand. In the case of this public dataset, the average prediction MAPE of the prediction model of the present invention can reach below 5%, the highest prediction MAPE is below 10%, and the lowest prediction MAPE can reach 1.56%, indicating that the model has better robustness after being optimized by multi-source data fusion compared with the existing methods and has good prediction effects on DMAs of different user types.

[0100] References: [1] Perelman G, Romano Y, Ostfeld A. Optimizing Time Series Models for Water Demand Forecasting[J]. Engineering Proceedings, 2024, 69(1): 9.

[0101] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many improvements and changes without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A method for predicting the water demand of a water supply network based on multi-source data fusion and a hybrid model, characterized in that, The hybrid model is a combination of multiple models, including a clustering model, a daily type classification model, and a prediction model. The method for predicting the water demand of a water supply network includes the following steps: Step 1: Collect and preprocess multi-source data; The multi-source data includes historical water consumption data, meteorological data, and holiday information data. Among them, the granularity of the historical water consumption data is one number per hour, and the time span is more than 1 year; the meteorological data includes temperature, rainfall, humidity, and wind speed, with a granularity of one number per hour and a time span of more than 1 year; the holiday information includes working days, weekends, legal holidays, and daylight saving time information, with a granularity of one number per day and a time span of more than one year; Preprocess the collected historical water consumption data and meteorological data above, use the z-score method to detect and filter outliers, and use the K-nearest neighbor method to implement missing value imputation; Step 2: Use the clustering model to perform daily type clustering based on the daily historical water consumption data; Use the clustering method to extract features from the preprocessed historical water consumption data. The historical water consumption data is sliced into several daily water consumption samples on a daily basis. The k-means algorithm is used to cluster the daily water consumption samples to obtain the clustering results of the daily historical water consumption data, thereby forming the clustering result vector y cluster ; Step 3: Make a feature matrix for training the classification model and train the daily type classification model; A feature matrix for training a classification model is made based on holiday information, preprocessed historical water consumption data, and meteorological data, denoted as feature matrix A, and label Y A is the clustering result vector y cluster ; Input the feature matrix A into the XGBoost algorithm for training to form a daily type classification model; Input the feature information of the feature matrix A corresponding to the day to be predicted into the daily type classification model, and the output of the daily type classification model is the daily type classification result of the day to be predicted; The XGBoost algorithm combines multiple weak learners through the additive model therein. During the process of training to form the daily type classification model, the weak learner is a decision tree; The feature matrix A and the output label Y A are structured as follows: Y A = y cluster Among them: Among them, x se is the seasonal factor vector, where the daylight saving time and winter time to which the day belongs are represented by 0 and 1 respectively; x mon is the monthly factor vector, where the months are encoded in sequence, and the encoding result of the month to which the day belongs is used as the value of the feature; x day is the date factor vector, where the dates are encoded in sequence, and the encoding corresponding to the day is used as the value of the feature. The meteorological factors include four-dimensional vectors, namely rainfall x rain , temperature x temp , humidity x hum , and wind speed x wind , and the average value of the daily meteorological factors is used as the value of the feature; x holi is the holiday factor vector, where the working days and rest days are encoded by 0 and 1 respectively according to the local calendar information, and the encoding result is used as the value of the feature, where the legal holidays and weekends are both recognized as rest days; The lagged daily water consumption information includes the average water consumption of the lagged day and the standard deviation of the water consumption where the lagged daily water consumption refers to the water consumption a certain number of days before the current day. For example, the water consumption on the 1st lagged day is the water consumption 1 day before the current day; respectively include the average hourly water consumption feature vectors from the 1st lagged day to the nth lagged day; respectively include the standard deviation feature vectors of the hourly water consumption from the 1st lagged day to the nth lagged day; is the average hourly water consumption feature vector for the kth lagged day; is the standard deviation feature vector of the hourly water consumption for the kth lagged day; x ik is the water consumption at the ith hour of the kth lagged day; Use the Z-score standardization method to standardize the features of each dimension of the feature matrix A, and its calculation formula is as follows: Among them, x is the feature value, μ is the mean of the feature vector, and σ is the standard deviation of the feature vector; Step 4: Make a multi-source data fusion feature matrix for training the prediction model and train the water demand prediction model Combine the daily historical water consumption data clustering result vector obtained by the clustering model with the multi-source data to make a multi-source data fusion feature matrix for training the prediction model, denoted as feature matrix B; input this feature matrix B into the XGBoost algorithm for training to form a water demand prediction model; The XGBoost algorithm combines multiple weak learners through the additive model therein. During the process of training to form the water demand prediction model, the weak learner is a regression tree; The feature matrix B and the output label Y B are structured as follows: Among them, the steps for obtaining the long-term lagged water consumption of the same type of day include: First, according to the clustering model, obtain the clustering result y of the daily historical water consumption data cluster , group the daily historical water consumption and rearrange it in chronological order to obtain the historical water consumption sequence of each type of day; then, according to the historical water consumption sequence of the same type of day to which the current water consumption belongs, perform lag processing on the current water consumption to obtain the long-term lagged water consumption of the same type of day; After obtaining the daily type classification model and the water demand prediction model, the water demand at a certain moment on a certain future day can be predicted, which specifically includes the following steps: 4-1) According to the structure of the feature matrix A, make the features of the day to be predicted into a feature sample of the day to be predicted and input it into the daily type classification model to obtain the daily type classification result of the day to be predicted; 4-2) According to the structure of the feature matrix B, combine the daily type classification result of the day to be predicted to make a feature sample of the moment to be predicted on the day to be predicted and input it into the water demand prediction model. The output of the water demand prediction model is the predicted value of the water demand at that moment on that day.

2. The water demand prediction method for water supply networks based on multi-source data fusion and hybrid model according to claim 1, characterized in that According to the feature importance of the short-term lag water consumption feature and the long-term lag water consumption feature in the water demand prediction model, characterize the user's water use behavior. Among them, the feature importance of the short-term lag water consumption feature characterizes the inertia of the user's water use behavior, and the feature importance of the long-term lag water consumption feature characterizes the periodicity of the user's water use behavior. The larger the feature importance value, the more obvious the corresponding water use behavior characteristics are.

3. The water demand prediction method for water supply pipe networks based on multi-source data fusion and hybrid model according to claim 1, characterized in that The feature importance is the degree of contribution of each dimension of data in the feature matrix to the prediction result after training the water demand prediction model using the XGBoost algorithm, and the result is represented by the Gain value.

Citation Information

Patent Citations

  • City short-term water consumption prediction method based on least square support vector machine model

    CN104715292A

  • Daily electricity consumption prediction method

    CN111178611A

  • Daily water consumption prediction method based on big data

    CN111210093A

  • Intelligent building user energy consumption behavior prediction method based on LSTM neural network

    CN113780684A

  • Residential water consumption prediction method based on MIC-XGBoost algorithm

    CN114792169A