Method and device for predicting water consumption of residential area based on multi-source data fusion
Through multi-source data fusion and gradient enhancement regression tree model, the problem of difficulty in considering the mutual influence between multiple influencing factors in the existing technology is solved, and high-precision prediction of daily water consumption in residential communities is achieved.
Patent Information
- Application Number
- CN202510288633.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively consider the mutual influence between multiple influencing factors in the prediction of water consumption in residential communities, resulting in low prediction accuracy.
Using a multi-source data fusion method, the factors that have the greatest impact on water consumption among meteorological factors are determined by collecting and processing the historical data of the daily water consumption in residential communities and its multi-source influencing factors, and a predictive regression model based on gradient enhancement regression tree is constructed.
The accurate prediction of the daily water consumption of residential communities is achieved, with the prediction accuracy within ±6%, and the overfitting problem is avoided.
Smart Images

Figure CN120218329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting water consumption in residential communities, and particularly to a method and device for predicting water consumption in residential communities based on multi-source data fusion. Background Art
[0002] The water demand of residential buildings is jointly affected by various conditions. These influencing factors include meteorology, season, policy, culture, and hot events, etc. Temperature, ground temperature, humidity, air pressure, wind speed, etc. in meteorological factors all affect the water consumption of residential buildings, and these factors jointly affect the water consumption of residential buildings. For example, the water consumption of residential buildings changes with the temperature. As the temperature rises, the water consumption for household bathing, laundry, and landscaping increases. The higher the temperature, the relatively higher the water consumption. According to the statistical data of residential water consumption in a certain city, it is found that the residential water consumption in summer is about 1.2 times that in winter, and the residential water consumption during high-temperature periods fluctuates with the temperature, thus reflecting the huge impact of temperature on residential water consumption.
[0003] Currently, water resources are facing problems of overall shortage and uneven distribution, which bring great challenges to the work of urban water service departments. With the development of artificial intelligence technology and the enhancement of computer computing power, water service departments in various cities have successively launched research on intelligent water supply projects, and the core content of which is the accurate prediction of residential water consumption.
[0004] The prediction of residential water consumption can be divided into medium- and long-term prediction and short-term prediction in terms of time dimension. Among them, short-term prediction mainly focuses on the prediction of daily and hourly water consumption. The prediction results can assist water service departments in the planning, management, and operation of water supply systems, and contribute to the sustainable development of cities.
[0005] There are many methods for predicting residential water consumption. Currently, they can be roughly divided into three categories:
[0006] The first category is the time series prediction method, which only relies on historical data for modeling and prediction, such as the autoregressive method, etc. Currently, this method generally targets one of many factors to establish a prediction model, without considering the mutual influence among various factors, resulting in a one-sided prediction model and low accuracy.
[0007] The second category is the structural analysis method. In addition to using historical data, it also needs to consider other factors related to water consumption. However, this method requires giving the explicit relationship between various influencing factors and water consumption, and such a relationship is not easy to obtain.
[0008] The third category is the system method. Similar to the structural analysis method, it uses various influencing factors of water consumption and historical data, and adopts non-linear models such as neural networks to establish a prediction system. This type of method requires a huge amount of historical data to train the model, and it is impossible to know the importance of factors such as air temperature, ground temperature, humidity, air pressure, and wind speed in meteorological factors relative to residential water consumption. Summary of the Invention
[0009] The present invention provides a method and device for predicting water consumption in residential communities based on multi-source data fusion to solve the technical problems existing in the prior art.
[0010] The technical solution adopted by the present invention to solve the technical problems existing in the prior art is:
[0011] A method for predicting water consumption in residential communities based on multi-source data fusion, the method comprising the following steps:
[0012] Step 1, collect historical data of the daily water consumption of a residential community and its multi-source influencing factors at the same time node, process the collected historical data and make a training sample set;
[0013] Step 2, determine the correlation between multi-source influencing factors and between multi-source influencing factors and the daily water consumption of the residential community using the data in the training sample set; determine the factor in meteorological factors that has the greatest impact on the daily water consumption of the residential area;
[0014] Step 3, construct a prediction regression model for the daily water consumption of the residential community based on multi-source influencing factors; use the data in the training sample set to perform fitting training and verification on the prediction regression model; during training, optimize and adjust the parameters of the prediction regression model;
[0015] Step 4, collect real-time data of multi-source influencing factors corresponding to the residential community, and predict the daily water consumption of the residential community by the prediction regression model that has completed fitting training.
[0016] Furthermore, use a python program to analyze the collected data of the daily water consumption of the residential community and / or multi-source influencing factors, and process abnormal data, missing data and duplicate data;
[0017] Process the abnormal data as follows: delete the abnormal data and interpolate the missing data after deletion, and the interpolation method is to use hot deck imputation.
[0018] Furthermore, the multi-source influencing factors include meteorological factors, water use pattern factors and historical water consumption factors.
[0019] Further, the meteorological factors include: minimum air pressure, minimum temperature, average relative humidity, daily precipitation, maximum wind speed, wind direction of the maximum wind speed, wind direction of the extreme wind speed, sunshine duration, average ground temperature, and perceived temperature; the perceived temperature is divided into general perceived temperature and segmented perceived temperature. Let AT be the general perceived temperature, and let T s be the segmented perceived temperature; the calculation formulas for AT and T s are as follows:
[0020] AT = 1.07T + 0.2u - 0.65V - 2.7;
[0021]
[0022] In the formula:
[0023] u is the vapor pressure;
[0024] RH is the relative humidity;
[0025] T is the average temperature of the day;
[0026] T max is the highest temperature of the day;
[0027] T min is the lowest temperature of the day;
[0028] V is the average wind speed of the day.
[0029] Further, in step 2, the data in the training sample set is classified by data name. Among any two types of data extracted from the training sample set, one type of data is called A, and the other type of data is called B;
[0030] The correlation between multi-source influencing factors and / or between multi-source influencing factors and the daily water consumption of residential communities is calculated according to the following formula:
[0031]
[0032] In the formula:
[0033] ρ A,B is the correlation between A and B;
[0034] cov(A, B) is the covariance between A and B;
[0035] σ A is the standard deviation of type A data;
[0036] σ B is the standard deviation of type B data;
[0037] k is the data serial number in each type of data set;
[0038] K is the number of data in each type of data set;
[0039] A k is the k-th data in Class A data;
[0040] is the average value of Class A data;
[0041] B k is the k-th data in Class B data;
[0042] is the average value of Class B data;
[0043] ρ A,B > 0 indicates a positive correlation between A and B; ρ A,B < 0 indicates a negative correlation between A and B; ρ A,B the closer the absolute value of ρ is to 1, the greater the degree of linear correlation between A and B; A,B the closer the absolute value is to zero, the less linear correlation there is between A and B; 0 < ρ A,B < 1 indicates a correlation between A and B; |ρ A,B |≥0.8 indicates a very strong correlation between A and B; 0.6 < |ρ A,B |< 0.8 indicates a strong correlation between A and B; 0.2 < |ρ A,B |< 0.6 indicates a medium correlation between A and B; 0 < |ρ A,B |< 0.2 indicates a weak correlation between A and B.
[0044] Furthermore, in step 3, a prediction regression model is constructed based on the gradient boosting regression tree algorithm as follows:
[0045] f t (X, ω t ) = ρ t h t (X, ω t );
[0046]
[0047] X = (x1, x2,..., x H ), j = 1, 2,..., H;
[0048] Let the daily water consumption of residential communities and their multi-source influencing factor data at the same time node be a set of samples. For N sets of samples, the optimal value of the prediction regression model is as follows:
[0049]
[0050] In the above formulas:
[0051] X is the input sample; it is a vector containing multi-source influencing factor data;
[0052] j is the serial number of the influencing factor in the input sample;
[0053] H is the number of influencing factors in the input sample;
[0054] x j is the j-th influencing factor in the input sample vector X;
[0055] t is the serial number of the regression tree;
[0056] T is the number of regression trees;
[0057] ω is the parameter vector of the algorithm;
[0058] ω t is the parameter to be adjusted for the t-th regression tree;
[0059] ρ t is the set parameter of the t-th regression tree;
[0060] C1, C2, and C3 are constants;
[0061] θ1 and θ2 are thresholds calculated according to ω t ;
[0062] F(X, ω) is the predicted water consumption value corresponding to X when the parameter is ω;
[0063] f t (X, ω t ) is the predicted water consumption value corresponding to X for the t-th regression tree;
[0064] F * is the optimal value of the prediction regression model;
[0065] i is the serial number of the sample group;
[0066] N is the number of sample groups;
[0067] y i is the daily water consumption of the residential community in the i-th group of samples;
[0068] X i is the multi-source influencing factor data vector in the i-th group of samples;
[0069] F(X i , ω) is the predicted water consumption value corresponding to the i-th group of samples when the parameter is ω;
[0070] The argmin function represents the variable value when the objective function reaches the minimum;
[0071] L() is the loss function.
[0072] Furthermore, in step 3, the genetic algorithm is used to adjust the following hyperparameters in the gradient boosting regression tree algorithm: the maximum number of iterations of the weak learner, the weight reduction coefficient of each weak learner, and the maximum depth of the regression tree; where:
[0073] The weight reduction coefficient of each weak learner, also known as the step size, has a value selection range of 0.01 - 1.00;
[0074] The maximum number of iterations of the weak learner, also known as the maximum number of weak learners. When the step size is 1, its value selection range is 10 - 150;
[0075] The weight reduction coefficient of each weak learner, also known as the step size. When the step size is 1, its value selection range is 2 - 100;
[0076] The maximum depth of the regression tree is used to control overfitting. When the model has a relatively large number of samples or a relatively large number of influencing factors, the specific value of the maximum depth of the regression tree depends on the data distribution.
[0077] Furthermore, the iterative process of the gradient boosting regression tree algorithm is as follows:
[0078] Step A1, set a group of data as a sample; define:
[0079]
[0080] Step A2, construct the training sample based on the regression tree as follows:
[0081]
[0082] Construct the objective function based on the regression tree as follows:
[0083] L(y i ,F(X i ))=(y i -F(X i )) 2 ;
[0084] Step A3, use the gradient descent direction to train the decision tree to obtain the following fitting data:
[0085] Let ω * be the best fitting data, and the calculation formula of ω * is as follows:
[0086]
[0087] Step A4, let ρ * be the best step size in the gradient descent direction, and the calculation formula of ρ * is as follows:
[0088]
[0089] Step A5, the weak learner of the t-th regression tree is obtained as:
[0090] f t = ρ * h t (X i , ω * );
[0091] Step A6, the predicted function after iteration is:
[0092] F t (X) = F t-1 (X) + f t ;
[0093] If the loss function satisfies the error convergence condition or the t value of the obtained regression tree reaches the preset value, the iteration terminates; if not, continue the iteration;
[0094] In the above formulas:
[0095] F(X) is an array of regression trees that have not been fully trained;
[0096] f i is the weak learner corresponding to the i-th group of samples;
[0097] D is the training sample based on the regression tree;
[0098] (X i y i ) represents an array that includes the independent variable X i and the dependent variable y i ;
[0099] D(X i ) is the regression prediction corresponding to the i-th group of samples;
[0100] F t (X) is the predicted water consumption value of the t-th regression tree;
[0101] F t-1 (X) is the predicted water consumption value of the (t - 1)-th regression tree;
[0102] F t-1 (X i ) is the predicted water consumption value of the (t - 1)-th regression tree corresponding to the i-th group of samples;
[0103] f t is the weak learner of the t-th regression tree;
[0104] h t (X i , ω* ) indicates that the t-th regression tree adopts the optimal fitting parameter ω * for the function value;
[0105] ρ t0 is the initial weight of the t-th regression tree.
[0106] Furthermore, in step 3, after training is completed, the importance of multi-source influencing factors is evaluated. By analyzing the importance of influencing factors, influencing factors are selected, and those with less impact on the prediction effect are removed, thereby simplifying the model;
[0107] The evaluation of the importance of multi-source influencing factors refers to the Gini index of each influencing factor; when evaluating the importance of multi-source influencing factors, first calculate the contribution of each influencing factor in each tree of the random forest, take the average of the contributions of all influencing factors, and finally compare the contribution sizes among influencing factors; the calculation formula is as follows:
[0108]
[0109] In the formula:
[0110] e is the serial number of the tree in the random forest;
[0111] h is the serial number of the influencing factor;
[0112] q, l, r are the node serial numbers of the tree in the random forest;
[0113] c is the class serial number of the tree in the random forest;
[0114] C is the number of classes of the tree in the random forest;
[0115] I is the number of trees in the random forest;
[0116] Q is the set of nodes where the h-th influencing factor appears in the e-th tree;
[0117] is the Gini index of the q-th node of the e-th tree;
[0118] is the proportion of the c-th class in the e-th tree;
[0119] is the Gini index of the l-th node of the e-th tree after branching;
[0120] is the Gini index of the r-th node of the e-th tree after branching;
[0121] is the importance of the h-th influencing factor at the q-th node of the e-th tree;
[0122] is the importance of the h-th influencing factor in the e-th tree;
[0123] is the importance of the h-th influencing factor.
[0124] The present invention also provides a device for a method of predicting water consumption in a residential community based on multi-source data fusion, including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the steps of the method of predicting water consumption in a residential community based on multi-source data fusion as described above.
[0125] The advantages and positive effects of the present invention are as follows: A method of predicting water consumption in a residential community based on multi-source data fusion according to the present invention constructs a prediction regression model of daily water consumption in a residential community based on multi-source influencing factors, with meteorological factors (such as temperature, ground temperature, humidity, air pressure, wind speed, etc.) and historical water consumption data of the residential area as the basis. By means of correlation analysis, the factor with the greatest influence on the daily water consumption in the residential area among meteorological factors is determined. Finally, the relationship between these factors and water consumption is quantified through gradient boosting regression trees.
[0126] The present invention utilizes multiple meteorological factors such as average ground temperature, perceived temperature, minimum temperature, minimum air pressure, average relative humidity, maximum wind speed, sunshine duration, daily precipitation, extreme wind speed and direction, and maximum wind speed and direction. Through the gradient boosting regression tree algorithm, the daily water consumption of a residential community with an accuracy within ±6% can be predicted. In addition, the present invention also analyzes that the factor with the greatest influence on the daily water consumption of a residential community is the average ground temperature (the feature importance accounts for 81% of the total), followed by the perceived temperature (the feature importance accounts for 7% of the total). This indicates that the hypothesis that meteorological conditions can be used to predict human water use behavior holds. The Pearson correlation coefficient of the training group predicted by the present invention is 0.987, and the Pearson correlation coefficient of the test group is 0.971. The Pearson correlation coefficients of the two groups are relatively close. Thus, it can be seen that the algorithm does not have the problem of overfitting.
[0127] During the use of the prediction regression model of daily water consumption in a residential community, the algorithm can continuously incorporate newly added water consumption data and meteorological factors into the historical water consumption data set and the meteorological factor data set. This enables the basic database of the algorithm to be updated and refined, thereby better predicting water consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0128] Figure 1 is a flowchart of the working process of a method of predicting water consumption in a residential community based on multi-source data fusion according to the present invention.
[0129] Figure 2 is a graph of one-day water consumption monitoring data that has not been processed.
[0130] Figure 3 It is a flowchart for imputation processing of missing data.
[0131] Figure 4 It is a graph of daily water consumption monitoring data processed by a Python program.
[0132] Figure 5 It is a statistical heat map for significance detection.
[0133] Figure 6 It is a scatter plot of the relationship between sunshine duration of weakly correlated meteorological factors and water consumption.
[0134] Figure 7 It is a scatter plot of the relationship between the maximum wind speed of weakly correlated meteorological factors and water consumption.
[0135] Figure 8 It is a convergence graph of genetic algorithm parameter tuning iteration.
[0136] Figure 9 It is a frequency statistical graph of the prediction relative error of the prediction regression model for the daily water consumption of residential communities.
[0137] Figure 10 It is a scatter plot of the relationship between the predicted water consumption using the prediction regression model for the daily water consumption of residential communities and the actual water consumption.
[0138] Figure 11 It is a pie chart of the feature importance of meteorological factors including average ground temperature with respect to water consumption.
[0139] Figure 12 It is a pie chart of the feature importance of meteorological factors with respect to water consumption after removing the average ground temperature.
[0140] Figure 3 In it: N represents no; Y represents yes; o represents the counting variable of water consumption data; g represents the counting variable of abnormal data; d represents the counting variable of water consumption data after imputation is completed. Specific implementation manners
[0141] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0142] The Chinese interpretations of the following English words, phrases and abbreviations are as follows:
[0143] Python program: It refers to a program written in the Python language. Python is a high-level programming language widely used in fields such as data analysis, artificial intelligence, and website development.
[0144] GBRT Algorithm: GBRT is the abbreviation of Gradient Boosting Regression Trees, which is translated as the "Gradient Boosting Regression Trees" algorithm. It is an ensemble learning algorithm that improves the prediction performance of the model by combining multiple weak regression models (usually decision trees).
[0145] Boosting Algorithm: An ensemble learning method that combines multiple weak classifiers into a strong classifier to gradually improve the model performance. Its basic idea is to make improvements based on the errors in the previous training each time. GBRT is an implementation of the Boosting method.
[0146] AdaBoost Algorithm: AdaBoost is the abbreviation of Adaptive Boosting, which is translated as the "Adaptive Boosting Algorithm". It is a common Boosting algorithm that improves the model performance by increasing the attention to misclassified samples in each round and is often used in classification problems.
[0147] KNN: KNN is the abbreviation of K-Nearest Neighbors, which is translated as the "K-Nearest Neighbors Algorithm". This is an algorithm for classification and regression. Its basic idea is: given a data point, find its K nearest neighbors, and then predict the target value based on the categories or values of these neighbors.
[0148] Ranking: Translated as "sorting". In machine learning and information retrieval, Ranking usually refers to sorting a set of objects according to certain criteria so that the sorted objects are more prioritized or relevant under certain metrics. For example, in a recommendation system, Ranking can sort items based on the user's historical behavior to recommend the most relevant content to the user.
[0149] Please refer to Figures 1 to 12 , a method for predicting the water consumption of residential communities based on multi-source data fusion, and the method includes the following steps:
[0150] Step 1, collect the historical data of the daily water consumption of the residential community and its multi-source influencing factors at the same time node, process the collected historical data and make it into a training sample set.
[0151] Step 2, use the data in the training sample set to determine the correlation between multi-source influencing factors and between multi-source influencing factors and the daily water consumption of the residential community; determine the factor in meteorological factors that has the greatest impact on the daily water consumption of the residential area.
[0152] Step 3: Construct a prediction regression model for the daily water consumption of residential communities based on multi-source influencing factors; use the data in the training sample set to fit, train, and validate the prediction regression model; during training, optimize and adjust the parameters of the prediction regression model.
[0153] Step 4: Collect real-time data on multi-source influencing factors for the corresponding residential community, and predict the daily water consumption of the residential community using the prediction regression model that has completed fitting training.
[0154] Preferably, a python program can be used to analyze the collected daily water consumption data and / or multi-source influencing factor data of the residential community, and process abnormal data, missing data, and duplicate data.
[0155] The abnormal data can be processed as follows: delete the abnormal data and interpolate the missing data after deletion. The interpolation method adopts hot deck imputation.
[0156] Preferably, the multi-source influencing factors can include meteorological factors, water usage pattern factors, and historical water consumption factors.
[0157] Preferably, the meteorological factors can include: minimum air pressure, minimum temperature, average relative humidity, daily precipitation, maximum wind speed, wind direction of the maximum wind speed, wind direction of the extreme wind speed, sunshine duration, average ground temperature, and perceived temperature.
[0158] Preferably, the perceived temperature can be divided into general perceived temperature and segmented perceived temperature. Let AT be the general perceived temperature, and let T s be the segmented perceived temperature; the calculation formulas for AT and T s are as follows:
[0159] AT = 1.07T + 0.2u - 0.65V - 2.7;
[0160]
[0161]
[0162] In the formula:
[0163] u is the vapor pressure;
[0164] RH is the relative humidity;
[0165] T is the average temperature of the day;
[0166] T max is the highest temperature of the day;
[0167] T min is the lowest temperature of the day;
[0168] V is the average wind speed of the day.
[0169] Preferably, in step 2, the data in the training sample set can be classified by data name. Among any two types of data extracted from the training sample set, one type of data can be called A, and the other type of data can be called B.
[0170] The correlation between multi-source influencing factors and / or between multi-source influencing factors and the daily water consumption of residential communities can be calculated according to the following formula:
[0171]
[0172] In the formula:
[0173] ρ A,B is the correlation between A and B;
[0174] cov(A,B) is the covariance between A and B;
[0175] σ A is the standard deviation of type A data;
[0176] σ B is the standard deviation of type B data;
[0177] k is the data serial number in each type of data set;
[0178] K is the number of data in each type of data set;
[0179] A k is the kth data in type A data;
[0180] is the average value of type A data;
[0181] B k is the kth data in type B data;
[0182] is the average value of type B data;
[0183] ρ A,B >0 indicates that A and B are positively correlated; ρ A,B <0 indicates that A and B are negatively correlated; ρ A,B The closer the absolute value of ρ is to 1, the greater the linear correlation degree between A and B. A,B The closer the absolute value is to zero, the less linear correlation there is between A and B; 0 < ρ A,B <1 indicates that there is a correlation between A and B; |ρ A,B |≥0.8 indicates that the correlation between A and B is a very strong correlation; 0.6 < |ρ A,B |<0.8 indicates that the correlation between A and B is a strong correlation; 0.2 < |ρ A,B |<0.6 indicates that the correlation between A and B is a medium correlation; 0 < |ρ A,B|<0.2 indicates that the correlation between A and B is weakly correlated.
[0184] Preferably, in step 3, a prediction regression model can be constructed based on the gradient boosting regression tree algorithm as follows:
[0185] f t (X, ω t ) = ρ t h t (X, ω t );
[0186]
[0187] X = (x1, x2,..., x H ), j = 1, 2,…, H;
[0188] Let the daily water consumption of residential communities and the data of their multi-source influencing factors at the same time node be a set of samples. For N groups of samples, the optimal values of the prediction regression model are as follows:
[0189]
[0190] In the above formulas:
[0191] X is the input sample; it is a vector containing multi-source influencing factor data;
[0192] j is the serial number of the influencing factor in the input sample;
[0193] H is the number of influencing factors in the input sample;
[0194] x j is the jth influencing factor in the input sample vector X;
[0195] t is the serial number of the regression tree;
[0196] T is the number of regression trees;
[0197] ω is the parameter vector of the algorithm;
[0198] ω t is the parameter to be adjusted for the tth regression tree;
[0199] ρ t is the set parameter of the tth regression tree;
[0200] C1, C2, C3 are constants;
[0201] θ1, θ2 are thresholds calculated according to ω t ;
[0202] F(X, ω) is the predicted water consumption value corresponding to X when the parameter is ω;
[0203] f t (X, ω t ) is the predicted water consumption value corresponding to X for the t-th regression tree;
[0204] F * is the optimal value of the predicted regression model;
[0205] i is the sample group serial number;
[0206] N is the number of sample groups;
[0207] y i is the daily water consumption of the residential community in the i-th group of samples;
[0208] X i is the multi-source influencing factor data vector in the i-th group of samples;
[0209] D(X i , ω) is the predicted water consumption value corresponding to the i-th group of samples with parameter ω;
[0210] The argmin function represents the variable value when the objective function reaches the minimum value;
[0211] L() is the loss function.
[0212] Preferably, in step 3, a genetic algorithm can be used to adjust the following hyperparameters in the gradient boosting regression tree algorithm: the maximum number of iterations of the weak learner, the weight reduction coefficient of each weak learner, the maximum depth of the regression tree; where:
[0213] The weight reduction coefficient of each weak learner, also known as the step size, has a value selection range of 0.01 - 1.00.
[0214] The maximum number of iterations of the weak learner, also known as the maximum number of weak learners, when the step size is 1, its value selection range is 10 - 150.
[0215] The weight reduction coefficient of each weak learner, also known as the step size, when the step size is 1, its value selection range is 2 - 100.
[0216] The maximum depth of the regression tree is used to control overfitting. When the model sample size is relatively large or the number of influencing factors is relatively large, the specific value of the maximum depth of the regression tree depends on the data distribution.
[0217] ω t is the parameter vector to be adjusted for the t-th regression tree; ω t includes the maximum number of iterations of the weak learner, the learning rate of each weak learner, and the maximum depth of the regression tree.
[0218] Preferably, the iterative process of the gradient boosting regression tree algorithm can be as follows:
[0219] Step A1, set a set of data as a sample; define:
[0220]
[0221] Step A2, construct the training sample based on the regression tree as follows:
[0222]
[0223] Construct the objective function based on the regression tree as follows:
[0224] L(y i ,F(X i ))=(y i -F(X i )) 2 ;
[0225] Step A3, use the gradient descent direction to train the decision tree to obtain the following fitting data:
[0226] Let ω * be the best fitting data, and the calculation formula of ω * is as follows:
[0227]
[0228] Step A4, let ρ * be the best step size in the gradient descent direction, and the calculation formula of ρ * is as follows:
[0229]
[0230] Step A5, obtain the weak learner of the t-th regression tree as:
[0231] f t =ρ * h t (X i ,ω * );
[0232] Step A6, the predicted function after iteration is:
[0233] F t (X)=F t-1 (X)+f t ;
[0234] If the loss function satisfies the error convergence condition or the t value of the obtained regression tree reaches the preset value, the iteration terminates; if not, continue the iteration;
[0235] In the above formulas:
[0236] F(X) is the regression tree group that has not completed training;
[0237] f i is the weak learner corresponding to the i-th group of samples;
[0238] D is the training sample based on the regression tree;
[0239] (X i y i ) represents an array that includes the independent variable X i and the dependent variable y i ;
[0240] D(X i ) is the regression prediction corresponding to the i-th group of samples;
[0241] F t (X) is the predicted water consumption value of the t-th regression tree;
[0242] F t-1 (X) is the predicted water consumption value of the (t - 1)-th regression tree;
[0243] F t-1 (X i ) is the predicted water consumption value of the (t - 1)-th regression tree corresponding to the i-th group of samples;
[0244] f t is the weak learner of the t-th regression tree;
[0245] h t (X i , ω * ) represents the function value of the t-th regression tree using the optimal fitting parameter ω * ;
[0246] ρ t0 is the initial weight of the t-th regression tree.
[0247] Preferably, in step 3, after training is completed, the importance of multi-source influencing factors can be evaluated. By analyzing the importance of influencing factors, influencing factors are selected, and those with less influence on the prediction effect are removed, thereby simplifying the model.
[0248] The evaluation of the importance of multi-source influencing factors refers to the Gini index of each influencing factor; when evaluating the importance of multi-source influencing factors, first calculate the contribution of each influencing factor in each tree of the random forest, take the average of the contributions of all influencing factors, and finally compare the contribution sizes among the influencing factors; the calculation formula is as follows:
[0249]
[0250] In the formula:
[0251] e is the serial number of the tree in the random forest;
[0252] h is the serial number of the influencing factor;
[0253] q, l, and r are the node serial numbers of the tree in the random forest;
[0254] c is the class serial number of the tree in the random forest;
[0255] C is the number of classes of the tree in the random forest;
[0256] I is the number of trees in the random forest;
[0257] Q is the set of nodes where the h-th influencing factor appears in the e-th tree;
[0258] is the Gini index of the q-th node of the e-th tree;
[0259] is the proportion of the c-th class in the e-th tree;
[0260] is the Gini index of the l-th node of the e-th tree after branching;
[0261] is the Gini index of the r-th node of the e-th tree after branching;
[0262] is the importance of the h-th influencing factor at the q-th node of the e-th tree;
[0263] is the importance of the h-th influencing factor in the e-th tree;
[0264] is the importance of the h-th influencing factor.
[0265] The present invention also provides a device for a method of predicting water consumption in residential communities based on multi-source data fusion, including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the steps of the method of predicting water consumption in residential communities based on multi-source data fusion as described above.
[0266] The working process and working principle of the present invention will be further described below with reference to the preferred embodiments of the present invention:
[0267] A method for predicting the water consumption of residential communities based on multi-source data fusion, which is based on meteorological factors (such as temperature, ground temperature, humidity, air pressure, wind speed, etc.) and historical water consumption data of residential areas. By means of correlation analysis, the factors that have the greatest impact on the daily water consumption of residential areas among meteorological factors are determined. Finally, the gradient boosting regression tree is used to quantify the relationship between these factors and water consumption.
[0268] This method includes the following steps:
[0269] Step 1: Collect historical data on the daily water consumption of residential communities and their multi-source influencing factors at the same time node, process the collected historical data and make a training sample set.
[0270] Step 2: Use the data in the training sample set to determine the correlation between multi-source influencing factors and between multi-source influencing factors and the daily water consumption of residential communities; determine the factors that have the greatest impact on the daily water consumption of residential areas among meteorological factors.
[0271] Step 3: Build a prediction regression model for the daily water consumption of residential communities based on multi-source influencing factors; use the data in the training sample set to perform fitting training and verification on the prediction regression model; during training, optimize and adjust the parameters of the prediction regression model.
[0272] Step 4: Collect real-time data on multi-source influencing factors of the corresponding residential community, and predict the daily water consumption of the residential community by the prediction regression model that has completed fitting training.
[0273] In Step 1, use a Python program to process the historical data of the daily water consumption of residential communities and the meteorological factors of the geographical location of the community: In the monitoring data, due to the influence of the monitoring equipment itself and monitoring conditions, some of the monitored data are abnormal. When analyzing the monitoring data under the conditions of adhering to the truth and conforming to the actual situation, the processing and repair of abnormal data are particularly important.
[0274] Regarding the identification of invalid monitoring data, the basic idea is to find the data in the monitoring data that obviously does not conform to the water use law. The manifestations of data that do not conform to the water use law are generally divided into three types: data missing, data duplication, and data fluctuation significantly exceeding the threshold. Typical representatives in the unprocessed water consumption monitoring data are shown in the appendix Figure 2 .
[0275] Data missing is manifested as the flow value recorded at the corresponding moment being a null value. The identification of this kind of error is relatively simple, and directly identify the null value in the data column. Data duplication is manifested as the same data continuously appearing within the same period of time. Data fluctuation significantly exceeding the threshold is manifested as a single data being an obvious outlier.
[0276] After deleting abnormal data, imputation of missing data is required. To accurately restore the variation pattern of the water consumption in the community, the imputation method is based on hot deck imputation and takes into account statistical laws. Hot deck imputation finds an object in the complete data that is most similar to it and fills the current value with the most similar value. Hot deck imputation is essentially a special form of using KNN for prediction. KNN refers to K, while hot deck filling refers to the nearest 1. After finding the set of nearest neighbor points using the idea of hot deck imputation, the water consumption data is imputed according to the statistical laws of the set of nearest neighbor points. The process is shown in Appendix Figure 3 The water consumption data after data cleaning and imputation is shown in Appendix Figure 4 .
[0277] In step 2, before determining the correlation between meteorological factors and water consumption, some meteorological factors that affect short-term water consumption are determined in advance according to experience, including the perceived temperature, ground temperature, wind speed, wind direction, air pressure, and sunshine.
[0278] The perceived temperature is different from the temperature measured at the meteorological station. Because it is affected by comprehensive conditions such as air temperature, air humidity, and wind speed, its variation is more complex. There are two main current calculation formulas for the perceived temperature:
[0279] 1. General formula for perceived temperature:
[0280] AT = 1.07T + 0.2u - 0.65V - 2.7;
[0281]
[0282] 2. Piecewise calculation formula for perceived temperature:
[0283]
[0284] In the formula:
[0285] AT is the perceived temperature, in °C;
[0286] T s is the perceived temperature, in °C;
[0287] u is the vapor pressure, in hPa;
[0288] RH is the relative humidity, in %;
[0289] T is the daily average temperature, in °C;
[0290] T max is the daily maximum temperature, in °C;
[0291] T min is the daily minimum temperature, in °C;
[0292] V is the average wind speed of the day, with the unit of m / s.
[0293] In previous studies, power grid dispatchers and scientific researchers found that the perceived temperature affects the use of air conditioners in summer and electric heaters in winter by humans. This change in electrical load can be explained as the change in electrical load caused by the change in human comfort due to the change in meteorological conditions. This meteorological change also causes changes in human water use behavior. Therefore, the water consumption can be analyzed and predicted based on the factor of perceived temperature. Wind speed, air pressure, and sunshine duration also affect water consumption by influencing human sensations.
[0294] In addition, since the geographical location of the city is determined and does not change, and the geographical features around the city do not change significantly in a short period of time, the wind blowing from different directions of the city will have different subjective feelings and impacts on people. The wind direction even directly guides the water use behavior of people in specific regions of China.
[0295] The ground temperature is one of the important indicators characterizing the thermal energy characteristics of the soil. It is the expression that the soil surface absorbs solar radiation and then transforms it into soil thermal energy and transfers it to the deeper layer. Its change is more conservative and lagging than the change in air temperature. There are seasonal changes in the heat of the soil deep inside, whether it is the influence of solar radiation or the change in the heat flow inside the earth itself. The ground temperature can reflect the historical meteorological effect and is closer to the temperature in buildings without temperature control devices where people live and work. Therefore, the ground temperature also affects water consumption.
[0296] When considering the influence of meteorological factors on water consumption, it is necessary to process daily characteristic meteorological factors. First, normalize the meteorological data. Next, the influence of meteorological factors on water consumption will be studied to analyze the correlation degree between meteorological factors and water consumption to help improve the accuracy of water consumption prediction.
[0297] Classify the data in the training sample set according to the data name. Among any two types of data extracted from the training sample set, one type of data is called A and the other type of data is called B.
[0298] The correlation between multi-source influencing factors and / or between multi-source influencing factors and the daily water consumption of residential communities is calculated according to the following formula:
[0299]
[0300] In the formula:
[0301] ρ A,B is the correlation between A and B;
[0302] cov(A,B) is the covariance between A and B;
[0303] σ A is the standard deviation of type A data;
[0304] σ B is the standard deviation of Class B data;
[0305] k is the data serial number in each class of dataset;
[0306] K is the number of data in each class of dataset;
[0307] A k is the k-th data in Class A data;
[0308] is the average value of Class A data;
[0309] B k is the k-th data in Class B data;
[0310] is the average value of Class B data;
[0311] ρ A,B > 0 indicates that A and B are positively correlated; ρ A,B < 0 indicates that A and B are negatively correlated; ρ A,B The closer the absolute value of ρ is to 1, the greater the linear correlation degree between A and B; A,B The closer the absolute value is to zero, the less linear correlation there is between A and B; 0 < ρ A,B < 1 indicates that there is a correlation between A and B; |ρ A,B |≥0.8 indicates that the correlation between A and B is extremely strong; 0.6 < |ρ A,B |< 0.8 indicates that the correlation between A and B is strong; 0.2 < |ρ A,B |< 0.6 indicates that the correlation between A and B is medium; 0 < |ρ A,B |< 0.2 indicates that the correlation between A and B is weak.
[0312] Calculate the Pearson correlation coefficients (Table 1.a) and significance (two-tailed) (Table 1.b) between all meteorological factors and water consumption, and draw a heat map (attached Figure 5 ).
[0313] Among them, in Table 1.a:
[0314] Marking "-" before the value indicates negative correlation.
[0315] Marking ** after the value indicates that the correlation is significant at the 0.01 level (two-tailed).
[0316] Table 1.a Pearson correlation coefficients between meteorological factors and between meteorological factors and water consumption
[0317]
[0318] Table 1.b Correlation Significance between Meteorological Factors and between Meteorological Factors and Water Consumption (Two-tailed)
[0319]
[0320]
[0321] By observing the Pearson correlation coefficient table, significance table, and statistical heat map, it can be seen that at a significance (two-tailed) level of 0.01, the correlations between the general perceived temperature, segmented perceived temperature, minimum air pressure, ground temperature, maximum air pressure, average relative humidity, daily precipitation, maximum wind speed and direction, and water consumption are significant. Among them, the correlation of ground temperature is the strongest, with a Pearson correlation coefficient of 0.930; the correlations of sunshine duration and maximum wind speed with water consumption are the weakest. However, in the scatter plots of sunshine duration and water consumption (attached Figure 6 ) and maximum wind speed and water consumption (attached Figure 7 ), it can be found that when the sunshine duration is long and the maximum wind speed is high, the water consumption remains at a relatively low level, indicating that extreme sunshine and strong wind weather have an inhibitory effect on human water use behavior. Although the significance levels between these two meteorological factors of long sunshine duration and high maximum wind speed and water consumption are not high, they are still two factors that cannot be ignored in the subsequent gradient boosting regression analysis.
[0322] The significances between the general perceived temperature and the segmented perceived temperature and water consumption are both very high, and both show a positive correlation with water consumption. However, the essence of both is the human sensitivity index to comprehensive meteorological factors. Therefore, it is incorrect to include both indicators as influencing factors in subsequent research. Therefore, a choice needs to be made between these two influencing factors. Compared with the general perceived temperature, the segmented perceived temperature has a smaller degree of dispersion. In addition, statistically, the Pearson correlation coefficient between the general perceived temperature and water consumption is 0.856, which is less than the Pearson correlation coefficient of 0.896 between the segmented perceived temperature and water consumption. Therefore, it is most appropriate to use the segmented perceived temperature as the representative of human intuitive perception of meteorological factors in subsequent calculations.
[0323] In Step 3, a predictive regression model is constructed based on the Gradient Boosting Regression Tree algorithm. The Gradient Boosting algorithm is a machine learning technique used for regression, classification, and ranking tasks and is part of the Boosting algorithm family. The GradientBoosting algorithm constructs a learner at each iteration that can reduce the loss along the steepest direction of the gradient to make up for the deficiencies of the existing model. The classic AdaBoost algorithm can only handle binary classification learning tasks using the exponential loss function, while the Gradient Boosting method can handle various learning tasks (multi-classification, regression, Ranking, etc.) by setting different differentiable loss functions, greatly expanding its application scope. The Gradient Boosting algorithm uses the negative gradient of the loss function as the way to fit the residuals. If the base function among them is a regression tree, the Gradient Boosting Regression Tree is obtained.
[0324] The predictive regression model constructed based on the Gradient Boosting Regression Tree algorithm is as follows:
[0325] f t (X, ω t ) = ρ t h t (X, ω t );
[0326]
[0327] X = (x1, x2,..., x H ), j = 1, 2,…, H;
[0328] Let the daily water consumption of residential communities and the data of their multi-source influencing factors at the same time node be a set of samples. For N sets of samples, the optimal values of the predictive regression model are as follows:
[0329]
[0330] In the above formulas:
[0331] X is the input sample; it is a vector containing multi-source influencing factor data;
[0332] j is the serial number of the influencing factor in the input sample;
[0333] H is the number of influencing factors in the input sample;
[0334] x j is the j-th influencing factor in the input sample vector X;
[0335] t is the serial number of the regression tree;
[0336] T is the number of regression trees;
[0337] ω is the parameter vector of the algorithm;
[0338] ωt is the parameter to be adjusted for the $t$-th regression tree;
[0339] $\rho$ t is the set parameter of the $t$-th regression tree;
[0340] $C_1$, $C_2$, $C_3$ are constants;
[0341] $\theta_1$, $\theta_2$ are thresholds calculated according to $\omega$ t ;
[0342] $F(X, \omega)$ is the predicted water consumption value of $X$ corresponding to the parameter $\omega$;
[0343] $f$ t $(X, \omega$ t ) is the predicted water consumption value of $X$ corresponding to the $t$-th regression tree;
[0344] $F$ * is the optimal value of the prediction regression model;
[0345] $i$ is the sample group serial number;
[0346] $N$ is the number of sample groups;
[0347] $y$ i is the daily water consumption of the residential community in the $i$-th group of samples;
[0348] $X$ i is the multi-source influencing factor data vector in the $i$-th group of samples;
[0349] $F(X$ i , $\omega)$ is the predicted water consumption value of the $i$-th group of samples corresponding to the parameter $\omega$;
[0350] The argmin function represents the variable value when the objective function reaches the minimum;
[0351] $L()$ is the loss function.
[0352] The iterative process of the above gradient boosting regression tree algorithm is as follows:
[0353] Step A1, set a set of data as a sample; define:
[0354]
[0355] Step A2, construct the training samples based on the regression tree as follows:
[0356]
[0357] Construct the objective function based on the regression tree as follows:
[0358] $L(y$ i , $F(X$ i)) = (y i -F(X i )) 2 ;
[0359] Step A3: Train the decision tree in the gradient descent direction to obtain the following fitting data:
[0360] Let ω * be the best fitting data. The calculation formula for ω * is as follows:
[0361]
[0362] Step A4: Let ρ * be the optimal step size in the gradient descent direction. The calculation formula for ρ * is as follows:
[0363]
[0364] Step A5: Obtain the weak learner of the t-th regression tree as:
[0365] f t = ρ * h t (X i , ω * );
[0366] Step A6: The predicted function after iteration is:
[0367] F t (X) = F t-1 (X) + f t ;
[0368] If the loss function satisfies the error convergence condition or the t value of the obtained regression tree reaches the preset value, the iteration terminates; otherwise, continue the iteration.
[0369] In the above formulas:
[0370] F(X) is an array of regression trees that have not completed training;
[0371] f i is the weak learner corresponding to the i-th group of samples;
[0372] D is the training sample based on the regression tree;
[0373] (X i y i ) represents an array that includes the independent variable X i and the dependent variable y i ;
[0374] F(X i) is the regression prediction corresponding to the i-th group of samples;
[0375] F t (X) is the predicted water consumption value of the t-th regression tree;
[0376] F t-1 (X) is the predicted water consumption value of the (t - 1)-th regression tree;
[0377] F t-1 (X i ) is the predicted water consumption value of the (t - 1)-th regression tree corresponding to the i-th group of samples;
[0378] f t is the weak learner of the t-th regression tree;
[0379] h t (X i ,ω * ) represents the function value of the t-th regression tree using the optimal fitting parameter ω * ;
[0380] ρ t0 is the initial weight of the t-th regression tree.
[0381] In step 3, the genetic algorithm is used to adjust the following hyperparameters in the gradient boosting regression tree algorithm: the maximum number of iterations of the weak learner, the weight reduction coefficient of each weak learner, and the maximum depth of the regression tree; the gradient boosting regression tree algorithm is also known as the GBRT algorithm. There are three hyperparameters that need to be adjusted in this algorithm. These three hyperparameters are: the maximum number of iterations of the weak learner (learning_rate), the weight reduction coefficient of each weak learner (n_estimators), and the maximum depth of the regression tree (max_depth).
[0382] The maximum number of iterations of the weak learner, also known as the maximum number of weak learners, if the maximum number of iterations of the weak learner is too small, it is easy to underfit, and if the maximum number of iterations of the weak learner is too large, it is easy to overfit. Therefore, generally, a moderate value is selected, and the default is 100. In the actual process of parameter tuning, the weight reduction coefficient of each weak learner should be considered together with the maximum number of iterations of the weak learner.
[0383] The weight reduction coefficient of each weak learner, also known as the step size, for the same training set fitting effect, a smaller step size means that more iterations of the weak learner are required. Usually, the step size and the maximum number of iterations are used together to determine the fitting effect of the algorithm. Therefore, the weight reduction coefficient of each weak learner and the maximum number of iterations of the weak learner should be adjusted together. Generally speaking, it can be adjusted starting from a smaller value, and the default is 1.
[0384] The maximum depth of the regression tree, which can control overfitting because the deeper the classification tree, the more likely it is to overfit. Generally speaking, when the model has a large number of samples and many influencing factors, it is recommended to limit this maximum depth, and the specific value depends on the data distribution.
[0385] In the present invention, the parameter adjustment range of the maximum number of iterations of the weak learner is set to 0.01 - 1.00 (step size 0.01); the parameter adjustment range of the weight reduction coefficient of each weak learner is set to 10 - 150 (step size 1); the parameter adjustment range of the maximum depth of the regression tree is set to 2 - 100 (step size 1). Therefore, the number of parameter combinations of this algorithm is 1,358,280. Due to the huge number of parameter combinations, it is not easy to obtain the optimal parameter combination by using the traditional parameter tuning method, so the genetic algorithm is used for parameter tuning in the present invention.
[0386] The specific idea of parameter tuning is to regard each parameter tuning process as a process of generating new offspring, and save the optimal individuals of the offspring and the parent generation each time. Using the Pearson correlation coefficient of the test group as the objective function, the convergence process is as shown in the appendix Figure 8 As shown. It can be seen that after 97 iterations, the Pearson correlation coefficient of the test group converges from 0.650 before parameter tuning to 0.971, and the parameter adjustment effect is relatively good.
[0387] Analyzing the results of the gradient boosting regression tree, it is obtained that the Pearson correlation coefficient of the training group of the prediction regression model of the daily water consumption of residential communities based on multi-source influencing factors is 0.987, and the Pearson correlation coefficient of the test group is 0.971. The Pearson correlation coefficients of the two groups are relatively close. Thus, it can be seen that the prediction regression model does not overfit.
[0388] Appendix Figure 9 shows the relative error between the predicted value and the actual value. By observing, it can be found that the relative error between all predicted values and the actual value is within 6%. Among them, the proportion of predicted value errors within ±2% accounts for 34.18% of the total, and the proportion within ±4% accounts for 83.54% of the total. Appendix Figure 10 is the comparison between the water consumption predicted according to meteorological factors and the actual water consumption. By observing this figure, it can be found that the fitting effect of this algorithm is better.
[0389] In the gradient boosting tree algorithm adopted by the present invention, a total of 10 features are included, namely the lowest air pressure, the lowest temperature, the average relative humidity, the daily precipitation, the maximum wind speed, the wind direction of the maximum wind speed, the wind direction of the extreme wind speed, the sunshine duration, the average ground temperature, and the perceived temperature. Next, it is necessary to determine the importance of these features. In this study, the evaluation of feature importance refers to the Gini index of each feature. The idea of feature importance evaluation is actually very simple. It mainly focuses on how much contribution each feature makes to each tree in the random forest, takes the average of these contributions, and finally compares the contribution sizes among features. The calculation formula is as follows:
[0390]
[0391] In the formula:
[0392] e is the serial number of the tree in the random forest;
[0393] h is the serial number of the influencing factor;
[0394] q, l, and r are the node serial numbers of the tree in the random forest;
[0395] c is the category serial number of the tree in the random forest;
[0396] C is the number of categories of the tree in the random forest;
[0397] I is the number of trees in the random forest;
[0398] Q is the set of nodes where the h-th influencing factor appears in the e-th tree;
[0399] is the Gini index of the q-th node of the e-th tree;
[0400] is the proportion of the c-th category in the e-th tree;
[0401] is the Gini index of the l-th node of the e-th tree after branching;
[0402] is the Gini index of the r-th node of the e-th tree after branching;
[0403] is the importance of the h-th influencing factor at the q-th node of the e-th tree;
[0404] is the importance of the h-th influencing factor in the e-th tree;
[0405] is the importance of the h-th influencing factor.
[0406] Table 2 shows the importance of each feature in the gradient boosting regression tree algorithm obtained through calculation. Attached Figure 11 and attached Figure 12 are the statistical attached drawings of the feature importance percentages drawn based on Table 2. Observing attached Figure 11 it can be found that the feature importance of the average ground temperature accounts for 81% of the total. Subsequently, it is the perceived temperature (7%), the minimum air temperature (5%), and the minimum air pressure (2%). Observing attached Figure 12 it can be found that after removing the influence of the average ground temperature, the feature importance of the perceived temperature has the highest proportion, accounting for 36% of the total. Subsequently, it is the minimum air temperature (24%), the minimum air pressure (10%), the average relative humidity (8%), the maximum wind speed (7%), the sunshine duration (6%), the daily precipitation (4%), the direction of the extreme wind speed (3%), and the direction of the maximum wind speed (2%).
[0407] Table 2. Feature Importance of Climate Features
[0408]
[0409] Affected by factors such as air temperature, precipitation, and solar radiation, the average ground temperature can reflect comprehensive meteorological factors. In addition, the average ground temperature is also affected by historical meteorological factors, and its change is more conservative and lagging than the change in air temperature. Since most human production and life activities are carried out in various forms of buildings, compared with other single meteorological factors, the average ground temperature can better represent the temperature felt by humans in structures. This can explain why the feature importance of the average ground temperature is the highest in this model. In addition, because people will inevitably engage in outdoor activities due to factors such as production, life, or entertainment, the perceived temperature becomes the most directly felt meteorological condition at this time. Therefore, in this model, the feature importance of the perceived temperature is second only to the average ground temperature. Other meteorological factors will also have a certain impact on people's water use behaviors. For example, the average relative humidity, sunshine duration, wind direction, and wind speed will affect people's bathing behaviors, and the daily precipitation, sunshine duration, and wind speed will affect people's laundry behaviors, etc. Incorporating these meteorological factors into the model can improve the accuracy of water consumption prediction.
[0410] The above-mentioned algorithm software such as the Python program, gradient boosting regression tree algorithm, Boosting algorithm, AdaBoost algorithm, GBRT algorithm, regression tree, decision tree, hot deck imputation, and KNN can all adopt the applicable algorithm software in the prior art, or adopt the algorithm software in the prior art and be constructed by conventional technical means.
[0411] The embodiments described above are only used to illustrate the technical idea and features of the present invention. The purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The patent scope of the present invention cannot be limited only by these embodiments. That is, any equivalent changes or modifications made in accordance with the spirit disclosed by the present invention still fall within the patent scope of the present invention.
Claims
1. A method for predicting water consumption in residential areas based on multi-source data fusion, characterized in that: The method comprises the following steps: Step 1: Collect historical data of daily water consumption and its multi-source influencing factors of residential areas at the same time node, process the collected historical data and make them into a training sample set; Step 2, using the data in the training sample set to determine the correlation between the multi-source influencing factors and between the multi-source influencing factors and the daily water consumption of the residential area; Determine the meteorological factors that have the greatest impact on daily water consumption in residential areas; Step 3: construct a prediction regression model for daily water consumption in residential areas based on multi-source influencing factors; use data from the training sample set to perform fitting training and verification on the prediction regression model; during training, optimize and adjust the parameters of the prediction regression model; Step 4: Collect real-time data of multi-source influencing factors of the corresponding residential area, and use the prediction regression model that has completed fitting training to predict the daily water consumption of the residential area.
2. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 1 is characterized in that: The python program is used to analyze the collected data of daily water consumption and / or multi-source influencing factors in residential areas, and to process abnormal data, missing data and duplicate data; The abnormal data are processed as follows: the abnormal data are deleted and the missing data are interpolated after deletion, and the interpolation method is hot card interpolation.
3. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 1 is characterized in that: Multi-source influencing factors include meteorological factors, water use pattern factors and historical water consumption factors.
4. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 3 is characterized in that: Meteorological factors include: minimum air pressure, minimum temperature, average relative humidity, daily precipitation, maximum wind speed, wind direction of maximum wind speed, wind direction of maximum wind speed, sunshine hours, average ground temperature, and body temperature. The body temperature is divided into general body temperature and segmented body temperature. AT is the general body temperature and T is the segmented body temperature. s AT and T are segmented body temperature. s The calculation formula is as follows: AT=1.07T+0.2u-0.65V-2.7; Where: u is the water vapor pressure; RH is relative humidity; T is the average temperature of the day; T max is the highest temperature of the day; T min The lowest temperature of the day; V is the average wind speed for the day.
5. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 1 is characterized in that: In step 2, the data in the training sample set are classified according to the data name. Among any two types of data extracted from the training sample set, one type of data is called A and the other type of data is called B. The correlation between multiple influencing factors and / or between multiple influencing factors and daily water consumption in residential areas is calculated according to the following formula: Where: ρ A,B is the correlation between A and B; cov(A,B) is the covariance between A and B; σ A is the standard deviation of Class A data; σ B is the standard deviation of Class B data; k is the data sequence number in each type of data set; K is the number of data in each type of data set; A k is the kth data in category A; is the average value of Class A data; B k is the kth data in category B; is the average value of category B data; ρ A,B >0, indicating that A and B are positively correlated; ρ A,B <0, indicating that A and B are negatively correlated; ρ A,B The closer the absolute value of is to 1, the greater the linear correlation between A and B. A,B The closer the absolute value is to zero, the less linear correlation there is between A and B; 0<ρ A,B <1, indicating that A and B are correlated; |ρ A,B |≥0.8, indicating that the correlation between A and B is extremely strong; 0.6<|ρ A,B |<0.8, indicating that the correlation between A and B is strong; 0.2<|ρ A,B |<0.6, indicating that the correlation between A and B is moderate; 0<|ρ A,B |<0.2 indicates that the correlation between A and B is weak.
6. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 1 is characterized in that: In step 3, the prediction regression model is constructed based on the gradient boosting regression tree algorithm as follows: f t (X,ω t )=ρ t h t (X,ω t ); x=(x1,x2,…,x H ),j=1,2,...,H: Assume that the daily water consumption of residential areas and its multi-source influencing factors at the same time node are a group of samples. For M groups of samples, the optimal value of the prediction regression model is as follows: In the above formulas: X is the input sample; it is a vector containing multi-source influencing factor data; j is the serial number of the influencing factor in the input sample; H is the number of influencing factors in the input sample; x j is the jth influencing factor in the input sample vector X; t is the regression tree number; T is the number of regression trees; ω is the parameter vector of the algorithm; ω t is the parameter to be adjusted for the tth regression tree; ρ t Set parameters for the tth regression tree; C1, C2, and C3 are constants; θ1 and θ2 are based on ω t The calculated threshold value; F(X,ω) is the predicted value of water consumption corresponding to X when the parameter is ω; f t (X,ω t ) is the predicted value of water consumption corresponding to X of the tth regression tree; F * is the optimal value of the predictive regression model; i is the sample group number; N is the number of sample groups; y i is the daily water consumption of the residential area in the i-th group of samples; X i is the data vector of multi-source influencing factors in the i-th group of samples; F(X i ,ω) is the predicted value of water consumption corresponding to the i-th group of samples when the parameter is ω; The argmin function represents the variable value when the objective function reaches the minimum value; L() is the loss function.
7. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 6 is characterized in that: In step 3, a genetic algorithm is used to adjust the following hyperparameters in the gradient boosting regression tree algorithm: the maximum number of iterations of the weak learner, the weight reduction coefficient of each weak learner, and the maximum depth of the regression tree; where: The weight reduction coefficient of each weak learner, also known as the step size, is selected in the range of 0.01-1.00; The maximum number of iterations of the weak learner, also known as the maximum number of weak learners, is 10-150 when the step size is 1; The weight reduction coefficient of each weak learner is also called the step size. When the step size is 1, its value range is 2-100; The maximum depth of the regression tree is used to control overfitting. When the model has a large number of samples or many influencing factors, the specific value of the maximum depth of the regression tree depends on the distribution of the data.
8. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 6 is characterized in that: The iterative process of the gradient boosting regression tree algorithm is as follows: Step A1, assume a set of data is a sample; define: Step A2, construct the training samples based on regression tree as follows: The objective function based on regression tree is constructed as follows: L(y i ,F(X i ))=(y i -F(X i )) 2 ; Step A3, train the decision tree using the gradient descent direction to obtain the following fitting data: Assume ω * is the best fit to the data, ω * The calculation formula is as follows: Step A4: Assume ρ * is the optimal step size in the gradient descent direction, ρ * The calculation formula is as follows: Step A5, the weak learner of the tth regression tree is obtained as: f t =ρ * h t (X i ,oh * ); Step A6, the prediction function after iteration is: F t (X)=F t-1 (X)+f t ; If the loss function meets the error convergence condition or the t value of the obtained regression tree reaches the preset value, the iteration is terminated; if not, the iteration continues; In the above formulas: F(X) is the set of regression trees that have not completed training; f i is the weak learner corresponding to the i-th group of samples; D is the training sample based on regression tree; (X i y i ) represents an array, which includes the independent variable X i With the dependent variable y i ; F(X i ) is the regression prediction corresponding to the i-th group of samples; F t (X) is the water consumption prediction value of the tth regression tree; F t-1 (X) is the water consumption prediction value of the t-1th regression tree; F t-1 (X i ) is the predicted value of water consumption of the ith group of samples corresponding to the t-1th regression tree; f t is the weak learner of the tth regression tree; h t (X i ,ω * ) indicates that the tth regression tree uses the best fitting parameter ω * The function value of ρ t0 is the initial weight of the tth regression tree.
9. The method for predicting water consumption in residential areas based on multi-source data fusion according to claim 1 is characterized in that: In step 3, after the training is completed, the importance of multi-source influencing factors is evaluated. By analyzing the importance of influencing factors, influencing factors are selected and factors with little impact on the prediction effect are removed, thereby simplifying the model; The importance of multi-source influencing factors is evaluated by referring to the Gini index of each influencing factor. When evaluating the importance of multi-source influencing factors, the contribution of each influencing factor to each tree in the random forest is first calculated, the contribution of all influencing factors is averaged, and finally the contribution between the influencing factors is compared. The calculation formula is as follows: Where: e is the ordinal number of the tree in the random forest; h is the serial number of the influencing factor; q, l, r are the node numbers of the trees in the random forest; c is the category number of the tree in the random forest; C is the number of tree categories in the random forest; I is the number of trees in the random forest; Q is the set of nodes where the hth influencing factor appears in the eth tree; is the Gini index of the qth node of the eth tree; is the proportion of the cth category in the eth tree; is the Gini index of the lth node of the eth tree after branching; is the Gini index of the rth node of the eth tree after branching; is the importance of the h-th influencing factor at the q-th node of the e-th tree; is the importance of the h-th influencing factor in the e-th tree; is the importance of the hth influencing factor.
10. A device for predicting water consumption in residential areas based on multi-source data fusion, comprising a memory and a processor, characterized in that: The memory is used to store computer programs; the processor is used to execute the computer program and implement the steps of the method for predicting water consumption in residential areas based on multi-source data fusion as described in any one of claims 1 to 9 when executing the computer program.