Water supply network single-node flow prediction method based on wavelet real-time decomposition and iTransform model
Through the combination of wavelet real-time decomposition and iTransformer model, data leakage and feature capture problems in the prediction of short-term fluctuations of water supply pipelines are solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202510370855.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the prior art, when predicting high-frequency data of short-term fluctuations in water supply networks, overall decomposition can easily lead to data leakage, and it is difficult to accurately capture data characteristics, affecting the prediction accuracy.
The data is decomposed into high-frequency detail components and low-frequency approximate components by using wavelet real-time decomposition, and combined with the iTransformer model, a prediction model of approximate components and detail components is constructed separately to improve prediction accuracy.
Through the combination of wavelet decomposition and iTransformer model, data leakage is avoided, the most important features in the data are focused, and the accuracy of single-node flow prediction in the water supply pipeline network is improved.
Smart Images

Figure CN120234615A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water demand prediction for water supply networks, and particularly relates to a method for predicting the single-node flow rate of a water supply network based on wavelet real-time decomposition and the iTransformer model. Background Art
[0002] With the continuous acceleration of the urbanization process, the operating load of water supply networks is increasing day by day. Water demand prediction plays a leading role in the operation management and dispatching of water supply networks. However, water usage behavior is affected by various factors such as seasonal changes, weather conditions, and residents' habits, which makes water demand data highly random and uncertain. At the same time, water demand data often exhibits complex non-linear and chaotic characteristics, posing a huge challenge to the prediction task. Therefore, establishing an accurate prediction model is of great significance for improving the accuracy of water demand prediction.
[0003] Water demand prediction models for water supply networks are mainly divided into three categories: physical models, statistical models, and machine learning methods. Physical models abstract influencing factors such as water pressure, flow rate, and flow velocity into corresponding mathematical models, and real-time monitor these parameters through sensors deployed in the water supply network, thereby realizing real-time prediction of water demand. The statistical model method uses statistical knowledge to estimate future water demand by analyzing, fitting, and trend predicting historical water usage data. Common statistical models include autoregressive integrated moving average models, etc. However, due to the non-stationarity and seasonal changes of water demand data, it may be difficult to obtain accurate predictions relying solely on traditional time series models. In addition to traditional physical and statistical models, although machine learning algorithms have improved the prediction accuracy to a certain extent, their processing ability for high-frequency data with short-term fluctuations and adaptability to real-time data still need to be improved. Especially when facing large-scale water supply networks, existing methods often have difficulty in accurately predicting the flow rate of a single node, and cannot meet the refined requirements of network operation management. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for predicting the single-node flow rate of a water supply network based on wavelet real-time decomposition and the iTransformer model, so as to solve the problems in the prior art that when predicting high-frequency data with short-term fluctuations in a water supply network, the overall decomposition is prone to data leakage and it is difficult to accurately capture data characteristics, thus affecting the prediction accuracy.
[0005] The present invention achieves the above purpose through the following technical solutions:
[0006] The present invention proposes a method for predicting the single-node flow rate of a water supply network based on wavelet real-time decomposition and the iTransformer model, and the method includes the following steps:
[0007] S1. Obtain the historical flow data of all nodes in the water supply network;
[0008] S2. Select a single node and calculate the correlation between this node and other nodes to obtain a similar node cluster;
[0009] S3. Divide the flow data within the similar node cluster into a training set and a test set;
[0010] S4. Use the sliding window and wavelet decomposition technology to decompose the training set and the test set to obtain the local approximation components and detail components of all nodes, and perform overall wavelet decomposition on the data of the single node in the training set to obtain the overall approximation component and overall detail component of the single node;
[0011] S5. Construct a single-feature iTransformer model, use the local and overall approximation components of the single node in the training set as training sample inputs to predict the approximation components; and construct a multi-feature iTransformer model, use the local detail components of all nodes and the overall detail component of the single node in the training set as training sample inputs to predict the detail components;
[0012] S6. Sum the predicted values of the approximation components and the predicted values of the detail components to obtain the flow prediction value of the single node.
[0013] Further, step S2 includes:
[0014] S201. Randomly select a single node and sequentially calculate the Pearson correlation coefficients between the historical flow data of this single node and all other nodes;
[0015] S202. Set a correlation coefficient threshold, and combine the nodes whose correlation coefficients with the single node are greater than the threshold with the single node as the center to obtain the historical flow data of the similar node cluster of the single node.
[0016] Further, step S4 includes:
[0017] S401. Define a sliding window with a length of l, and use the sliding window to obtain multiple groups of data on the training set and test set of all nodes within the similar node cluster respectively;
[0018] S402. Adopt the db2-based wavelet and determine its low-pass filtering coefficient h0 and high-pass filtering coefficient g0;
[0019] S403. Perform convolution operations to calculate the approximation components and detail components where i represents the current position index of the input signal in the convolution operation, k represents the index of the filtering coefficient, and * represents the convolution operation;
[0020] S404. Downsample the approximate component cA[i] and the detail component cD[i]. The formula is as follows:
[0021] cA[j] = cA[2j + 1], cD[j] = cD[2j + 1]
[0022] where j represents the index after downsampling;
[0023] S405. After performing one decomposition of the above S402 - S404, obtain the local approximate components and detail components of all nodes, as well as the overall approximate component and overall detail component of a single node.
[0024] Further, in step S5, when constructing the single - feature iTransformer model, using the local and overall approximate components of a single node in the training set as training sample inputs to predict the approximate components, it includes:
[0025] S501. Construct a single - feature neural network model including an embedding layer, an encoding layer, and a projection layer connected in sequence;
[0026] S502. Use the local approximate components of a single node in the training set as training sample inputs and the overall approximate component of a single node as the training sample label to train the single - feature neural network model;
[0027] S503. Input the local approximate components of a single node in the test set into the trained single - feature neural network model to obtain the predicted value of the approximate component of the single node.
[0028] Further, in step S5, when constructing the multi - feature iTransformer model, using the local detail components of all nodes and the overall detail component of a single node in the training set as training sample inputs to predict the detail components, it includes:
[0029] S511. Construct a multi - feature neural network model including an embedding layer, an encoding layer, and a projection layer connected in sequence;
[0030] S512. Use the local detail components of all nodes and the overall detail component of a single node in the training set as training sample inputs and the overall detail component of a single node as the training sample label to train the multi - feature neural network model;
[0031] S513. Input the local detail components of all nodes in the test set into the trained multi - feature neural network model to obtain the predicted value of the detail component of the single node.
[0032] Further, in the single-feature neural network model, constructing the encoding layer includes constructing a self-attention layer, a residual connection layer, a normalization layer, a feed-forward network layer, and a residual connection layer. Constructing the self-attention layer includes taking the output of the embedding layer as the input of the self-attention layer:
[0033]
[0034] output2 = softmax(scores1) * Values1
[0035] Where Queries1, Keys1, and Values1 are output1, representing the input data of the self-attention layer, and output2 represents the output of the self-attention layer.
[0036] Further, in the multi-feature neural network model, constructing the encoding layer includes constructing a multi-head self-attention layer, a residual connection layer, a normalization layer, a feed-forward network layer, and a residual connection layer. The multi-head self-attention layer is used to capture the associations between different feature variables. Constructing the multi-head self-attention layer includes:
[0037] Taking the output of the embedding layer as the input of the multi-head self-attention layer;
[0038] Calculating the attention scores for each head and calculating the output:
[0039]
[0040] V = softmax(scores 1 ) * Values 1
[0041] Finally, combining the inputs of all heads:
[0042] MultiHead = Concat(head1, head2..., head8) * W
[0043] Where Queries 1 , Keys 1 , and Values 1 represent the input data of each head, which is output 2 , V represents the weighted value matrix, that is, the output of a single head, W represents the weight matrix, and MultiHead represents the output of the multi-head self-attention layer.
[0044] The beneficial effects of the present invention are as follows:
[0045] (1) Compared with the traditional overall decomposition prediction method, the training strategy adopted by the present invention avoids the data leakage problem caused by the possible exposure of information in the test dataset in advance due to the use of specific wavelet decompositions. It is possible to subdivide the data into subsets with independent features while maintaining data integrity, thereby processing and using the data more securely.
[0046] (2) The present invention uses a single wavelet decomposition to decompose the sequence into high-frequency detail components and low-frequency approximation components, and then establishes a prediction model. The high-frequency part usually contains the short-term fluctuations of the data, while the low-frequency part represents the long-term trend and periodic components of the data. Constructing prediction models separately can enable the prediction model to focus on the most important features in the data, thereby improving the prediction accuracy.
[0047] (3) Aiming at the problem of difficult prediction of high-frequency data containing short-term fluctuations, the present invention starts from the spatial correlation, takes similar nodes in the same water supply network as the input for the node to be predicted, and uses the advanced iTransformer model to capture the correlation of water use behaviors existing between different nodes, thereby improving the prediction accuracy of the high-frequency detail components.
[0048] (4) The iTransformer model adopted by the present invention is improved based on the Transformer model. Different node inputs are considered separately, and the input of each node is independently encoded. Then, the multi-head self-attention mechanism is used to model the correlation between different node inputs. Finally, the feed-forward network is used to model the temporal correlation of different node inputs. Compared with the LSTM that relies on recursive calculation or the TCN with local connections, the model adopted by the present invention can capture the complex interdependent relationships between nodes and is more suitable for predicting the water demand of different nodes in the water supply network. Description of the Drawings
[0049] Figure 1 A flowchart of a single-node flow prediction method for a water supply network based on wavelet real-time decomposition and iTransformer model provided by an embodiment of the present application;
[0050] Figure 2 Another flowchart of a single-node flow prediction method for a water supply network based on wavelet real-time decomposition and iTransformer model provided by an embodiment of the present application;
[0051] Figure 3 A single-feature iTransformer architecture diagram of a single-node flow prediction method for a water supply network based on wavelet real-time decomposition and iTransformer model provided by an embodiment of the present application;
[0052] Figure 4Multi - feature iTransformer architecture diagram based on wavelet real - time decomposition and iTransformer model provided by the embodiments of the present application;
[0053] Figure 5 Water demand prediction result diagram provided for the case part of the specific implementation manner of the present application. Specific implementation manner
[0054] The following further describes the present application in detail with reference to the accompanying drawings. It is necessary to point out here that the following specific implementation manners are only used to further illustrate the present application and cannot be understood as limiting the protection scope of the present application. Those skilled in the art can make some non - essential improvements and adjustments to the present application according to the above application content.
[0055] Embodiment 1
[0056] As Figure 1-2 shown, this embodiment proposes a single - node flow prediction method for water supply pipe networks based on wavelet real - time decomposition and the iTransformer model (a deep learning model based on the self - attention mechanism). The method includes the following steps:
[0057] S1. Obtain the historical flow data of all n nodes at m moments in the water supply pipe network where a represents the a - th node in the water supply pipe network, j represents the j - th moment, n represents that there are n nodes in the water supply pipe network, and m represents the m - th moment of the node.
[0058] S2. Randomly select a node Xi=(x i1 ,x i2 ,x ij ,…,x im ) as a single node, and calculate the Pearson correlation coefficient between this single node and the other n - 1 nodes in turn. Set the correlation coefficient threshold k, and combine the high - correlation - coefficient nodes with the single node that are greater than the coefficient threshold k centered on the single node to obtain the historical flow data of the similar - node cluster of the single node where c represents the c - th node in the cluster, and z represents all z nodes in the cluster.
[0059] S3. Divide the flow data in the similar - node cluster into a training set and a test set;
[0060] Specifically, in the historical flow data of all nodes in the similar - node cluster of the obtained single node, divide the historical flow data from the 1 - q moment into the training set The historical flow data at the
[0061] Among them, the data of a single node in the training set is \(X_{i\_train}=(x i1 ,x i2 ,x ij1 ,…,x iq ), and the data of a single node in the test set is \(X_{i\_test}=(x i(q+1) ,x i(q+2) ,x ij2 ,…,x im )
[0062] S4. Use the sliding window and wavelet decomposition technology to decompose the training set and the test set, obtain the local approximation components and detail components of all nodes, and perform the overall wavelet decomposition on the data of a single node in the training set to obtain the overall approximation component and overall detail component of the single node;
[0063] In a specific embodiment, step S4 includes:
[0064] Create sliding windows of length \(l\) on the training sets and test sets of all nodes in the cluster respectively to obtain multiple groups of data of the training sets of all nodes in the cluster Multiple groups of data of the test set Among them, it includes multiple groups of data of the training set of a single node (\(L i1 ,L i2 ,L iw ,…,L i(q-l) ), multiple groups of data of the test set (\(L i(q+1) ,L i(q+2) ,L iu ,…,L i(m-l) )
[0065] Use the wavelet decomposition technology to decompose the multiple groups of data of the training set and the test set respectively to obtain the local approximation components and detail components of all nodes in the training set and detail components of all nodes in the test set i1 ,A i2 ,A iw ,…,A i(q-l)}, the local approximation component \(X_{i\_Atest}=\{A i(q+1) ,A i(q+2) ,A iu ,…,A i(m-l )\) of a single node in the test set. Subsequently, use the wavelet decomposition technology to decompose the data \(X_{i\_train}\) of a single node in the training set as a whole to obtain the overall approximation component \(Y_{i\_Atrain}=(a1,a2,a j1 ,…,aq ), the overall detailed component Yi_Dtrain of a single node = (d1, d2, d j1 ,…,d q )(j1 = 1, 2, …, q).
[0066] Using wavelet decomposition technology, decompose multiple groups of data in the training set and the test set respectively. The specific steps are as follows:
[0067] A. Determine the type of the base wavelet and its low-pass filtering coefficients and high-pass filtering coefficients. In the present invention, the db2 base wavelet is selected, and its low-pass filtering coefficient h0 and high-pass filtering coefficient g0 are as follows:
[0068]
[0069] B. Perform convolution operations to calculate the approximate component and detailed component of each group of data. The formula is as follows:
[0070] Approximate component
[0071] Detailed component
[0072] Where i represents the current position index of the input signal in the convolution operation, k represents the index of the filtering coefficient, and * represents the convolution operation.
[0073] C. Perform downsampling on the approximate component and detailed component. The formula is as follows:
[0074] cA[j] = cA[2j + 1], cD[j] = cD[2j + 1]
[0075] Where j represents the index after downsampling.
[0076] After one decomposition, obtain the local approximate components and local detailed components of all nodes in the training set and local detailed components of all nodes in the test set i1 , A i2 , A iw ,…, A i(q-l)}, the local approximate component Xi_Atrain of a single node in the test set = {A i(q+1) , A i(q+2) , A iu ,…, A i(m-l)}.
[0077] Use wavelet decomposition technology to decompose the data Xi_train of a single node in the training set as a whole. The specific steps are as follows:
[0078] A. Determine the type of the base wavelet and its low-pass filtering coefficients and high-pass filtering coefficients. In the present invention, the db2 base wavelet is selected, and its low-pass filtering coefficients h0 and high-pass filtering coefficients g0 are as follows:
[0079]
[0080] B. Perform a convolution operation to calculate the approximation component and detail component of each data. The formula is as follows:
[0081] Approximation component
[0082] Detail component
[0083] Where i represents the current position index of the input signal in the convolution operation, k represents the index of the filtering coefficient, and * represents the convolution operation.
[0084] C. Perform downsampling on the approximation component and detail component. The formula is as follows:
[0085] cA[j] = cA[2j + 1], cD[j] = cD[2j + 1]
[0086] Where j represents the index after downsampling.
[0087] Obtain the overall approximation component of a single node Yi_Atrain = (a1, a2, a j1 , …, a q ), the overall detail component of a single node Yi_Dtrain = (d1, d2, d j1 , …, d q ).
[0088] S5. Construct a single-feature iTransformer neural network model. Use the local approximation component of a single node in the obtained training set as the training sample input of the single-feature iTransformer neural network model. Use the last q - l elements in the obtained overall approximation component Yi_Atrain = (a1, a2, a j1 , …, a q ) of a single node as the training sample label to train the above single-feature iTransformer neural network model. Input the local approximation component Xi_Atest of a single node in the test set into the trained single-feature iTransformer model to obtain the predicted value A of the approximation component of a single node = (a i(q+1) , a i(q+2) , a iv , …, a im )(v = q + 1, q + 2, v, …, m);
[0089] and constructing a multi-feature iTransformer model, and using the local detail components of all nodes in the obtained training set as the training sample input of the multi-feature iTransformer neural network model, and using the obtained overall detail component Yi_Dtrain = (d1, d2, d j1 , …, d q ) as the training sample labels of the last q-l elements of the multi-feature model to train the above multi-feature iTransformer neural network model. Input the local detail components Xc_Dtest of all nodes in the test set into the trained multi-feature iTransformer model to obtain the detail component prediction value D of a single node = (d i(q+1) , d i(q+2) , d iv , …, di m )(v = q + 1, q + 2, v, …, m).
[0090] S6. Sum the obtained approximate component prediction value A and the detail component prediction value D obtained in step 4 to obtain the flow prediction value I of a single node = (p i(q+1) , p i(q+2) , p iv , …, p im )(v = q + 1, q + 2, v, …, m).
[0091] In a specific embodiment, in step 1, obtain the historical flow data of all n nodes at m moments in the water supply network Randomly select a single node Xi = (x i1 , x i2 , x ij , …, x im ) from the n nodes, and sequentially calculate the correlation coefficients between the single node and the other n-1 nodes. The calculation formula is as follows:
[0092]
[0093] where X j is the observed value of the historical flow data of the single node, Y j is the observed values of the other n-1 nodes, is the mean value of the node historical flow data. Calculate the correlation coefficient array r ia = (r i1 , r i2 , r ib , …, r i(n-1) )(b = 1, 2, …, n-1), set the correlation coefficient threshold k, and combine the highly correlated nodes with the single node that are greater than the coefficient threshold k with the single node as the center, that is, r ib≥k to obtain the historical traffic data of the similar node clusters of a single node where c represents the c-th node in the cluster, and z represents all z nodes in the cluster.
[0094] In a specific embodiment, in step S5, a single-feature iTransformer model is constructed, and the local and global approximation components of a single node in the training set are used as training sample inputs to predict the approximation components, including:
[0095] S501. Construct a single-feature neural network model including an embedding layer, an encoding layer, and a projection layer connected in sequence;
[0096] S502. Use the local approximation component of a single node in the training set as the training sample input, and the global approximation component of the single node as the training sample label to train the single-feature neural network model;
[0097] S503. Input the local approximation component of a single node in the test set into the trained single-feature neural network model to obtain the predicted value of the approximation component of the single node.
[0098] In a specific embodiment, in step S5, a multi-feature iTransformer model is constructed, and the local detail components of all nodes in the training set and the global detail component of a single node are used as training sample inputs to predict the detail components, including:
[0099] S511. Construct a multi-feature neural network model including an embedding layer, an encoding layer, and a projection layer connected in sequence;
[0100] S512. Use the local detail components of all nodes in the training set and the global detail component of a single node as the training sample input, and the global detail component of the single node as the training sample label to train the multi-feature neural network model;
[0101] S513. Input the local detail components of all nodes in the test set into the trained multi-feature neural network model to obtain the predicted value of the detail component of the single node.
[0102] In a specific embodiment, the single-feature neural network model includes an embedding layer, an encoding layer, and a projection layer connected in sequence, and it includes the following steps:
[0103] A. Construct the embedding layer:
[0104] The embedding layer is a linear layer, and the dimension of the linear mapping is 64.
[0105] ouput1 = A1 T *input1 + b1
[0106] Among them, A1 represents the weight matrix, b1 represents the bias vector, input1 represents the input vector of the linear layer, and output1 represents the output of the linear layer.
[0107] B. Construct the encoding layer, which includes a self-attention layer, a residual connection layer, a normalization layer, a feed-forward network layer, a residual connection layer, and a normalization layer connected in sequence.
[0108] 1). Construct the self-attention layer:
[0109] Construct a fully-connected self-attention mechanism layer, and use the output of the embedding layer as the input of the self-attention layer:
[0110]
[0111] output2 = softmax(scores1) * Values1
[0112] Among them, Queries1, Keys1, and Values1 are output1, representing the input data of the self-attention layer, and output2 represents the output of the self-attention layer.
[0113] 2). Construct the residual connection layer and the normalization layer:
[0114] L1 regularization is introduced during the residual connection, and the regularization factor is 0.001.
[0115] output3 = output2 + drop(output1)
[0116] output1 and output2 represent the inputs of the residual connection layer, and output3 represents the output of the residual connection layer.
[0117] Normalization layer:
[0118] ouput4 = Laynorm(output3)
[0119] output3 represents the input of the normalization layer, and output4 represents the output of the normalization layer.
[0120] 3). Construct the feed-forward network layer
[0121] The feed-forward network layer includes a first conv1d layer, an activation layer, and a second conv1d layer. The input channels and output channels of the first conv1d layer are 32 and 64, the size of the convolutional kernel is 1, the activation function of the activation layer is Relu, the input channels and output channels of the second conv1d layer are 64 and 32, the size of the convolutional kernel is 1, and L1 regularization is introduced for the two convolutional layers, and the regularization factor is 0.001.
[0122]
[0123] output6 = Relu(output5)
[0124]
[0125] where w represents the weight matrix of the convolutional kernel, with dimensions (c out , c in , k), output4 represents the input of the first conv1d layer, b represents the bias term, with dimensions (c out ), k is 0, output5 represents the output of the first conv1d layer, output6 represents the output of the activation layer, and output7 represents the output of the second conv1d layer.
[0126] 4). Construct the residual connection layer:
[0127] Introduce L1 regularization during the residual connection, with the regularization factor being 0.001.
[0128] output8 = output7 + drop(output4)
[0129] output4 and output7 represent the inputs of the residual connection layer, and output8 represents the output of the residual connection layer.
[0130] C. Construct the projection layer
[0131] The projection layer includes two linear layers, and the dimensions of the linear mapping are 64 and 1.
[0132] output9 = A T * output8 + b
[0133] where A represents the weight matrix, b represents the bias vector, output8 represents the input of the projection layer, and output9 represents the output of the projection layer.
[0134] In a specific embodiment, the multi - feature neural network model includes an embedding layer, an encoding layer, and a projection layer connected in sequence, and includes the following steps:
[0135] A. Construct the embedding layer:
[0136] The embedding layer includes a transpose layer and a linear layer, and the dimension of the linear mapping is 64.
[0137] Transpose layer:
[0138] output 1 = permute(input 1 )
[0139] where input 1 represents the input of the transpose operation, and output 1 represents the output of the transpose operation.
[0140] Linear layer:
[0141] ouput 2 = A 1T * output 1 + b 1
[0142] where A 1 represents the weight matrix, and b 1 represents the bias vector, and output 1 represents the input vector of the linear layer, and output 2 represents the output of the linear layer.
[0143] B. Construct two encoding layers. A single encoding layer includes a multi - head self - attention layer, a residual connection layer, a normalization layer, a feed - forward network layer, and a residual connection layer connected in sequence.
[0144] 1). Construct the multi - head self - attention layer:
[0145] Construct a fully - connected multi - head self - attention layer. Take the output of the embedding layer as the input of the multi - head self - attention layer, and the number of heads is 8. The iTransformer model cancels the decoding layer in the Transformer model to simplify the model structure. Different from the multi - head attention mechanism in the Transformer model that acts on the time dimension, the multi - head attention mechanism of the multi - head self - attention layer is used to capture the associations between different feature variables, that is, it acts on the feature dimension.
[0146] Calculate the attention scores for each head and calculate the output:
[0147]
[0148] V = softmax(scores 1 ) * Values 1
[0149] Finally, combine the inputs of all heads:
[0150] MultiHead = Concat(head1, head2..., head8) * W
[0151] where Queries 1 (queries), Keys 1 (keys), and Values 1 (values) represent the input data for each head and are output2 , V represents the weighted value matrix, i.e., the output of a single head, W represents the weight matrix, and MultiHead represents the output of the multi-head self-attention.
[0152] 2). Construct the residual connection layer and the normalization layer:
[0153] L1 regularization is introduced during the residual connection, and the regularization factor is 0.001.
[0154] output 3 = MultiHead + drop(input 3 )
[0155] input 3 represents the input of the residual connection layer, V represents the output of the self-attention layer, and output 3 represents the output of the residual connection layer.
[0156] Normalization layer:
[0157] output 4 = Laynorm(output 3 )
[0158] 3). Construct the feed-forward network layer:
[0159] The feed-forward network layer includes the first conv1d layer, the activation layer, and the second conv1d layer. The input channels and output channels of the first conv1d layer are 64 and 512 respectively, the size of the convolutional kernel is 1, the activation function of the activation layer is Relu, the input channels and output channels of the second conv1d layer are 512 and 64 respectively, the size of the convolutional kernel is 1, and L1 regularization is introduced for both convolutional layers, with the regularization factor being 0.001.
[0160]
[0161] ouput 6 = Re lu(output 5 )
[0162]
[0163] where w represents the weight matrix of the convolutional kernel, with the dimension (c out , c in , k), output 4 represents the input of the first conv1d layer, b represents the bias term, with the dimension (c out ), k is 0, and output 5 represents the output of the first conv1d layer, output 6Represents the output of the activation layer, output 7 Represents the output of the second conv1d layer.
[0164] 4). Construct the residual connection layer:
[0165] Introduce L1 regularization during the residual connection, and the regularization factor is 0.001.
[0166] output 8 = output 7 + drop(output 4 )
[0167] output4 and output7 represent the inputs of the residual connection layer, and output8 represents the output of the residual connection layer.
[0168] C. Construct the projection layer
[0169] The projection layer includes two linear layers, and the dimensions of the linear mapping are 64 and 1.
[0170] ouput 9 = A T * output 8 + b
[0171] Where A represents the weight matrix, b represents the bias vector, output 7 represents the input of the projection layer, and output 9 represents the output of the projection layer.
[0172] It should be noted that the training sample input and training sample label input of the multi-feature model are input into the constructed multi-feature iTransformer model. Select MSE as the loss function of the model, use the Adam optimizer to optimize the model parameters, the batch_size is 32, the epochs is 50, and the learning rate is 0.001 to train the multi-feature iTransformer model.
[0173] To more clearly illustrate the present invention and its advantages, the following will further explain the method provided by the present invention in combination with specific embodiments and relevant parts of the drawings.
[0174] The present invention is a short-term flow prediction method for water supply networks based on wavelet real-time decomposition and the iTransformer model. The deep water cloud brain "Residential Community Secondary Water Supply Demand Prediction" competition dataset is selected as the experimental research object. This dataset provides hourly flow data for 20 communities in the same water supply network, that is, the cumulative flow in the past hour is collected every hour at 20 flow monitoring points in the same area. A total of 2714 moments from 1:00 on January 1, 2022 to 1:00 on April 24, 2022 at each monitoring point are selected as the experimental dataset.
[0175] This implementation discloses a short-term flow prediction method for water supply networks based on wavelet real-time decomposition and the iTransformer model, which specifically includes the following steps:
[0176] Step 1, obtain the historical flow data of all 20 nodes at 2714 moments in the water supply network
[0177] Select a single node X1=(29.7, 21.9, 16.9, ……, 28.6) among the 20 nodes, and calculate the correlation coefficients between the single node and the remaining 19 nodes in turn to obtain a correlation coefficient array: (0.88, 0.86, 0.83, 0.86, 0.88, 0.79, 0.41, 0.87, 0.24, 0.41, 0.26, 0.39, 0.49, 0.40, 0.40, 0.41, 0.25, 0.25, 0.75)
[0178] Set the correlation coefficient threshold to 0.85, and combine the high-correlation monitoring points with a coefficient greater than 0.80 with the single node centered on the single node, that is, monitoring points 2, 3, 5, 6, 9, to obtain the historical flow data of the similar node cluster of the single node:
[0179] Step 2, among the historical flow data of all nodes in the similar node cluster of monitoring point 1 obtained in step 1, divide the historical flow data from 1 to 2171 moments into the training set:
[0180] Divide the historical flow data from 2172 to 2714 moments into the test set Among them, the data of the single node in the training set is Xi_train=(29.7, 21.9, 16.8, …, 49.1), and the data of the single node in the test set is Xi_test=(48.1, 46.8, 40.1, …, 28.6). Create sliding windows with a length of 24 on the training sets and test sets of all nodes in the cluster to obtain multiple groups of data for the training sets of all nodes in the cluster Multiple groups of data for the test set It contains multiple sets of data for the training set with single nodes \(\{(29.7,\cdots,38.9),(21.9,\cdots,25.4),\cdots,(47.2,\cdots,51.1)\}\) and multiple sets of data for the test set \(\{(48.1,\cdots,45.5),(46.8,\cdots,45.1),\cdots,(33.2,\cdots,35.4)\}\). Select the db2 basis wavelet and use wavelet decomposition technology to decompose the multiple sets of data in the training set and the test set respectively. Perform one decomposition to obtain the local approximation components of all nodes in the training set Local detail components Local approximation components of all nodes in the test set Local detail components It includes the local approximation components of single nodes in the training set \(X_{i\_A_{train}}=\{(25.5,\cdots,40.1),(19.1,\cdots,21.9),\cdots,(47.7,\cdots,53.5)\}\) and the local approximation components of single nodes in the test set \(X_{i\_A_{test}}=\{(47.6,\cdots,46.1),(42.9,\cdots,44.7),\cdots,(17.7,\cdots,24.4)\}\). Subsequently, use wavelet decomposition technology to decompose the data \(X_{i\_train}\) of single nodes in the training set as a whole to obtain the overall approximation component \(Y_{i\_A_{train}}=(25.5,23.9,17.7,\cdots,35.2)\) and the overall detail component \(Y_{i\_D_{train}}=(4.2, - 1.9,-0.8,\cdots,0.46)\).
[0183] Step 3, construct a single - feature iTransformer model to predict the approximation component. The specific steps are as follows:
[0184] S1. Use the last 2147 elements in the local approximation components \(X_{i\_A_{train}}=\{(25.5,\cdots,40.1),(19.1,\cdots,21.9),\cdots,(47.7,\cdots,53.5)\}\) and the overall approximation component \(Y_{i\_A_{train}}=(25.5,23.9,17.7,\cdots,35.2)\) obtained in step 2 as the training input and training label of the single - feature iTransformer model.
[0185] It includes the following steps:
[0186] A. Construct an embedding layer:
[0187] The embedding layer is a linear layer, and the dimension of the linear mapping is 64.
[0188] output1 = A T *input1 + b
[0189] Among them, A represents the weight matrix, b represents the bias vector, input1 represents the input vector of the linear layer, and output1 represents the output of the linear layer.
[0190] B. Construct the encoding layer, which includes a self-attention layer, a residual connection layer, a normalization layer, a feed-forward network layer, and a residual connection layer connected in sequence.
[0191] 1). Construct the self-attention layer:
[0192] Construct a fully-connected self-attention mechanism layer, and use the output of the embedding layer as the input of the self-attention layer:
[0193]
[0194] output2 = softmax(scores) * Values
[0195] Among them, Queries, Keys, and Values are output1, representing the input data of the self-attention layer, and output2 represents the output of the self-attention layer.
[0196] 2). Construct the residual connection layer and the normalization layer:
[0197] Introduce L1 regularization during the residual connection, and the regularization factor is 0.001.
[0198] output3 = output2 + drop(output1)
[0199] output1 and output2 represent the inputs of the residual connection layer, and output3 represents the output of the residual connection layer.
[0200] Normalization layer:
[0201] ouput4 = Laynorm(output3)
[0202] output3 represents the input of the normalization layer, and output4 represents the output of the normalization layer.
[0203] 3). Construct the feed-forward network layer
[0204] The feed-forward network layer includes a first conv1d layer, an activation layer, and a second conv1d layer. The input channels and output channels of the first conv1d layer are 32 and 64, the size of the convolution kernel is 1, the activation function of the activation layer is Relu, the input channels and output channels of the second conv1d layer are 64 and 32, the size of the convolution kernel is 1, and L1 regularization is introduced for the two convolutional layers, and the regularization factor is 0.001.
[0205]
[0206] output6 = Relu(output5)
[0207]
[0208] where w represents the weight matrix of the convolutional kernel, with dimensions (c out , c in , k), output4 represents the input of the first conv1d layer, b represents the bias term, with dimensions (c out ), k is 0, output5 represents the output of the first conv1d layer, output6 represents the output of the activation layer, and output7 represents the output of the second conv1d layer.
[0209] 5). Construct the residual connection layer:
[0210] L1 regularization is introduced during the residual connection, and the regularization factor is 0.001.
[0211] output8 = output7 + drop(output4)
[0212] output4 and output7 represent the inputs of the residual connection layer, and output8 represents the output of the residual connection layer.
[0213] C. Construct the projection layer
[0214] The projection layer consists of two linear layers, and the dimensions of the linear mapping are 64 and 1.
[0215] ouput9 = A T *output8 + b
[0216] where A represents the weight matrix, b represents the bias vector, output8 represents the input of the projection layer, and output9 represents the output of the projection layer.
[0217] S3. Input the training sample input and training sample label of the single - feature model into the single - feature iTransformer model described above. Select MSE as the loss function of the model, use the Adam optimizer to optimize the model parameters, with batch_size being 32, epochs being 50, and the learning rate being 0.001.
[0218] S4. Input the local approximation component Xi_Atest of a single node in the test set into the trained single - feature model to obtain the predicted value of the approximation component of the single node A = (45.2, 44.3, 35.9, …, 28.2).
[0219] Step 4. Build a multi-feature iTransformer model to predict the detail components. The specific steps are as follows:
[0220] S1. Use the detail components of all nodes within the cluster in the training set obtained in Step 2 and the last 2,174 elements in the overall detail component Yi_Dtrain = (4.2, -1.9, -0.8, …, 0.46) of a single node as the training sample input and training sample label of the multi-feature iTransformer model.
[0221] S1. Build a multi-feature neural network model including an embedding layer, an encoding layer, and a projection layer connected in sequence, including the following steps:
[0222] A. Build the embedding layer:
[0223] The embedding layer includes a transpose layer and a linear layer, and the dimension of the linear mapping is 64.
[0224] Transpose layer:
[0225] output 1 = permute(input 1 )
[0226] where input 1 represents the input of the transpose operation, and output 1 represents the output of the transpose operation.
[0227] Linear layer:
[0228] ouput 2 = A T * output 1 + b
[0229] where A represents the weight matrix, b represents the bias vector, output 1 represents the input vector of the linear layer, and output 2 represents the output of the linear layer.
[0230] D. Build two encoding layers. A single encoding layer includes a multi-head self-attention layer, a residual connection layer, a normalization layer, a feed-forward network layer, and a residual connection layer connected in sequence.
[0231] 1). Build the multi-head self-attention layer:
[0232] Build a fully connected multi-head self-attention layer, and use the output of the embedding layer as the input of the multi-head self-attention layer, with the number of heads being 8.
[0233] Calculate the attention scores of each head and calculate the output:
[0234]
[0235] V = softmax(scores) * Values
[0236] Finally, merge the inputs of all heads:
[0237] MultiHead = Concat(head1, head2..., head8) * W
[0238] where Queries, Keys, and Values represent the input data of each head, and are the output 2 , V represents the weighted value matrix, which is the output of a single head, W represents the weight matrix, and MultiHead represents the output of the multi - head self - attention.
[0239] 2). Construct the residual connection layer and the normalization layer:
[0240] Introduce L1 regularization during the residual connection, and the regularization factor is 0.001.
[0241] output 3 = MultiHead + drop(input 3 )
[0242] input 3 represents the input of the residual connection layer, V represents the output of the self - attention layer, and output 3 represents the output of the residual connection layer.
[0243] Normalization layer:
[0244] output 4 = Laynorm(output 3 )
[0245] 3). Construct the feed - forward network layer:
[0246] The feed - forward network layer includes the first conv1d layer, the activation layer, and the second conv1d layer. The number of input channels and output channels of the first conv1d layer are 64 and 512 respectively, the size of the convolutional kernel is 1, the activation function of the activation layer is Relu, the number of input channels and output channels of the second conv1d layer are 512 and 64 respectively, the size of the convolutional kernel is 1, and L1 regularization is introduced for both convolutional layers, with the regularization factor being 0.001.
[0247]
[0248] ouput 6 = Re lu(output 5 )
[0249]
[0250] where w represents the weight matrix of the convolutional kernel, with dimensions (c out , c in , k), output 4 represents the input of the first conv1d layer, b represents the bias term, with dimensions (c out ), k is 0, and output 5 represents the output of the first conv1d layer, output 6 represents the output of the activation layer, and output 7 represents the output of the second conv1d layer.
[0251] 5). Construct the residual connection layer:
[0252] Introduce L1 regularization during the residual connection, with a regularization factor of 0.001.
[0253] output 8 = output 7 + drop(output 4 )
[0254] output4 and output7 represent the inputs of the residual connection layer, and output8 represents the output of the residual connection layer.
[0255] E. Construct the projection layer
[0256] The projection layer includes two linear layers, with the dimensions of the linear mapping being 64 and 1.
[0257] ouput 9 = A T * output 8 + b
[0258] where A represents the weight matrix, b represents the bias vector, output 7 represents the input of the projection layer, and output 9 represents the output of the projection layer.
[0259] S3. Input the training sample input and training sample label of the multi-feature model into the multi-feature iTransformer model described above, select MSE as the loss function of the model, use the Adam optimizer to optimize the model parameters, with batch_size being 32, epochs being 50, and the learning rate being 0.001.
[0260] S4. Input the local detail component Xc_Dtest of all nodes in the test set into the trained multi-feature model to obtain the predicted value D of the detail component of a single node: D = (-5.3, 2.9, -0.1, …, 0.5).
[0261] Step 5. Sum the predicted value A of the approximate component obtained in Step 3 and the predicted value D of the detail component obtained in Step 4 to obtain the predicted value I of the flow rate of a single node: I = (39.9, 47.2, 35.8, …, 28.7). Use performance evaluation metrics such as MAE (Mean Absolute Error) and R 2 to evaluate the results.
[0262] Figure 5 is the water demand prediction result graph provided for the case part. In the graph, the abscissa is the time point of single-step prediction, the scale division value is 1 h, the ordinate is the interval flow rate of a single node, the scale division value is 1 L, the mean absolute error MAE between the predicted flow rate and the actual flow rate is 2.92, and the correlation coefficient R 2 is 0.90. The predicted flow rate curve can fit well with the actual flow rate curve.
[0263] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0264] In addition, in each embodiment of this application, each functional module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0265] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application.
Claims
1. A single-node flow prediction method for water supply network based on wavelet real-time decomposition and iTransformer model, characterized in that: The method comprises the following steps: S1. Obtain historical flow data of all nodes in the water supply network; S2, select a single node and calculate the correlation between the node and other nodes to obtain a similar node cluster; S3, dividing the traffic data in the similar node cluster into a training set and a test set; S4, using sliding window and wavelet decomposition technology to decompose the training set and test set to obtain local approximate components and detail components of all nodes, and perform overall wavelet decomposition on the data of a single node in the training set to obtain the overall approximate component and overall detail component of the single node; S5, constructing a single-feature iTransformer model, taking the local and global approximate components of a single node in the training set as training sample inputs, and predicting the approximate components; and constructing a multi-feature iTransformer model, taking the local detail components of all nodes in the training set and the global detail components of a single node as training sample inputs, and predicting the detail components; S6. Sum the approximate component prediction value and the detailed component prediction value to obtain the traffic prediction value of a single node.
2. According to claim 1, a single-node flow prediction method for a water supply network based on wavelet real-time decomposition and iTransformer model is characterized in that: Step S2 includes: S201, randomly selecting a single node, and calculating the Pearson correlation coefficient of the historical traffic data between the single node and all other nodes in turn; S202 , setting a correlation coefficient threshold, taking the single node as the center and combining nodes whose correlation coefficients with the single node are greater than the threshold, to obtain historical traffic data of similar node clusters of the single node.
3. According to claim 1, a single-node flow prediction method for a water supply network based on wavelet real-time decomposition and iTransformer model is characterized in that: Step S4 includes: S401, defining a sliding window of length l, and using the sliding window to obtain multiple sets of data on the training set and the test set of all nodes in the similar node cluster; S402, using db2-based wavelet and determining its low-pass filter coefficient h0 and high-pass filter coefficient g0; S403: Perform convolution operation to calculate the approximate components of each set of data and detail weight Where i represents the current position index of the input signal in the convolution operation, k represents the index of the filter coefficient, and * represents the convolution operation; S404, downsampling the approximate component cA[i] and the detail component cD[i], the formula includes: cA[j]=cA[2j+1], cD[j]=cD[2j+1] Where j represents the index after downsampling; S405. After performing the above S402-S404 decomposition once, local approximate components and detail components of all nodes, as well as overall approximate components and overall detail components of a single node are obtained.
4. According to claim 1, a single-node flow prediction method for a water supply network based on wavelet real-time decomposition and iTransformer model is characterized in that: In step S5, the single-feature iTransformer model is constructed, the local and overall approximate components of a single node in the training set are used as training sample inputs, and the approximate components are predicted, including: S501, constructing a single feature neural network model including an embedding layer, a coding layer and a projection layer connected in sequence; S502, using the local approximate component of a single node in the training set as a training sample input, and the overall approximate component of the single node as a training sample label, to train a single-feature neural network model; S503: Input the local approximate component of the single node in the test set into the trained single-feature neural network model to obtain the predicted value of the approximate component of the single node.
5. According to claim 1, a method for predicting single-node flow in a water supply network based on real-time wavelet decomposition and iTransformer model, characterized in that: In step S5, the multi-feature iTransformer model is constructed, the local detail components of all nodes in the training set and the overall detail components of a single node are used as training sample inputs, and the detail components are predicted, including: S511, constructing a multi-feature neural network model including an embedding layer, a coding layer and a projection layer connected in sequence; S512, using the local detail components of all nodes in the training set and the overall detail components of a single node as training sample inputs, and the overall detail components of a single node as training sample labels, to train a multi-feature neural network model; S513: Input the local detail components of all nodes in the test set into the trained multi-feature neural network model to obtain the predicted value of the detail component of the single node.
6. The method for predicting single-node flow in a water supply network based on real-time wavelet decomposition and iTransformer model according to claim 4 is characterized in that: In the single-feature neural network model, constructing the encoding layer includes constructing a self-attention layer, a residual connection layer, a normalization layer, a feedforward network layer, and a residual connection layer, and constructing the self-attention layer includes using the output of the embedding layer as the input of the self-attention layer: output2=softmax(scores1)*Values1 Among them, Queries1, Keys1 and Values1 are output1, which represents the input data of the self-attention layer, and output2 represents the output of the self-attention layer.
7. The method for predicting single-node flow in a water supply network based on real-time wavelet decomposition and iTransformer model according to claim 5, characterized in that: In the multi-feature neural network model, constructing the encoding layer includes constructing a multi-head self-attention layer, a residual connection layer, a normalization layer, a feedforward network layer and a residual connection layer. The multi-head self-attention layer is used to capture the association between different feature variables. Constructing the multi-head self-attention layer includes: Use the output of the embedding layer as the input of the multi-head self-attention layer; Compute the attention scores for each head and calculate the output: V=softmax(scores 1 )*Values 1 Finally merge the inputs of all headers: MultiHead=Concat(head1,head2...,head8)*W Queries 1 、Keys 1 , and Values 1 Represents the input data of each header, which is output 2 , V represents the weighted value matrix, that is, the output of a single head, W represents the weight matrix, and MultiHead represents the output of the multi-head self-attention layer.
Citation Information
Patent Citations
Time sequence data prediction method and device based on wavelet time-frequency domain
CN119167283A
Coordinated optimization multi-step prediction method based on target correlation space-time coding-task decoding
CN119599189A
Cited By
Pipeline leakage detection method and device
CN120667655A