Stack model-based water level prediction system and method, medium and computer program product
By combining stacked models with multiple deep learning models and methods of optimizing hyperparameters, the problem of water level prediction model capturing complex features and adapting to different conditions is solved, achieving higher accuracy and stable water level prediction.
Patent Information
- Application Number
- CN202510696435.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing water level prediction model is difficult to fully capture the multi-faceted characteristics of the data, lacks generalization, and cannot adapt to the complex and changing water level prediction needs, resulting in large prediction errors.
A water level prediction system based on stacked models is adopted, including a base model building module, a stack model building module, a base model training module and a stack model training module. Through the combination of multiple deep learning models, grid search and early stop strategies are used to optimize hyperparameters to build a multi-layer neural network to improve prediction accuracy.
It improves the accuracy and stability of water level prediction, enhances the generalization ability of the model, ensures that new data can be better processed in practical applications, reduces waste of computing resources, and improves training efficiency.
Smart Images

Figure CN120355036A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and deep learning, and particularly relates to a water level prediction system, method, medium and computer program product based on a stacked model. Background Art
[0002] With the continuous intensification of global climate change, the importance of hydrological problems such as water resource management and flood disaster warning has become increasingly prominent. As one of the core tasks of water conservancy, waterway management and water resource planning, the accurate prediction of water level is directly related to the effectiveness of flood control and disaster reduction and the rational allocation of water resources. In this context, it is particularly important to improve the public service ability and scientific decision-making level of waterway water level prediction and forecasting.
[0003] Traditional water level prediction models based on statistical methods usually have difficulty in capturing the complex non-linear characteristics of water level changes, resulting in relatively large prediction errors. To make up for the deficiencies of traditional methods and make full use of the information in big data, introducing data-driven deep learning models has become one of the key ways to improve the accuracy and precision of water level prediction. In recent years, scholars at home and abroad have conducted a large number of studies on the water level prediction problem based on machine learning and deep learning models, and the machine learning methods and deep learning methods for water level prediction have developed rapidly.
[0004] Although these deep learning models have made certain progress in water level prediction, a single prediction model often can only capture a single feature of the time series and is difficult to comprehensively capture the multi-faceted features of the data. In addition, for specific water level prediction problems, it is difficult for people to pre-judge which method is the most applicable. Therefore, there is a lack of a water level prediction model framework with generalization ability and adaptability to complex changes to meet the water level prediction needs under different conditions. Summary of the Invention
[0005] In order to improve the prediction ability of water level changes and provide more reliable scientific support for water conservancy management and water resource planning, the present invention proposes a water level prediction system, method, medium and computer program product based on a stacked model.
[0006] The water level prediction system based on a stacked model for achieving one of the purposes of the present invention includes: Base model construction module: used to construct a plurality of base models for predicting water level changes; the input data of each base model is historical water level data, and the output data is the initial prediction value of the water level change in the future time period; Stacked model construction module: used to construct a stacked model for predicting water level changes; the input data of the stacked model is the initial prediction value of the water level change in the future time period output by each base model, and the output value is the final prediction value of the water level change in the future time period; Base model training module: used to train the base model using historical water level data to obtain each trained base model; Stacked model training module: used to train the stacked model using the initial predicted values for the future time period output by each trained base model as training data to obtain a water level prediction model for predicting the water level in the future time period; Water level prediction module: used to predict the water level change in the future time period using the water level prediction model.
[0007] Furthermore, when training each base model or stacked model, the grid search method is used to determine the optimal values of the hyperparameters of each base model or stacked model, and each base model is evaluated by adopting cross-validation with a set multiple on the validation set; when the error on the validation set fails to improve in consecutive multiple training cycles, the training is stopped. The technical effects include: the grid search method can find the optimal hyperparameter settings that perform best on the given dataset, thus fully exerting the performance potential of the model; by finding the optimal hyperparameters, the performance of the model on the training set and the validation set is more balanced, avoiding overfitting or underfitting problems caused by improper hyperparameter settings, and then improving the generalization ability of the model to unseen data, enabling it to make more accurate predictions or classifications in practical applications; multiple-fold cross-validation can make more full use of the validation set data compared to only using the validation set once for evaluation, reducing the evaluation result deviation caused by different data partitioning methods, making the evaluation results more stable and reliable; by conducting multiple cross-validation evaluations on the validation set, the performance of different base models can be more accurately compared, so as to select the most suitable base model for the current task, providing a better foundation for the subsequent construction of the stacked model. When the error on the validation set fails to improve in consecutive multiple training cycles, it means that the model may have reached the optimal state or started to overfit. Stopping the training at this time can avoid wasting a large amount of computing resources and time in the ineffective training process and improve the training efficiency. The early stopping strategy can timely terminate the training of the model, prevent the model from learning too deeply on the training set, thus preventing the model from overfitting to the training set data, maintaining the generalization ability of the model on the validation set and the test set, and enabling the model to better process new data in practical applications; it helps to optimize the robustness and prediction accuracy of the model, ensuring that the potential of the model is fully utilized to improve the performance.
[0008] Further, the stacking model includes two hidden layers and an output layer. Each hidden layer contains multiple neurons. The first hidden layer is used to map the initial predicted values of water level changes output by each base model into a multi-dimensional space, so as to extract preliminary features from the initial predicted values of water level changes output by each base model and obtain the first output feature values. The second hidden layer is used to map the multi-dimensional space corresponding to the first output feature values into another multi-dimensional space, perform a non-linear transformation on the first output feature values, and obtain the second output feature values. The output layer is used to map the second output feature values extracted by the second hidden layer into the final predicted values of water level changes.
[0009] The technical effects of the above technical solutions include: The two hidden layers of the stacking model can extract more complex features from the initial predicted values of the base model through non-linear mapping, thereby improving the prediction accuracy. The first hidden layer extracts preliminary features, and the second hidden layer performs non-linear transformation. This hierarchical structure can better capture the complex patterns of water level changes. By mapping the features of the second hidden layer to the final predicted values through the output layer, the accuracy and stability of the prediction results are ensured.
[0010] Further, the calculation of each layer of the stacking model includes:
[0011] In the formula: h 1, h 2, and h 3 are the first output feature values, the second output feature values, and the final predicted values of water level changes respectively; σ is the RELU activation function; W 1, W 2, and W 3 are the weight matrices of the first hidden layer, the second hidden layer, and the output layer respectively; b 1, b 2, and b 3 are the bias vectors of the first hidden layer, the second hidden layer, and the output layer respectively.
[0012] The technical effects of the above technical solutions include: By clearly describing the calculation process of each layer with mathematical formulas, the interpretability and reproducibility of the model are ensured; Using the RELU activation function avoids the problem of gradient disappearance and accelerates model training; By optimizing the weight matrix and bias vector, the prediction performance of the model is further improved.
[0013] Furthermore, the stacking model uses the mean squared error (MSE) and the mean absolute error (MAE) as the loss function.
[0014] Furthermore, the base model includes a first base model based on RNN, including two RNN layers and a fully connected layer, the hidden state output by the first RNN layer is used as the input value of the second RNN layer; the number of neurons in the first RNN layer is greater than the number of neurons in the second RNN layer; the fully connected layer is used to convert the output value of the second RNN layer into a final water level change prediction value to obtain an initial prediction value of the water level change.
[0015] Furthermore, in each RNN layer t The calculation formula for the output information in the RNN of the step includes: h t =σ ( W h h t-1 + W x x t + b h ) Where: h t is the time step t The hidden state at the time step contains all the input information up to the current time step. The hidden state of the first RNN layer outputs h t As the input value of the second RNN layer; W h , W x is a weight matrix, which is used to connect the hidden state of the previous moment and the input features of the current moment; b h is the bias, used to adjust the input of the activation function; x t is the current time step t The input feature of the current time step t Water level data; h t-1 The previous time step in memory t -1 stored status information; σ is the activation function.
[0016] The technical effects of the above technical solution include: through two RNN layers and one fully connected layer, it is possible to better capture the time dependence of water level data; the number of neurons in the first RNN layer is more than that in the second RNN layer, and this design can gradually compress information and extract more important features; the output of the RNN layer is converted into an initial prediction value of the water level change to ensure that the output of the base model can be effectively utilized by the stacked model.
[0017] Furthermore, the base model includes a second base model based on LSTM, which includes two LSTM layers, and the number of neurons in one LSTM layer is half of that in the other LSTM layer.
[0018] Even further, the calculation formula for the output information in the LSTM layer includes:
[0019] In the formula: f t is the forget gate, which determines which information is discarded from the cell state C t-1 ; is the hidden state at the previous time step, and the hidden state contains all the input information up to the current time step. For the first time step, it is usually initialized as a zero vector or a random vector; represents the measured value of the water level data at time step t ; i t is the input gate, which determines which new information will be stored in the cell state C t ; O t is the output gate, which determines which information will be output from the cell state C t ; is the candidate memory, which generates new candidate values that may be added to the cell state; C t is the update cell state, which is used to update the current cell state; σ is the sigmoid activation function; W f 、W i and W o are weights; bf 、b i and b o are biases, corresponding to the forget gate, input gate, and output gate respectively; W c is the weight of the candidate memory unit; b c is the bias of the candidate memory unit.
[0020] Furthermore, the base model includes a third base model based on GRU, including two GRU layers and a fully connected layer. The calculation formula for the output information in each GRU layer includes:
[0021] In the formula: z t is the update gate, used to control the degree to which the hidden state information at the previous moment is passed to the current moment; r t is the reset gate, used to control how much information in the hidden state at the previous moment is forgotten; represents the input feature at time step t, that is, the water level data at the current time step t; is the candidate memory; and represent the hidden states at the current time step t and the previous time step t-1 respectively; W z 、 W r 、 W h represent the weight matrices of the update gate, reset gate, and calculation of the candidate hidden state, b z 、 b r 、 b h are the corresponding biases respectively; tanh is the hyperbolic tangent activation function, used to generate the candidate memory.
[0022] The technical effects of the above technical solution include: through two LSTM layers, the long-term dependence relationship in the water level data can be better captured. The number of neurons in one LSTM layer is half of that in the other, and this design can gradually extract more abstract features while reducing the computational complexity.
[0023] Further, the multiple base models include: a first base model based on RNN, a second base model based on LSTM, and a third base model based on GRU; each base model contains two layers, and each layer contains a first set number of neurons.
[0024] The technical effects of the above technical solution include: through two GRU layers and a fully connected layer, the time dependence of the water level data can be efficiently captured; the introduction of the update gate and the reset gate can dynamically control the transmission and forgetting of information, thereby improving the prediction ability of the model; by clearly describing the calculation process of the GRU layer with mathematical formulas, the interpretability and reproducibility of the model are ensured; using the tanh activation function to generate candidate memories ensures the non-linear transformation ability of information.
[0025] Furthermore, each base model uses the mean square error (MSE) and the mean absolute error (MAE) as loss functions.
[0026] Further, the system further includes a data preprocessing module for normalizing the historical water level data to improve the stability and efficiency of model training.
[0027] The water level prediction method based on the stacked model for achieving the second object of the present invention includes: Constructing multiple base models for predicting water level changes; the input data of each base model is the historical water level data, and the output data is the initial predicted value of the water level change in the future time period; Constructing a stacked model for predicting water level changes; the input data of the stacked model is the initial predicted value of the water level change in the future time period output by each base model, and the output value is the final predicted value of the water level change in the future time period; Training the base models with the historical water level data to obtain each trained base model; Using the initial predicted values of the future time period output by each trained base model as training data to train the stacked model to obtain a water level prediction model for predicting the water level in the future time period.
[0028] When training the model, the grid search method is used to determine the optimal values of the hyperparameters of each base model or the stacked model, and each base model is evaluated by adopting cross-validation with a set multiple on the validation set; when the error on the validation set fails to improve in consecutive multiple training cycles, the training is stopped.
[0029] A non-transitory computer-readable storage medium for achieving the third object of the present invention, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the water level prediction method based on the stacking model are implemented.
[0030] A computer program product for achieving the fourth object of the present invention, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the water level prediction method based on the stacking model are implemented.
[0031] The beneficial effects of the present invention include: The present invention provides a water level prediction model framework that combines the advantages of multiple stacking models based on deep learning and has generalization. By integrating the advantages of multiple deep learning models, the prediction ability for water level changes is improved, providing more reliable scientific support for water conservancy management and water resource planning, and enhancing the public service ability and scientific decision-making level of waterway water level prediction and forecasting. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of the stacking model architecture; Figures 2 to 5 They are respectively the comparison charts of water levels predicted by RNN, GRU, LSTM, and RGL-Stacking stacking models at 4 different stations during the dry season (December - April of the following year); Figure 6 and Figure 7 They are respectively the mean absolute error (MAE) and root mean square error (RMSE) of water levels predicted by RNN, GRU, LSTM, and RGL-Stacking stacking models at different stations. Detailed Embodiments
[0033] The following detailed embodiments are used to explain the technical solutions of the claims of the present invention so that those skilled in the art can understand the claim book. The protection scope of the present invention is not limited to the following specific implementation structures. Those made by those skilled in the art that include the technical solutions of the claim book of the present invention and are different from the following detailed embodiments are also within the protection scope of the present invention.
[0034] The embodiment of the present invention provides a water level prediction method based on a stacking model, including: (1) Data preprocessing Use the water level data of the Yangtze River Basin in the past ten years. Clean, supplement, smooth, and denoise the data, and then perform normalization processing so that its values are distributed between 0 and 1 to improve the stability and efficiency of model training, and divide it into a training set and a test set.
[0035] (2) Base model construction Three base models based on RNN, GRU, and LSTM are constructed respectively. Each base model contains two layers, and each layer contains 128 neurons. The Adam optimizer is used, with the learning rate set to 0.001, the batch size to 32, and the number of training epochs to 100. Among them: Base model 1: The first base model based on RNN; RNN is a basic recurrent neural network model. It maintains the memory of previous input information through a recurrent structure (hereinafter referred to as the memory body), which enables RNN to consider the time sequence and dependence of data when processing sequence data and is suitable for processing short sequence data. A model with multiple hidden layers is built to capture the temporal characteristics of water level data.
[0036] The first RNN layer of the first base model based on RNN has a larger number of neurons. More neurons mean that the model can learn and represent more information to enhance the model's representation ability; the second RNN layer has a smaller number of neurons, which helps to further extract more abstract features from the features extracted by the first layer. Finally, a fully connected layer is added to convert the output of the second RNN layer into the final predicted value of water level change, and the final predicted value of water level change is output. The calculation formula for the output information of the RNN at each step in each RNN layer is as follows: t Step of the RNN in each step of the RNN is as follows: h t =σ ( W h h t-1 + W x x t + b h ) Where h t is the hidden state at the time step, which contains all the input information up to the current time step. The hidden state t output by the first RNN layer h t is used as the input value of the second RNN layer; W h , W x are weight matrices, which are used to connect the hidden state of the previous moment and the input features of the current moment respectively; b h is the bias, which is used to adjust the input of the activation function; x t is the input feature at the current time step t , that is, the water level data at the current time step t ;h t-1 The state information stored at the previous time step in the memory, where σ is the activation function (using t , and the same applies hereinafter). tanh
[0037] Assume that the length of the input sequence is T. After calculation through two RNN layers, the hidden state h of the second RNN layer at time step T is obtained T ; Take the hidden state h of the second RNN layer at time step T T as the input of the fully connected layer. Assume that the weight matrix of the fully connected layer is W1 fc , and the bias is b 1 fc , then the calculation of the fully connected layer is as follows:
[0038] Y predRNN is the predicted water level change value of the first base model based on RNN.
[0039] Base model 2: The second base model based on LSTM; LSTM is a recurrent neural network model that performs excellently in processing long sequence data. It is a model proposed to solve the vanishing gradient problem in RNN and can better capture long-distance dependencies in the sequence. The second base model based on LSTM contains two LSTM layers. The input of the first LSTM layer is the input feature at each time step x t , such as the water level data at the current time step, and the output values are the hidden state and cell state at each time step; the input of the second LSTM layer is the hidden state output by the first LSTM layer, and the output values are the hidden state and cell state; where the number of neurons in the second layer is half of that in the first layer. Through the LSTM layer, the model can effectively learn the temporal features and trends of the water level data. The calculation formula for the output information in LSTM is as follows:
[0040] where f t is the forget gate, which determines which information is discarded from the cell state C t-1 ; is the hidden state at the previous time step. The hidden state contains all the input information up to the current time step. For the first time step, it is usually initialized as a zero vector or a random vector; represents the measured value of the water level data at time step t ; i t is the input gate, which determines which new information will be stored in the cell state C t ; O t is the output gate, which determines which information will be output from the cell state C t ; is the candidate memory, which generates new candidate values that may be added to the cell state; C t is the update cell state, which is used to update the current cell state; σ is the sigmoid activation function, W f 、W i and W o are weights, b f 、b i and b o are biases, corresponding to the forget gate, input gate, and output gate respectively; W c is the weight of the candidate memory cell; b c is the bias of the candidate memory cell.
[0041] After passing through two LSTM layers, the hidden state ( T is the last time step of the input sequence) output by the second LSTM layer is used as the input to the fully connected layer. Assuming the weight matrix of the fully connected layer is W2 fc , and the bias is b2 fc , then the calculation of the fully connected layer is as follows:
[0042]
[0043] Y predLSTM is the final predicted water level change value based on the second base model of LSTM. Through the two-layer LSTM layer, the model can effectively learn the temporal characteristics and trends of the water level data.
[0044] Base model 3: The third base model based on GRU; GRU is a recurrent neural network model between RNN and LSTM, which can reduce the computational cost while maintaining good performance. Compared with LSTM, GRU cancels the output gate, and sets the reset gate and update gate to jointly control how to calculate the new hidden state from the previous hidden state, which helps the model capture the important change characteristics of the water level data. The calculation formula for the output information in GRU is as follows:
[0045] where z t is the update gate, which is used to control the degree to which the hidden state information at the previous moment is passed to the current moment; is passed to the current moment; r t is the reset gate, which is used to control how much information of the hidden state at the previous moment is forgotten; is forgotten; represents the input feature at time step t, that is, the water level data at the current time step t; is the candidate memory; and represent the hidden states at the current time step t and the previous time step t-1 respectively; W z 、 W r 、 W h represent the weight matrices for the update gate, the reset gate, and calculating the candidate hidden state, b z 、 b r 、 b h are the corresponding biases respectively; tanh is the hyperbolic tangent activation function, which is used to generate the candidate memory.
[0046] Assume that the length of the input water level data sequence is T. After calculation by the GRU layer, the hidden state at time step T is obtained . Assume that the weight matrix of the fully connected layer is W3 fc , and the bias is b3 fc , then the calculation of the fully connected layer is as follows:
[0047] Y predGRU is the final predicted water level change value of the third base model based on GRU.
[0048] (3) Model output fusion Take the prediction results Y predRNN , y predLSTM , y predGRU of the above three base models as meta-features and input them into a multi-layer feedforward neural network with two hidden layers (hereinafter referred to as the stacked model), and each hidden layer contains 64 neurons.
[0049] (4) Stacked model construction and training Use the prediction results output by each trained base model as training data to train the stacking model. The stacking model is a multi-layer perceptron (MLP) with three layers, which uses the prediction results from each base model to perform further learning, thereby enhancing the prediction stability and robustness across different sites. Combine the prediction results of all base models into a new feature matrix X meta as the input to the stacking model:
[0050] The calculation of each layer of the stacking model is as follows:
[0051] The first layer: It is a hidden layer used to input the meta-features X meta into the first-layer perceptron to calculate the first output feature value h 1, where σ is the RELU activation function; W 1 is the weight matrix of the first-layer perceptron, with a dimension of 3×64 because the meta-features have 3 elements; b 1 is the bias vector of the first-layer perceptron, with a dimension of 64×1.
[0052] The second layer: It is a hidden layer used to input the first output feature value h 1 into the second-layer perceptron to calculate the second output feature value h 2; where W 2 is the weight matrix of the second-layer perceptron, with a dimension of 64×64; b 2 is the bias vector of the second-layer perceptron, with a dimension of 64×1.
[0053] The third layer: It is the output layer used to input the second output feature value h 2 into the third-layer perceptron to calculate the final feature value h final , which is also the water level change value finally predicted by the stacking model; where W 3 is the weight matrix of the third-layer perceptron, with a dimension of 64×1; b 3 is the bias vector of the third-layer perceptron, with a dimension of 1×1.
[0054] For each base model (RNN, GRU, LSTM), the training process includes the following steps:
[0055] Finally, the stacking model M meta uses X meta as the input, Yfinal (i.e., the aforementioned h final ) as the final output of the stacked model:
[0056] In both the base model and the stacked model, the rectified linear unit (RELU) is used as the activation function. RELU is well-known for its computational efficiency and ability to mitigate the vanishing gradient problem, thereby improving the training efficiency and overall performance of the model.
[0057] During the entire training process, an early stopping mechanism is implemented to prevent overfitting. If the error on the validation set fails to show improvement over five consecutive training epochs, the training stops. This strategy is crucial for maintaining the performance of the model on the validation set and ensuring that it does not degrade due to overtraining.
[0058] To determine the most effective hyperparameter configuration for each base model and the stacked model, a grid search method for exploring the best combination of hyperparameters for different site models is used to determine the most effective hyperparameter configuration for each model. This involves defining a series of potential hyperparameters, including the learning rate, the number of neurons, and the batch size. Subsequently, we evaluate the efficacy of each model configuration by adopting up to 5-fold cross-validation on the validation set. Under different hyperparameters, this meticulous examination greatly contributes to optimizing the robustness of the model and its prediction accuracy, ensuring that we fully utilize the potential of the model to improve performance.
[0059] In the embodiments of the present invention, each base model and the stacked model adopt the mean squared error (MSE) and the mean absolute error (MAE) as loss functions to control errors and optimize model parameters: Mean absolute error (MAE):
[0060] where T is the number of time steps in the test set, is the predicted value output by the model, x t is the actual value.
[0061] Root mean squared error (RMSE):
[0062] (5) Water level prediction The trained stacked model is used as a water level prediction model for predicting the water level change value. The water level prediction model is used to predict the water level data for a future time period, and the water level change for the future time period is obtained. The prediction result is compared with the actual water level of the test set to evaluate the prediction performance of the water level prediction model.
[0063] Let the water level time series data be { x t}, where t = 1, 2, …, T, and T is the length of the time series. The prediction result of the base model can be expressed as:
[0064] where n is the time lag step considered by the model, f RNN , f GRU and f LSTM respectively represent the prediction functions of the first model based on RNN, the second model based on LSTM, and the third model based on GRU.
[0065] The meta - features of the stacking model are the set of the prediction results of the base models:
[0066] The stacking model can be expressed as:
[0067] where g ( X t ) represents the prediction function of the stacking model, which takes the prediction results of the base models as input and outputs the final water level prediction value .
[0068] It can be seen from Figures 2 to 5 that the accuracy of the water level change values predicted by the water level prediction model based on the stacking model (RGL - Stacking) of the present invention during the dry season at 4 different stations is the highest compared to using RNN, GRU, or LSTM alone. At the same time, it can also be seen from Figure 6 and Figure 7 that the mean absolute error (MAE) and root mean square error (RMSE) of the water level prediction model based on the stacking model (RGL(RNN - GRU - LSTM) - Stacking) of the present invention are both the lowest compared to the models using RNN, GRU, and LSTM alone.
[0069] It should be understood that the magnitudes of the sequence numbers of the steps in the above - mentioned embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0070] The embodiments of the present invention also provide a water level prediction system based on a stacking model, including: Base model construction module: used to construct multiple base models for predicting water level changes; the input data of each base model is historical water level data, and the output data is the initial predicted value of the water level change in the future time period; Stacked model construction module: used to construct a stacked model for predicting water level changes; the input data of the stacked model is the initial predicted value of the water level change in the future time period output by each base model, and the output value is the final predicted value of the water level change in the future time period; Base model training module: used to train the base models using historical water level data to obtain each trained base model; Stacked model training module: used to train the stacked model using the initial predicted values of the future time period output by each trained base model as training data to obtain a water level prediction model for predicting the water level in the future time period; Water level prediction module: used to predict the water level change value in the future time period using the water level prediction model.
[0071] When training the model, the grid search method is used to determine the optimal values of the hyperparameters of each base model or stacked model, and each base model is evaluated by adopting 5-fold cross-validation on the validation set; when the error on the validation set fails to improve in consecutive multiple training cycles, the training is stopped.
[0072] In some embodiments, the stacked model includes two hidden layers and an output layer, each hidden layer contains multiple neurons, the first hidden layer is used to map the initial predicted value of the water level change output by each base model into a multi-dimensional space, so as to extract preliminary features from the initial predicted value of the water level change output by each base model to obtain the first output feature value; the second hidden layer is used to map the multi-dimensional space corresponding to the first output feature value into another multi-dimensional space, perform a non-linear transformation on the first output feature value to obtain the second output feature value; the output layer is used to map the second output feature value extracted by the second hidden layer to the final predicted value of the water level change.
[0073] In some embodiments, the calculation of each layer of the stacked model includes:
[0074] Where: h 1, h 2 and h 3 are the first output feature value, the second output feature value and the final predicted value of the water level change respectively; σ is the RELU activation function; W 1, W 2 and W3 are the weight matrices of the first hidden layer, the second hidden layer, and the output layer respectively; b 1, b 2, and b 3 are the bias vectors of the first hidden layer, the second hidden layer, and the output layer respectively.
[0075] In some embodiments, the base model includes a first base model based on RNN, which includes two RNN layers and one fully connected layer. The hidden state output by the first RNN layer is used as the input value of the second RNN layer; the number of neurons in the first RNN layer is more than that in the second RNN layer; the fully connected layer is used to convert the output value of the second RNN layer into the final predicted water level change value, and an initial predicted value of the water level change is obtained.
[0076] In some embodiments, the base model includes a second base model based on LSTM, which includes two LSTM layers, and the number of neurons in one LSTM layer is half of that in the other LSTM layer.
[0077] In some embodiments, the base model includes a third base model based on GRU, which includes two GRU layers and one fully connected layer. The calculation formula for the output information in each GRU layer includes:
[0078] Where: z t is the update gate, which is used to control the degree to which the hidden state information at the previous moment is passed to the current moment; r t is the reset gate, which is used to control how much information of the hidden state at the previous moment is forgotten; represents the input feature at time step t, that is, the water level data at the current time step t; is the candidate memory; and respectively represent the hidden states at the current time step t and the previous time step t - 1; W z , W r and W h represent the weight matrices of the update gate, the reset gate, and the calculation of the candidate hidden state, b z , b r andb h are the corresponding offsets respectively; tanh is the hyperbolic tangent activation function, which is used to generate candidate memories.
[0079] In some embodiments, the multiple base models include: a first base model based on RNN, a second base model based on LSTM, and a third base model based on GRU; each base model includes two layers, and each layer includes a first set number of neurons.
[0080] An embodiment of the present invention further provides a water level prediction method based on a stacked model that uses the water level prediction method based on a stacked model, including: Construct multiple base models for predicting water level changes; the input data of each base model is historical water level data, and the output data is the initial predicted value of the water level change in a future time period; Construct a stacked model for predicting water level changes; the input data of the stacked model is the initial predicted value of the water level change in a future time period output by each base model, and the output value is the final predicted value of the water level change in a future time period; Use historical water level data to train the base models to obtain each trained base model; Use the initial predicted values of the future time period output by each trained base model as training data to train the stacked model to obtain a water level prediction model for predicting the water level in the future time period; Use the water level prediction model to predict the water level change in a future time period.
[0081] When training the model, use the grid search method to determine the optimal values of the hyperparameters of each base model or stacked model, and evaluate each base model by using 5-fold cross-validation on the validation set; when the error on the validation set fails to improve in consecutive multiple training cycles, stop training.
[0082] An embodiment of the present invention further provides a non-transitory computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, each step of the method of the present invention is implemented, which will not be elaborated here.
[0083] A computer-readable storage medium may be an internal storage unit of the data transmission device or the computer device provided in any of the foregoing embodiments, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device.
[0084] Furthermore, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data to be output or already output.
[0085] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, thereby providing instructions executed on the computer or other programmable device for implementing the process Figure 1 in one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.
[0089] An embodiment of the present invention also provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the water level prediction method based on the stacking model are implemented.
[0090] Contents not detailed in this specification belong to the prior art well known to those skilled in the art.
Claims
1. A water level prediction system based on a stacked model, characterized in that Comprising: Base model construction module: used to construct multiple base models for predicting water level changes; the input data of each base model is historical water level data, and the output data is the initial prediction value of the water level change in the future time period; Stacked model construction module: used to construct a stacked model for predicting water level changes; the input data of the stacked model is the initial prediction value of the water level change in the future time period output by each base model, and the output value is the final prediction value of the water level change in the future time period; Base model training module: used to train the base model using historical water level data to obtain each trained base model; Stacked model training module: used to use the initial prediction value of the future time period output by each trained base model as training data to train the stacked model to obtain a water level prediction model for predicting the water level in the future time period; Water level prediction module: used to predict the water level change in the future time period using the water level prediction model.
2. The water level prediction system based on the stacked model according to claim 1, wherein The stacked model includes two hidden layers and an output layer. Each hidden layer contains multiple neurons. The first hidden layer is used to map the initial prediction value of the water level change output by each base model into a multi-dimensional space, so as to extract preliminary features from the initial prediction value of the water level change output by each base model to obtain the first output feature value; The second hidden layer is used to map the multi-dimensional space corresponding to the first output feature value into another multi-dimensional space, perform a non-linear transformation on the first output feature value to obtain the second output feature value; The output layer is used to map the second output feature value extracted by the second hidden layer to the final water level change prediction value.
3. The water level prediction system based on the stacked model according to claim 2, wherein, The calculation of each layer of the stacked model includes: ; Where: h1, h2, and h3 are the first output feature value, the second output feature value, and the final water level change prediction value respectively; σ is the RELU activation function; W1, W2, and W3 are the weight matrices of the first hidden layer, the second hidden layer, and the output layer respectively; b1, b2, and b3 are the bias vectors of the first hidden layer, the second hidden layer, and the output layer respectively.
4. The water level prediction system based on the stacked model according to claim 1, wherein The base model includes a first base model based on RNN, including two RNN layers and a fully connected layer. The hidden state output by the first RNN layer is used as the input value of the second RNN layer; the number of neurons in the first RNN layer is more than that in the second RNN layer; the fully connected layer is used to convert the output value of the second RNN layer into the final water level change prediction value to obtain the initial prediction value of the water level change.
5. The water level prediction system based on the stacked model according to any one of claims 1 or 4, characterized in that, The base model includes a second base model based on LSTM, including two LSTM layers, and the number of neurons in one LSTM layer is half of that in the other LSTM layer.
6. The water level prediction system based on a stacked model according to any one of claims 1 or 4, characterized in that The base model includes a third base model based on GRU, including two GRU layers and a fully connected layer. The calculation formula for the output information in each GRU layer includes: ; Where: z t For updating the gate, which is used to control the degree to which the hidden state information at the previous moment is passed to the current moment; To the current moment; r t For resetting the door, used to control the hidden state at the previous moment How much information is forgotten; Denote the input feature at time step t, i.e., the water level data at the current time step t; is a candidate memory; and represent the hidden states at the current time step \(t\) and the previous time step \(t - 1\), respectively; W z , W r , W h represent the update gate, the reset gate, and the weight matrix for calculating the candidate hidden state, b z , b r , b h are the corresponding biases respectively; tanh is the hyperbolic tangent activation function, used to generate candidate memories.
7. The water level prediction system based on the stacked model according to claim 1, wherein The multiple base models include: a first base model based on RNN, a second base model based on LSTM, and a third base model based on GRU; each base model contains two layers, and each layer contains the first set number of neurons.
8. The water level prediction method based on a stacked model using the system according to claim 1, characterized in that Comprising: Construct multiple base models for predicting water level changes; the input data of each of the base models is historical water level data, and the output data is the initial predicted value of the water level change in a future time period; Construct a stacking model for predicting water level changes; the input data of the stacking model is the initial predicted value of the water level change in a future time period output by each base model, and the output value is the final predicted value of the water level change in a future time period; Train the base models using the historical water level data to obtain each trained base model; Use the initial predicted values of the future time periods output by each trained base model as training data to train the stacking model, and after training is completed, obtain a water level prediction model for predicting the water level in a future time period; Use the water level prediction model to predict the water level change in a future time period.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the water level prediction method based on a stacking model as described in claim 8.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, it implements the steps of the water level prediction method based on a stacking model as described in claim 8.
Citation Information
Patent Citations
Intelligent water level prediction method based on recurrent neural network and convolutional neural network
CN111242344A
Water flooded waterwheel chamber early warning method and system based on knowledge graph
CN111736636A
Road traffic flow prediction method based on stacked regression
CN117116042A
Hybrid clustering and stacking integrated deep learning photovoltaic output prediction method and system
CN119382063A
Multi-target intelligent hydraulic regulation and control method for tidal river network sluice group
CN119863016A
Cited By
Prediction method based on pyramid decomposition framework and multi-scale stacked LSTMs
CN120822666A
Warehouse-out flow prediction method based on Stacking framework
CN121146176A