Water level prediction method and device, electronic equipment and storage medium
The integration of an improved LSTM neural network with self-attention mechanisms in water level prediction methods addresses the challenge of capturing complex dynamics, resulting in enhanced prediction accuracy and robustness by retaining local details and understanding broader temporal patterns.
Patent Information
- Application Number
- CN202510796197.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing water level prediction methods rely too much on manual experience, poor generalization capabilities of the model, and lack dynamic timing relationship capture capabilities, resulting in insufficient prediction accuracy.
Using an improved LSTM neural network and self-attention mechanism model, the input sequence and output sequence are constructed through a sliding window, combined with input gate, forget gate, output gate and encoder decoder, to capture the time dependence and global information in water level changes, and trained using quantile loss function and Adam optimizer.
It improves the accuracy and robustness of water level prediction, can better reflect the true laws of water level changes, and adapt to complex water level changes.
Smart Images

Figure CN120315071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water level prediction, and particularly to a water level prediction method, device, electronic device, and storage medium. Background Art
[0002] Water level prediction is of great significance in water resource management, flood control and drought relief, and ecological balance maintenance. Traditional water level prediction methods, such as regression analysis and time series prediction models (such as ARIMA), although can provide certain prediction ability in simple scenarios, are difficult to accurately capture the complex water level change rules due to their inherent limitations. These traditional methods usually assume that the data follows a specific distribution or linear relationship, while the actual water level change is affected by various non-linear factors, including rainfall, evaporation, upstream inflow, etc., so it is difficult to meet the requirements of high-precision prediction. In recent years, with the development of deep learning technology, especially the application of recurrent neural networks (RNN) and their variants (such as LSTM, GRU) and self-attention mechanism, new possibilities have been provided for the modeling of complex sequence data. These technologies can automatically extract non-linear features and temporal dependencies in the data without relying on strict mathematical assumptions, thus significantly improving the accuracy of predicting dynamic systems.
[0003] However, when applying deep learning technology to the field of water level prediction, it is overly dependent on manual experience, the model has poor generalization ability, and lacks the ability to capture dynamic temporal relationships, resulting in insufficient prediction accuracy. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a water level prediction method, device, electronic device, and storage medium.
[0005] In a first aspect, an embodiment of the present invention provides a water level prediction method, which includes: Obtain the historical water level data and historical rainfall data of the area to be predicted; Use a sliding window to construct the historical water level data and historical rainfall data into an input sequence and an output sequence; Input the input sequence into a pre-trained water level prediction model to output a prediction result; the prediction result corresponds to the output sequence; Wherein, the water level prediction model includes an improved LSTM neural network and a self-attention mechanism model; the improved LSTM neural network includes an input gate, a forget gate, and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
[0006] Combined with the first aspect, the step of inputting the input sequence into a pre-trained water level prediction model to output a prediction result includes: Input the input sequence into the improved LSTM neural network, and successively pass through the input gate, forget gate, and output gate for local feature extraction, and output the hidden state and memory; Take the hidden state as the eigenvalue and input it into the self-attention mechanism model. After passing through the encoder to capture the global dependency relationship between the input data, the prediction result is output through the decoder.
[0007] Combined with the first aspect, the steps of constructing the input sequence and output sequence from the historical water level data and rainfall data using a sliding window include: For each prediction point in the area to be predicted, align the water level data and rainfall data at the prediction point at the same time frequency; Set the input window length to the data of the first consecutive number of time points in the past, and the output window length to the water level data of the second consecutive number of time points in the future; Move the window with a preset sliding step to generate a training sample set; Among them, the input sequence of each sample in the training sample set contains the water level data and rainfall data of the first number of time points, and the output sequence contains the water level data of the second number of time points; the first number is greater than or equal to the second number.
[0008] Combined with the first aspect, the encoder includes a masked multi-head attention layer and a feed-forward neural network; The steps of taking the hidden state as the eigenvalue and inputting it into the self-attention mechanism model, passing through the encoder to capture the global dependency relationship between the input data, and then outputting the prediction result through the decoder include: Add positional encoding to the input hidden state to preserve the timing information, and obtain the positional encoding data; among them, the positional encoding is generated by sine and cosine functions; Input the positional encoding data into the multi-head attention layer, calculate the query vector, key vector, and value vector, and capture the global dependency relationship through the scaled dot self-attention mechanism to obtain the multi-head attention output; Perform a non-linear transformation on the output of the multi-head attention layer through the feed-forward neural network to obtain the encoder output; Input the encoder output into the decoder to output the prediction result.
[0009] Combined with the first aspect, the steps of adding positional encoding to the input hidden state to preserve the timing information and obtaining the positional encoding data include: Calculate the positional encoding with the following formula: ; ; Among them, is the positional encoding function, is the position index, is the a time point, is the dimension of the self-attention mechanism model.
[0010] Combined with the first aspect, the decoder includes a masked multi-head attention layer; The step of inputting the encoder output into the decoder and outputting the prediction result includes: Taking the encoder output as the input of the masked multi-head attention layer, masking future time step information through a lower triangular matrix to prevent prediction leakage, and outputting the masked attention result; Interacting the masked attention result with the encoder output, fusing features again through the multi-head attention layer, generating the final prediction sequence through the feed-forward neural network and the fully connected layer, and outputting the final prediction sequence as the prediction result.
[0011] Combined with the first aspect, when training the water level prediction model, the quantile loss is used as the loss function and the Adam optimizer is used for parameter update.
[0012] In the second aspect, the present application provides a water level prediction device, and the device includes: An acquisition module for acquiring historical water level data and historical rainfall data of the area to be predicted; A data preprocessing module for constructing the historical water level data and historical rainfall data into an input sequence and an output sequence by using a sliding window; A prediction module for inputting the input sequence into a pre-trained water level prediction model and outputting a prediction result, and the prediction result corresponds to the output sequence; Wherein, the water level prediction model includes an improved LSTM neural network and a self-attention mechanism model; the improved LSTM neural network includes an input gate, a forget gate and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
[0013] In the third aspect, the present application provides an electronic device, and the electronic device includes a memory and a processor. The memory is used for storing a computer program, and the processor runs the computer program to enable the electronic device to execute the above method.
[0014] In the fourth aspect, the present application provides a readable storage medium, and computer program instructions are stored in the readable storage medium. When the computer program instructions are read and run by a processor, the above method is executed.
[0015] The embodiments of the present invention bring the following beneficial effects: The water level prediction method, device, electronic device, and storage medium provided by this application. The method includes obtaining historical water level data and historical rainfall data of the area to be predicted; using a sliding window to construct the historical water level data and historical rainfall data into an input sequence and an output sequence; inputting the input sequence into a pre-trained water level prediction model to output a prediction result; the prediction result corresponds to the output sequence; wherein, the water level prediction model includes an improved LSTM neural network and a self-attention mechanism model; the improved LSTM neural network includes an input gate, a forget gate, and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
[0016] The water level prediction method provided by this application combines an improved LSTM neural network and a self-attention mechanism model, which can effectively capture the time-dependent relationship in water level changes, that is, the understanding of global information, so as to capture a wider time series relationship without losing local details, thereby better reflecting the true law of water level changes and improving prediction accuracy, robustness, and adaptability.
[0017] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0018] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flowchart of the water level prediction method provided by the embodiments of the present invention; Figure 2 It is a schematic structural diagram of the self-attention mechanism model in the water level prediction method provided by the embodiments of the present invention; Figure 3 It is a combined schematic diagram of the water level prediction device provided by the embodiments of the present invention; Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present invention.
[0021] Reference Signs: 10 - Acquisition module, 20 - Data preprocessing module, 30 - Prediction module; 130 - Processor, 131 - Memory, 132 - Bus, 133 - Communication interface. Detailed implementation manners
[0022] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] To facilitate the understanding of this embodiment, the application scenario and design concept of the embodiments of this application will be briefly introduced below.
[0024] When the deep learning technology provided by the prior art is applied to the field of water level prediction, it overly relies on manual experience, has poor model generalization ability, and lacks the ability to capture dynamic time series relationships, resulting in insufficient prediction accuracy.
[0025] Based on this, the embodiments of this application provide a water level prediction method, device, electronic device, and storage medium.
[0026] Embodiment 1 This application provides a water level prediction method, which combines Figure 1 As shown, this method includes: S110, acquiring historical water level data and historical rainfall data of the area to be predicted.
[0027] S120, using a sliding window to construct the historical water level data and historical rainfall data into an input sequence and an output sequence.
[0028] S130, inputting the input sequence into a pre-trained water level prediction model to output a prediction result; the prediction result corresponds to the output sequence.
[0029] Among them, the water level prediction model includes an improved LSTM neural network-based and self-attention mechanism model; the improved LSTM neural network includes an input gate, a forget gate, and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
[0030] The water level prediction method provided by this application, through the combination of an improved LSTM neural network and a self-attention mechanism model, can effectively capture the time-dependent relationship in the water level change, that is, the understanding of global information, so as to capture a wider time series relationship without losing local details, thereby better reflecting the true law of water level change to improve prediction accuracy, robustness, and adaptability.
[0031] Combined with the first aspect, step S130 includes: S131, input the input sequence into the improved LSTM neural network, and sequentially pass through the input gate, forget gate, and output gate to perform local feature extraction, and output the hidden state and memory.
[0032] S132, use the hidden state as the eigenvalue to input into the self-attention mechanism model. After passing through the encoder to capture the global dependency relationship between the input data, the prediction result is output through the decoder.
[0033] Among them, the forget gate is used to determine which information needs to be discarded from the memory unit (cell state) at the previous moment (such as outdated water level data or irrelevant rainfall records), and its mathematical formula is as follows: ; Among them, is the output of the forget gate, used to control the discarding ratio of old information, indicating the degree of forgetting; is the Sigmoid function, and the output range is [0, 1]; is the weight matrix of the forget gate; is the concatenation of the hidden state at the previous moment and the current input (water level, rainfall); is the bias term of the forget gate.
[0034] In water level prediction, if the upstream inflow suddenly decreases (such as the gate is closed), the forget gate may output , discarding the memory of the past "high water level" to avoid misleading the prediction.
[0035] The input gate is used to determine which new information needs to be stored in the memory unit, and its mathematical formula is as follows: Input gate activation value: ; This step is used to determine which information is important; among them, is the input gate activation value, that is, the output of the input gate, indicating the opening degree of the input gate; is the Sigmoid function, with an output range of [0, 1], determining the retention ratio of information; is the weight matrix corresponding to the input gate; is the hidden state at the previous moment; is the input at the current moment (water level, rainfall, etc.); is the bias term.
[0036] Candidate memory content: ; This step is used to generate new information to be added; among them, is the candidate memory content, indicating new information; is the weight matrix of the candidate state; is the hyperbolic tangent function, with an output range of [-1, 1], generating potential values of new information; Memory cell update: ; where, is the state of the memory cell at the current moment, used to update the memory cell; is the output of the forget gate, used to control the proportion of old information discarded; is the state of the memory cell at the previous moment; is the filtered new information.
[0037] The input gate determines the storage of new information in two steps: First, filter information: Determine which information needs to be retained through the Sigmoid function (0 to 1, close to 1 indicates importance).
[0038] Subsequently, generate candidate information: Generate candidate memory content (-1 to 1, representing possible values of new information) through the function.
[0039] Finally, the input gate adds the filtered candidate information to the memory cell.
[0040] An application example of the input gate in water level prediction is as follows: When a rainstorm occurs suddenly, , it is considered that the rainfall data is crucial, generate a candidate memory of "the water level will rise rapidly" and update it to .
[0041] The output gate is used to determine which information needs to be output from the memory cell and the hidden state ( ), for final prediction. Its mathematical formula is as follows: Output gate activation value: ; where, is the output of the output gate, indicating the opening degree of the output gate; is the Sigmoid function, with an output range of [0, 1]; is the weight matrix of the output gate; is the bias term of the output gate.
[0042] Hidden state generation: ; where, is the hidden state at the current moment. An application example of the output gate in water level prediction is as follows: If it is necessary to predict the water level in the next 10 days: The output gate filters information related to the long-term trend (such as the rainy season cycle) in the memory cell and ignores short-term fluctuations (such as daily evaporation).
[0043] In this application, the improved LSTM neural network includes a forget gate, an input gate, and an output gate. The three gates work together to handle the suddenness (such as heavy rain) and periodicity (such as tides) of water level changes. At the same time, the forget gate and the input gate jointly maintain the memory unit to retain important information (such as seasonal rainfall patterns) in the long term while ignoring irrelevant noise and avoiding the vanishing gradient problem of traditional RNNs. In addition, the output gate ensures that the model only transmits features related to future water levels, improving prediction stability. In this way, the improved LSTM neural network can effectively handle complex temporal dependencies in water level prediction and outperform traditional statistical models.
[0044] In this embodiment, the weight matrix is initialized randomly, such as using a Gaussian distribution or a normal distribution; the bias term is usually initialized to zero or a small constant (0.01).
[0045] Combined with the first aspect, step S120 includes: S121, for each prediction point in the area to be predicted, align the water level data and rainfall data of the prediction point at the same time frequency.
[0046] S122, set the input window length to the data of the first consecutive number of time points in the past, and the output window length to the water level data of the second consecutive number of time points in the future.
[0047] S123, move the window with a preset sliding step to generate a training sample set.
[0048] Among them, the input sequence of each sample in the training sample set includes the water level data and rainfall data of the first number of time points, and the output sequence includes the water level data of the second number of time points; the first number is greater than or equal to the second number.
[0049] After obtaining the historical water level data and historical rainfall data of the area to be predicted in step S110, in step S120, a sequence is constructed. Specifically, in step S121, the historical water level data and historical rainfall data are aligned at the same time frequency to ensure the consistency of the water level data and rainfall data in the time dimension, so that they can be accurately used for subsequent model training. The specific operations are as follows: Determine the time frequency: First, determine a suitable time frequency (for example, every hour, every day, etc.), and this frequency should be determined according to the data characteristics and prediction requirements.
[0050] Data alignment: For each time point, check whether there is corresponding water level data and rainfall data.
[0051] If a certain time point lacks a certain type of data, it needs to be filled in by interpolation or other data filling methods. Ensure that each time point has complete water level and rainfall data.
[0052] After these steps are completed, two sets of data will be obtained - historical water level data and historical rainfall data, both of which are arranged according to the selected time frequency and correspond one by one at each time point. The processed data can be directly used to construct the input sequence and output sequence, and then for model training and prediction.
[0053] Subsequently, step S122 sets two key parameters: the input window length and the output window length. The input window length is the data of the past consecutive "first quantity" time points. The output window length is defined as the water level data of the future consecutive "second quantity" time points, which is the target value that the model needs to predict.
[0054] Because the model needs sufficient historical data to accurately predict future water level changes, generally, the "first quantity" is greater than or equal to the "second quantity".
[0055] In this embodiment, the first quantity is 30 days, and the water level data and rainfall data of the past consecutive 30 days are used as the input of the model. The second quantity is 7 days, indicating that it is expected that the model predicts the water level data for the next 7 days.
[0056] Subsequently, step S123 moves the window with a preset sliding step to generate a training sample set; the training sample includes the water level data and rainfall data of the corresponding previous 30 days within the window.
[0057] It can be understood that the water level prediction model needs to be trained before step S110. The training sample set required in this training process is similar to that in step S120. Preferably, the original data can be collected first, a local data warehouse can be constructed, and then the method similar to step S120 can be used to slide the window to select data. Subsequently, the selected data is divided into a training set and a validation set. The water level prediction model is trained through the training set, and the trained water level prediction model is verified through the validation set until the trained water level prediction model is obtained.
[0058] Among them, the steps of constructing the local data warehouse include: Access and save the real-time water level data and historical water level data (at least 5 years) of the water level monitoring sites to be predicted and their upstream tributary water level monitoring sites, with a time interval of once a day / once an hour / once every half hour.
[0059] Access and save the rainfall data in the control area of the monitoring site to be predicted and the predicted rainfall data for the next 7 days, with a time interval of once a day or once an hour or once every half hour, and the frequency is the same as that of the water level data.
[0060] Combined with the first aspect, the encoder includes a masked multi-head attention layer and a feed-forward neural network. Step S132 includes: S1321. Add positional encoding to the input hidden state to retain temporal information, obtaining positional encoding data. The positional encoding is generated through sine and cosine functions.
[0061] S1322. Input the positional encoding data into the multi-head attention layer, calculate the query vector, key vector, and value vector to capture global dependencies through the scaled dot-product self-attention mechanism, obtaining the multi-head attention output.
[0062] S1323. Perform a non-linear transformation on the output of the multi-head attention layer through a feed-forward neural network to obtain the encoder output.
[0063] S1324. Input the encoder output into the decoder to output the prediction result.
[0064] The core of the self-attention mechanism model is the encoder-decoder architecture. The main components include: the encoder - processing the time series features of the input; the decoder - outputting the future time series. The self-attention mechanism is used to capture the global dependencies between the input data.
[0065] Combined with the first aspect, step S1321 includes: Calculate the positional encoding with the following formula: ; ; where is the positional encoding function, is the position index, is the th time point, is the dimension of the self-attention mechanism model, is the dimension of the self-attention mechanism model, is the sine function, is the cosine function.
[0066] Use the hidden state at the last moment of the improved neural network in step S131 as the feature and input it into the self-attention mechanism model, and add positional encoding through the above formula.
[0067] Subsequently, as shown in Figure 2 , in step S1322, the multi-head attention mechanism is adopted in the encoder part to enable the self-attention mechanism model to focus on the information of all positions in the sequence when calculating the representation of a certain position. Let the input matrix be , where is the sequence length, and obtain the query vector, key vector, and value vector through linear transformation. Specifically: ; ; ; Among them, is the query vector, which is used to clarify the information required for the prediction target; is the key vector, which is used to filter the key patterns in the historical data; is the value vector, which is used to provide specific numerical contributions, are the weight matrices for learning respectively.
[0068] Calculate the attention weights: ; Among them, is and of the dimension; is used for scaling to prevent the inner product value from being too large; is the normalization process, which is used to ensure the weight normalization; is the transpose of the vector.
[0069] The feed-forward neural network is used to further extract features, and the calculation formula is as follows: ; Among them, is the input matrix, is the preset first weight matrix, is the preset first bias term, is the preset second weight matrix, is the preset second bias term.
[0070] In step S130, first add positional encoding to the input to retain the temporal information to obtain the positional encoding data; then use the output of step S131 (i.e., the positional encoding data) as the input of the multi-head attention layer, and calculate the query vector , key vector and value vector to capture the global dependency output; then, perform a non-linear transformation through the feed-forward neural network to obtain the encoder output.
[0071] Combined with the first aspect, the decoder includes a masked multi-head attention layer. Step S133 includes: S1331, use the encoder output as the input of the masked multi-head attention layer, and shield the future time step information through the lower triangular matrix to prevent prediction leakage, and output the masked attention result.
[0072] The masked multi-head attention layer (Masked Multi-Head Attention) uses the encoder output (historical water level / rainfall characteristics) as the key vector and value vector , the decoder input (the predicted partial sequence) is used as a query vector .
[0073] Specifically: ; Among them, is a lower triangular mask matrix. For example: .
[0074] S1332. Interact the masked attention result with the encoder output. After fusing features through the multi-head attention layer again, generate the final prediction sequence through the feed-forward neural network and the fully connected layer, and output the final prediction sequence as the prediction result.
[0075] Use the output of step S1331 (the masked attention result) as the new query vector , and use the encoder output as the key vector again and the value vector to align the historical features (encoder) with the current prediction state (decoder).
[0076] Subsequently, perform secondary fusion through the multi-head attention layer to enhance the model's attention to historical key events; then further extract non-linear features through the feed-forward neural network, and map the high-dimensional features to the output dimension through the fully connected layer (combining the prediction target in this application: water level data for the next 7 days).
[0077] Combined with Figure 2 as shown, the input embedding (Input Embedding) is used to convert the original input (such as time series data such as water level and rainfall) into a dense vector representation. Specifically, map features such as daily or hourly water level values and rainfall into high-dimensional vectors, retaining the potential relationships between time series and variables; the positional encoding (Positional Encoding) is used to add position information to the input sequence to make up for the self-attention mechanism's neglect of the time series order. For example, distinguish the impact differences of water level data on "the first day" and "the 30th day" on predicting future water levels; the encoder (Encoder) is stacked by N identical layers, and each layer contains the following sub-modules with residual connections: multi-head attention layer (Multi-Head Attention), layer normalization (LayerNormalization), feed-forward neural network (Feed Forward Network).
[0078] Among them, the multi-head attention layer is used to calculate multiple groups of query vectors, key vectors, and value vectors in parallel to capture the global dependencies at different positions in the sequence. For example, it can simultaneously focus on the collaborative effects of multiple factors such as "upstream inflow", "local rainfall", and "evaporation" on the water level; layer normalization is used to stabilize the model prediction process by normalizing the output of each layer; the feed-forward neural network further extracts non-linear features through fully connected layers; the residual connection is used to add the output of each sub-module to the input (residual connection) to alleviate the vanishing gradient problem.
[0079] The decoder is also stacked by N identical layers. In addition to the modules of the encoder, it also includes: a masked multi-head attention layer.
[0080] Among them, the masked multi-head attention layer is used to set the weights of future positions to negative infinity through a lower triangular mask matrix to prevent the decoder from "peeking" at future information during prediction (ensuring that the time step only depends on historical data). For example, when predicting the water level on the 7th future day, only known historical data is used to avoid information leakage; the multi-head attention layer is used to interact the query vectors of the decoder with the key vectors and value vectors of the encoder to align the key features of the input and output. For example, it dynamically associates the historical water level (encoder output) with the future prediction target (decoder input).
[0081] The output part of the self-attention mechanism model includes a linear layer and a Softmax layer. The linear layer is used to map the high-dimensional vector output by the decoder to the target dimension (such as the water level values for the next 7 days); the Softmax layer is used to perform probability transformation on the output result of the linear layer. If the task is classification (such as predicting the water level warning level), it outputs a probability distribution. If it is a regression task (such as predicting the water level value), it may also be replaced by a linear activation.
[0082] Combined with the first aspect, the quantile loss is used as the loss function during the training of the water level prediction model, and the Adam optimizer is used to update the parameters.
[0083] In this embodiment, the water level prediction model needs to be trained before actual application to enable the model to accurately predict the water level. Define the hyperparameters of the model: the number of hidden layers of the improved LSTM neural network is a constant greater than 1; the number of hidden layers of the self-attention mechanism module is a constant greater than 1; the number of layers of the self-attention mechanism module is a constant greater than 1; the learning rate is between 0 and 1; the number of iterations is a constant greater than 1.
[0084] Among them, the number of layers of the hidden layer of the LSTM neural network usually ranges from 2 to 4. A deep network can capture complex time series features (such as the rainfall-water level lag effect); the number of hidden units of the LSTM neural network is 64 to 256. The more hidden units there are, the larger the model capacity, but overfitting needs to be prevented; the number of hidden layers of the self-attention mechanism module is usually 128 - 512, and this value affects the expression ability of feature interaction; the learning rate controls the step size of parameter update. If it is too large, it is easy to cause oscillation, and if it is too small, it will lead to slow convergence; the number of iterations usually ranges from 50 to 200, and it is not easy to choose a too large value to cause overfitting.
[0085] For example, define the number of layers of the hidden layer of the LSTM neural network as 3, the number of hidden units of the LSTM neural network as 64, the number of hidden layers of the self-attention mechanism module as 128, and the number of layers of the self-attention mechanism module as 8; the learning rate is 0.01, and the number of iterations is 100.
[0086] In addition, define the quantile loss as the loss function, and use the Adam optimizer to update the model parameters. Record the loss value of each epoch during the model training process to observe whether the model converges gradually.
[0087] The specific loss function can be expressed as:
[0088]
[0089]
[0090] Among them, is the predicted water level; is the measured water level; is the quantile loss value; is the number of training samples; is the quantile threshold, ranging between 0 and 1 (e.g., 0.5 corresponds to the median); is the error; is the indicator function. When is true, = 1; when is true, = 0.
[0091] Adopt the optimizer (Adam) to adaptively adjust the learning rate and combine momentum to accelerate convergence.
[0092] After constructing the local data warehouse, the training dataset and the validation dataset are selected in the manner of step S120. The training set is input into the water level prediction model for model training, and the validation is carried out through the validation set, and iterative optimization is performed. After each training, the loss function is calculated to improve the model prediction accuracy.
[0093] Preferably, a comparison chart is drawn by combining historical water level data (combining the above example, the water level data of the past 30 days) and prediction results (combining the above example, that is, the water level trend predicted for the next 7 days) with different color or line representation forms to facilitate the observation of the predicted water level.
[0094] In a second aspect, the present application provides a water level prediction device, in combination with Figure 3 As shown, the device includes: an acquisition module 10, a data preprocessing module 20, and a prediction module 30.
[0095] The acquisition module 10 is used to acquire historical water level data and historical rainfall data of the area to be predicted.
[0096] The data preprocessing module 20 is used to construct the historical water level data and historical rainfall data into an input sequence and an output sequence by using a sliding window.
[0097] The prediction module 30 is used to input the input sequence into the pre-trained water level prediction model and output a prediction result, and the prediction result corresponds to the output sequence.
[0098] Among them, the water level prediction model includes an improved LSTM neural network and a self-attention mechanism model; the improved LSTM neural network includes an input gate, a forget gate, and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
[0099] In a third aspect, an embodiment of the present application provides an electronic device, in combination with Figure 4 As shown, the electronic device includes a memory 131 and a processor 130. The memory 131 is used to store a computer program, and the processor 130 runs the computer program to enable the electronic device to execute the above method.
[0100] Furthermore, the electronic device shown in combination with Figure 4 also includes a bus 132 and a communication interface 133. The processor 130, the communication interface 133, and the memory 131 are connected through the bus 132.
[0101] Among them, the memory 131 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is implemented through at least one communication interface 133 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 132 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a bidirectional arrow is used in Figure 4 , but it does not mean that there is only one bus or one type of bus.
[0102] The processor 130 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 130 or the instructions in software form. The above-mentioned processor 130 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 131, and the processor 130 reads the information in the memory 131 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0103] In a fourth aspect, an embodiment of the present application provides a readable storage medium, in which computer program instructions are stored. When the computer program instructions are read and run by a processor, the above-mentioned method is executed.
[0104] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0105] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0106] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0107] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0108] Finally, it should be noted that the above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A water level prediction method, characterized in that, The method includes: Obtaining historical water level data and historical rainfall data of the area to be predicted; Using a sliding window to construct the historical water level data and the historical rainfall data into an input sequence and an output sequence; Inputting the input sequence into a pre-trained water level prediction model to output a prediction result; the prediction result corresponds to the output sequence; Wherein, the water level prediction model includes an improved LSTM neural network and a self-attention mechanism model; the improved LSTM neural network includes an input gate, a forget gate, and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
2. The method according to claim 1, wherein The step of inputting the input sequence into the pre-trained water level prediction model to output a prediction result includes: Inputting the input sequence into the improved LSTM neural network, and successively passing through the input gate, the forget gate, and the output gate to perform local feature extraction, and outputting a hidden state and a memory; Taking the hidden state as a feature value and inputting it into the self-attention mechanism model, after capturing the global dependence relationship among the input data through the encoder, the prediction result is output through the decoder.
3. The method according to claim 1, wherein The step of using a sliding window to construct the historical water level data and the rainfall data into an input sequence and an output sequence includes: For each prediction point in the area to be predicted, aligning the water level data and the rainfall data of the prediction point at the same time frequency; Setting the input window length as the data of the first consecutive number of time points in the past, and the output window length as the water level data of the second consecutive number of time points in the future; Moving the window with a preset sliding step length to generate a training sample set; Wherein, the input sequence of each sample in the training sample set includes the water level data and the rainfall data of the first number of time points, and the output sequence includes the water level data of the second number of time points; the first number is greater than or equal to the second number.
4. The method according to claim 2, characterized in that, The encoder includes a masked multi-head attention layer and a feed-forward neural network; The step of taking the hidden state as a feature value and inputting it into the self-attention mechanism model, after capturing the global dependence relationship among the input data through the encoder, and outputting the prediction result through the decoder includes: Adding position encoding to the input hidden state to retain the timing information, obtaining position-encoded data; wherein, the position encoding is generated by sine and cosine functions; Inputting the position-encoded data into the multi-head attention layer, calculating query vectors, key vectors, and value vectors to capture the global dependence relationship through the scaled dot self-attention mechanism, and obtaining the multi-head attention output; Performing a non-linear transformation on the output of the multi-head attention layer through a feed-forward neural network to obtain the encoder output; Inputting the encoder output into the decoder to output the prediction result.
5. The method according to claim 4, wherein The step of adding position encoding to the input hidden state to retain the timing information and obtaining position-encoded data includes: Calculating the position encoding with the following formula: ; ; Among them, is the position encoding function, is the position index, is the th time point, is the dimension of the self-attention mechanism model.
6. The method according to claim 4, characterized in that The decoder includes a masked multi-head attention layer; The step of inputting the encoder output into the decoder to output the prediction result includes: Taking the encoder output as the input of the masked multi-head attention layer, shielding the future time step information through a lower triangular matrix to prevent prediction leakage, and outputting the masked attention result; Interact the masked attention result with the encoder output. After fusing features through the multi-head attention layer again, generate a final prediction sequence through the feed-forward neural network and the fully connected layer, and output the final prediction sequence as the prediction result.
7. The method according to claim 1, characterized in that When training the water level prediction model, the quantile loss is used as the loss function and the Adam optimizer is used for parameter update.
8. A water level prediction device, characterized in that, The device includes: An acquisition module, configured to acquire historical water level data and historical rainfall data of the area to be predicted; A data preprocessing module, configured to construct the historical water level data and the historical rainfall data into an input sequence and an output sequence by using a sliding window; A prediction module, configured to input the input sequence into a pre-trained water level prediction model and output a prediction result, where the prediction result corresponds to the output sequence; Wherein, the water level prediction model includes an improved LSTM neural network and a self-attention mechanism model; the improved LSTM neural network includes an input gate, a forgetting gate, and an output gate, and the self-attention mechanism model includes an encoder and a decoder.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, Computer program instructions are stored in the storage medium, and when the computer program instructions are read and run by a processor, the method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Underground water level prediction method, equipment and medium
CN118014391A
Method for constructing intersection pedestrian-vehicle trajectory prediction model based on heterogeneous graph network
CN118261051A
Charging station cluster load prediction method and system based on deep fusion of exogenous variables
CN118504792A
Password guessing method and device, electronic equipment and medium
CN119646798A
Spare part inventory data prediction method, system, equipment and medium
CN119886428A