Multi-step Prediction Method of Groundwater Level Based on Convolutional Attention Long Short-Term Neural Network

Through the convolutional attention long-term neural network model, multi-sensor data is analyzed and feature extraction is solved, and the traditional method's insufficient accuracy of groundwater level prediction in complex engineering environments is achieved, and high-precision prediction of groundwater level for a long time in the future is achieved, supporting groundwater disaster decision-making.

CN117520784BActive Publication Date: 2025-07-25CSIC INTERNATIONAL ENGINEERING CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311622198.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-07-25
Estimated Expiration
2043-11-30

AI Technical Summary

Technical Problem

The existing groundwater level prediction methods rely on statistical or physical models and are difficult to adapt to complex engineering environments. The diversity and long-term dependence of groundwater level data increase the difficulty of prediction. Traditional methods have shortcomings in accuracy and stability.

Method used

A hybrid neural network model based on convolutional attention long and short-term neural network is adopted, combining convolutional neural network, attention mechanism and long-term memory network to analyze and feature extraction of multi-sensor monitoring timing data. Through the combination of one-dimensional convolutional neural network and attention mechanism, accurate prediction of the groundwater level for a long time in the future can be achieved.

Benefits of technology

It improves the prediction accuracy of groundwater levels for a long time in the future, can cope with the diversity and complexity of groundwater level data, overcome the impact of lack and noise, provide more reliable data support, and provide a basis for groundwater disaster decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117520784B_ABST
    Figure CN117520784B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-step prediction method for groundwater levels based on a convolutional attention long short-term neural network, comprising: S10, obtaining monitoring data related to groundwater levels to obtain multi-feature monitoring time series data; S20, constructing a multi-feature monitoring data sample tensor based on the multi-feature monitoring time series data; S30, establishing a convolutional attention long short-term neural network model; S40, training the convolutional attention long short-term neural network model based on the multi-feature monitoring data sample tensor to obtain an optimal model; S50, substituting the monitoring data for prediction into the optimal model, and outputting multi-step prediction data through model calculation. The present invention can combine a convolutional neural network, an attention mechanism and a long short-term memory neural network, and is a hybrid neural network model, which can greatly improve the prediction accuracy of future long-term groundwater levels and provide more reliable data support for groundwater disaster decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of research on disaster prevention and control in civil engineering and water conservancy projects, and particularly to a method for predicting groundwater levels based on deep learning, specifically a multi-step prediction method for groundwater levels based on a convolutional attention long short-term neural network. Background Art

[0002] Groundwater is an important factor affecting the safety of underground projects. Groundwater can cause various engineering disasters, such as soil liquefaction, settlement or water ingress of underground tunnels, pipelines, and other underground facilities. In addition, groundwater can also induce geological disasters such as landslides and collapses. Therefore, accurately predicting groundwater level changes can better prevent and manage potential engineering disasters.

[0003] Traditional groundwater level prediction methods usually rely on statistical models or physical models. These methods are limited to a certain extent by data quality and model complexity, and the prediction accuracy cannot adapt to complex engineering environments.

[0004] With the development of deep learning technology, especially the rise of neural network structures such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), some new methods have emerged to solve the groundwater level prediction problem. These methods use neural networks to automatically capture temporal patterns and feature extraction in data, and have shown good performance in some groundwater level prediction tasks.

[0005] However, groundwater level data is affected by various factors such as meteorology, geology, and hydrology. The water level data has a high degree of diversity and complexity, and multi-step prediction of groundwater levels is still challenging. Groundwater level data usually has long-term dependencies, so long-term future predictions need to be considered, which increases the difficulty of groundwater level prediction. Groundwater level data may also be affected by missing values, noise, or outliers, and effective data processing methods are required. Since accurate prediction of groundwater levels is crucial for engineering safety, more effective, accurate, and stable groundwater level prediction methods are needed to improve prediction accuracy to meet the needs of engineering disaster prevention. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention establishes a multi-step prediction method for groundwater levels based on a convolutional attention long short-term neural network. This method can combine convolutional neural networks, attention mechanisms, and long short-term memory neural networks, and is a hybrid neural network model that can greatly improve the prediction accuracy of future long-term groundwater levels.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0008] The present invention provides a multi-step prediction method for groundwater level based on a convolutional attention long short-term neural network, comprising the following steps: S10, obtaining monitoring data related to the groundwater level to obtain multi-feature monitoring time series data; S20, constructing a multi-feature monitoring data sample tensor based on the multi-feature monitoring time series data; S30, establishing a convolutional attention long short-term neural network model; S40, training the convolutional attention long short-term neural network model based on the multi-feature monitoring data sample tensor to obtain an optimal model; S50, substituting the monitoring data for prediction into the optimal model, and calculating through the model to output multi-step prediction data.

[0009] As a preferred embodiment, in step S10, the monitoring data is multi-sensor data collected according to time series, including: groundwater level height, rainfall, temperature sensor data, local reservoir drainage volume, local river water level height.

[0010] As a preferred embodiment, in step S20, the constructing of the multi-feature monitoring data sample tensor includes:

[0011] S201, preprocessing the multi-feature monitoring time series data, filling in the missing data, and deleting the abnormal data;

[0012] S202, forming the multi-feature monitoring time series data into a multi-feature monitoring time series data matrix X * , where each row represents a time series sample and each column represents different features;

[0013] S203, performing normalization conversion on each column data X * of the multi-feature monitoring time series data matrix X * i to obtain a normalized sample matrix X;

[0014] S204, processing the monitoring time data, converting the date into a timestamp, and incorporating the timestamp into the normalized sample matrix X as a column of the matrix;

[0015] S205, dividing the normalized sample matrix X into a training set, a validation set, and a test set according to the number of rows;

[0016] S206, respectively dividing the training set, the validation set, and the test set matrices into three three-dimensional tensors x_train, x_val, x_test, and the dimensions of the three-dimensional tensors are: batch size, time step, and number of feature channels.

[0017] As a preferred embodiment, in step S203, the following formula is used to perform normalization conversion on each column data X * of the multi-feature monitoring time series data matrix X * i :

[0018]

[0019] Wherein, k is the serial number of each column of data, that is represents the multi-feature monitoring time-series data matrix X * the data value of the k-th time-series data in the i-th column of is the time-series data the normalized function value, and it is also the k-th time-series data in the i-th column that constitutes the normalized sample matrix X. The value range of the i-th column of data is , that is and are respectively the minimum value and the maximum value of the i-th column of data X* in the multi-feature monitoring time-series data matrix X* i Each column of data X in the normalized sample matrix X i represents a kind of monitoring data i.

[0020] As a preferred embodiment, in step S204, the date is converted into a timestamp X by using the following formula t , and the timestamp X t is incorporated into the normalized sample matrix X:

[0021]

[0022] Wherein, t represents the date, and X t represents the converted timestamp.

[0023] As a preferred embodiment, in step S30, the convolutional attention long short-term neural network model includes: a convolutional layer, a multi-head self-attention mechanism layer, a long short-term recurrent neural network layer, and a fully connected layer.

[0024] As a preferred embodiment, the convolutional layer includes a one-dimensional convolutional layer and a one-dimensional max pooling layer, wherein:

[0025] The one-dimensional convolutional layer is represented by the following formula:

[0026]

[0027] Wherein, ReLU represents the rectified linear unit activation function, conv1d represents the one-dimensional convolution operation, W is the convolution kernel, b is the bias of the convolutional layer, and x cov represents a three-dimensional tensor;

[0028] The one-dimensional max pooling layer is represented by the following formula:

[0029] MaxPool1d ( x pool ) = max(Conv( x pool )[i:i + kernel_size ])

[0030] In the formula, MaxPool1d represents a one-dimensional max pooling operation, i represents the starting position of the pooling kernel on the input sequence, kernel_size represents the pooling kernel size, and x pool represents the output tensor of the one-dimensional convolutional layer.

[0031] As a preferred embodiment, the multi-head self-attention mechanism layer is composed of a self-attention score layer and a multi-head attention layer. The multi-head attention layer is represented by the following formula:

[0032]

[0033] In the formula, Q represents the query, which is the representation used to calculate the attention score, K represents the key, which is the representation used to compare with the query, V represents the value, which is the representation multiplied by the attention weight to generate the final output. Q, K, and V are obtained through linear transformation by the following self-attention score layer. i represents the i-th attention head in the multi-head self-attention mechanism:

[0034]

[0035]

[0036]

[0037] In the formula, X is the input tensor of the self-attention score layer, W Qi and W Ki and W Vi are trainable parameter matrices.

[0038] As a preferred embodiment, the long short-term memory neural network layer includes: a forget gate, an input gate, and an output gate, as well as a cell state and a hidden state; the steps for the long short-term memory neural network layer to process data are as follows:

[0039] (1), At each time step, calculate the values of the forget gate, the input gate, and the output gate according to the hidden state of the previous time step and the input of the current time step;

[0040] (2), Use the forget gate to determine the information to be forgotten from the cell state;

[0041] (3), Use the input gate and the new candidate value to update the cell state and add new information to the cell state;

[0042] (4), Use the output gate to determine the information extracted from the cell state and generate the hidden state of the current time step;

[0043] (5), The cell state and the hidden state are passed to the next long short-term recurrent neural network layer or used to generate the final output at the next time step;

[0044] (6), Connect the output data of the long short-term recurrent neural network layer to a fully connected layer to achieve feature combination and non-linear transformation, and output a three-dimensional tensor with dimensions: batch size, time step, and number of prediction days.

[0045] As a preferred implementation, in step S40, the training of the convolutional attention long short-term neural network model includes the following steps:

[0046] S401, Select the time window size according to requirements, and adjust and evaluate it;

[0047] S402, Use the convolutional attention long short-term neural network model for training, use the training set data and apply the backpropagation algorithm to adjust the model parameters to reduce the loss function;

[0048] S403, Use the Adam optimizer and an appropriate learning rate to optimize the model;

[0049] S404, Monitor the model performance metrics, loss function, and accuracy, and use the validation set to evaluate the model performance to avoid overfitting;

[0050] S405, Dynamically adjust the learning rate using the cosine annealing strategy;

[0051] S406, When the loss value on the validation set reaches the lowest point, save the model parameters to obtain the optimal model.

[0052] As a preferred implementation, the model performance metrics are measured by the mean squared error SME and the coefficient of determination R 2 value, and the model with the optimal values of the mean squared error SME and the coefficient of determination R 2 is taken for multiple trainings.

[0053] The beneficial effects of the present invention compared with the prior art are as follows: The multi-step prediction method of groundwater level based on convolutional attention long short-term neural network proposed by the present invention uses convolutional neural network and attention mechanism in deep learning to analyze and extract features from multi-sensor monitoring time series data, and then inputs the extracted data into a long short-term neural network for model training to achieve multi-step prediction of groundwater level. By combining one-dimensional convolutional neural network and attention mechanism, the accuracy of feature extraction of both in the neural network model is fully exerted, and accurate prediction of groundwater level in the future for a long time is realized. This technological innovation aims to address the complexity of groundwater level monitoring and the need for long-term prediction of groundwater level in the future, and can provide more reliable data support for decision-making related to groundwater disasters.

[0054] The idea of the present invention is a multi-step prediction method of groundwater level based on deep learning technology. The specific advantages include at least one or more of the following:

[0055] (1) The multi-step prediction of groundwater level based on the improved long short-term recurrent neural network layer LSTM model of the present invention can achieve accurate prediction of groundwater level in the future for a long time.

[0056] (2) The present invention collects multiple monitoring indicators of the physical field, such as groundwater level height, rainfall, temperature sensor data, local reservoir drainage volume, local river water level height, etc., covering comprehensive data, and can cope with the high diversity and complexity of groundwater level-related data, improving the accuracy of model prediction.

[0057] (3) The present invention overcomes the influence that groundwater level data may be affected by missing values, noise or outliers, and adopts a suitable mean substitution method, improving the accuracy of model prediction.

[0058] (4) The present invention uses technologies such as CNN, Attention and LSTM in deep learning to analyze and extract features from multi-sensor monitoring time series data, and realizes multi-step prediction of groundwater level.

[0059] (5) By adopting the method of combining one-dimensional convolutional neural network and attention mechanism, the present invention improves the accuracy of feature extraction and realizes more accurate prediction of groundwater level in the future for multiple days.

[0060] (6) The present invention aims to address the complexity of groundwater level monitoring and prediction, and the prediction results can provide more reliable data support for decision-making.

[0061] It should be understood that the implementation of any embodiment of the present invention does not mean that multiple or all of the above beneficial effects need to be simultaneously possessed or achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.

[0063] The structures, proportions, sizes, etc. depicted in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical substantive significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0064] Figure 1 It is a schematic diagram of the overall process of the groundwater level multi-step prediction method based on the convolutional attention long short-term neural network of the present invention;

[0065] Figure 2 It is a schematic diagram of the cell structure of the long short-term memory network (LSTM) of the present invention;

[0066] Figure 3 It is a schematic diagram of the convolutional attention long short-term neural network structure of the present invention;

[0067] Figure 4 It is a comparison diagram of prediction results, where (a) is the comparison between the prediction results of the ordinary double-layer LSTM multi-step prediction method and the true values, and (b) is the comparison between the prediction results of the convolutional attention long short-term neural network prediction method and the true values.

[0068] In each drawing, the same or corresponding reference numerals represent the same or corresponding parts. Specific Embodiments

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following will further elaborate on the embodiments of the present invention in combination with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0070] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "connection", "linkage", "fixation" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral one; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0071] It should be understood that the terms "including / comprising", "consisting of" or any other variants are intended to cover non-exclusive inclusion, so that a product, device, process or method including a series of elements not only includes those elements, but may also include other elements not explicitly listed when necessary, or further includes elements inherent to such product, device, process or method. Without further limitations, the elements defined by the statement "including / comprising..." or "consisting of..." do not exclude the existence of additional identical elements in the product, device, process or method including the said elements.

[0072] It should also be understood that terms such as "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating orientation or positional relationship are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device, component or structure referred to must have a specific orientation, be constructed or operated in a specific orientation, and should not be construed as a limitation to the present invention.

[0073] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality" is two or more unless otherwise clearly and specifically defined.

[0074] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the drawings and specific embodiments.

[0075] The present invention provides a multi-step prediction method for groundwater level based on a convolutional attention long short-term neural network. See Figure 1The flowchart shown includes the following steps: S10, obtaining monitoring data related to the groundwater level to obtain multi-feature monitoring time-series data; S20, constructing a multi-feature monitoring data sample tensor based on the multi-feature monitoring time-series data; S30, establishing a convolutional attention long short-term neural network model; S40, training the convolutional attention long short-term neural network model based on the multi-feature monitoring data sample tensor to obtain an optimal model; S50, substituting the monitoring data for prediction into the optimal model, and calculating through the model to output multi-step prediction data. The idea of the present invention is a multi-step prediction method for groundwater level based on deep learning technology. The convolutional neural network and attention mechanism in deep learning are used to analyze and extract features from multi-sensor time-series data, and then the extracted data is input into the long short-term neural network for model training, ultimately realizing multi-step prediction of the groundwater level, providing more reliable data support for decision-making related to groundwater disasters.

[0076] It should be noted that the so-called multi-step prediction means that by substituting the monitoring data into the optimal model, it is possible to predict and output data for multiple days. For example, given 30 days of monitoring data, the usual method uses the 30-day data to predict the data for the 31st day, and can only predict the groundwater level data for 1 day. However, the multi-step prediction method can use the 30-day data to predict the data from the 31st day to the 40th day, which is equivalent to predicting the groundwater level data for 10 days. Of course, through the method provided by the present invention, it is also possible to predict data for more days when the complexity of the monitoring data is low and the accuracy of the final prediction result is within the required range.

[0077] In step S10, monitoring data related to the groundwater level is obtained to obtain multi-feature monitoring time-series data.

[0078] The monitoring data related to the groundwater level referred to here is multi-sensor data collected according to time series, which means data containing multiple features, specifically including but not limited to the following features: groundwater level height, rainfall, temperature sensor data, local reservoir drainage volume, local river water level height, etc., which can be collected according to the time series. These data are recorded in real time by daily monitoring instruments or sensors, including two dimensions of time and specific values, such as the number of millimeters of rainfall corresponding to a specific date. Of course, due to the real-time recording of the monitoring instruments or sensors, the time dimension can be divided more finely according to the needs of the model, such as every hour, every minute, etc.

[0079] After the data collection is completed, multi-feature monitoring time-series data is obtained through the monitoring data related to the groundwater level and organized in the form of a database for later use.

[0080] In step S20, a multi-feature monitoring data sample tensor is constructed based on the multi-feature monitoring time-series data.

[0081] In this step, data preprocessing and data segmentation are mainly performed on multi-feature monitoring time series data, which specifically includes the following sub-steps:

[0082] In step S201, the multi-feature monitoring time series data is preprocessed. Due to the complexity of the data collection process, the multi-feature monitoring time series data collected often have missing values, erroneous values, and extreme values. The entry of erroneous data into the model will affect the accuracy of the model, so the erroneous data needs to be preprocessed before modeling. In some embodiments, the mean fitting method can be used to replace or delete the erroneous data. It is easy to understand that the mean fitting method replaces the erroneous data with the mean of the relevant data according to the actual situation.

[0083] In step S202, multi-feature monitoring time series data, such as groundwater level, rainfall, temperature sensor data, local reservoir discharge, local river water level, etc., are combined into a multi-feature monitoring time series data matrix X * (i.e., sample data matrix), where one column of data represents a monitoring data feature i, represented by X * i express, Represents the multi-feature monitoring time series data matrix X * The data value of the kth time series data in the i-th column, where k is the time series number of each column of data.

[0084] In step S203, the multi-feature monitoring time series data matrix X * Each column of data X * i Perform normalization transformation to obtain the normalized sample matrix X. Due to the diversity of monitoring time series data features, the features of the data set have different value ranges. Therefore, before using the data, the data needs to be normalized and converted into dimensionless scalars so that different features have the same measurement scale. At the same time, it can also eliminate the adverse effects caused by erroneous data, which is convenient for subsequent model training and prediction.

[0085] In some embodiments, the multi-feature monitoring time series data matrix X is calculated using the following formula: * Each column of data X * i Perform normalization:

[0086]

[0087] In the formula, k is the serial number of each column of data, that is Represents the multi-feature monitoring time series data matrix X * The data value of the kth time series data in the i-th column; It is time series data The normalized function value is also the kth time series data of the ith column of the normalized sample matrix. The value range of the ith column data is [ ],Right now and are the i-th column data X* in the multi-feature monitoring time series data matrix X* i The minimum and maximum values of each column of data in the normalized sample matrix X i Represents a type of monitoring data i.

[0088] As a step in data preprocessing, normalization transforms the interval of multi-feature monitoring time series data into [0,1], so as to better adapt to the training of deep learning models and accelerate convergence. When the evaluation criteria of sample data are different, it is necessary to dimensionalize them. Normalization can eliminate the impact of the dimension on the evaluation results and make different indicators comparable.

[0089] After the above normalization process, each feature data sequence is converted into the same measurement scale, and all data sequences are mapped to [0,1] to facilitate subsequent model calculations.

[0090] In step S204, the monitoring time data is processed and the date is converted into a timestamp.

[0091] In some embodiments, the conversion is performed using the following formula:

[0092]

[0093] In the formula, t represents the date, Xt represents the converted timestamp, and the obtained Xt is incorporated into the normalized sample matrix X as a column of the sample matrix X, which is convenient for calling data in the modeling process.

[0094] Taking January 10 as an example, when t=10, Xt is 0.59. Through timestamp conversion, the date data is also mapped to [0,1] to facilitate the subsequent model calculation.

[0095] It should be noted that the monitoring time data here refers to the date data column. For example, the continuous date data column from January 1 to January 31 has a total of 31 rows, each row is one day, and is arranged in date order.

[0096] By using the formula to map the date to a range between 0 and 1, as a periodic feature, it is helpful to capture the seasonal trend of groundwater level. Furthermore, the sine function can make the converted result smoother.

[0097] In step S205, the normalized sample matrix X is divided into a training set, a validation set, and a test set according to the number of rows. Among them, 70% of the data is used as the training set data, 10% of the data is used as the validation set data, and 20% of the data is used as the test set data. In machine learning, the training set, the validation set, and the test set are three important parts of the data set, which are used to train, evaluate, and test the performance of the machine learning model.

[0098] The training set is the data set used by the machine learning model for training and learning, and is used to train the parameters of the model. The validation set is the data set used to evaluate the performance of the model, and is used to adjust the parameters of the model during the training process to improve the performance of the model and avoid overfitting or underfitting of the model. The test set is the data set used to evaluate the final performance of the model, which does not overlap with the training set and the validation set, and determines whether the model is accurate. Generally speaking, the proportion of the training set is relatively large, usually accounting for 60%-80% of the total data set, while the proportion of the validation set or the test set is relatively small, usually accounting for 10%-20% of the total data set.

[0099] In step S206, the training set, the validation set, and the test set matrices are respectively split into three three-dimensional tensors x_train, x_val, and x_test as the input tensors of the neural network model. The dimensions are: batch size, time step, and number of feature channels. These three dimensions are the input dimensions of models such as CNN and LSTM. Splitting the data into these three dimensions facilitates input into the neural network for training. In some embodiments, the batch size is set according to the video memory size of the computer used to train the neural network model, and is generally 32. The time step can be custom-set to the optimal value, such as using a time step of 10. The number of feature channels is the same as the number of features in the normalized sample matrix X of the multi-feature monitoring time series data. For example, if there are five data features, the number of feature channels is 5. The training data unit x_train after data processing is in the form of Tensor[32,10,5].

[0100] In step S30, a convolutional attention long short-term neural network model is established. See Figure 3 , the neural network model includes: a convolutional layer, a multi-head self-attention mechanism layer, a long short-term recurrent neural network layer, and a fully connected layer.

[0101] Specifically, the convolutional layer is composed of a one-dimensional convolutional layer and a one-dimensional max pooling layer.

[0102] The one-dimensional convolutional layer is represented by the following formula:

[0103]

[0104] In the formula, ReLU represents the rectified linear unit activation function, conv1d represents the one-dimensional convolution operation, W is the convolution kernel, b is the bias of the convolutional layer, and x cov is the three-dimensional tensor constructed in step S206.

[0105] The one-dimensional max pooling layer is represented by the following formula:

[0106] MaxPool1d ( x pool ) = max(Conv( x pool )[i:i + kernel_size ])

[0107] Where MaxPool1d represents the one-dimensional max pooling operation, i represents the starting position of the pooling kernel on the input sequence, kernel_size represents the size of the pooling kernel, and x pool represents the output tensor of the one-dimensional convolutional layer (the previous neural network layer).

[0108] At each position i of the input sequence through the one-dimensional max pooling layer, the maximum value is selected from the subsequence from position i to i + kernel_size as the output. This can effectively reduce the length of the input sequence while retaining the most important features.

[0109] The convolutional layer realizes the following key functions through one-dimensional convolutional operations and pooling operations: capturing local features from the input data, which helps to identify patterns, structures, and associated information in the data; reducing the data scale, reducing network parameters and computational complexity, thereby alleviating the overfitting problem; achieving translational invariance through the weight sharing mechanism, that is, detecting the same features at different positions, which improves the robustness of the network to changes in the position of the input data.

[0110] Specifically, the multi-head self-attention mechanism layer consists of a self-attention score layer and a multi-head attention layer. The multi-head attention layer is represented by the following formula:

[0111] Specifically, the multi-head attention layer is represented by the following formula:

[0112]

[0113] Where Q represents the query, which is the representation used to calculate the attention score, and the query represents what you are interested in. K represents the key, which is the representation used to compare with the query, and the key represents the information in the input. V represents the value, which is the representation multiplied by the attention weights to produce the final output, and the value represents what you hope to output. Q, K, and V are obtained through the following linear transformation of the self-attention score layer. i represents the i-th attention head in the multi-head self-attention mechanism:

[0114]

[0115]

[0116]

[0117] In the formula, X is the input tensor of the self-attention score layer, and W Qi , W Ki , W Vi are trainable parameter matrices.

[0118] is the normalization factor, which is used to scale to ensure that they are within a suitable range, which is helpful for the training and stability of the model.

[0119] The attention mechanism is usually applied to natural language processing and machine translation. At the same time, the effect of the attention mechanism in time series data prediction is also very prominent. Its core functions include optimizing the model's attention to the input features at different time steps, realizing dynamic weight allocation, enabling the model to focus more on the information related to the current task, which is beneficial to solving the long-term dependence relationship in time series data and overcoming the gradient problems that traditional RNNs may face. In addition, the attention mechanism provides interpretability for the model's decisions, enabling the network to clearly understand which information is crucial for specific predictions. The attention mechanism also helps to fuse multi-modal data, thereby effectively integrating different data sources according to the task requirements.

[0120] Specifically, the long short-term memory neural network layer (LSTM) is a variant of the recurrent neural network (RNN) commonly used to process sequence data, which can effectively solve the long sequence problem and is used to solve problems such as gradient vanishing and gradient explosion. The cell structure of the LSTM neuron is shown in Figure 2 .

[0121] Figure 2 In it, Xt represents the input at the current time step (t), ht represents the hidden state at the current time step (t), σ represents the Sigmoid function, tanh is the tangent function, whose role is to map the input value to the range between -1 and 1, and X represents the element-wise multiplication operation.

[0122] In LSTM, there are three gates: the forget gate, the input gate, and the output gate. The gating mechanism in the LSTM structure can better control the flow of information and effectively avoid the interference of irrelevant information and the problem of gradient vanishing. In addition, LSTM also has a cell state and a hidden state. The working process of LSTM is as follows:

[0123] (1) At each time step, calculate the values of the forget gate, the input gate, and the output gate according to the hidden state of the previous time step and the input of the current time step;

[0124] (2) Use the forget gate to determine the information to be forgotten from the cell state;

[0125] (3) Update the cell state using the input gate and the new candidate value, and add new information to the cell state;

[0126] (4) Use the output gate to determine the information extracted from the cell state and generate the hidden state at the current time step;

[0127] (5) The cell state and the hidden state are passed to the next layer of LSTM at the next time step or used to generate the final output;

[0128] (6) Connect the LSTM output data to the fully connected layer to achieve feature combination and non-linear transformation.

[0129] The fully connected layer is a basic layer in a neural network, also known as the dense layer or the multi-layer perceptron layer. Its main function is to connect all neurons in the previous layer of the neural network to each neuron in the current layer, thereby achieving feature combination and non-linear transformation.

[0130] The output result of the final neural network model is a three-dimensional tensor, with dimensions: batch size, time step, number of prediction days, and the dimension is [32, 10, 30].

[0131] In step S40, based on the multi-feature monitoring data sample tensor, train the convolutional attention long short-term neural network model to obtain the optimal model. In this step, the established model needs to be trained, using the multi-feature monitoring data sample tensor for model training, monitoring the model performance metrics, loss function, and accuracy. Use the validation set and the early stopping method to evaluate the model performance to avoid overfitting.

[0132] When setting hyperparameters, the following hyperparameters need to be concerned: time window size, batch size, number of features, number of LSTM hidden layers, number of prediction days, learning rate, and the patience value in the early stopping method. Among them, the specific size of the time window should be selected according to the problem requirements and needs to be adjusted and evaluated; the patience value is the number of epochs used to monitor the continuous non-decrease of the validation set loss function value during model training. When the number of epochs of continuous non-decrease of the validation set loss exceeds the patience value, the training process will be terminated. Setting the patience value helps to stop training when the model reaches the best performance, to avoid overfitting and save computing resources.

[0133] Step S40 specifically includes the following sub-steps:

[0134] In step S401, the specific size of the time window should be selected according to the problem requirements, that is, select the time window size according to the requirements and adjust and evaluate it;

[0135] In step S402, after preprocessing the multi-feature monitoring time-series data, the Attention-LSTM model is used for training. Using the training set data, the model is trained using the backpropagation algorithm. The backpropagation algorithm, abbreviated as the BP algorithm, is a learning algorithm suitable for multi-layer neural networks. It is based on the gradient descent method, and its information processing ability comes from the multiple compositions of simple non-linear functions. Therefore, it has a strong function reproduction ability.

[0136] The model parameters are continuously adjusted to reduce the loss function. The loss function is used to measure the deviation between the predicted value and the true value of the model. Generally, the smaller it is, the better. In this step, the model parameters need to be adjusted to continuously reduce the loss function.

[0137] In step S403, the Adam optimizer and an appropriate learning rate are used to optimize the model. The Adam optimizer is an adaptive optimization algorithm that can adjust the learning rate according to historical gradient information. In actual use, the learning rate can be adjusted according to the data scale and data complexity, and the optimal range is generally in [0.0001, 0.01].

[0138] In step S404, the model performance metrics, loss function, and accuracy are monitored, and the validation set is used to evaluate the model performance to avoid overfitting. The model performance is measured by the mean squared error SME and the coefficient of determination R 2 value. The model with the optimal values of SME and the coefficient of determination R 2 is obtained through multiple trainings. The loss function is an operation function used to measure the difference between the predicted value and the true value of the model. It is a non-negative real-valued function. The smaller the loss function, the better the robustness of the model. In the process of parameter tuning, the model is solved and evaluated by minimizing the loss function. The accuracy of the model refers to the number of all correctly predicted samples in the model / the total number of observed samples for a given test set. In the process of parameter tuning, the accuracy of the model should be continuously improved as much as possible.

[0139] In step S405, the cosine annealing strategy is used to dynamically adjust the learning rate. The cosine annealing strategy adjusts the learning rate to the value of the cosine function between the minimum and maximum values within a specified range. The learning rate gradually decreases from the maximum value to the minimum value during training and then gradually increases. By using the cosine annealing strategy to smoothly and dynamically adjust the learning rate, local optima can be effectively avoided, the learning rate can be adaptively adjusted, and the model training process can be simplified.

[0140] In step S406, when the loss function value on the validation set reaches the lowest point, the model parameters are saved to obtain the optimal model.

[0141] By executing the above steps, the optimal prediction model is obtained and used to predict the groundwater level.

[0142] In step S50, the monitoring data for prediction is substituted into the optimal model, and multi-step prediction data is output through model calculation.

[0143] When predicting the groundwater level, the input data has the same structure as the input data of the training model. A three-dimensional tensor is input, and the tensor dimensions are batch size, time step, and number of feature channels. The predicted output result is a three-dimensional tensor, and the tensor dimensions are batch size, time step, and number of predicted days.

[0144] The present invention is a multi-step prediction of groundwater level based on an improved convolutional attention long short-term neural network model (LSTM model). Technologies such as CNN, Attention, and LSTM in deep learning are used to analyze and extract features from multi-sensor time series data to achieve multi-step prediction of groundwater level. By adopting a method of combining a one-dimensional convolutional neural network with an attention mechanism, the accuracy of feature extraction is improved, and more accurate prediction of the groundwater level in the future for multiple days is realized. This technological innovation aims to address the complexity of groundwater level monitoring and prediction and provide more reliable data support for decision-making.

[0145] See Figure 4 , in this embodiment, the groundwater level data of Petrignano, Italy is used as an example. The characteristics of the example data include the groundwater level height, rainfall, temperature sensor data, local reservoir drainage volume, and local river water level height in Petrignano, Italy from 2009 to 2016. The prediction time is the groundwater level in the next 30 days.

[0146] Figure 4 In (a), the comparison between the multi-step prediction result of the ordinary double-layer LSTM and the true value is shown. The mean square error value SME is 0.367, and the coefficient of determination R 2 is 0.710; in (b), the comparison between the prediction result of the convolutional attention long short-term neural network and the true value is shown. The mean square error SME is 0.289, and the coefficient of determination R 2 is 0.822. It can be seen that through the multi-step prediction method of groundwater level based on the convolutional attention long short-term neural network of the present invention, the error of the prediction model is reduced, the accuracy of model prediction is improved, and accurate prediction of the groundwater level in the future for a long time is realized.

[0147] Although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present invention. Certain features described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.

[0148] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A multi-step prediction method for groundwater level based on convolutional attention long short-term neural network, characterized in that It includes the following steps: S10. Obtain the monitoring data related to the groundwater level to get the multi-feature monitoring time-series data; S20. Based on the multi-feature monitoring time-series data, construct a multi-feature monitoring data sample tensor. Among them, a three-dimensional tensor is input, and the tensor dimensions are: batch size, time step, and number of feature channels; S30. Establish a convolutional attention long short-term neural network model; the convolutional attention long short-term neural network model includes: a convolutional layer, a multi-head self-attention mechanism layer, a long short-term recurrent neural network layer, and a fully connected layer; S40. Based on the multi-feature monitoring data sample tensor, train the convolutional attention long short-term neural network model to obtain the optimal model; S50. Substitute the monitoring data for prediction into the optimal model, and calculate and output multi-step prediction data through the model; among them, the output result is a three-dimensional tensor, and the tensor dimensions are: batch size, time step, and number of prediction days.

2. The prediction method according to claim 1, characterized in that, In step S10, the monitoring data is multi-sensor data collected in time series, including: groundwater level height, rainfall, temperature sensor data, local reservoir drainage volume, and local river water level height.

3. The prediction method according to claim 1, characterized in that, In step S20, the construction of the multi-feature monitoring data sample tensor includes: S201. Preprocess the multi-feature monitoring time-series data, fill in the missing data, and delete the abnormal data; S202, compose the multi-feature monitoring time-series data into a multi-feature monitoring time-series data matrix X * , where each row represents a time-series sample and each column represents a different feature; S203, the multi-feature monitoring time series data matrix X * Each column of data X * i Perform normalization transformation to obtain the normalized sample matrix X; S204. Process the monitoring time data, convert the date to a timestamp, and incorporate the timestamp into the normalized sample matrix X as a column of the matrix; S205. Divide the normalized sample matrix X into a training set, a validation set, and a test set according to the number of rows; S206. Divide the training set, validation set, and test set matrices into three three-dimensional tensors x_train, x_val, and x_test respectively.

4. The prediction method according to claim 3, wherein In step S203, each column of data \(X\) in the multi-feature monitoring time series data matrix \(X\) is normalized using the following formula: * of the matrix \(X\) * i is normalized as follows: ; where k is the serial number of each column of data, that is represents the multi-feature monitoring time series data matrix X * the data value of the k-th time series data in the i-th column of is the time series data the normalized function value, and it is also the k-th time series data in the i-th column that constitutes the normalized sample matrix X. The value range of the i-th column of data is , that is and are respectively the * minimum and maximum values of the data X in the i-th column of the multi-feature monitoring time series data matrix X * i . Each column of data X in the normalized sample matrix X i represents a kind of monitoring data i.

5. The prediction method according to claim 3, characterized in that, In step S204, the date is converted into a timestamp X using the following formula t , and the timestamp X t is incorporated into the normalized sample matrix X: ; where t represents the date, and X t represents the converted timestamp.

6. The prediction method according to claim 1, wherein The convolutional layer includes a one-dimensional convolutional layer and a one-dimensional max pooling layer, where: The one-dimensional convolutional layer is represented by the following formula: ; Wherein, ReLU represents the rectified linear unit activation function, conv1d represents the one-dimensional convolution operation, W is the convolution kernel, b is the bias of the convolutional layer, and x cov represents a three-dimensional tensor; The one-dimensional max pooling layer is represented by the following formula: ; Wherein, MaxPool1d represents a one-dimensional maximum pooling operation, i represents the starting position of the pooling kernel on the input sequence, kernel_size represents the pooling kernel size, and x pool represents the output tensor of the one-dimensional convolutional layer.

7. The prediction method according to claim 1, characterized in that The multi-head self-attention mechanism layer is composed of a self-attention score layer and a multi-head attention layer. The multi-head attention layer is represented by the following formula: ; In the formula, Q represents the query, which is used to calculate the representation of the attention score, K represents the key, which is used to compare with the query, V represents the value, which is multiplied by the attention weight to generate the final output representation. Q, K, and V are obtained through the following linear transformation of the self-attention score layer. i represents the i-th attention head in the multi-head self-attention mechanism: ; ; ; where X is the input tensor of the self-attention score layer, W Qi , W Ki , W Vi are trainable parameter matrices.

8. The prediction method according to claim 1, wherein The long short-term recurrent neural network layer includes: a forget gate, an input gate, and an output gate, as well as a cell state and a hidden state; the steps for the long short-term recurrent neural network layer to process data include the following: (1). At each time step, calculate the values of the forget gate, input gate, and output gate according to the hidden state of the previous time step and the input of the current time step; (2). Use the forget gate to determine the information forgotten from the cell state; (3). Use the input gate and the new candidate value to update the cell state and add new information to the cell state; (4). Use the output gate to determine the information extracted from the cell state and generate the hidden state of the current time step; (5), the cell state and the hidden state are passed to the next long short-term recurrent neural network layer at the next time step or used to generate the final output; (6), the output data of the long short-term recurrent neural network layer is connected to a fully connected layer to achieve feature combination and non-linear transformation, and a three-dimensional tensor is output, with dimensions: batch size, time step, and number of prediction days.

9. The prediction method according to claim 1, wherein In step S40, the training of the convolutional attention long short-term neural network model includes the following steps: S401, select the time window size according to requirements, and adjust and evaluate it; S402, use the convolutional attention long short-term neural network model for training, use the training set data and apply the backpropagation algorithm to adjust the model parameters to reduce the loss function; S403, optimize the model using the Adam optimizer and an appropriate learning rate; S404, monitor the model performance metrics, loss function, and accuracy, and use the validation set to evaluate the model performance to avoid overfitting; S405, dynamically adjust the learning rate using the cosine annealing strategy; S406, when the loss value on the validation set reaches the lowest point, save the model parameters to obtain the optimal model.

10. The prediction method according to claim 9, characterized in that The model performance metrics are measured by the mean squared error SME and the coefficient of determination R 2 value, and the model with the optimal values of the mean squared error SME and the coefficient of determination R 2 is taken after multiple trainings.

Citation Information

Patent Citations

  • Underground water level prediction method based on time sequence convolution feature filtering neural network model

    CN118966567A