Multivariable underground water level time sequence prediction method, device, equipment and medium
By combining the random forest regression model and the CNN-LSTM network, the feature selection and model structure are optimized, and the problem of insufficient accuracy and stability of groundwater level prediction in the prior art is solved, and more efficient multivariate groundwater level time series prediction is achieved.
Patent Information
- Application Number
- CN202510451125.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to effectively deal with complex nonlinear relationships and dependencies in long-time series data, resulting in insufficient accuracy and stability of groundwater level prediction, and poor generalization ability and training speed.
A multivariate groundwater level time series prediction method is adopted, combining the random forest regression model, convolutional neural network (CNN) and long and short-term memory network (LSTM), and optimize the model structure and parameters through the calculation of feature variable contribution ratio, feature screening and normalization processing, and optimize the model performance using evaluation indicators.
The applicability, generalization ability and training speed of the model are improved, the convergence of the model is optimized, the accuracy and stability of multivariate groundwater level time series prediction is enhanced, and the sustainable utilization efficiency of groundwater resources is improved.
Smart Images

Figure CN120234776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a multi-variable groundwater level time series prediction method, device, equipment and medium. Background Art
[0002] The accurate prediction of groundwater level is of great significance for water resource management, ecological environment protection and disaster warning, etc. Traditional groundwater level prediction methods usually rely on statistical models or physical models, but these methods cannot effectively handle complex non-linear relationships and dependencies in long time series data. Therefore, in recent years, prediction methods based on deep learning have gradually attracted attention, especially the advantages of CNN (Convolutional Neural Networks) and LSTM (Long Short-Term Memory) in time series prediction. However, the computational complexity of LSTM is higher, and it requires more parameters and computational volume. The complexity of LSTM makes its internal operation mechanism less intuitive and difficult to explain the decision-making process of the network. At the same time, LSTM has more parameters that need to be trained, so more data is required to avoid overfitting. If the training data is insufficient, LSTM may face the problem of insufficient generalization ability.
[0003] As can be seen from the above, how to improve the applicability, generalization ability, training speed of the model, optimize the convergence of the model, improve the accuracy and stability of multi-variable groundwater level time series prediction, and increase the sustainable utilization efficiency of groundwater resources is an issue to be solved in this field. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a multi-variable groundwater level time series prediction method, device, equipment and medium, which can improve the applicability, generalization ability, training speed of the model, optimize the convergence of the model, improve the accuracy and stability of multi-variable groundwater level time series prediction, and increase the sustainable utilization efficiency of groundwater resources. The specific solutions are as follows:
[0005] In a first aspect, the present application discloses a multi-variable groundwater level time series prediction method, including:
[0006] Obtain groundwater data of multi-variable groundwater level to be predicted, and perform verification analysis and screening processing on the groundwater data according to the time series to obtain processed data;
[0007] Based on the random forest regression model and using the mean impurity reduction method, calculate the contribution ratio of feature variables, perform feature screening and normalization processing on the processed data to obtain target data;
[0008] Construct a prediction model for multi-variable groundwater level time series prediction based on convolutional neural network and long short-term memory network, and perform structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain the target prediction model;
[0009] Input the target data into the target prediction model to obtain a predicted value, generate an evaluation index based on the predicted value and the actual measured value of the multi-variable groundwater level to be predicted, and use the evaluation index to optimize the target prediction model.
[0010] Optionally, the method for obtaining groundwater data of the multi-variable groundwater level to be predicted, performing verification analysis and screening processing on the groundwater data according to the time series, and obtaining the processed data includes:
[0011] Screen the multi-variable groundwater level to be predicted that meets the preset conditions, and collect the corresponding groundwater data;
[0012] Use Kriging interpolation method and spatial distribution analysis method, combine with the statistical analysis tool in professional geographic information software, and perform cross-validation, analysis, missing value elimination, and screening on the groundwater data according to the time series to obtain the processed data.
[0013] Optionally, the method for calculating the contribution ratio of feature variables, performing feature screening, and normalizing the processed data based on the random forest regression model and using the mean impurity reduction method to obtain the target data includes:
[0014] Construct a random forest regression model using multiple decision trees;
[0015] Based on the random forest regression model and using the mean impurity reduction method, calculate the contribution ratio of feature variables for the processed data, and screen out the feature variables corresponding to the largest contribution ratio of feature variables;
[0016] Use the Min-Max normalization method to normalize the feature variables to obtain the target data.
[0017] Optionally, before constructing the prediction model for multi-variable groundwater level time series prediction based on convolutional neural network and long short-term memory network, it further includes:
[0018] Construct a convolutional neural network based on an image input layer, a convolutional layer, a batch normalization layer, a ReLU activation layer, a max pooling layer, a fully connected layer, and a regression layer;
[0019] Construct a long short-term memory network based on long short-term memory network units, a ReLU activation layer, a fully connected layer, a sequence input layer, and an output layer.
[0020] Optionally, the structural setting, cross-validation, parameter setting, and model training of the prediction model to obtain a target prediction model include:
[0021] Set the convolutional neural network in the prediction model to extract data features, and set the long short-term memory network to receive sequential data to complete the structural setting;
[0022] Use the time series segmentation method to perform cross-validation on the prediction model after structural setting and set the target hyperparameters;
[0023] Perform parameter setting and model training on the prediction model after setting the target hyperparameters to obtain a target prediction model; the parameter setting includes learning rate setting and batch setting.
[0024] Optionally, the generation of evaluation indicators based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted includes:
[0025] Calculate the root mean square error, mean absolute error, and coefficient of determination between the predicted value and the actual measured value of the multivariate groundwater level to be predicted;
[0026] Generate evaluation indicators based on the root mean square error, the mean absolute error, and the coefficient of determination.
[0027] Optionally, the optimization of the target prediction model using the evaluation indicators includes:
[0028] Use the Akaike information criterion and the Bayesian information criterion to determine the lag order;
[0029] Adjust and optimize the parameters of the convolutional neural network and the long short-term memory network in the target prediction model based on the evaluation indicators and the lag order.
[0030] In a second aspect, the present application discloses a multivariate groundwater level time series prediction device, including:
[0031] A verification analysis and screening processing module, configured to obtain groundwater data of the multivariate groundwater level to be predicted, process the groundwater data according to a time series to obtain processed data;
[0032] A calculation screening and normalization processing module, configured to calculate the contribution ratio of feature variables, perform feature screening, and perform normalization processing on the processed data based on a random forest regression model using the mean impurity reduction method to obtain target data;
[0033] A model construction and training module, which is used to construct a prediction model for multivariate groundwater level time series prediction based on a convolutional neural network and a long short-term memory network, perform structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain a target prediction model;
[0034] A prediction module, which is used to input the target data into the target prediction model to obtain a predicted value, generate an evaluation index based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and optimize the target prediction model by using the evaluation index.
[0035] In a third aspect, the present application discloses an electronic device, including:
[0036] A memory, which is used to store a computer program;
[0037] A processor, which is used to execute the computer program to implement the foregoing multivariate groundwater level time series prediction method.
[0038] In a fourth aspect, the present application discloses a computer storage medium, which is used to store a computer program; wherein, when the computer program is executed by a processor, the steps of the foregoing disclosed multivariate groundwater level time series prediction method are implemented.
[0039] It can be seen that the present application provides a multi-variable groundwater level time series prediction method, which includes obtaining groundwater data of the multi-variable groundwater level to be predicted, performing verification analysis and screening processing on the groundwater data according to the time series to obtain processed data; calculating the contribution ratio of feature variables, feature screening, and normalization processing on the processed data based on the random forest regression model and using the average impurity reduction method to obtain target data; constructing a prediction model for multi-variable groundwater level time series prediction based on the convolutional neural network and the long short-term memory network, performing structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain a target prediction model; inputting the target data into the target prediction model to obtain a predicted value, generating an evaluation index based on the predicted value and the actual measured value of the multi-variable groundwater level to be predicted, and optimizing the target prediction model using the evaluation index. The present application obtains the groundwater data of the multi-variable groundwater level to be predicted, performs verification analysis and screening processing on the groundwater data according to the time series, uses the random forest regression model and the average impurity reduction method to calculate the contribution ratio of feature variables, feature screening, and normalization processing on the processed data to obtain target data, improves the prediction accuracy of the model, can retain variables with higher contribution degrees as input features through the random forest regression model, improves the effectiveness of model training, constructs a prediction model for multi-variable groundwater level time series prediction based on the convolutional neural network and the long short-term memory network, enhances the feature extraction ability, performs structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain a target prediction model, reduces the scale influence between different variables, stabilizes the training process and accelerates model convergence, inputs the target data into the target prediction model to obtain a predicted value, generates an evaluation index based on the predicted value and the actual measured value, improves the applicability and generalization ability of the model, enables it to adapt to the groundwater level prediction requirements of different regions, optimizes the target prediction model using the evaluation index, effectively improves the training speed, optimizes the convergence of the model, reduces the computational complexity, and at the same time ensures the applicability and generalization ability of the model on large-scale data sets, improves the accuracy and stability of multi-variable groundwater level time series prediction, and increases the sustainable utilization efficiency of groundwater resources. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0041] Figure 1 Flowchart of a multi-variable groundwater level time series prediction method disclosed in the present application;
[0042] Figure 2 Another flowchart of the multivariate groundwater level time series prediction method disclosed in this application;
[0043] Figure 3 A system structure diagram of the multivariate groundwater level time series prediction disclosed in this application;
[0044] Figure 4 A characteristic contribution ratio diagram disclosed in this application;
[0045] Figure 5 (a) An interpolation result diagram for March disclosed in this application;
[0046] Figure 5 (b) An interpolation result diagram for April disclosed in this application;
[0047] Figure 5 (c) An interpolation result diagram for May disclosed in this application;
[0048] Figure 6 A standardized error diagram for different sampling points disclosed in this application;
[0049] Figure 7 A schematic diagram of the long short-term neural memory network principle disclosed in this application;
[0050] Figure 8 A CNN-LSTM neural network structure diagram disclosed in this application;
[0051] Figure 9 A loss diagram corresponding to the training with different parameters disclosed in this application;
[0052] Figure 10 An effect diagram of deep training disclosed in this application;
[0053] Figure 11 An analysis diagram of the lag order optimization disclosed in this application;
[0054] Figure 12 A CNN-LSTM optimization program diagram disclosed in this application;
[0055] Figure 13 An effect diagram of the optimized training disclosed in this application;
[0056] Figure 14 A summary flowchart of the multivariate groundwater level time series prediction disclosed in this application;
[0057] Figure 15 A schematic diagram of the structure of the multivariate groundwater level time series prediction device disclosed in this application;
[0058] Figure 16 A structural diagram of an electronic device provided for this application. Specific implementation manner
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] The accurate prediction of the groundwater level is of great significance for water resource management, ecological environment protection, disaster warning, etc. Traditional groundwater level prediction methods usually rely on statistical models or physical models, but these methods cannot effectively handle complex non-linear relationships and dependencies in long time series data. Therefore, in recent years, prediction methods based on deep learning have gradually received attention, especially the advantages of CNN and LSTM in time series prediction. However, LSTM has a higher computational complexity, requires more parameters and computational volume. The complexity of LSTM makes its internal operation mechanism less intuitive and difficult to explain the decision-making process of the network. At the same time, LSTM has more parameters to be trained, so more data is needed to avoid overfitting. If the training data is insufficient, LSTM may face the problem of insufficient generalization ability. As can be seen from the above, how to improve the applicability, generalization ability, training speed of the model, optimize the convergence of the model, improve the accuracy and stability of multi-variable groundwater level time series prediction, and increase the sustainable utilization efficiency of groundwater resources are problems to be solved in this field.
[0061] See Figure 1 As shown, the embodiments of the present invention disclose a multi-variable groundwater level time series prediction method, which may specifically include:
[0062] Step S11: Obtain groundwater data of multi-variable groundwater levels to be predicted, and perform verification analysis and screening processing on the groundwater data according to the time series to obtain processed data.
[0063] In this embodiment, multi-variable groundwater levels to be predicted that meet preset conditions are screened, and corresponding groundwater data is collected; using Kriging interpolation method and spatial distribution analysis method, combined with statistical analysis tools in professional geographic information software, and performing cross-validation, analysis, missing value elimination, and screening on the groundwater data according to the time series to obtain processed data.
[0064] Based on Kriging interpolation and spatial distribution characteristics, this application selects points that can better reflect the changes in regional groundwater levels. The Kriging interpolation method can visually represent the spatial distribution pattern of groundwater levels. Combining with the Geostatistical Analyst tool (statistical analysis tool) of ArcGIS (professional geographic information software), missing data in the groundwater data of the multivariate groundwater levels to be predicted is removed, and the missing data is interpolated with the mean value to reduce the prediction error. The residuals and prediction errors of the interpolation results are analyzed, and data points with smaller residuals and prediction errors are selected; when the residuals are not much different, those with smaller standardized errors are selected as representative experimental data to obtain the processed data.
[0065] Step S12: Based on the random forest regression model and using the mean decrease in impurity method, calculate the contribution ratio of feature variables, perform feature screening, and normalization processing on the processed data to obtain the target data.
[0066] In this embodiment, a random forest regression model is constructed using multiple decision trees; based on the random forest regression model and using the mean decrease in impurity method, calculate the contribution ratio of feature variables for the processed data, and screen out the feature variables corresponding to the largest contribution ratio of feature variables; use the Min - Max normalization method to perform normalization processing on the feature variables to obtain the target data.
[0067] This application trains an RF (Random Forest) regression model, selects the feature variables with the largest contribution ratio to groundwater level prediction, and then uses the Min - Max normalization method to normalize the feature variables to avoid the influence of differences between feature values on the model. The specific steps are as follows:
[0068] (1) A random forest consists of multiple Decision Trees. Each tree samples from the training data through Bootstrap Sampling, and randomly selects some features for splitting nodes when constructing the tree. The node splitting of a decision tree depends on a certain feature, making the data more pure. This change in "purity" is called Impurity Reduction, and in this paper, the method based on MDI (Mean Decrease in Impurity) is used to calculate feature importance.
[0069] (2) Impurity measurement: Decision trees often use MSE (Mean Squared Error) or Gini Index to measure the splitting quality. In regression tasks (such as groundwater level prediction), MSE is usually adopted:
[0070] ;
[0071] Among them, N is the number of samples, is the true value, is the predicted value; when a feature is used in a splitting node of the decision tree, this split will reduce the mean squared error MSE, that is:
[0072] ;
[0073] Among them, is the impurity of the parent node, and are the impurities of the left and right child nodes, and are the data proportions of the left and right child nodes;
[0074] (3) Feature importance calculation; in the random forest, each tree will be divided based on a certain feature resulting in a reduction in impurity. Calculate the average contribution of this feature in all trees to obtain its importance:
[0075] ;
[0076] Among them, T is the number of trees in the random forest, is the set of all nodes that use the feature for division, is the reduction in impurity when using for splitting at the s-th node of the t-th tree.
[0077] Finally, the importance of all features will be normalized:
[0078] ;
[0079] Among them, M is the total number of features;
[0080] (4) Set a threshold to screen out the corresponding feature variables with a high contribution ratio;
[0081] (5) Use the min-max normalization method to normalize the feature variables to obtain the target data, making all time series data take values between 0 and 1, avoiding the impact of differences between feature values on the model. Divide the data into a training set and a test set in chronological order, using the first 70% of the data as the training set and the last 30% of the data as the test set:
[0082] ;
[0083] Among them, is the normalized value, is the minimum value of the data set, is the maximum value of the data set.
[0084] Step S13: Construct a prediction model for multivariate groundwater level time series prediction based on a convolutional neural network and a long short-term memory network, perform structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain a target prediction model.
[0085] Step S14: Input the target data into the target prediction model to obtain a predicted value, generate an evaluation index based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and optimize the target prediction model using the evaluation index.
[0086] In this embodiment, the target data is input into the target prediction model to obtain a predicted value, the root mean square error, mean absolute error, and coefficient of determination between the predicted value and the actual measured value of the multivariate groundwater level to be predicted are calculated, an evaluation index is generated based on the root mean square error, the mean absolute error, and the coefficient of determination, the lag order is determined using the Akaike information criterion and the Bayesian information criterion, and the parameters of the convolutional neural network and the long short-term memory network in the target prediction model are adjusted and optimized based on the evaluation index and the lag order.
[0087] This application selects appropriate evaluation indexes according to the nature of the time series problem and the specific application scenario, and uses RMSE (Root Mean Square Error), MAE (Mean Absolute Error), and coefficient of determination as the indexes for evaluating the groundwater level prediction effect of the model. Their calculation formulas are as follows:
[0088] ;
[0089] where n is the number of samples, is the actual observed value of the i-th sample, is the model predicted value of the i-th sample. RMSE represents the average error between the predicted value and the actual observed value, and MAE represents the average absolute error between the predicted value and the actual observed value:
[0090] ;
[0091] where, is the sum of squared residuals, is the total sum of squares, is the coefficient of determination.
[0092] Based on the evaluation results of CNN-LSTM, calculate the lag order of the trained model. The lag order is selected based on AIC (Akaike information criterion) and BIC (Bayesian Information Criterion) to choose the appropriate value, and the lag order corresponding to the calculated minimum value is selected as the optimal parameter. The calculation formula is as follows:
[0093] ;
[0094] where, ln(L) is the log-likelihood function of the model, k is the number of parameters of the model, and n is the size of the sample.
[0095] Combined with the optimal step size and the number of neurons, the number of convolutional kernels in the CNN layer is adjusted, and the number of convolutional layers and pooling layers is adjusted according to the optimal lag order to further extract data features, enabling this model to better complete the groundwater level prediction task.
[0096] See Figure 2 As shown, an embodiment of the present invention discloses a multi-variable groundwater level time series prediction method, which may specifically include:
[0097] Step S21: Obtain groundwater data of the multi-variable groundwater level to be predicted, and perform verification analysis and screening processing on the groundwater data according to the time series to obtain processed data.
[0098] Step S22: Based on the random forest regression model and using the average impurity reduction method, calculate the contribution ratio of feature variables, perform feature screening, and normalization processing on the processed data to obtain target data.
[0099] Step S23: Construct a convolutional neural network based on an image input layer, convolutional layer, batch normalization layer, ReLU activation layer, max pooling layer, fully connected layer, and regression layer, and construct a long short-term memory network based on long short-term memory network units, ReLU activation layer, fully connected layer, sequence input layer, and output layer.
[0100] In this embodiment, the network structure settings of the convolutional neural network and the long short-term memory network include: for the CNN layer, specify the size of the input data through the image input layer, and then sequentially add a convolutional layer, batch normalization layer, ReLU activation layer, and max pooling layer to extract the features of the input data; for the LSTM layer, receive sequential data by setting a sequence input layer, which contains the number of features of the training data. The CNN layer extracts the features of the input data, and the LSTM layer receives sequential data.
[0101] In this embodiment, the CNN layer includes: in the convolutional layer, the convolutional kernel size is set to 5x1, and 8 convolutional kernels are generated; the batch normalization layer is used to accelerate the training process of the model and improve the generalization ability of the model, perform normalization processing on the input data, and help avoid the problems of gradient disappearance or explosion; the ReLU activation function is used to perform a non-linear transformation on the output of the convolutional layer, introduce non-linear factors, and enhance the expression ability of the network; the pooling window size of the max-pooling layer is 2x1, and the stride is 1, which is used to reduce the size of the feature map while retaining the most significant features; a fully connected layer and a regression layer are connected to the output end of the network to generate the output of the regression task.
[0102] The LSTM layer includes: a specified number of LSTM units for capturing long-term dependencies in sequential data; the ReLU activation layer introduces non-linear factors; the fully connected layer maps the output of the LSTM layer to the final output dimension; the regression layer is used to calculate the loss of the regression task; the dimensions of the output layer and the number of neurons in the hidden layer are additionally set.
[0103] Step S24: Construct a prediction model for multi-variable groundwater level time series prediction based on a convolutional neural network and a long short-term memory network. Set the convolutional neural network in the prediction model to extract data features, and set the long short-term memory network to receive sequential data to complete the structure setting. Use the time series segmentation method to perform cross-validation on the prediction model after the structure setting, and set the target hyperparameters. Perform parameter setting and model training on the prediction model after setting the target hyperparameters to obtain the target prediction model; the parameter setting includes learning rate setting and batch setting.
[0104] In this embodiment, the features of the input data are extracted through the CNN layer in the grid structure, specifically including:
[0105] 1. Convolutional layer: Extract local temporal features in the water level data through the Conv1D layer (one-dimensional convolutional layer). When there are multi-variable inputs, the convolutional kernel will act on all channels (feature channels) simultaneously. After passing through the multi-variable input features, the joint temporal patterns of variables such as water level, rainfall, temperature, and evaporation can be learned, rather than just the individual trend of the water level; this cross-variable feature extraction ability helps to improve the prediction accuracy; then the ReLU activation function is used for non-linear transformation; that is:
[0106] ;
[0107] where C is the number of input variables (number of channels, such as water level, rainfall, temperature, evaporation), is the convolutional kernel parameter of channel c, is the c-th feature at time step t;
[0108] 2. Pooling layer: Use MaxPooling1D (maximum pooling layer) to reduce the feature dimension while maintaining important temporal features. Let the pooling window size be p, then the maximum pooling output at time position t is:
[0109] ;
[0110] Then input the data features into the LSTM layer to perform sequence modeling and capture long-term dependencies on these features.
[0111] This application selects the best window size and number of hidden layers through cross-validation, and retrains the final model based on the optimal parameters to improve the generalization ability and prediction accuracy: First, divide the time series dataset into multiple training-validation subsets, and set a series of combinations of candidate window sizes and numbers of hidden layers. Adopt the TimeSeriesSplit cross-validation method to construct a CNN-LSTM model for each combination of window size and number of hidden layers, and train and evaluate it on different training-validation subsets. Use metrics such as RMSE and MAE to measure the model performance, and compare the cross-validation results of different hyperparameter combinations. Finally, select the window size and number of hidden layers that make the model prediction performance optimal, and retrain the final model based on the optimal parameters to improve the generalization ability and prediction accuracy.
[0112] This application can help improve the model's performance, generalization ability, and training efficiency by learning the features and patterns of the data, reasonably adjusting parameters, and configuring the deep learning model. The training parameter settings include: training in the GPU (Graphics Processing Unit) environment, selecting the Adam optimization algorithm for parameter optimization, setting the maximum number of training iterations to 4000 times, controlling the gradient threshold to 1 to prevent gradient explosion, setting the initial learning rate to 0.001, and adopting a segmented learning rate adjustment strategy to adjust the learning rate every 50 epochs and update the learning rate by multiplying it by 0.5. At the same time, add L2 regularization to prevent overfitting.
[0113] Step S25: Input the target data into the target prediction model to obtain a predicted value, generate an evaluation index based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and use the evaluation index to optimize the target prediction model.
[0114] The optimization process for the CNN layer in the target prediction model in this application is as follows: (1) Adjust the number of convolutional kernels: Verify the impact of different numbers of convolutional kernels on feature extraction ability through experiments, and select the optimal configuration to ensure sufficient extraction of temporal features and improve the model's representation ability; (2) Optimize the pooling layer stride: Conduct comparative experiments under different stride settings, and select the optimal stride to reduce the data dimension while retaining key features, enabling the LSTM layer to capture sequence patterns more effectively; (3) Set the optimal convolutional kernel size (Kernel Size = 3): Adjust the convolutional kernel size through multiple experiments, and finally select Kernel Size = 3 to extract the most representative local temporal correlation features; (4) Increase the number of CNN layers: On the basis of the original two convolutional layers, verify the effectiveness of the three-layer CNN structure through experiments, further optimize the feature representation, enabling the LSTM layer to obtain higher-quality feature information and improve the ability to capture time series patterns.
[0115] LSTM layer optimization: (1) Optimize the number of neurons in the hidden layer: Verify the impact of different hidden layer scales on prediction performance through experiments, and finally determine to optimize the number of neurons in the hidden layer according to the optimal lag order (Len): when Len = 4 and hidden_size = 64, the coefficient of determination = 0.71; when Len = 6 and hidden_size = 128, the coefficient of determination = 0.64; when Len = 8 and hidden_size = 256, the coefficient of determination = 0.92; indicating that this parameter combination has the best effect in this instance. (2) Optimize the number of LSTM layers. On the basis of the two-layer LSTM structure, appropriately increase or decrease the LSTM layer to optimize the time series modeling ability.
[0116] In addition, for cross-validation and AIC / BIC selection, (1) Adopt the cross-validation (TimeSeriesSplit) method to train and evaluate for different lag orders and hidden layer parameters, and select the hyperparameter combination with the minimum loss; (2) Calculate AIC and BIC, evaluate the impact of different model complexities on prediction performance, and finally select the lag order with the minimum AIC / BIC.
[0117] The system structure of this application is as Figure 3 shown, including: a data preprocessing module, a data screening module, a data processing module, a network structure setting module, a training parameter setting module, a cross-validation module, an evaluation index selection module, and a performance optimization module.
[0118] Taking the groundwater data in the central urban area of a certain city as an example, in the data preprocessing module, all groundwater data are first classified and sorted according to the time dimension to ensure the integrity and temporal consistency of the data. Before constructing the prediction model, it is necessary to screen the input features to eliminate environmental factors that are irrelevant to or have low contribution to the deformation prediction, thereby improving the accuracy and generalization ability of the prediction model. This application uses a random forest regression model to calculate the importance of each environmental feature based on the impurity reduction method and sets a threshold (taking 0.125 as an example) to screen high-contribution features. The feature contribution ratios are as Figure 4 shown. Among them, features such as temperature (°C), evaporation (mm), and rainfall (mm) have higher importance, while the influence of some features on the prediction is relatively small. To reduce redundant information and avoid interference of similar variables on the model stability, this application selects to retain the temperature (°C) with a higher contribution value and remove the minimum temperature (°C) with a relatively small contribution. After feature screening, the final input variables mainly include groundwater level data, and key external environmental factors such as rainfall, temperature, and evaporation are introduced to ensure that the model can fully learn the influencing factors of water level changes and improve the prediction accuracy; at the same time, interpolation analysis is performed using ArcGIS software, and the Kriging interpolation method is selected for the interpolation method. The interpolation result in March is as Figure 5 shown in (a), the lowest water level is 199.973m, and the highest water level is 675.541m. The water level gradually increases in an almost concentric circle shape from the center to the boundary of the study area. The interpolation result in April is as Figure 5 shown in (b), and the interpolation result in May is as Figure 5 shown in (c). The water level change also gradually increases in an almost concentric circle shape from the center to the boundary of the study area. The lowest water levels are 199.828m and 200.078m respectively, and the highest water levels are 675.73m and 670.866m respectively; seasonal rainfall may be the main reason for the water level change. Among all the data, it is usually desirable to select data points with smaller residuals and prediction errors because they can better reflect the overall data pattern and spatial variation law. In this example, the Geostatistical Analyst tool of ArcGIS is used to perform cross-validation on the Kriging interpolation result and analyze the residuals and prediction errors of the interpolation result. Among the data collected in different regions within the area, those with smaller residuals are selected as representative experimental data; when the residuals are not much different, those with smaller standardized errors are selected as representative experimental data. Figure 6 is the standardized error map for different sampling points. Among them, the monitoring station numbered 500103220001 has a smaller standardized error, so this number is selected as the representative experimental data.
[0119] In the data processing module, the daily average data of the groundwater level is adopted. A suitable monitoring site (taking the number 500103220001 as an example) is selected, and the data quality is ensured. For records with a water level of 0, interpolation filling can be used. To make the data converge more easily during the training process and avoid the impact of differences between eigenvalue on the model, the data is uniformly normalized. The data is divided into a training set and a test set in chronological order. The first 70% of the data is used as the training set, and the last 30% of the data is used as the test set.
[0120] In the network structure setting module, the network structure is divided into a CNN layer and an LSTM layer for setting. In the CNN layer, first, the size of the input data is specified through the image input layer, and then a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer are added in sequence to extract the features of the input data. In the convolutional layer, the convolutional kernel size is set to 5x1, and 8 convolutional kernels are generated. The batch normalization layer is used to accelerate the training process of the model and improve the generalization ability of the model. It normalizes the input data, which helps to avoid the problem of gradient disappearance or explosion. The ReLU activation function is used to perform a non-linear transformation on the output of the convolutional layer, introducing non-linear factors and enhancing the expression ability of the network. The pooling window size of the max pooling layer is 2x1, and the stride is 1, which is used to reduce the size of the feature map while retaining the most significant features. A fully connected layer and a regression layer are connected to the output end of the network to generate the output of the regression task. In the LSTM layer, a sequence input layer is set to receive sequential data, which contains the number of features of the training data. The LSTM layer contains a specified number of LSTM units to capture the long-term dependencies in the sequential data. The ReLU activation layer introduces non-linear factors. The fully connected layer maps the output of the LSTM layer to the final output dimension. The regression layer is used to calculate the loss of the regression task, and the dimension of the output layer and the number of hidden layer neurons are additionally set. The principle of the long short-term neural memory network is as Figure 7 shown.
[0121] In the training parameter setting module, the CNN-LSTM model is trained using the training set described in the data processing module to learn the features and patterns of the data. The CNN-LSTM neural network structure is as Figure 8 shown. Training is carried out in a GPU environment. The Adam optimization algorithm is selected for parameter optimization. The maximum number of training iterations is set to 4000 times, the gradient threshold is controlled to be 1 to prevent gradient explosion, the initial learning rate is set to 0.01, and a piecewise learning rate adjustment strategy is adopted. The learning rate is adjusted every 50 epochs, and the learning rate is updated by multiplying it by 0.5. At the same time, L2 regularization is added to prevent overfitting. The losses corresponding to the training of different parameters are as Figure 9 shown. The effect of deep training is as Figure 10 shown.
[0122] In the cross-validation module, TimeSeriesSplit is used to implement time series cross-validation to ensure the time order of the training set and the test set. The average evaluation metrics (such as RMSE) are calculated through cross-validation for different window sizes and different hidden layers, and then the optimal window size (i.e., the best lag order) and the best number of hidden layers are selected. The window size is dynamically selected through cross-validation instead of using a fixed window. For example, multiple window sizes (such as 6, 8, 10) are traversed, and the RMSE is calculated for each window size using cross-validation, and the window size with the smallest RMSE is selected as the best window.
[0123] In the evaluation metric selection module, there are many evaluation metrics for time series problems. In this application, the root mean square error, mean absolute error, and coefficient of determination are used as the metrics to evaluate the groundwater level prediction effect of the model. RMSE represents the average error between the predicted value and the actual observed value, with the same unit as the actual value. The smaller the RMSE, the better the prediction effect of the model; MAE represents the average absolute error between the predicted value and the actual observed value, with the same unit as the actual value. The smaller the MAE, the better the prediction effect of the model. The value range of the coefficient of determination is between 0 and 1. When the coefficient of determination is close to 1, it means that the model fits the data well and most of the data variance is explained by the model; when the coefficient of determination is close to 0, it means that the fitting effect of the model is poor.
[0124] In the performance optimization module, the lag order (Len) and the network structure layer of this instance are optimized to improve the prediction accuracy of the model. The optimization analysis of the lag order is as Figure 11 shown. According to the Akaike Information Criterion (AIC), when selecting a model, it is necessary to consider both the complexity and the fitting degree of the model, so that the model still has good fitting ability without overly increasing the complexity. The smaller the AIC value, the better the model performance. On the other hand, the Bayesian Information Criterion (BIC) punishes the model complexity more strictly in the case of large samples. Therefore, in this application, the lag order (Len) corresponding to the minimum value of AIC or BIC is selected as the optimal value. To determine the best lag order (Len), training sets are constructed based on different lag orders, and the CNN-LSTM model is used for training. TimeSeriesSplit is used for cross-validation, and the AIC and BIC values are calculated according to the formula to select the Len that minimizes AIC or BIC as the best lag order. In this instance, through cross-validation, it is found that when the lag order Len = 8 and the hidden layer size hidden_size = 256, the corresponding loss value is the smallest. That is, the best lag order is obtained by judging the loss values trained with different lag orders; therefore, Len = 8 is selected as the best lag order, and on this basis, the hyperparameter configuration of the CNN and LSTM layers is further optimized.
[0125] Once the optimal lag order is determined, the CNN and LSTM layers are further optimized to enhance the model's feature extraction ability and time series modeling ability. The optimization methods mainly include hyperparameter adjustment, structural improvement, and computational optimization to ensure that the model can more accurately capture the spatio-temporal variation patterns of the groundwater level.
[0126] The CNN-LSTM optimization process is as Figure 12 shown. After a series of hyperparameter adjustments and model training, a relatively optimal prediction effect is finally obtained. The training effect after optimization is as Figure 13 shown. Through multivariate analysis and deep learning techniques, the present invention realizes accurate prediction of the change trend and volatility of the groundwater level, and can be widely applied to fields such as groundwater resource management, environmental monitoring, and water conservancy projects, improving the sustainable utilization efficiency of groundwater resources.
[0127] The abstract process of this application is as Figure 14 shown. Based on the CNN-LSTM combined model, the present invention is optimized by integrating multiple evaluation indicators. The model performance is evaluated through RMSE, MAE, and to ensure the accuracy and stability of the prediction results. At the same time, AIC and BIC are used to select the most lag order, and combined with the TimeSeriesSplit cross-validation method, the optimal model parameters with the smallest loss are selected. The two are combined to achieve fine optimization of the window size and the number of hidden layers, improving the applicability and generalization ability of the model, so that it can adapt to the groundwater level prediction needs of different regions; the present invention uses a random forest regression model to optimize feature selection and improve the model prediction accuracy. By calculating the contribution ratio of different features to the water level prediction through the RF model, only the variables with higher contribution degrees are retained as input features, improving the effectiveness of model training. At the same time, in terms of the structural optimization of the CNN layer and the LSTM layer, the present invention reasonably designs the convolution kernel size and the pooling layer stride to enhance the feature extraction ability. In the data processing stage, Kriging interpolation is used to select the monitoring points with the smallest error to ensure the representativeness of the experimental data, and further Min-Max normalization processing is carried out to reduce the scale impact between different variables. In addition, the present invention introduces a batch normalization (BatchNormalization) layer and a ReLU activation function to stabilize the training process and accelerate model convergence. To improve the training efficiency and computational performance, the present invention uses GPU acceleration training, combined with the Adam optimization algorithm and the segmented learning rate adjustment strategy, effectively improving the training speed, optimizing the convergence of the model, reducing the computational complexity, and at the same time ensuring the applicability and generalization ability of the model on large-scale datasets.
[0128] In this embodiment, groundwater data of multi-variable groundwater levels to be predicted is obtained, and the groundwater data is verified, analyzed, and screened according to the time series to obtain processed data; based on the random forest regression model and using the average impurity reduction method, the contribution ratio of feature variables is calculated, feature screening, and normalization processing are performed on the processed data to obtain target data; a prediction model for multi-variable groundwater level time series prediction is constructed based on the convolutional neural network and the long short-term memory network, and the structure setting, cross-validation, parameter setting, and model training of the prediction model are performed to obtain the target prediction model; the target data is input into the target prediction model to obtain a predicted value, and an evaluation index is generated based on the predicted value and the actual measured value of the multi-variable groundwater level to be predicted, and the target prediction model is optimized using the evaluation index. In this application, groundwater data of multi-variable groundwater levels to be predicted is obtained, the groundwater data is verified, analyzed, and screened according to the time series, the random forest regression model is used and the average impurity reduction method is used to calculate the contribution ratio of feature variables, perform feature screening, and perform normalization processing on the processed data to obtain target data, improving the prediction accuracy of the model. Through the random forest regression model, variables with higher contribution degrees can be retained as input features, improving the effectiveness of model training. A prediction model for multi-variable groundwater level time series prediction is constructed based on the convolutional neural network and the long short-term memory network, enhancing the feature extraction ability. The structure setting, cross-validation, parameter setting, and model training of the prediction model are performed to obtain the target prediction model, reducing the scale influence between different variables, stabilizing the training process and accelerating model convergence. The target data is input into the target prediction model to obtain a predicted value, and an evaluation index is generated based on the predicted value and the actual measured value, improving the applicability and generalization ability of the model, enabling it to adapt to the groundwater level prediction requirements of different regions. The target prediction model is optimized using the evaluation index, effectively improving the training speed, optimizing the convergence of the model, reducing the computational complexity, and at the same time ensuring the applicability and generalization ability of the model on large-scale datasets, improving the accuracy and stability of multi-variable groundwater level time series prediction, and increasing the sustainable utilization efficiency of groundwater resources.
[0129] See Figure 15 As shown, an embodiment of the present invention discloses a multi-variable groundwater level time series prediction device, which may specifically include:
[0130] A verification, analysis, and screening processing module 11, configured to obtain groundwater data of multi-variable groundwater levels to be predicted, and process the groundwater data according to the time series to obtain processed data;
[0131] A calculation, screening, and normalization processing module 12, configured to calculate the contribution ratio of feature variables, perform feature screening, and perform normalization processing on the processed data based on the random forest regression model and using the average impurity reduction method to obtain target data;
[0132] The model construction and training module 13 is used to construct a prediction model for multivariate groundwater level time series prediction based on a convolutional neural network and a long short-term memory network, and perform structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain a target prediction model;
[0133] The prediction module 14 is used to input the target data into the target prediction model to obtain a predicted value, generate an evaluation index based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and optimize the target prediction model using the evaluation index.
[0134] In this embodiment, groundwater data of the multivariate groundwater level to be predicted is obtained, and the groundwater data is verified, analyzed, and screened according to the time series to obtain processed data; based on the random forest regression model and using the mean impurity reduction method, the contribution ratio of feature variables is calculated, feature screening, and normalization processing are performed on the processed data to obtain target data; a prediction model for multivariate groundwater level time series prediction is constructed based on a convolutional neural network and a long short-term memory network, and structure setting, cross-validation, parameter setting, and model training are performed on the prediction model to obtain a target prediction model; the target data is input into the target prediction model to obtain a predicted value, an evaluation index is generated based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and the target prediction model is optimized using the evaluation index. In this application, groundwater data of the multivariate groundwater level to be predicted is obtained, the groundwater data is verified, analyzed, and screened according to the time series, the random forest regression model is used and the mean impurity reduction method is used to calculate the contribution ratio of feature variables, perform feature screening, and normalization processing on the processed data to obtain target data, improving the prediction accuracy of the model. Through the random forest regression model, variables with higher contribution degrees can be retained as input features, improving the effectiveness of model training. A prediction model for multivariate groundwater level time series prediction is constructed based on a convolutional neural network and a long short-term memory network, enhancing the feature extraction ability. Structure setting, cross-validation, parameter setting, and model training are performed on the prediction model to obtain a target prediction model, reducing the scale influence between different variables, stabilizing the training process and accelerating model convergence. The target data is input into the target prediction model to obtain a predicted value, an evaluation index is generated based on the predicted value and the actual measured value, improving the applicability and generalization ability of the model, enabling it to adapt to the groundwater level prediction needs of different regions. The target prediction model is optimized using the evaluation index, effectively improving the training speed, optimizing the convergence of the model, reducing the computational complexity, and at the same time ensuring the applicability and generalization ability of the model on large-scale datasets, improving the accuracy and stability of multivariate groundwater level time series prediction, and increasing the sustainable utilization efficiency of groundwater resources.
[0135] In some specific embodiments, the verification analysis and screening processing module 11 may specifically include:
[0136] A screening module, configured to screen multivariate groundwater levels to be predicted that meet preset conditions, and collect corresponding groundwater data;
[0137] A processing module, configured to perform cross-validation, analysis, missing value elimination, and screening on the groundwater data according to a time series by using the Kriging interpolation method and the spatial distribution analysis method, in combination with statistical analysis tools in professional geographic information software, to obtain processed data.
[0138] In some specific embodiments, the calculation screening and normalization processing module 12 may specifically include:
[0139] A random forest regression model construction module, configured to construct a random forest regression model by using multiple decision trees;
[0140] A feature variable contribution ratio calculation module, configured to calculate the contribution ratio of feature variables to the processed data based on the random forest regression model by using the average impurity reduction method, and screen out the feature variables corresponding to the largest contribution ratio of feature variables;
[0141] A normalization processing module, configured to perform normalization processing on the feature variables by using the Min-Max normalization method to obtain target data.
[0142] In some specific embodiments, the model construction and training module 13 may specifically include:
[0143] A convolutional neural network construction module, configured to construct a convolutional neural network based on an image input layer, a convolutional layer, a batch normalization layer, a ReLU activation layer, a max pooling layer, a fully connected layer, and a regression layer;
[0144] A long short-term memory network construction module, configured to construct a long short-term memory network based on long short-term memory network units, a ReLU activation layer, a fully connected layer, a sequence input layer, and an output layer.
[0145] In some specific embodiments, the model construction and training module 13 may specifically include:
[0146] A data feature extraction module, configured to set the convolutional neural network in the prediction model to extract data features, and set the long short-term memory network to receive sequence-type data to complete the structure setting;
[0147] A target hyperparameter setting module, configured to perform cross-validation on the prediction model after structure setting by using the time series segmentation method, and set target hyperparameters;
[0148] A parameter setting and model training module is used to perform parameter setting and model training on the prediction model after setting target hyperparameters to obtain a target prediction model; the parameter setting includes learning rate setting and batch setting.
[0149] In some specific embodiments, the prediction module 14 may specifically include:
[0150] An error and coefficient of determination calculation module is used to calculate the root mean square error, mean absolute error, and coefficient of determination between the predicted value and the actual measured value of the multivariate groundwater level to be predicted;
[0151] An evaluation index generation module is used to generate evaluation indexes based on the root mean square error, the mean absolute error, and the coefficient of determination.
[0152] In some specific embodiments, the prediction module 14 may specifically include:
[0153] A lag order determination module is used to determine the lag order by using the Akaike information criterion and the Bayesian information criterion;
[0154] An adjustment and optimization module is used to adjust and optimize the parameters of the convolutional neural network and the long short-term memory network in the target prediction model based on the evaluation indexes and the lag order.
[0155] Figure 16 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the multivariate groundwater level time series prediction method executed by the electronic device disclosed in any of the foregoing embodiments.
[0156] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is made here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0157] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon include an operating system 221, a computer program 222, and data 223, etc., and the storage method can be short-term storage or permanent storage.
[0158] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, so as to realize the operation and processing of the data 223 in the memory 22 by the processor 21. It can be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the multi-variable groundwater level time series prediction method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks. In addition to the data that can include the data transmitted by external devices received by the multi-variable groundwater level time series prediction device, the data 223 may also include the data collected by its own input / output interface 25, etc.
[0159] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be implemented directly by hardware, software modules executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0160] Furthermore, the embodiment of the present application also discloses a computer-readable storage medium. When the computer program stored in the storage medium is loaded and executed by a processor, the steps of the multi-variable groundwater level time series prediction method disclosed in any of the foregoing embodiments are realized.
[0161] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0162] The above has introduced in detail a multi-variable groundwater level time series prediction method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multivariate groundwater level time series prediction method, characterized in that: include: Acquire groundwater data of multivariate groundwater levels to be predicted, perform verification analysis and screening processing on the groundwater data according to a time series, and obtain processed data; Based on the random forest regression model and using the average impurity reduction method, the processed data is subjected to characteristic variable contribution ratio calculation, feature screening, and normalization processing to obtain target data; A prediction model for multivariate groundwater level time series prediction is constructed based on a convolutional neural network and a long short-term memory network, and the prediction model is subjected to structure setting, cross-validation, parameter setting, and model training to obtain a target prediction model; The target data is input into the target prediction model to obtain a predicted value, an evaluation index is generated based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and the target prediction model is optimized using the evaluation index.
2. The multivariate groundwater level time series prediction method according to claim 1, characterized in that: The method of obtaining groundwater data of a multivariate groundwater level to be predicted and performing verification analysis and screening processing on the groundwater data according to a time series to obtain processed data includes: Screen the multivariate groundwater levels to be predicted that meet the preset conditions and collect the corresponding groundwater data; By using the Kriging interpolation method and the spatial distribution analysis method, combined with the statistical analysis tools in the professional geographic information software, the groundwater data were cross-validated, analyzed, missing values were eliminated, and screened according to the time series to obtain the processed data.
3. The multivariate groundwater level time series prediction method according to claim 1, characterized in that: The method of calculating the contribution ratio of feature variables, screening features, and normalizing the processed data based on the random forest regression model and using the average impurity reduction method to obtain target data includes: Use multiple decision trees to build a random forest regression model; Based on the random forest regression model and using the average impurity reduction method, the characteristic variable contribution ratio of the processed data is calculated, and the characteristic variable corresponding to the characteristic variable contribution ratio with the largest value is screened out; The Min-Max normalization method is used to normalize the characteristic variables to obtain target data.
4. The multivariate groundwater level time series prediction method according to claim 1, characterized in that: Before constructing a prediction model for multivariate groundwater level time series prediction based on a convolutional neural network and a long short-term memory network, the method further includes: Construct a convolutional neural network based on image input layer, convolution layer, batch normalization layer, ReLU activation layer, maximum pooling layer, fully connected layer, and regression layer; A long short-term memory network is constructed based on long short-term memory network units, ReLU activation layers, fully connected layers, sequence input layers, and output layers.
5. The multivariate groundwater level time series prediction method according to claim 1, characterized in that: The prediction model is subjected to structure setting, cross-validation, parameter setting, and model training to obtain a target prediction model, including: The convolutional neural network in the prediction model is set to extract data features, and the long short-term memory network is set to receive sequence data to complete the structure setting; Cross-validate the prediction model after structure setting using time series segmentation method and set target hyperparameters; The prediction model after setting the target hyperparameters is subjected to parameter setting and model training to obtain a target prediction model; the parameter setting includes learning rate setting and batch setting.
6. The multivariate groundwater level time series prediction method according to claim 1, characterized in that: The generating of the evaluation index based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted comprises: Calculating the root mean square error, mean absolute error, and coefficient of determination between the predicted value and the actual measured value of the multivariate groundwater level to be predicted; An evaluation index is generated based on the root mean square error, the mean absolute error, and the determination coefficient.
7. The multivariate groundwater level time series prediction method according to any one of claims 1 to 6, characterized in that: The optimizing the target prediction model by using the evaluation index includes: The Akaike information criterion and the Bayesian information criterion were used to determine the lag order; Based on the evaluation index and the lag order, the parameters of the convolutional neural network and the long short-term memory network in the target prediction model are adjusted and optimized.
8. A multivariate groundwater level time series prediction device, characterized in that: include: A verification analysis and screening processing module is used to obtain groundwater data of multivariate groundwater levels to be predicted, and to process the groundwater data according to a time series to obtain processed data; A calculation screening and normalization processing module is used to calculate the contribution ratio of feature variables, screen features, and normalize the processed data based on a random forest regression model and using an average impurity reduction method to obtain target data; A model building and training module is used to build a prediction model for multivariate groundwater level time series prediction based on a convolutional neural network and a long short-term memory network, and to perform structure setting, cross-validation, parameter setting, and model training on the prediction model to obtain a target prediction model; The prediction module is used to input the target data into the target prediction model to obtain a predicted value, generate an evaluation index based on the predicted value and the actual measured value of the multivariate groundwater level to be predicted, and optimize the target prediction model using the evaluation index.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the multivariate groundwater level time series prediction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the multivariate groundwater level time series prediction method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Deep foundation pit intelligent precipitation method and system based on Beidou communication and water level switch linkage
CN120803086A
Performance prediction method, computer program product, equipment and computer medium
CN120892309A
Key soil monitoring point screening method based on multiple machine learning methods
CN121030287A
Key soil monitoring point screening method based on multiple machine learning methods
CN121030287B
Multi-source fusion underground water level prediction system and method under dual-model machine learning
CN122133873A