A method and device for constructing a saltwater tide upstream salinity prediction model
By constructing a salt tide forward salinity forecasting method based on long and short-term memory model and sapli additive interpretation model, the calculation time-consuming problem in the prior art is solved, and rapid salinity forecasting is achieved.
Patent Information
- Application Number
- CN202411599928.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-11-11
AI Technical Summary
The existing salt tide upward salinity forecasting technology cannot achieve rapid forecasting, and requires a large amount of detailed information on boundaries or initial conditions such as runoff, tide, wind, riverbed terrain, etc., resulting in a long calculation time.
By obtaining the actual hydrological meteorological measurement data set, pre-processing, and constructing the hydrological meteorological training data set and test data set, using the long-term and short-term memory model and the Shapril additive interpretation model, the salt tide upward salinity forecast model is screened out, reducing the dependence on detailed boundary conditions, and achieving rapid salinity forecast.
Fast salt tide upward salinity forecast is achieved, reducing calculation time, and can make salinity forecasts without relying on a large number of detailed boundary conditions.
Smart Images

Figure CN119474682B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrology and water resources, and in particular to a method and device for constructing a saltwater tide upstream salinity forecasting model. Background Art
[0002] Saltwater upstreaming is a natural phenomenon that occurs in estuaries, where seawater flows upstream due to tidal action, causing salinity to rise at or near the mouth of a freshwater river. This phenomenon is caused by a variety of factors, including reduced river flow, tidal fluctuations, wind direction and strength, and rising sea levels. Especially during the dry season, when river freshwater flow decreases and high tides rise, saltwater can easily flow back into the river, creating a saltwater upstreaming phenomenon.
[0003] The upstream movement of saltwater can have a significant negative impact on freshwater resource utilization, industrial and agricultural production, and aquatic habitats in estuaries. According to national water supply standards, the chloride content in drinking water sources should not exceed 250 mg / L. Long-term consumption of water with excessive chloride levels can seriously harm human health.
[0004] Existing upstream saltwater salinity forecasting technologies are mostly based on numerical hydrodynamic salinity models based on the dynamics of salinity changes. Model construction and parameter verification require extensive detailed information on boundary or initial conditions, such as runoff, tides, wind, and riverbed topography. The process is complex and computationally time-consuming, making rapid salinity forecasting impossible. Summary of the Invention
[0005] The present invention provides a method and device for constructing a saltwater tide upstream salinity forecasting model, which are used to solve the technical problem that the existing saltwater tide upstream salinity forecasting technology cannot achieve rapid forecasting.
[0006] A first aspect of the present invention provides a method for constructing a saltwater tide upstream salinity prediction model, comprising:
[0007] Acquiring a hydrological and meteorological measured data set, and preprocessing the hydrological and meteorological measured data set to generate a target hydrological and meteorological measured data set;
[0008] Constructing a hydrometeorological training dataset, a hydrometeorological test dataset, and a hydrometeorological observation dataset based on multiple preset lag times, multiple preset lead times, and the target hydrometeorological measured dataset;
[0009] Using the hydrological and meteorological training data set and the hydrological and meteorological observation data set to perform model training on an initial long-short-term memory model, and determine an intermediate long-short-term memory model;
[0010] Calculating Shapley values corresponding to a subset of hydrometeorological test data of multiple time series in the hydrometeorological test data set based on the hydrometeorological test data set using a preset Shapley additive interpretation model and the intermediate long-short-term memory model;
[0011] constructing a plurality of target long-short-term memory models according to the plurality of time series hydrometeorological test data subsets, the hydrometeorological observation data set, and the Shapley values corresponding to the plurality of time series hydrometeorological test data subsets;
[0012] The hydrological and meteorological test data set and the hydrological and meteorological observation data set are used to screen the target long-short term memory models to determine the saltwater tide upstream salinity forecast model.
[0013] Optionally, the target hydrometeorological measured dataset includes multiple salinity datasets and multiple water-gas coupling datasets; and constructing the hydrometeorological training dataset, the hydrometeorological test dataset, and the hydrometeorological observation dataset based on multiple preset lag times, multiple preset lead times, and the target hydrometeorological measured dataset includes:
[0014] Based on a plurality of preset lag times, performing lag processing on the plurality of salinity data sets, and outputting a plurality of lag salinity data sets corresponding to the respective preset lag times;
[0015] Calculating, based on the plurality of lag-time salinity data sets corresponding to the respective preset lag times, the partial correlation coefficient associated with the plurality of lag-time salinity data sets corresponding to the respective preset lag times;
[0016] Based on multiple preset advance times, a salinity dataset associated with a lagged salinity dataset corresponding to a maximum partial correlation coefficient is used as a target salinity dataset, and the target salinity dataset is time-advanced to output a time-advanced salinity dataset corresponding to each of the preset advance times;
[0017] Based on the plurality of preset advance times, time-advanced the plurality of water-gas coupling data sets, and outputting a plurality of initial time-advanced water-gas coupling data sets corresponding to the respective preset advance times;
[0018] Calculating a Pearson correlation coefficient between the multiple salinity data sets and the multiple initial time-advanced water-gas coupling data sets corresponding to the preset advance times;
[0019] Based on the plurality of preset advance times, time-advance the water-gas coupling dataset associated with the initial time-advance water-gas coupling dataset corresponding to the largest Pearson correlation coefficient, and output a target time-advance water-gas coupling dataset corresponding to each of the preset advance times;
[0020] performing data sorting and standardization on the plurality of target time-advanced water-gas coupled data sets and the plurality of time-advanced salinity data sets, and outputting a plurality of standardized time-advanced water-gas coupled data sets and a plurality of standardized time-advanced salinity data sets;
[0021] Dividing the plurality of the standardized time-advanced water-gas coupling data sets and the plurality of the standardized time-advanced salinity data sets to generate a hydrometeorological training data set and a hydrometeorological test data set;
[0022] The target salinity dataset is used as a hydrological and meteorological observation dataset.
[0023] Optionally, the using a preset Shapley additive interpretation model and the intermediate long-short-term memory model to calculate, based on the hydrological and meteorological test dataset, Shapley values corresponding to hydrological and meteorological test data subsets of multiple time series in the hydrological and meteorological test dataset, includes:
[0024] Using the intermediate long-short term memory model, outputting salinity prediction values corresponding to the hydrological and meteorological test data subsets of each time series respectively according to the hydrological and meteorological test data subsets of each time series;
[0025] The preset Shapley additive interpretation model is used to calculate the Shapley value corresponding to each of the time series hydrological and meteorological test data subsets according to the salinity prediction value corresponding to each of the time series hydrological and meteorological test data subsets.
[0026] Optionally, constructing multiple target long-short-term memory models based on multiple subsets of the hydrological and meteorological test data of the time series, the hydrological and meteorological observation dataset, and the Shapley values corresponding to multiple subsets of the hydrological and meteorological test data of the time series includes:
[0027] sorting the Shapley values corresponding to the plurality of time series hydrological and meteorological test data subsets in descending order, selecting a preset number of time series hydrological and meteorological test data subsets as a first training data subset, and sequentially adding the remaining time series hydrological and meteorological test data subsets to the first training data subset to generate a plurality of second training data subsets;
[0028] The initial long short-term memory model is trained using the first training data subset, multiple second training data subsets, and the hydrological and meteorological observation data set to generate multiple target long short-term memory models.
[0029] Optionally, the using the hydrometeorological test dataset and the hydrometeorological observation dataset to screen each of the target long-short term memory models to determine the saltwater tide upstream salinity forecast model includes:
[0030] Using the hydrological and meteorological test data sets as inputs of the target long-short-term memory models, and outputting a plurality of salinity prediction values corresponding to the target long-short-term memory models;
[0031] Performing a model accuracy evaluation on each target long-short-term memory model using a plurality of salinity prediction values corresponding to each target long-short-term memory model and the hydrological and meteorological observation dataset, and outputting a Nash efficiency coefficient and a root mean square error corresponding to each target long-short-term memory model;
[0032] The target long-short-term memory model corresponding to the largest Nash efficiency coefficient and the smallest root mean square error is selected as the upstream salinity prediction model of saltwater tide.
[0033] Optionally, after the step of using the hydrological and meteorological test dataset and the hydrological and meteorological observation dataset to screen each of the target long-short-term memory models to determine the saltwater tide upstream salinity forecast model, the method further includes:
[0034] When the hydrological and meteorological measured data to be detected is received, the hydrological and meteorological measured data to be detected is input into the salt tide upstream salinity prediction model to perform salinity forecasting, and an initial salinity forecast result is output;
[0035] The initial salinity forecast result is denormalized to generate a target salinity forecast result.
[0036] A second aspect of the present invention provides a device for constructing a saltwater tide upstream salinity forecast model, comprising:
[0037] An acquisition module is used to acquire a hydrological and meteorological measured data set, and preprocess the hydrological and meteorological measured data set to generate a target hydrological and meteorological measured data set;
[0038] Based on the module, it is used to construct a hydrometeorological training dataset, a hydrometeorological test dataset and a hydrometeorological observation dataset based on multiple preset lag times, multiple preset lead times and the target hydrometeorological measured dataset;
[0039] An adoption module is used to use the hydrological and meteorological training data set and the hydrological and meteorological observation data set to perform model training on the initial long-short-term memory model and determine an intermediate long-short-term memory model;
[0040] a calculation module, configured to calculate, based on the hydrological and meteorological test data set, Shapley values corresponding to a subset of hydrological and meteorological test data of multiple time series in the hydrological and meteorological test data set using a preset Shapley additive interpretation model and the intermediate long-short-term memory model;
[0041] a construction module, configured to construct a plurality of target long-short-term memory models based on a plurality of the time series hydrometeorological test data subsets, the hydrometeorological observation dataset, and the Shapley values corresponding to the plurality of the time series hydrometeorological test data subsets;
[0042] The screening module is used to screen each of the target long-short term memory models using the hydrological and meteorological test data set and the hydrological and meteorological observation data set to determine a saltwater tide upstream salinity forecast model.
[0043] A third aspect of the present invention provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the method for constructing a saltwater upstream salinity forecast model as described in any one of the above items.
[0044] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the steps of the method for constructing a saltwater upstream salinity forecast model as described in any one of the above items.
[0045] A fifth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the steps of the method for constructing a saltwater upstream salinity forecast model as described in any one of the above items.
[0046] It can be seen from the above technical solutions that the present invention has the following advantages:
[0047] The above technical solution of the present invention provides a method for constructing a saltwater tide upstream salinity prediction model. First, a hydrometeorological measured data set is obtained, and the hydrometeorological measured data set is preprocessed to generate a target hydrometeorological measured data set; then, based on multiple preset lag times, multiple preset lead times, and the target hydrometeorological measured data set, a hydrometeorological training data set, a hydrometeorological test data set, and a hydrometeorological observation data set are constructed; the hydrometeorological training data set and the hydrometeorological observation data set are used to train an initial long-short-term memory model, and an intermediate long-short-term memory model is determined; a preset Shapley additive interpretation model and an intermediate long-short-term memory model are used to calculate the Shapley values corresponding to the hydrometeorological test data subsets of multiple time series in the hydrometeorological test data set according to the hydrometeorological test data set; The Shapley values corresponding to the subsets of hydrometeorological test data of a time series are used to construct multiple target long-short-term memory models; finally, the hydrometeorological test data set and the hydrometeorological observation data set are used to screen the target long-short-term memory models and determine the salt tide upstream salinity forecast model; based on the above scheme, the acquired hydrometeorological measured data set is processed to obtain the hydrometeorological training data set, the hydrometeorological test data set and the hydrometeorological observation data set, and the preset Shapley additive interpretation model and the long-short-term memory model are combined to obtain the salt tide upstream salinity forecast model according to the hydrometeorological training data set, the hydrometeorological test data set and the hydrometeorological observation data set. The process does not require a large amount of detailed information on boundaries or initial conditions such as runoff, tide, wind, riverbed topography, etc., and the obtained salt tide upstream salinity forecast model can be directly used to make the corresponding salinity forecast, which can reduce the calculation time and thus achieve rapid salinity forecast. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A flowchart of a method for constructing a saltwater tide upstream salinity forecasting model provided in Example 1 of the present invention;
[0050] Figure 2 A schematic diagram of the partial autocorrelation coefficients of salinity series with different lags provided in Example 1 of the present invention;
[0051] Figure 3A schematic diagram of the Pearson correlation coefficients between the tide level factor sequence, runoff factor sequence, wind speed and direction factor sequence, and salinity factor sequence at different lead times provided in Example 1 of the present invention;
[0052] Figure 4 A basic structural diagram of a long short-term memory neural network neuron provided in Example 1 of the present invention;
[0053] Figure 5 A schematic diagram of the results of a SHAP analysis based on the constructed long short-term memory model provided in Example 1 of the present invention;
[0054] Figure 6 The accuracy changes of the model training set (a) and test set (b) with different numbers of features provided in Example 1 of the present invention;
[0055] Figure 7 A schematic diagram of the prediction results of Model G provided in Example 1 of the present invention;
[0056] Figure 8 This is a structural block diagram of a device for constructing a saltwater upstream salinity forecast model provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0057] The embodiment of the present invention provides a method and device for constructing a saltwater tide upstream salinity forecasting model, which is used to solve the technical problem that the existing saltwater tide upstream salinity forecasting technology cannot achieve rapid forecasting.
[0058] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0059] See also Figure 1 , Figure 1 This is a flowchart of the steps of a method for constructing a saltwater upstream salinity forecast model provided in Example 1 of the present invention.
[0060] The present invention provides a method for constructing a saltwater tide upstream salinity prediction model, comprising:
[0061] Step 101: Acquire a hydrological and meteorological measured dataset, and preprocess the hydrological and meteorological measured dataset to generate a target hydrological and meteorological measured dataset.
[0062] It should be noted that, through the analysis of the influencing factors of the physical process of saltwater upwelling and data query, the hydrological and meteorological measured data sets of the dry season in the target area are collected. The hydrological and meteorological measured data sets include multiple salinity data sets and multiple water-gas coupling data sets that have not been preprocessed. The salinity data sets include multiple salinity data, and the water-gas coupling data sets include multiple tide level data sets, multiple flow data sets, and multiple wind speed and direction data sets.
[0063] Furthermore, the data preprocessing process includes data cleaning, such as removing duplicate data, correcting or deleting outliers, processing missing values, etc., to ensure the integrity and continuity of the data. In addition, if the scale of the collected data is inconsistent with the forecast scale, data integration must be performed first, such as aggregating hourly data into daily data.
[0064] Step 102: Based on multiple preset lag times, multiple preset lead times, and a target hydrometeorological measured data set, a hydrometeorological training data set, a hydrometeorological test data set, and a hydrometeorological observation data set are constructed.
[0065] Furthermore, step 102 may include the following sub-steps S21-S29:
[0066] Step S21: performing time lag processing on multiple salinity data sets based on multiple preset time lags, and outputting multiple time lag salinity data sets corresponding to each preset time lag;
[0067] Step S22: calculating the partial correlation coefficients of the multiple lag salinity data sets corresponding to the preset lag times according to the multiple lag salinity data sets corresponding to the preset lag times;
[0068] Step S23: Based on multiple preset advance times, the salinity dataset associated with the lagged salinity dataset corresponding to the largest partial correlation coefficient is used as the target salinity dataset, and the target salinity dataset is time-advanced to output the time-advanced salinity dataset corresponding to each preset advance time;
[0069] It should be noted that the partial correlation coefficients between salinity series at different lags are calculated and visualized. This reflects the net correlation between two salinity series at different lags, while controlling for one or more other variables. The salinity series corresponding to the lag with the highest (largest) partial correlation coefficient is selected as the input data series for model training.
[0070] Specifically, the multiple preset lag times are multiple different lag times that are set, and the multiple preset advance times are multiple different advance times that are set. Based on each lag time, each salinity data set is subjected to lag processing, and multiple lag salinity data sets corresponding to each preset lag time are output; for example, assuming that the preset lag time is a lag of 24 hours and a lag of 25 hours, and the salinity data set includes a salinity data set for January and a salinity data set for February, the salinity data set for January with a lag of 24 hours, the salinity data set for February with a lag of 24 hours, the salinity data set for January with a lag of 25 hours, and the salinity data set for February with a lag of 25 hours can be obtained.
[0071] Furthermore, among all the partial correlation coefficients associated with the lagged salinity data sets, the salinity data set associated with the lagged salinity data set corresponding to the largest partial correlation coefficient is selected. For example, if the lagged salinity data set corresponding to the largest partial correlation coefficient is the salinity data set of January that is lagged by 24 hours, then the salinity data set of January is used as the target salinity data set, and then the target salinity data set is time-advanced by setting multiple different advance times, and the time-advanced salinity data sets corresponding to each preset advance time are output. For example, assuming that the preset advance time is 24 hours in advance and 25 hours in advance, and the target salinity data set is the salinity data set of January, then the time-advanced salinity data sets corresponding to the preset advance time can be the salinity data set of January that is 24 hours in advance and the salinity data set of January that is 25 hours in advance.
[0072] For example, see Figure 2 , assuming that the hydrological and meteorological measured data set is the measured hydrological and meteorological historical data collected from the dry season of Modaomen Waterway from 2020 to 2023, including the hourly average salinity data of Pinggang Station, the hourly tide data of Sanzao Station, the hourly flow data of Makou Station, and the wind direction and speed data. Among them, the wind speed and direction data of Macau, China are derived from the ERA5-Land hourly data from 1950 to present of the European Centre for Medium-Range Weather Forecasts (ECMWF). The wind speed in the u and v directions at 10m in the data set is used as the two candidate prediction factors. Based on this hydrological and meteorological measured data set, the partial autocorrelation coefficients of the salinity series at different lags are calculated as follows: Figure 2 shown.
[0073] Step S24: based on the multiple preset advance times, time-advance the multiple water-gas coupling data sets, and output multiple initial time-advance water-gas coupling data sets corresponding to the preset advance times;
[0074] Step S25, calculating the Pearson correlation coefficient of the multiple initial time advance water-vapor coupling data sets corresponding to each preset advance time based on the multiple salinity data sets and the multiple initial time advance water-vapor coupling data sets corresponding to each preset advance time;
[0075] Step S26: Based on the multiple preset advance times, time-advance the water-gas coupling dataset associated with the initial time-advance water-gas coupling dataset corresponding to the largest Pearson correlation coefficient, and output the target time-advance water-gas coupling dataset corresponding to each preset advance time;
[0076] It should be noted that for other collected data sequences except the salinity factor, the Pearson correlation coefficient between the tide level factor sequence, runoff factor sequence, wind speed and direction factor sequence and salinity factor sequence at different lead times is first calculated. The calculation formula of the Pearson correlation coefficient can be expressed as:
[0077] ;
[0078] Among them, R is the Pearson correlation coefficient, which ranges from [-1, 1] and indicates the degree of linear correlation. The closer its absolute value is to 1, the greater the correlation is, and the better the fit is. is the salinity data of the i-th hour; is the average value of the salinity series; is the value corresponding to the water-gas coupling data of the ith hour in the water-gas coupling data sequence with different preset advance times. The water-gas coupling data includes tide level data, flow data, and wind speed and direction data; is the average value of the water-air coupling data series; n is the number of data contained in the data series.
[0079] Furthermore, for each prediction factor sequence, the exogenous driving item of the advance time series with the larger (largest) correlation coefficient is taken as the sequence to be input to the model; for example, assuming that the preset advance time is 24 hours in advance and 25 hours in advance, the initial time-advanced water-gas coupling dataset corresponding to the largest Pearson correlation coefficient is the water-gas coupling dataset of January 24 hours in advance, then based on 24 hours in advance and 25 hours in advance, the water-gas coupling dataset of January is time-advanced, and the output is the tide level dataset of January 24 hours in advance, the flow data of January 24 hours in advance, the wind speed and direction dataset of January 24 hours in advance, the tide level dataset of January 25 hours in advance, the flow data of January 25 hours in advance, and the wind speed and direction dataset of January 25 hours in advance.
[0080] For example, see Figure 3, the Pearson correlation coefficients between the tide level factor sequence, runoff factor sequence, wind speed and direction factor sequence and salinity factor sequence at different lead times were calculated. The Pearson correlation coefficients of some characteristic variables are as follows: Figure 3 As shown in Table 1, taking the hourly scale forecast of 24-hour salinity as an example, the selected forecast factors (i.e., water-gas coupling data and salinity data of different time series) are shown in Table 1, where S is the hourly salinity of Pinggang station (mg / L); F is the hourly flow of Makou (m 3 / s); T is the hourly tide level at Sanzao station (m); W1 is the wind speed in the u direction of Macao, China (m / s); W2 is the wind speed in the v direction of Macao, China (m / s); S_24 represents salinity data 24 hours ahead of the salinity forecast value (salinity data 24 hours ahead), S_25 represents salinity data 25 hours ahead of the salinity forecast value, T_24 represents tide data 24 hours ahead of the tide forecast value, F_24 represents flow data 24 hours ahead of the flow forecast value, W1_24 represents wind speed data in the u direction of Macao, China 24 hours ahead of the wind speed forecast value in the u direction of Macao, China, and W2_24 represents wind speed data in the v direction of Macao, China 24 hours ahead of the wind speed forecast value in the v direction of Macao, China.
[0081]
[0082] Step S27: performing data sorting and standardization on the multiple target time-leading water-gas coupling data sets and the multiple time-leading salinity data sets, and outputting multiple standardized time-leading water-gas coupling data sets and multiple standardized time-leading salinity data sets;
[0083] It should be noted that all the selected prediction factor sequences (target time-lead water-gas coupling dataset and time-lead salinity dataset) are organized into a two-dimensional dataset, one dimension is the time series dimension, and the other dimension is the input feature type, and the input data are standardized; among them, in order to avoid the dimensional influence between the various prediction factors before all data are input, all data need to be standardized.
[0084] Step S28: dividing the multiple standardized time-advanced water-gas coupling data sets and the multiple standardized time-advanced salinity data sets to generate a hydrometeorological training data set and a hydrometeorological test data set;
[0085] It should be noted that the hydrological and meteorological training dataset and the hydrological and meteorological test dataset are taken as a whole dataset and divided into a training set and a test set according to the ratio of 6:4 according to the time series, namely the hydrological and meteorological training dataset and the hydrological and meteorological test dataset; the hydrological and meteorological training dataset includes a plurality of salinity datasets, tide level datasets, flow datasets, and wind speed and direction datasets of different time series used for the initial model training, and the hydrological and meteorological test dataset includes a plurality of salinity datasets, tide level datasets, flow datasets, and wind speed and direction datasets of different time series used for shap value calculation, etc.; for example, assuming that the time series is 24 hours in advance and 25 hours in advance, the hydrological and meteorological training dataset includes a salinity dataset, a tide level dataset, a flow dataset, and a wind speed and direction dataset in January 24 hours in advance, and a salinity dataset, a tide level dataset, a flow dataset, and a wind speed and direction dataset in January 25 hours in advance.
[0086] Step S29: Use the target salinity dataset as a hydrological and meteorological observation dataset.
[0087] It should be noted that the data in the hydrometeorological observation dataset corresponds to the data in the hydrometeorological training dataset and the hydrometeorological test dataset. For example, assuming that the hydrometeorological training dataset includes the salinity dataset and tide level dataset of January 24 hours in advance, the hydrometeorological observation dataset also includes the salinity dataset and tide level dataset of January.
[0088] Step 103: Use the hydrological and meteorological training data set and the hydrological and meteorological observation data set to train the initial long-short-term memory model and determine the intermediate long-short-term memory model.
[0089] It should be noted that the Python library required for prediction is imported, the preprocessed data set is imported, and the data set is converted into a three-dimensional data set to obtain a three-dimensional data set (number of features, time series, number of samples).
[0090] Further, see Figure 4 , build the LSTM model and set the corresponding model parameters. The LSTM (Long Short Term Memory) neural network model, also known as the long short-term memory model, is a special recurrent neural network (RNN) architecture proposed by Hochreiter and Schmidhuber in 1997. It can alleviate the gradient vanishing or gradient exploding problems encountered by standard RNNs when processing long sequence data by selectively retaining and forgetting input memory. It can also solve the dependency problem of different time series periods. Its basic structure is as follows: Figure 4As shown. The LSTM memory unit contains three gates, namely forget, input, and output gates. The formula for its output is as follows:
[0091] ;
[0092] ;
[0093] ;
[0094] ;
[0095] ;
[0096] ;
[0097] ;
[0098] ;
[0099] in, For the Gate of Forgetfulness; is a nonlinear activation function; is the first weight vector corresponding to the forget gate; is the input value of the first layer; is the second weight vector corresponding to the forget gate; is the activation vector of the block at time t-1; is the basis vector corresponding to the forget gate; is the input gate; is the first weight vector corresponding to the input gate; is the second weight vector corresponding to the input gate; is the basis vector corresponding to the input gate; is the candidate memory neuron state; is the tanh activation function; is the first weight matrix corresponding to the candidate memory unit state; is the second weight matrix corresponding to the candidate memory unit state; is the bias vector of the candidate memory cell state; is the neuron state at time t; is the neuron state at time t-1; is the output gate; is the first weight matrix corresponding to the output gate; is the second weight matrix corresponding to the output gate; is the bias vector of the output gate; is the activation vector of the block at time t; is the output value of the second layer; is the weight matrix of the hidden state; is the bias vector of the hidden state.
[0100] Furthermore, the model accuracy is preliminarily verified based on the hydrological and meteorological observation dataset, and the model hyperparameters are adjusted. If the model accuracy is poor, the model parameters need to be further adjusted and retrained until the model accuracy reaches a certain standard.
[0101] In this embodiment, the above-mentioned preliminarily selected prediction factor datasets (hydrometeorological training dataset and hydrometeorological observation dataset) are used to construct an LSTM model, set corresponding model parameters, and build a suitable model by training the model with different parameter combinations.
[0102] Step 104 : Using a preset Shapley additive explanatory model and an intermediate long short-term memory model, the Shapley values corresponding to the hydrological and meteorological test data subsets of multiple time series in the hydrological and meteorological test data set are calculated based on the hydrological and meteorological test data set.
[0103] The hydrological and meteorological test data subsets of multiple time series include salinity data sets of multiple time series, tide level data sets of multiple time series, flow data sets of multiple time series, and wind speed and direction data sets of multiple time series.
[0104] Furthermore, step 104 may include the following sub-steps:
[0105] Step S41: using the intermediate long short-term memory model to output the salinity prediction value corresponding to the hydrological and meteorological test data subset of each time series according to the hydrological and meteorological test data subset of each time series;
[0106] Step S42: using a preset Shapley additive interpretation model to calculate the Shapley value corresponding to the hydrological and meteorological test data subset of each time series according to the salinity prediction value corresponding to the hydrological and meteorological test data subset of each time series.
[0107] It should be noted that the SHAP library (pre-configured with the Shapley Additive Explanation Model) in Python is used to calculate SHAP values for multiple time series hydrometeorological test data subsets. Specifically, the SHAP values corresponding to each hydrometeorological test data subset with different time series and data types are calculated. For example, if the hydrometeorological test data subsets include a 24-hour-ahead salinity dataset, a 25-hour-ahead salinity dataset, and a 24-hour-ahead flow dataset, the SHAP values corresponding to the 24-hour-ahead salinity dataset, the SHAP values corresponding to the 25-hour-ahead salinity dataset, and the 24-hour-ahead flow dataset can be calculated. The SHAP (SHapley Additive ExPlanations) model was proposed by Lundberg et al. in 2017. The SHAP method calculates the marginal contribution of a feature to the model's prediction results, known as the SHAP value (Shapley value). This method draws on the concept of Shapley value in game theory and can be applied to various machine learning models, including decision trees, random forests, gradient boosting machines, and deep learning models. Using SHAP values can help us better understand how features affect the model's prediction results and the extent of the impact, thereby improving the transparency and credibility of the model.
[0108] Furthermore, the calculation formula of the Shapley value can be expressed as:
[0109] ;
[0110] in, is the Shapley value corresponding to the feature subset of the i-th feature. The feature subset represents the salinity dataset, tide dataset, flow dataset, or wind speed and direction dataset of any time series input to the intermediate long short-term memory model. When it is greater than 0, it means that the feature improves the prediction value and plays a positive role. Otherwise, it means that the feature reduces the prediction value and plays a negative role. S is the feature subset that does not contain the i-th feature, that is, the feature subset that does not contain the salinity dataset, tide dataset, flow dataset, or wind speed and direction dataset of any time series input to the intermediate long short-term memory model. N is the set of all features, that is, the data set composed of all time series hydrological and meteorological test data subsets. The salinity prediction value of feature subset S plus feature subset i; It is the salinity prediction value output by the intermediate long short-term memory model when the input is the feature subset S.
[0111] Furthermore, the SHAP model decomposes the model's prediction results into the contributions of each feature. Each feature has a corresponding SHAP value, which represents the impact of the feature on the model's prediction results. For each feature, the SHAP value is obtained by calculating the marginal contribution of the feature in all possible feature combinations.
[0112] For example, see Figure 5 , call the shap library in Python to calculate the shap value of the test set of the long short-term memory model. The results of the SHAP analysis of the initially constructed long short-term memory model are as follows Figure 5 As shown in Figure 5, for the convenience of representation, the present invention names the salinity prediction factors as S1, S2, S3, S4, and S5, and the other variables are deduced by analogy, that is, Si represents salinity data of different time series, Fi represents flow data of different time series, Ti represents tide data of different time series, W1i represents wind speed in the u direction of different time series, and W2i represents wind speed in the v direction of different time series. As can be seen from Figure 5, in the prediction process of the model, the feature importance ranking of S1, S2, and F1 is at the top, that is, the salinity series (salinity data set) 24 hours in advance and 25 hours in advance play a greater role in the model prediction process, which is consistent with the research results of previous researchers who used physical models to explore the physical mechanism of saltwater tide traceability.
[0113] Step 105: Construct multiple target long-short-term memory models based on the hydrometeorological test data subsets of multiple time series, the hydrometeorological observation data set, and the Shapley values corresponding to the hydrometeorological test data subsets of multiple time series.
[0114] Furthermore, step 105 may include the following sub-steps S51-S52:
[0115] S51, sorting the Shapley values corresponding to the hydrological and meteorological test data subsets of the multiple time series in descending order, selecting a preset number of hydrological and meteorological test data subsets of the time series as the first training data subset, and gradually adding the hydrological and meteorological test data subsets of the remaining time series to the first training data subset in order to generate multiple second training data subsets;
[0116] S52: Using the first training data subset, multiple second training data subsets, and the hydrological and meteorological observation data set to perform model training on the initial long short-term memory model to generate multiple target long short-term memory models.
[0117] It should be noted that according to the calculation results of the shap value, the features included in the input data set of the model are adjusted, the predictor sequence with the feature importance ranked at the top is input first, and the number of input features, that is, the preset number generally starts from 3 and increases one by one, and models containing different numbers and types of features are established and trained respectively. For example, assuming that the Shapley values corresponding to the hydrological and meteorological test data subsets of all time series are sorted in descending order, the salinity sequence 24 hours in advance is ranked first, the salinity sequence 25 hours in advance is ranked second, the salinity sequence 26 hours in advance is ranked third, the tide sequence 24 hours in advance is ranked fourth, and the tide sequence 25 hours in advance is ranked fifth, then the salinity sequence 24 hours in advance, the salinity sequence 25 hours in advance, and the salinity sequence 26 hours in advance are used to construct the first training data subset, and the tide sequence 24 hours in advance and the tide sequence 25 hours in advance are gradually added to the first training data subset, that is, the first training data subset and the tide sequence 24 hours in advance are used for data construction, and the first training data subset, the tide sequence 24 hours in advance, and the tide sequence 25 hours in advance are used for data construction to obtain two second training data subsets. Then, the first training data subset is used to train the initial long short-term memory model, and the two second training data subsets are used to train the initial long short-term memory model respectively, and three target long short-term memory models are output.
[0118] For example, based on the SHAP calculation results, the features included in the model's input dataset are adjusted, with the predictor sequence with the highest feature importance being prioritized. The number of input features increases gradually from 3, and models containing different numbers of features are established and trained, respectively, named A, B, C, D, E, F, G, H, I, and J. The feature factors input to each model are shown in Table 2.
[0119]
[0120] Step 106: Use the hydrometeorological test data set and the hydrometeorological observation data set to screen the target long-short term memory models and determine the saltwater tide upstream salinity forecast model.
[0121] Furthermore, step 106 may include the following sub-steps:
[0122] S61, using the hydrological and meteorological test data set as the input of each target long-short-term memory model, and outputting multiple salinity prediction values corresponding to each target long-short-term memory model;
[0123] S62, using multiple salinity prediction values and hydrological and meteorological observation data sets corresponding to each target long and short-term memory model to evaluate the model accuracy of each target long and short-term memory model, and outputting the Nash efficiency coefficient and root mean square error corresponding to each target long and short-term memory model;
[0124] S63. Select the target long-short-term memory model corresponding to the largest Nash efficiency coefficient and the smallest root mean square error as the saltwater upstream salinity prediction model.
[0125] It should be noted that the model with different features (target long-short-term memory model) established based on the above steps uses the Nash efficiency coefficient (NSE) and root mean square error (RMSE) to evaluate the salinity forecast of the model. The calculation methods of NSE and RMSE are as follows:
[0126] ;
[0127] ;
[0128] Among them, NSE is the Nash efficiency coefficient, and the range of NSE is (-∞,1]. When NSE=1, it means that the model forecast is completely consistent; when NSE=0, it means that the explanatory power of the forecast result is the same as the mean of the observed value; when NSE<0, it means that the model forecast effect is lower than the mean forecast method; i is the predicted value of salinity i; P i is the i-th target salinity data (salinity observation value); RMSE is the root mean square error, which indicates the difference between the predicted value and the observed value. When the difference increases, it gradually increases from 0 to +∞; is the average value of the target salinity data (the average value of the salinity observations); n is the total number of samples.
[0129] For example, see Figure 6-Figure 7 , using NSE and RMSE to evaluate the model prediction accuracy, the model accuracy changes with different number of features are obtained as follows Figure 6 As shown in Figure 3, based on the calculation results of each model accuracy evaluation index, the model with the largest NSE and the smallest RMSE is selected as the optimal model. Based on the analysis of Table 3, it can be seen that when the number of model input features is 9, the model accuracy is the best, that is, model G is the optimal model. The prediction results of the model are visualized as follows Figure 7 As shown, the NSE of the training set and the test set can reach 0.764 and 0.738 respectively. Compared with the model of the initial prediction factor, the accuracy has been improved and can basically meet the forecast needs. Among them, the forecast results and the accuracy evaluation of model G show that the method of the present invention for salt tide retrospective forecasting based on interpretable machine learning is combined with the LSTM model and the SHAP model to perform refined salinity forecasting. It not only improves the shortcomings of the initial prediction factors that are too many and more troublesome to apply, improves the rationality of the model, but also further improves the forecast accuracy of the model. Furthermore, the current optimal model can be saved, and when the latest data of the study area is collected, efficient forecasts of future salinity can be made.
[0130]
[0131] Optionally, after the step of screening each target long-short term memory model using the hydrometeorological test dataset and the hydrometeorological observation dataset to determine the saltwater tide upstream salinity forecast model, the method includes:
[0132] When the hydrological and meteorological measured data to be detected is received, the hydrological and meteorological measured data to be detected is input into the salt tide upstream salinity prediction model to perform salinity forecasting and output the initial salinity forecast result;
[0133] The initial salinity forecast results are denormalized to generate the target salinity forecast results.
[0134] It should be noted that the best model (salinity forecast model for upstream saltwater tide) is saved in pb or h5 format. When necessary, the forecast factor with the required lag time is directly input to call the model (salinity forecast model for upstream saltwater tide) to make upstream saltwater tide forecast for the corresponding forecast period. Each time the saltwater forecast model for upstream saltwater tide is used for forecasting, the forecast result data needs to be denormalized.
[0135] As a comparison of technical effects, it can be combined with existing technologies for reference. Salt tide upstream is a natural hydrological phenomenon, mainly caused by factors such as reduced river flow, tidal changes, wind direction and force, and rising sea levels. The flow of fresh water in rivers decreases during the dry season, and the tide rises during high tides, making it easy for seawater to flow back into the river, forming a salt tide. In addition, when the wind force and direction are consistent with the tidal direction, the advancement of the salt tide will be exacerbated. The rise in sea levels caused by global warming has also expanded the scope of the impact of salt tides. In general, the upstream of salt tides will have a huge negative impact on the utilization of freshwater resources, industrial and agricultural production, and aquatic habitat systems in estuaries. Long-term drinking of water with excessive chloride content will seriously endanger human health.
[0136] Recurrent neural networks (RNNs) are deep learning models suitable for processing sequential data. RNNs can retain memory of previous information when processing sequences, capturing dynamic features in time series, making them well-suited for time series forecasting. The LSTM model effectively addresses the vanishing gradient problem of traditional RNNs by introducing a gating mechanism. It captures long-range dependencies in sequential data with fewer parameters and higher training efficiency, and has been widely used in the forecasting field. However, simulation studies on upstream salinity forecasting based on LSTM models are currently rare. In hydrology, LSTM models are primarily used for runoff forecasting.
[0137] In response to the above problems, the present invention proposes a method for predicting the time series of saltwater tides by using a long short-term memory network (LSTM), thereby realizing the prediction of saltwater tides at multiple time scales. The introduction of the SHAP model allows us to intuitively understand the contribution of the input features in the model and the importance ranking of the features. Combined with the physical mechanism of the saltwater tide, the rationality of the prediction model can be further verified. Therefore, combining the LSTM model and the SHAP model to quickly and effectively predict the saltwater tide in the estuary during the dry season is of great significance for ensuring the water resource security and ecological environmental protection of coastal areas.
[0138] Specifically, the present invention first collects and preprocesses data from the study area, then preliminarily selects predictors with appropriate lead times based on partial autocorrelation analysis and the Pearson correlation coefficient. The preliminarily selected predictors are used to construct an LSTM (Long Short-Term Memory) model. The established model is further analyzed using the SHAP model, and the predictors input into the model are screened again based on importance features, thereby constructing an optimal salinity prediction model to achieve the best forecasting effect. The present invention can effectively and finely predict the salinity value of saltwater tides in estuaries during the dry season. The introduction of the SHAP model improves the interpretability of the machine learning black box model, thereby providing a scientific basis for the precise prevention and control of saltwater tides in estuaries.
[0139] In an embodiment of the present invention, the present invention provides a method for constructing a saltwater tide upstream salinity forecast model. First, a hydrometeorological measured data set is obtained, and the hydrometeorological measured data set is preprocessed to generate a target hydrometeorological measured data set; then, based on multiple preset lag times, multiple preset lead times, and the target hydrometeorological measured data set, a hydrometeorological training data set, a hydrometeorological test data set, and a hydrometeorological observation data set are constructed; the hydrometeorological training data set and the hydrometeorological observation data set are used to train an initial long-short-term memory model, and an intermediate long-short-term memory model is determined; a preset Shapley additive interpretation model and an intermediate long-short-term memory model are used to calculate the Shapley values corresponding to the hydrometeorological test data subsets of multiple time series in the hydrometeorological test data set according to the hydrometeorological test data set; The Shapley value corresponding to the subset of the hydrometeorological test data of the sequence is used to construct multiple target long-short-term memory models; finally, the hydrometeorological test data set and the hydrometeorological observation data set are used to screen the target long-short-term memory models and determine the salt tide upstream salinity forecast model; based on the above scheme, the acquired hydrometeorological measured data set is processed to obtain the hydrometeorological training data set, the hydrometeorological test data set and the hydrometeorological observation data set. The preset Shapley additive interpretation model and the long-short-term memory model are combined to obtain the salt tide upstream salinity forecast model according to the hydrometeorological training data set, the hydrometeorological test data set and the hydrometeorological observation data set. The process does not require a large amount of detailed information on the boundaries or initial conditions such as runoff, tide, wind, riverbed topography, etc. The obtained salt tide upstream salinity forecast model can be directly used to make the corresponding salinity forecast, which can reduce the calculation time and simplify the process mechanism, thereby realizing rapid salinity forecast.
[0140] See also Figure 8 , Figure 8 A structural block diagram of a saltwater tide upstream salinity forecast model construction device provided in Example 2 of the present invention
[0141] The present invention provides a device for constructing a saltwater tide upstream salinity forecast model, comprising:
[0142] The acquisition module 801 is used to acquire a hydrological and meteorological measured data set, and pre-process the hydrological and meteorological measured data set to generate a target hydrological and meteorological measured data set;
[0143] Based on module 802, it is used to construct a hydrometeorological training dataset, a hydrometeorological test dataset, and a hydrometeorological observation dataset based on multiple preset lag times, multiple preset lead times, and a target hydrometeorological measured dataset;
[0144] Adopting module 803, for using the hydrological and meteorological training data set and the hydrological and meteorological observation data set to perform model training on the initial long short-term memory model, and determine the intermediate long short-term memory model;
[0145] A calculation module 804 is configured to calculate Shapley values corresponding to a subset of hydrometeorological test data of multiple time series in the hydrometeorological test data set based on the hydrometeorological test data set using a preset Shapley additive interpretation model and an intermediate long short-term memory model;
[0146] A construction module 805 is used to construct multiple target long-short-term memory models based on the hydrological and meteorological test data subsets of multiple time series, the hydrological and meteorological observation data set, and the Shapley values corresponding to the hydrological and meteorological test data subsets of multiple time series;
[0147] The screening module 806 is used to screen each target long-short term memory model using the hydrological and meteorological test data set and the hydrological and meteorological observation data set to determine the saltwater tide upstream salinity forecast model.
[0148] Furthermore, the target hydrological and meteorological measured data set includes multiple salinity data sets and multiple water-gas coupling data sets; based on module 802, specifically used for:
[0149] Based on multiple preset lag times, multiple salinity data sets are subjected to lag processing, and multiple lag salinity data sets corresponding to each preset lag time are output;
[0150] Calculating, based on the multiple lag salinity data sets corresponding to the respective preset lag times, the partial correlation coefficients associated with the multiple lag salinity data sets corresponding to the respective preset lag times;
[0151] Based on multiple preset advance times, the salinity dataset associated with the lagged salinity dataset corresponding to the largest partial correlation coefficient is used as the target salinity dataset, and the target salinity dataset is time-advanced to output the time-advanced salinity dataset corresponding to each preset advance time;
[0152] Based on multiple preset advance times, multiple water-gas coupling data sets are time-advanced, and multiple initial time-advanced water-gas coupling data sets corresponding to each preset advance time are output;
[0153] Calculating a Pearson correlation coefficient of the multiple salinity data sets and the multiple initial time advance water-gas coupling data sets corresponding to the preset advance time;
[0154] Based on multiple preset advance times, the water-gas coupling dataset associated with the initial time-advance water-gas coupling dataset corresponding to the largest Pearson correlation coefficient is time-advanced, and the target time-advance water-gas coupling dataset corresponding to each preset advance time is output;
[0155] Perform data collation and standardization on multiple target time-advance water-air coupled datasets and multiple time-advance salinity datasets, and output multiple standardized time-advance water-air coupled datasets and multiple standardized time-advance salinity datasets;
[0156] Divide multiple standardized time-advanced water-air coupling datasets and multiple standardized time-advanced salinity datasets to generate hydrometeorological training datasets and hydrometeorological test datasets;
[0157] The target salinity dataset is used as the hydrological and meteorological observation dataset.
[0158] Furthermore, the calculation module 804 is specifically configured to:
[0159] The intermediate long short-term memory model is used to output the salinity prediction value corresponding to the hydrological and meteorological test data subset of each time series according to the hydrological and meteorological test data subset of each time series;
[0160] The preset Shapley additive interpretation model is used to calculate the Shapley value corresponding to the subset of hydrometeorological test data of each time series based on the salinity prediction value corresponding to the subset of hydrometeorological test data of each time series.
[0161] Furthermore, the construction module 805 is specifically configured to:
[0162] Sort the Shapley values corresponding to the hydrological and meteorological test data subsets of multiple time series in descending order, select a preset number of hydrological and meteorological test data subsets of the time series as the first training data subset, and gradually add the hydrological and meteorological test data subsets of the remaining time series to the first training data subset in order to generate multiple second training data subsets;
[0163] The initial long short-term memory model is trained using a first training data subset, multiple second training data subsets, and a hydrological and meteorological observation data set to generate multiple target long short-term memory models.
[0164] Furthermore, the screening module 806 is specifically configured to:
[0165] The hydrological and meteorological test datasets are used as inputs of each target long-short-term memory model, and multiple salinity prediction values corresponding to each target long-short-term memory model are output;
[0166] The model accuracy of each target long-short-term memory model is evaluated using multiple salinity prediction values and hydrological and meteorological observation datasets corresponding to each target long-short-term memory model, and the Nash efficiency coefficient and root mean square error corresponding to each target long-short-term memory model are output;
[0167] The target long-short-term memory model corresponding to the largest Nash efficiency coefficient and the smallest root mean square error is selected as the upstream salinity prediction model of saltwater tide.
[0168] In an optional embodiment, the apparatus further comprises:
[0169] The first module is used for inputting the hydrological and meteorological measured data to be detected into the salt tide upstream salinity prediction model to perform salinity forecasting and outputting the initial salinity forecast result when receiving the hydrological and meteorological measured data to be detected;
[0170] The second module is used to denormalize the initial salinity forecast results and generate target salinity forecast results.
[0171] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0172] An embodiment of the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the method for constructing a saltwater upstream salinity prediction model as described in any of the above embodiments.
[0173] An embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the method for constructing a saltwater upstream salinity prediction model as described in any of the above embodiments are implemented.
[0174] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the method for constructing a saltwater upstream salinity prediction model as described in any of the above embodiments.
[0175] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0176] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0177] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a saltwater tide upstream salinity forecast model, characterized in that: include: Acquiring a hydrological and meteorological measured data set, and preprocessing the hydrological and meteorological measured data set to generate a target hydrological and meteorological measured data set; Constructing a hydrometeorological training dataset, a hydrometeorological test dataset, and a hydrometeorological observation dataset based on multiple preset lag times, multiple preset lead times, and the target hydrometeorological measured dataset; Using the hydrological and meteorological training data set and the hydrological and meteorological observation data set to perform model training on an initial long-short-term memory model, and determine an intermediate long-short-term memory model; Calculating Shapley values corresponding to a subset of hydrometeorological test data of multiple time series in the hydrometeorological test data set based on the hydrometeorological test data set using a preset Shapley additive interpretation model and the intermediate long-short-term memory model; constructing a plurality of target long-short-term memory models according to the plurality of time series hydrometeorological test data subsets, the hydrometeorological observation data set, and the Shapley values corresponding to the plurality of time series hydrometeorological test data subsets; Using the hydrological and meteorological test data set and the hydrological and meteorological observation data set to screen each of the target long-short term memory models, and determine a saltwater tide upstream salinity forecast model; The target hydrometeorological measured data set includes multiple salinity data sets and multiple water-gas coupling data sets; constructing a hydrometeorological training data set, a hydrometeorological test data set, and a hydrometeorological observation data set based on multiple preset lag times, multiple preset lead times, and the target hydrometeorological measured data set includes: Based on a plurality of preset lag times, performing lag processing on the plurality of salinity data sets, and outputting a plurality of lag salinity data sets corresponding to the respective preset lag times; Calculating, based on the plurality of lag salinity data sets corresponding to the respective preset lag times, the partial correlation coefficients associated with the plurality of lag salinity data sets corresponding to the respective preset lag times; Based on multiple preset advance times, a salinity dataset associated with a lagged salinity dataset corresponding to a maximum partial correlation coefficient is used as a target salinity dataset, and the target salinity dataset is time-advanced to output a time-advanced salinity dataset corresponding to each of the preset advance times; Based on the plurality of preset advance times, the plurality of water-gas coupling data sets are time-advanced, and a plurality of initial time-advanced water-gas coupling data sets corresponding to the respective preset advance times are output; Calculating a Pearson correlation coefficient between the multiple salinity data sets and the multiple initial time-advanced water-gas coupling data sets corresponding to the preset advance times; Based on the plurality of preset advance times, time-advance the water-gas coupling dataset associated with the initial time-advance water-gas coupling dataset corresponding to the largest Pearson correlation coefficient, and output a target time-advance water-gas coupling dataset corresponding to each of the preset advance times; performing data sorting and standardization on the plurality of target time-advanced water-gas coupled data sets and the plurality of time-advanced salinity data sets, and outputting a plurality of standardized time-advanced water-gas coupled data sets and a plurality of standardized time-advanced salinity data sets; Dividing the plurality of the standardized time-advanced water-gas coupling data sets and the plurality of the standardized time-advanced salinity data sets to generate a hydrometeorological training data set and a hydrometeorological test data set; The target salinity dataset is used as a hydrological and meteorological observation dataset.
2. The method for constructing a saltwater tide upstream salinity forecast model according to claim 1, wherein: The method of using the preset Shapley additive interpretation model and the intermediate long-short-term memory model to calculate the Shapley values corresponding to the hydrological and meteorological test data subsets of multiple time series in the hydrological and meteorological test data set according to the hydrological and meteorological test data set includes: Using the intermediate long-short term memory model, outputting salinity prediction values corresponding to the hydrological and meteorological test data subsets of each time series respectively according to the hydrological and meteorological test data subsets of each time series; The preset Shapley additive interpretation model is used to calculate the Shapley value corresponding to each of the time series hydrological and meteorological test data subsets according to the salinity prediction value corresponding to each of the time series hydrological and meteorological test data subsets.
3. The saltwater tide upstream salinity forecast model construction method according to claim 1, wherein The constructing of multiple target long-short term memory models based on the multiple time series hydrometeorological test data subsets, the hydrometeorological observation data set, and the Shapley values corresponding to the multiple time series hydrometeorological test data subsets includes: sorting the Shapley values corresponding to the plurality of time series hydrological and meteorological test data subsets in descending order, selecting a preset number of time series hydrological and meteorological test data subsets as a first training data subset, and sequentially adding the remaining time series hydrological and meteorological test data subsets to the first training data subset to generate a plurality of second training data subsets; The initial long short-term memory model is trained using the first training data subset, multiple second training data subsets, and the hydrological and meteorological observation data set to generate multiple target long short-term memory models.
4. The method for constructing a saltwater tide upstream salinity forecast model according to claim 1, wherein: The method of using the hydrological and meteorological test data set and the hydrological and meteorological observation data set to screen each of the target long-short term memory models and determine a saltwater tide upstream salinity forecast model includes: Using the hydrological and meteorological test data sets as inputs of the target long-short-term memory models, and outputting a plurality of salinity prediction values corresponding to the target long-short-term memory models; Performing a model accuracy evaluation on each target long-short-term memory model using a plurality of salinity prediction values corresponding to each target long-short-term memory model and the hydrological and meteorological observation dataset, and outputting a Nash efficiency coefficient and a root mean square error corresponding to each target long-short-term memory model; The target long-short-term memory model corresponding to the largest Nash efficiency coefficient and the smallest root mean square error is selected as the upstream salinity prediction model of saltwater tide.
5. The method for constructing a saltwater tide upstream salinity forecast model according to claim 1, wherein: After the step of using the hydrological and meteorological test data set and the hydrological and meteorological observation data set to screen each of the target long-short term memory models to determine the saltwater tide upstream salinity forecast model, the method includes: When the hydrological and meteorological measured data to be detected is received, the hydrological and meteorological measured data to be detected is input into the salt tide upstream salinity prediction model to perform salinity forecasting, and an initial salinity forecast result is output; The initial salinity forecast result is denormalized to generate a target salinity forecast result.
6. A saltwater tide upstream salinity forecast model construction device, applied to the saltwater tide upstream salinity forecast model construction method according to claim 1, characterized in that: include: An acquisition module is used to acquire a hydrological and meteorological measured data set, and preprocess the hydrological and meteorological measured data set to generate a target hydrological and meteorological measured data set; Based on the module, it is used to construct a hydrometeorological training dataset, a hydrometeorological test dataset and a hydrometeorological observation dataset based on multiple preset lag times, multiple preset lead times and the target hydrometeorological measured dataset; An adoption module is used to use the hydrological and meteorological training data set and the hydrological and meteorological observation data set to perform model training on the initial long-short-term memory model and determine an intermediate long-short-term memory model; a calculation module, configured to calculate, based on the hydrological and meteorological test data set, Shapley values corresponding to a subset of hydrological and meteorological test data of multiple time series in the hydrological and meteorological test data set using a preset Shapley additive interpretation model and the intermediate long-short-term memory model; a construction module, configured to construct a plurality of target long-short-term memory models based on a plurality of the time series hydrometeorological test data subsets, the hydrometeorological observation dataset, and the Shapley values corresponding to the plurality of the time series hydrometeorological test data subsets; The screening module is used to use the hydrological and meteorological test data set and the hydrological and meteorological observation data set to screen each of the target long-short term memory models to determine a saltwater tide upstream salinity forecast model.
7. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing a saltwater upstream salinity prediction model according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the method for constructing a saltwater upstream salinity forecast model according to any one of claims 1 to 5 is implemented.
9. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to execute the method for constructing a saltwater upstream salinity prediction model according to any one of claims 1 to 5.
Citation Information
Patent Citations
Calculation method and device for contribution degree of training data set, equipment and storage medium
CN111325353A