A Method for Predicting Wave Parameters of Ecological Coastlines Based on Sparse Gaussian Processes and SHAP Interpretation
By constructing a three-level cascaded sparse Gaussian process regression model and combining it with the SHAP interpreter, the problems of high computational complexity and poor interpretability of wave prediction methods in complex environments are solved, and efficient and interpretable wave parameter prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DALIAN MARITIME UNIVERSITY
- Filing Date
- 2025-11-18
- Publication Date
- 2026-07-31
AI Technical Summary
Existing wave prediction methods have limited ability to characterize local topographic changes, tidal currents, and nonlinear processes in complex nearshore areas or environments disturbed by artificial structures. They also have large parameterization errors, making it difficult to meet the requirements for fast and stable prediction. Furthermore, deep learning models lack a clear physical interpretation mechanism, which affects the interpretability and reliability of the models.
An ecological shoreline wave parameter prediction method based on sparse Gaussian processes and SHAP interpretation is adopted. By constructing a three-level cascaded sparse Gaussian process regression model and combining it with the SHAP interpreter, the prediction and interpretability analysis of wave parameters are realized.
It significantly reduces computational complexity, improves training and inference efficiency, provides transparent prediction results, enhances model interpretability and reliability, and enables high-precision, interpretable wave prediction in complex marine environments.
Smart Images

Figure CN121503276B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary integration of marine engineering and artificial intelligence, and in particular to a method for predicting ecological shoreline wave parameters based on sparse Gaussian processes and SHAP interpretation. Background Technology
[0002] Wave prediction is one of the core technologies in marine engineering and ocean dynamics research, and it has significant scientific and engineering application value in areas such as offshore construction safety, port operation scheduling, ship navigation and risk avoidance, and nearshore disaster prevention and mitigation. The generation and evolution of waves are influenced by multiple factors, including wind field, tidal current, topography, and seabed friction, and their changes exhibit significant nonlinear and multi-scale characteristics. Therefore, achieving high-precision, real-time wave prediction has always been a key challenge in marine science research.
[0003] Currently, wave prediction mainly relies on numerical simulation methods, such as third-generation spectral models SWAN (Simulating WAVEs Nearshore) and WW3 (Wave Watch III). These models, based on the spectral energy balance equation, accurately simulate the dynamic behaviors of wave propagation, shallowing, and breaking by solving the spatial and temporal transport, generation, and dissipation processes of wave energy. These methods have a solid physical theoretical foundation and exhibit high accuracy under deep-sea and regular boundary conditions. However, their computation is highly dependent on initial boundary conditions and physical parameters; numerical integration and grid discretization are complex, resulting in large computational loads and long processing times. In complex nearshore areas or environments disturbed by artificial structures, the models have limited ability to characterize local topographic changes, tidal currents, and nonlinear processes. Parameterization errors can easily lead to significant deviations in prediction results, making it difficult to meet the needs of practical engineering for rapid and stable predictions.
[0004] With the rapid development of artificial intelligence and big data technologies, data-driven deep learning methods have been gradually introduced into the field of wave prediction. Commonly used models include Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), and Graph Neural Networks (GNN). These methods can automatically extract feature patterns from historical observation data, avoiding explicit solutions to complex physical equations, significantly improving computational efficiency and reducing reliance on prior knowledge, making them particularly suitable for handling high-dimensional, nonlinear, and non-stationary time-series features. However, deep learning models generally suffer from "black box" characteristics, lacking a clear physical interpretation mechanism. Their internal feature weights and decision-making processes are difficult to explain intuitively, thus affecting the interpretability and reliability of the models. Furthermore, under the multi-source dynamic conditions of the marine environment, deep learning models often struggle to stably capture the coupling relationships between high-dimensional input features, easily leading to overfitting or insufficient generalization ability, resulting in insufficient physical consistency of prediction results.
[0005] In recent years, the rise of probabilistic modeling and Bayesian machine learning methods has provided new insights into wave prediction. Gaussian Process Regression (GPR) can naturally introduce uncertainty quantification into the prediction process, providing both the prediction mean and confidence intervals to reflect the model's reliability. However, the computational complexity of traditional Gaussian processes is... In the context of large-scale ocean observation data, it is difficult to apply directly. In addition, existing studies based on GPR models mostly focus on improving prediction performance, and lack systematic interpretive analysis methods. They cannot intuitively reveal the influence mechanism of environmental factors such as wind speed, tide level, and current velocity on wave characteristic changes, nor can they establish a transparent correlation between physics and data-driven models. Summary of the Invention
[0006] This invention provides a method for predicting ecological shoreline wave parameters based on sparse Gaussian processes and SHAP interpretation, in order to overcome the above-mentioned technical problems.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] A method for predicting ecological shoreline wave parameters based on sparse Gaussian processes and SHAP interpretation specifically includes the following steps:
[0009] S1: Collect and acquire historical time-series data of corresponding sample points at each observation station;
[0010] Furthermore, the historical time-series data includes environmental driving variables and their corresponding target outputs; the environmental driving variables include at least flow velocity and water level; the target outputs include at least water depth, wave height, and peak period.
[0011] S2: A sparse Gaussian process regression model constructed based on historical time series data to establish an ecological shoreline wave parameter prediction model based on sparse Gaussian processes. It includes an input module, a three-level cascaded prediction module, and an output module.
[0012] The input module is used to input historical time series data into the three-level cascaded prediction module;
[0013] The three-level cascaded prediction module includes a first SGPR stage model, a second SGPR stage model, and a third SGPR stage model connected in sequence. The first SGPR stage model is used to obtain predicted water depth features by taking environmental driving variables as feature inputs. The second SGPR stage model is used to obtain predicted wave height features by taking environmental driving variables and predicted water depth features as inputs. The third SGPR stage model is used to obtain predicted peak periodic features by taking environmental driving variables, predicted water depth features, and predicted wave height features as feature inputs.
[0014] The output module is used to output predicted output features that include predicted water depth features, predicted wave height features, and predicted peak period features.
[0015] S3: Train the ecological shoreline wave parameter prediction model based on historical time series data to obtain the optimal ecological shoreline wave parameter prediction model, and obtain the prediction output features through the optimal ecological shoreline wave parameter prediction model;
[0016] S4: The importance of SHAP features for each observation station and the global interpretation are calculated based on the predicted output features by the pre-set SHAP interpreter, so as to enhance the interpretability of the ecological shoreline wave parameter prediction model and realize the prediction of ecological shoreline wave parameters based on sparse Gaussian process and SHAP interpretation.
[0017] Furthermore, the first SGPR stage model, the second SGPR stage model, and the third SGPR stage model in S2 are sparse Gaussian process regression models constructed based on historical time series data.
[0018] Furthermore, the method for constructing a sparse Gaussian process regression model specifically includes the following steps:
[0019] S100: Establish the state function model of the Gaussian process;
[0020] S101: Observational data assuming a sparse Gaussian process regression model The latent function value f follows a Gaussian likelihood distribution. Its expression is:
[0021]
[0022] In the formula: y represents the observed data; f represents the latent function value; This represents the observation noise variance, also known as the hyperparameter. Represents the unit noise matrix; Represents a Gaussian probability distribution;
[0023] Based on the Gaussian process state function model combined with Gaussian likelihood distribution Obtain the joint distribution function of the sparse Gaussian process regression model, whose expression is:
[0024]
[0025] In the formula: This represents the set of input sample points, which contains the feature vectors of all observation points; The corresponding observation output vector represents the observed value of the target output quantity actually measured at each input sample point; This indicates that the model is based on the input sample set. The corresponding latent function value; Indicates the th in the training input sample set The latent function values corresponding to each input sample point and = ; This represents the observed values given a set of input samples. With latent function value The joint probability distribution of ; This indicates that the output of the latent function is Generate observation data under the circumstances The observed likelihood distribution of the probability relationship; Indicates that in a given set of inputs The prior probability distribution of the latent function value and satisfying ; Representing kernel function The covariance matrix formed and
[0026] ;
[0027] S102: Based on Bayesian principles and the joint distribution function, obtain the prediction points. posterior distribution on Its expression is
[0028]
[0029] In the formula: Indicates at the test point The output of the latent function at that location; This represents the predicted mean; Indicates the prediction variance; This represents the covariance matrix between the test point and the given induced point; This represents the covariance value of the test point itself; express transpose; Indicates the inclusion of test points A subset of features;
[0030] S103: By randomly selecting a limited number of input sample points from historical time-series data as induction points, and approximating the predicted points based on the limited number of induction points. The posterior distribution on the input sample points, i.e., the latent function values corresponding to the input sample points. With the induced point function value joint distribution Its expression is
[0031]
[0032] In the formula: This represents the latent function value corresponding to the input sample point; This represents the potential function value at the induced point location; Describe the set of induced points and ; K represents the covariance matrix between each input sample point; XZ K represents the covariance matrix between the input sample points and the induced points; ZX K represents XZ The transpose of K; ZZ This represents the covariance matrix between the induced points; Represents a Gaussian probability distribution;
[0033] This leads to the construction of a sparse Gaussian process regression model.
[0034] Furthermore, the Gaussian process state function model constructed in S100 has the following expression:
[0035]
[0036] In the formula: Indicates the sign for Gaussian process regression; These represent the latent function values corresponding to a certain input feature vector x, namely water depth, significant wave height, and peak period, respectively. This represents the input feature vector corresponding to another sample point; These represent the mean functions of water depth, significant wave height, and peak period, respectively. Let represent the kernel functions for water depth, significant wave height, and peak period, respectively.
[0037]
[0038] In the formula: The hyperparameter representing the control of the kernel function signal variance; These represent the characteristic scale factor hyperparameters for water depth, significant wave height, and peak period, respectively. These represent the Euclidean distances between input sample points for water depth, significant wave height, and peak period, respectively.
[0039] Furthermore, the method for obtaining the optimal ecological shoreline wave parameter prediction model in S3 specifically includes the following steps:
[0040] S31: Randomly divide the historical time series data into training and test sets according to a preset ratio;
[0041] S32: The Adam optimization algorithm, combined with the constructed loss function, trains the ecological shoreline wave parameter prediction model through the training set. The model training process is carried out within the preset maximum number of iterations, and the hyperparameters of the model are continuously updated through the backpropagation algorithm until the predetermined number of training rounds is reached, thereby obtaining the optimal ecological shoreline wave parameter prediction model.
[0042] S33: Using the optimal ecological shoreline wave parameter prediction model, the ecological shoreline wave parameters of the test set are predicted, and the prediction output features are obtained.
[0043] Furthermore, the method for constructing the loss function in S33 specifically includes:
[0044] S331: Based on the joint distribution Define the conditional probability distribution of input sample points under the condition of the induction point. Its expression is
[0045]
[0046] In the formula: Indicates the first The latent function value corresponding to each input sample point is... The amount; Indicates the first The conditional probability distribution of an input sample point under a given induced point function value;
[0047] S332: Based on the conditional probability distribution Obtain the marginal likelihood function;
[0048] And the marginal likelihood function The formula for obtaining it is
[0049]
[0050] In the formula: Represents the prior distribution of the induction point;
[0051] S333: Marginal likelihood function Apply variational distribution And using Jensen's inequality, we get:
[0052]
[0053] S334: Based on S333, construct a loss function with the variational evidence lower bound ELBO as the optimization objective, and its expression is:
[0054]
[0055] In the formula: KL divergences, representing water depth, significant wave height, and peak period respectively, are used to constrain the distance between the variational distribution and the true prior. represents the variational lower bounds of water depth, significant wave height, and peak period, respectively; y represents the actual observed value; These represent the variational distributions of water depth, significant wave height, and peak period, respectively. These represent the prior distributions of the induction points for water depth, significant wave height, and peak period, respectively. express ; Indicates information about water depth ; Indicates the significance of wave height ; Indicates the period of peak value .
[0056] Beneficial Effects: This invention provides a method for predicting ecological shoreline wave parameters based on sparse Gaussian processes and SHAP interpretation. It constructs an ecological shoreline wave parameter prediction model by introducing a sparsity-inducing point mechanism on top of Gaussian process regression. This effectively reduces the computational complexity of traditional GPR models in large-scale data processing. Traditional Gaussian processes require inverting the complete covariance matrix, and the computational complexity increases cubically with the number of samples, making it difficult to apply in big data environments. This invention uses a sparse Gaussian process regression (SGPR) structure, introducing a limited number of inducing points and combining them with a variational inference strategy to reduce the computational complexity from... Reduce to This invention significantly improves training and inference efficiency while maintaining prediction accuracy, and exhibits good scalability. By introducing a SHAP interpreter into the prediction framework, this invention makes the prediction results of the sparse Gaussian process regression model more transparent and traceable based on the SHAP interpretability analysis mechanism. Attached Figure Description
[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a flowchart of the ecological shoreline wave parameter prediction method based on sparse Gaussian process and SHAP interpretation of the present invention.
[0059] Figure 2 This is a schematic diagram showing the location of each observation station in this embodiment;
[0060] Figure 3 This is a schematic diagram of the ecological shoreline wave parameter prediction model in this embodiment;
[0061] Figure 4 This is a graph showing the comparison between the three-level variable prediction time series and the 95% confidence interval of the ecological shoreline wave parameter prediction model in this embodiment;
[0062] Figure 5 This is a scatter plot comparing the prediction accuracy of the ecological shoreline wave parameter prediction model in this embodiment;
[0063] Figure 6 This is a comparison chart of the importance of feature groups SHAP output by the model in this embodiment;
[0064] Figure 7 This is a graph showing the changes in the importance of features under different tidal conditions in this embodiment;
[0065] Figure 8 This is a graph showing the ranking of the importance of univariate features in the ecological shoreline wave parameter prediction model in this embodiment. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0067] This embodiment provides a method for predicting ecological shoreline wave parameters based on sparse Gaussian processes and SHAP interpretation. This method can achieve efficient modeling and probabilistic prediction of large-scale observation data, while also enabling the interpretability and visualization of the model decision-making process through SHAP, thus providing a more scientific basis for wave prediction and risk assessment in complex marine dynamic systems.
[0068] like Figure 1 As shown, the specific steps include:
[0069] S1: Collect and acquire historical time-series data of corresponding sample points at each observation station;
[0070] Furthermore, the historical time-series data includes environmental driving variables and their corresponding target outputs; the environmental driving variables include at least flow velocity and water level; the target outputs include at least water depth, wave height, and peak period.
[0071] Specifically, such as Figure 2As shown, multi-source ocean observation data were collected and loaded within the study area, including environmental driving variables such as current velocity and water level, as well as the corresponding target output—water depth. ), significant wave height ( ) and peak period ( The input features are derived from multiple observation stations to construct a comprehensive spatiotemporal input structure. This includes uniformly normalizing all input variables and scaling the data to the [-1,1] interval to eliminate the influence of different units and enhance the numerical stability of training. Furthermore, the training, validation, and test sets are subsequently divided according to preset indices to ensure consistent sample distribution and improve model generalization ability. This data preparation process provides a standardized and balanced input foundation for subsequent model learning.
[0072] S2: A sparse Gaussian process regression model constructed based on historical time series data;
[0073] The sparse Gaussian process regression model (SGPR) constructed in this embodiment uses... The kernel function (ν=2.5) is used as the covariance function, and the ARD (Automatic Relevance Determination) mechanism is introduced to realize the automatic discrimination of the correlation of different input dimensions;
[0074] Specifically, in this embodiment, we first assume that the observed data The latent function value f follows a Gaussian likelihood distribution: the Gaussian likelihood function is used to model noise in the observed data, and its expression is...
[0075] (1)
[0076] In the formula: y represents the observed data; f represents the latent function value; The hyperparameter represents the variance of observation noise (corresponding to random uncertainty); Let be the identity matrix, representing independent noise;
[0077] Based on this, a Gaussian process prior is introduced to construct and build a sparse Gaussian process regression model, specifically including the following steps:
[0078] S100: Establish the state function model of the Gaussian process, that is, the basic function of Gaussian process regression is:
[0079] (2)
[0080] (3)
[0081] (4)
[0082] In the formula: Indicates the sign for Gaussian process regression; These represent the latent function output (i.e., the true value function predicted by the model) corresponding to a certain input feature vector x, representing water depth, significant wave height, and peak period, respectively. This represents the input feature vector corresponding to another sample point (such as observed variables like water level, tide level, and flow velocity at each station). These represent the mean functions of water depth, significant wave height, and peak period, respectively. Let the kernel functions (covariance functions) represent water depth, significant wave height, and peak period, respectively, and let the kernel functions employ... The nucleus (ν=2.5) is in the form of
[0083] (5)
[0084] (6)
[0085] (7)
[0086] In the formula: The hyperparameter representing the control of the kernel function signal variance; These represent the characteristic scale factor hyperparameters for water depth, significant wave height, and peak period, respectively. These represent the Euclidean distances between input sample points for water depth, significant wave height, and peak period, respectively.
[0087] Based on the Gaussian process state function model combined with Gaussian likelihood distribution Obtain the joint distribution function of the sparse Gaussian process regression model, whose expression is:
[0088] (8)
[0089] In the formula: This represents the set of input sample points, which contains the feature vectors of all observation points; The corresponding observation output vector represents the observed value of the target output quantity (such as water depth, significant wave height, or peak period) actually measured at each input sample point; This indicates that the model is based on the input sample set. The corresponding latent function value, that is, the true function value generated by the Gaussian process under the condition of no observation noise; Indicates the th in the training input sample set The latent function values corresponding to each input sample point and = ; This represents the observed values given a set of input samples. With latent function value The joint probability distribution of ; This indicates that the output of the latent function is Generate observation data under the circumstances The observed likelihood function of probability relationships; Indicates that in a given set of inputs The prior probability distribution of the latent function value satisfies ; Representing kernel function The covariance matrix formed and
[0090] ;
[0091] S102: Based on Bayesian principles and the joint distribution function, obtain the predicted points through variational inference. posterior distribution on Its expression is
[0092]
[0093] In the formula: Indicates at the test point The output of the latent function at that location; This represents the predicted mean; Indicates the prediction variance; This represents the covariance matrix between the test point and the given induced point; This represents the covariance value of the test point itself; express transpose; Indicates that it contains a given test point A subset of features;
[0094] S103: By randomly selecting a limited number of input sample points from historical time series data as induction points, and approximating the posterior distribution of formula (9) based on the limited number of induction points, that is, the latent function value corresponding to the input sample points. With the induced point function value joint distribution Its expression is
[0095]
[0096] In the formula: This represents the latent function value corresponding to the input sample point; This represents the potential function value at the induced point location; Describe the set of induced points and ; K represents the covariance matrix between each input sample point; XZK represents the covariance matrix between the input sample points and the induced points; ZX K represents XZ The transpose of K; ZZ This represents the covariance matrix between the induced points; This represents a Gaussian probability distribution. In this embodiment, after obtaining the complete prediction distribution (Equations (9)-(11)), in order to reduce computational complexity and maintain the probability consistency of the Gaussian process, a finite set of induction points is introduced. By approximating the posterior distribution through a finite number of induced points, the prior definition of the Gaussian process (Equation (2)–(4)) can be obtained, and the function values on any set of input sample points all follow a multivariate Gaussian distribution.
[0097] S104: Based on steps S101 to S103, the sparse Gaussian process regression model (SGPR) is constructed. This embodiment selects a limited number of induction points and combines them with a variational inference strategy, enabling the model to... The SGPR model efficiently approximates the posterior distribution of the original Gaussian process with low complexity, thus significantly reducing the computational burden and adapting to the training requirements of large sample scenarios. The SGPR model can simultaneously output the predicted mean and the predicted variance, making the prediction results probabilistic.
[0098] Based on the sparse Gaussian process regression model, a prediction model for ecological shoreline wave parameters based on the sparse Gaussian process is constructed, such as... Figure 3 It includes an input module, a three-level cascaded prediction module, and an output module; the input module is used to input historical time series data into the three-level cascaded prediction module.
[0099] The three-level cascaded prediction module includes a first SGPR stage model, a second SGPR stage model, and a third SGPR stage model connected in sequence. The first SGPR stage model is used to obtain predicted water depth features by taking environmental driving variables as feature inputs. The second SGPR stage model is used to obtain predicted wave height features by taking environmental driving variables and predicted water depth features as inputs. The third SGPR stage model is used to obtain predicted peak periodic features by taking environmental driving variables, predicted water depth features, and predicted wave height features as feature inputs. In this embodiment, the first SGPR stage model, the second SGPR stage model, and the third SGPR stage model are formula models selected according to actual needs in the sparse Gaussian process regression model constructed based on historical time series data.
[0100] The output module is used to output predicted output features that include predicted water depth features, predicted wave height features, and predicted peak period features.
[0101] In this embodiment, a three-level cascaded prediction framework is designed: the first SGPR stage model uses environmental features such as flow velocity and water level as input to predict water depth; the second SGPR stage model uses the predicted water depth as an additional input feature, and inputs it together with environmental variables to predict significant wave height; the third SGPR stage model further uses the predicted water depth and wave height as input to predict peak period. Through this hierarchical structure, the dependency between different physical variables is explicitly reflected in the model, realizing causal modeling from local features to overall dynamics, and improving the physical consistency and scientific rationality of the prediction.
[0102] S3: Train the ecological shoreline wave parameter prediction model based on historical time series data to obtain the optimal ecological shoreline wave parameter prediction model, and obtain the prediction output features through the optimal ecological shoreline wave parameter prediction model;
[0103] The method for obtaining the optimal ecological shoreline wave parameter prediction model in this embodiment specifically includes the following steps:
[0104] S31: Randomly divide the historical time series data into training and test sets according to a preset ratio;
[0105] S32: The Adam optimization algorithm, combined with the constructed loss function, trains the ecological shoreline wave parameter prediction model using the training set. The training process will be carried out within the preset maximum number of iterations until the predetermined number of training rounds is reached. During this process, the hyperparameters of the model are continuously updated through the backpropagation algorithm. As the model training iterates, the loss value of the loss function gradually decreases and tends to stabilize. Finally, through this process, a model with optimal hyperparameters is obtained, thereby obtaining the optimal ecological shoreline wave parameter prediction model.
[0106] S33: Using the optimal ecological shoreline wave parameter prediction model, the ecological shoreline wave parameters of the test set are predicted, and the prediction output features are obtained.
[0107] Specifically, the loss function is constructed as follows:
[0108] S331: Based on the joint distribution Define the conditional probability distribution of input sample points under the condition of the induction point. Its expression is
[0109]
[0110] In the formula: Indicates the first The latent function value corresponding to each input sample point is... The amount; Indicates the first The conditional probability distribution of an input sample point under a given induced point function value; This indicates that the model is trained on the set of input sample points. The corresponding latent function output vector reflects the ideal function value under conditions of no observation noise; Indicates the first Input of training samples The corresponding single latent function value is specifically: The amount; Indicating the set of induction points The function outputs and ; This represents the number of inducing points, which is usually much smaller than the number of training samples. In this embodiment, the induced point is used to achieve a low-rank approximation of the posterior distribution while preserving the statistical properties of the Gaussian process, thereby reducing computational complexity. Indicates the first The conditional probability distribution of a training sample under the given function value of the inducing point describes the constraint relationship between the inducing point and the potential output of the sample.
[0111] During model training, the Adam optimization algorithm is used, with the variational evidence lower bound (ELBO) as the optimization objective for hyperparameter updates. In the optimization process, the SGPR model is trained by maximizing the evidence lower bound, and the loss function is optimized by minimizing the negative ELBO. This training strategy ensures stable convergence of the training. This embodiment combines the sparse approximation of equations (12) and (13) and employs a variational inference framework to analyze the posterior distribution of the induction point. To make an approximation, a variational distribution is introduced. It can transform the maximization of log marginal likelihood into an optimization problem of evidence lower bound (ELBO);
[0112] S332: Based on the conditional probability distribution That is, by combining the prior (12) and the sparse conditional distribution (13), the marginal likelihood function of the corresponding training data is obtained;
[0113] And the marginal likelihood function The formula for obtaining it is
[0114]
[0115] In the formula: Represents the prior distribution of the induction point;
[0116] S333: Marginal likelihood function Apply variational distribution And using Jensen's inequality, we get:
[0117]
[0118] S334: Based on formula (15), construct a loss function with the variational evidence lower bound ELBO as the optimization objective, and the ELBO of the three-stage model can be obtained, the expression of which is:
[0119]
[0120] In the formula: KL divergences, representing water depth, significant wave height, and peak period respectively, are used to constrain the distance between the variational distribution and the true prior. represents the variational lower bounds of water depth, significant wave height, and peak period, respectively; y represents the actual observed value; These represent the variational distributions of water depth, significant wave height, and peak period, respectively. These represent the prior distributions of the induction points for water depth, significant wave height, and peak period, respectively. express ; Indicates information about water depth ; Indicates the significance of wave height ; Indicates the period of peak value ;
[0121] This allows us to obtain the total loss function for training the ecological shoreline wave parameter prediction model, which is expressed as follows: In the formula: Lower bounds of strain respectively ;
[0122] S34: Using the optimal ecological shoreline wave parameter prediction model, predict the ecological shoreline wave parameters of the test set and obtain the prediction output features.
[0123] This embodiment also includes testing and performance evaluation of the trained model. Through inverse normalization, the prediction results are restored to the true physical dimensions, and multiple performance metrics are calculated, including root mean square error (RMSE) and coefficient of determination. Systematic bias and the index of dispersion (SI) are used to evaluate the model's accuracy and bias control capabilities from different perspectives. Meanwhile, the variance term in the model output is used to characterize prediction uncertainty. Based on the posterior distribution, cognitive uncertainty, noise uncertainty, and their sum are extracted separately. The former reflects the model's confidence level in the unknown parameters, while the latter reflects the noise impact of the data itself. Based on the variance term in the model output, the total uncertainty is calculated:
[0124]
[0125] In the formula: Indicates total uncertainty; This indicates cognitive uncertainty; This indicates random uncertainty;
[0126] In this embodiment, the 95% confidence interval (CI) is calculated based on the confidence coefficient z=1.96 of the standard normal distribution:
[0127]
[0128] In the formula: This represents the 95% confidence interval for the predicted result; 1.96 represents the confidence interval coefficient, corresponding to a 95% confidence level under a normal distribution. This represents the model's predicted mean;
[0129] Furthermore, the Prediction Interval Coverage (PIC) and Mean Prediction Interval Width (MPIW) are calculated. The formula for calculating the Prediction Interval Coverage is as follows:
[0130]
[0131] In the formula: N Indicates the total number of test samples; This indicates an indicator function that takes the value 1 if the condition is true and 0 otherwise. The numbers represent the water depth, significant wave height, and peak period, respectively. i The actual observed values of each sample; The numbers represent the water depth, significant wave height, and peak period, respectively. i The predicted mean of each sample; The numbers represent the water depth, significant wave height, and peak period, respectively. i The predicted standard deviation for each sample;
[0132] The formula for calculating the average prediction interval width (MPIW) is:
[0133]
[0134] like Figures 4 to 5 As shown, this comprehensively reflects the reliability, range width, and uncertainty level of the model's prediction results. A high PIC and a moderate MPIW indicate that the model not only predicts accurately but also reflects data fluctuation characteristics within a reasonable confidence range.
[0135] S4: The importance of SHAP features for each observation station and the global interpretation are calculated based on the predicted output features by the pre-set SHAP interpreter, so as to enhance the interpretability of the ecological shoreline wave parameter prediction model and realize the prediction of ecological shoreline wave parameters based on sparse Gaussian process and SHAP interpretation.
[0136] To reveal the mechanism by which internal features of the model influence prediction results, the SHAP method is used to perform interpretability analysis on the trained model. This method is based on game theory, calculating the influence of each feature on different input samples... The value is used to quantify its marginal contribution to the prediction results;
[0137] Specifically, the formula for calculating SHAP is as follows:
[0138] (27)
[0139] In the formula: Prediction represents the model's predicted output value; basevalue represents the model's average predicted value under the condition of no feature input, i.e., the baseline output; SHAP value of feature i Indicates the first i The marginal contribution of each feature to the prediction result, i.e. value;
[0140] The formula for calculating the value is:
[0141] (28)
[0142] (29)
[0143] (30)
[0144] In the formula: S Represents a subset of features; A This represents the input feature set (containing all features). These respectively represent features that do not include water depth, significant wave height, and peak period. i A subset of features; This represents the output of the model when only a subset of features S is used; These represent the characteristics of water depth, significant wave height, and peak period, respectively. i Join a subset S The output of the post-model;
[0145] In this embodiment, a prediction function corresponding to water depth, significant wave height, and peak period is constructed, and the SHAP value is calculated using a Kernel Explainer. For example... Figures 6 to 8For ease of explanation, this embodiment groups features according to physical location or function, such as Light, Ship, Delaware, Lewes, Cape, and tide level. By calculating the average absolute SHAP value of each group of features, the importance ranking of the groups is obtained, thereby revealing the overall impact of each station's features on the model output.
[0146] To study the model's response patterns under different marine environments, this embodiment performs hierarchical statistical analysis and visualization of the SHAP results. First, based on the 33% and 66% quantiles of tide levels in the training set, the samples are divided into three intervals: Low WL, Mid WL, and High WL. The mean SHAP values of different feature groups within each tide level interval are then calculated to reveal the changing trends in the model's sensitivity to various input variables under different tide levels. Second, when the data includes timestamp information, the system automatically divides the data into seasons (DJF, MAM, JJA, SON) by month, calculates and compares the average contribution of each feature group across the four seasons, and analyzes the seasonal response differences of the model. This multi-layered interpretation not only improves the model's transparency but also reveals the physical driving force of factors such as tide level and current velocity on wave evolution under different hydrodynamic conditions.
[0147] The method described in this embodiment integrates probabilistic prediction, uncertainty quantification, and interpretable analysis. The SGPR model, while ensuring prediction accuracy, effectively quantifies prediction reliability by outputting confidence intervals; while the SHAP framework provides a transparent interpretation path for the results, enabling the model to not only "predict accurately" but also "explain clearly." Theoretically, this method combines the Bayesian inference characteristics of Gaussian processes with the flexibility of data-driven models. In application, it can provide reliable scientific support for ecological shoreline wave early warning, storm surge disaster prevention and control, and port and waterway management, demonstrating significant engineering and academic research value. The method described in this embodiment enables wave prediction data-driven methods to maintain high efficiency in processing large-scale data while introducing uncertainty quantification mechanisms and predictive interpretability analysis. By combining the Sparse Gaussian Process Regression (SGPR) and SHAP (SHapley Additive Explanations) interpretation framework, it solves the problems of computational complexity, prediction accuracy, and interpretability in traditional wave prediction methods. This innovative method can not only provide high-precision wave predictions but also deeply reveal the physical mechanisms behind the predictions, promoting the advancement of intelligent and data-driven decision-making in the field of marine engineering.
[0148] In summary, the beneficial effects of the method described in this embodiment are as follows:
[0149] (1) Improved computational efficiency: The method described in this embodiment introduces a sparsity-induced point mechanism on the basis of Gaussian process regression, which effectively reduces the computational complexity of the traditional GPR model in large-scale data processing from the perspective of algorithm mechanism. The traditional Gaussian process requires the inversion of the complete covariance matrix, and the computational amount increases cubically with the number of samples, making it difficult to apply in a big data environment. This invention adopts the sparse Gaussian process regression (SGPR) structure, and by introducing a limited number of induced points and combining them with variational inference strategy, the computational complexity is reduced from... Reduce to This method significantly improves training and inference efficiency while maintaining prediction accuracy, and exhibits good scalability. Experimental results show that the method described in this embodiment maintains high computational performance and stable convergence characteristics when processing ocean observation data from multiple stations and at multiple time scales, making it suitable for large-scale modeling and real-time prediction tasks in complex ocean dynamic environments. (n: number of training samples, m: number of induction points)
[0150] (2) Quantification of Prediction Accuracy and Uncertainty: The method described in this embodiment utilizes the Bayesian inference characteristics of sparse Gaussian process regression to achieve high-precision prediction while simultaneously providing a quantitative description of prediction uncertainty. The model outputs the prediction mean and simultaneously provides a variance estimate, reflecting the confidence and reliability of the prediction results. Based on the posterior distribution, the model distinguishes between cognitive uncertainty and noise uncertainty; the former reflects the cognitive ambiguity of the model structure and parameters, while the latter reflects observation errors and external disturbances. Through the 95% confidence interval calculated using the standard confidence coefficient, as well as indicators such as prediction interval coverage and average prediction interval width, the reliability of the prediction results can be quantitatively assessed. This probabilistic output method enables the model to not only provide accurate point prediction results but also to provide the corresponding uncertainty range, offering reliable probabilistic support for risk assessment and scientific decision-making in the process of ocean wave evolution.
[0151] (3) Significantly Enhanced Model Interpretability and Physical Consistency: The SHAP interpretability analysis mechanism introduced in this embodiment into the prediction framework makes the prediction results of the sparse Gaussian process regression model more transparent and traceable. Through the SHAP framework, the model can quantify the marginal contribution of each input feature to the prediction results, revealing the driving role of physical factors such as flow velocity and water level on wave characteristics under different scenarios. This invention further performs statistical analysis on the SHAP results through feature grouping and hierarchical analysis (including division by tidal range and season), thereby revealing the response variation law of the model under different hydrodynamic and seasonal conditions. This interpretability mechanism not only improves the transparency and credibility of the model, but also ensures the rationality of the prediction results at the physical level, making the model output have a clear causal structure and scientific explanatory value. Compared with traditional black-box deep learning methods, the method described in this embodiment significantly enhances the interpretability and physical consistency of the model, making it easier for researchers and engineering decision-makers to understand and verify the prediction results.
[0152] (4) Improving Model Robustness: The method described in this embodiment fully considers the differences in multi-source input features and the impact of noise during data processing and model inference. Through unified normalization, input variables of different dimensions are standardized in numerical space, effectively reducing the scale bias between features and improving the stability of model training. In terms of model structure design, variational inference method is adopted, which can suppress overfitting under limited sample conditions and improve the model's adaptability to noise and abnormal samples. In addition, by simultaneously outputting the prediction mean and uncertainty index, the model can use dynamic indicators such as confidence interval and prediction interval coverage to reflect the prediction reliability in real time, thus maintaining stable performance even in complex marine environments such as input disturbances, missing observations, or drastic tidal changes. Experiments show that this method has excellent generalization ability and anti-interference characteristics at different stations and at different time scales, and can be stably applied in actual marine monitoring and forecasting tasks.
[0153] The method described in this embodiment can be widely applied in scenarios such as coastal zone management, disaster prevention and mitigation, meteorological route planning, offshore equipment operation and maintenance, and observation site optimization. By achieving accurate prediction of wave evolution and quantitative analysis of key influencing factors, it can provide scientific support for shoreline protection and disaster early warning, provide decision-making basis for route safety and scheduling optimization, and assist in assessing the operational risks of offshore equipment and their causes, guiding the layout and optimization of offshore monitoring networks. The widespread application of this method is of great significance for improving the comprehensive disaster prevention capabilities of nearshore waters, ensuring shipping safety and reliable facility operation, and also provides an embeddable intelligent prediction module for intelligent marine observation systems, offshore risk management platforms, and marine engineering decision support systems, possessing good promotional value and application potential.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An ecological shoreline sea wave parameter prediction method based on sparse Gaussian process and SHAP explanation, characterized in that, Specifically, the following steps are included: S1: Collect and acquire historical time-series data of corresponding sample points at each observation station; Furthermore, the historical time-series data includes environmental driving variables and their corresponding target outputs; Environmental driving variables include at least flow velocity and water level; target outputs include at least water depth, wave height, and peak period. S2: Construct a sparse Gaussian process regression model based on historical time series data to establish an ecological shoreline wave parameter prediction model based on sparse Gaussian process, which includes an input module, a three-level cascaded prediction module and an output module. The input module is used to input historical time series data into the three-level cascaded prediction module; The three-level cascaded prediction module includes a first SGPR stage model, a second SGPR stage model, and a third SGPR stage model connected in sequence. The first SGPR stage model is used to obtain predicted water depth features by taking environmental driving variables as feature inputs. The second SGPR stage model is used to obtain predicted wave height features by taking environmental driving variables and predicted water depth features as inputs. The third SGPR stage model is used to obtain predicted peak periodic features by taking environmental driving variables, predicted water depth features, and predicted wave height features as feature inputs. The output module is used to output predicted output features that include predicted water depth features, predicted wave height features, and predicted peak period features. S3: Train the ecological shoreline wave parameter prediction model based on historical time series data to obtain the optimal ecological shoreline wave parameter prediction model, and obtain the prediction output features through the optimal ecological shoreline wave parameter prediction model; S4: The importance of SHAP features for each observation station and the global interpretation are calculated based on the predicted output features by the pre-set SHAP interpreter, so as to enhance the interpretability of the ecological shoreline wave parameter prediction model and realize the prediction of ecological shoreline wave parameters based on sparse Gaussian process and SHAP interpretation.
2. The method of claim 1, wherein, In S2, the first SGPR stage model, the second SGPR stage model, and the third SGPR stage model are sparse Gaussian process regression models constructed based on historical time series data. Furthermore, the method for constructing a sparse Gaussian process regression model specifically includes the following steps: S100: Establish the state function model of the Gaussian process; S101: Assume the observation data of the sparse Gaussian process regression model satisfies the Gaussian likelihood distribution with the latent function value f The expression is: where y represents the observation data; f represents the latent function value; represents the observation noise variance, i.e., a hyperparameter; represents a unit noise matrix; represents a Gaussian probability distribution; According to the Gaussian process state function model combined with the Gaussian likelihood distribution , the joint distribution function of the sparse Gaussian process regression model is obtained, and the expression is In the formula: This represents the set of input sample points, which contains the feature vectors of all observation points; The corresponding observation output vector represents the observed value of the target output quantity actually measured at each input sample point; This indicates that the model is based on the input sample set. The corresponding latent function value; Indicates the first training input sample set. The latent function values corresponding to each input sample point and = ; These represent the latent function values corresponding to a certain input feature vector x, namely water depth, significant wave height, and peak period, respectively. This represents the observed values given a set of input samples. With latent function value The joint probability distribution of ; This indicates that the output of the latent function is Generate observation data under the circumstances The observed likelihood distribution of the probability relationship; Indicates that in a given set of inputs The prior probability distribution of the latent function value and satisfying ; Representing kernel function The covariance matrix formed and ; wherein: respectively represent the kernel functions for water depth, significant wave height, and peak period, respectively; S102: Obtain the posterior distribution on the prediction point based on the Bayesian principle according to the joint distribution function whose expression is In the formula: Indicates at the test point The output of the latent function at that location; This represents the predicted mean; Indicates the prediction variance; This represents the covariance matrix between the test point and the given induced point; This represents the covariance value of the test point itself; express transpose; Indicates that test points are included. A subset of features; S103: By randomly selecting a limited number of input sample points from historical time-series data as induction points, and approximating the predicted points based on the limited number of induction points. The posterior distribution on the input sample points, i.e., the latent function values corresponding to the input sample points. With the induced point function value joint distribution Its expression is In the formula: This represents the latent function value corresponding to the input sample point; This represents the potential function value at the induced point location; Describe the set of induced points and ; K represents the covariance matrix between each input sample point; XZ K represents the covariance matrix between the input sample points and the induced points; ZX K represents XZ The transpose of K; ZZ This represents the covariance matrix between the induced points; Represents a Gaussian probability distribution; This leads to the construction of a sparse Gaussian process regression model.
3. The method of claim 2, wherein the method is characterized by, The state function model of the Gaussian process constructed in S100 has the following expression: In the formula: Indicates the sign for Gaussian process regression; These represent the latent function values corresponding to a certain input feature vector x, namely water depth, significant wave height, and peak period, respectively. This represents the input feature vector corresponding to another sample point; These represent the mean functions of water depth, significant wave height, and peak period, respectively. Let represent the kernel functions for water depth, significant wave height, and peak period, respectively. In the formula: The hyperparameter representing the control of the kernel function signal variance; These represent the characteristic scale factor hyperparameters for water depth, significant wave height, and peak period, respectively. These represent the Euclidean distances between input sample points for water depth, significant wave height, and peak period, respectively.
4. The method of claim 3, wherein, The method for obtaining the optimal ecological shoreline wave parameter prediction model in S3 includes the following steps: S31: Randomly divide the historical time series data into training and test sets according to a preset ratio; S32: The Adam optimization algorithm, combined with the constructed loss function, trains the ecological shoreline wave parameter prediction model through the training set. The model training process is carried out within the preset maximum number of iterations, and the hyperparameters of the model are continuously updated through the backpropagation algorithm until the predetermined number of training rounds is reached, thereby obtaining the optimal ecological shoreline wave parameter prediction model. S33: Using the optimal ecological shoreline wave parameter prediction model, the ecological shoreline wave parameters of the test set are predicted, and the prediction output features are obtained.
5. The method of claim 4, wherein, The method for constructing the loss function in S33 specifically includes: S331: defining a conditional probability distribution of input sample points given the induced point condition whose expression is In the formula: Indicates the first The latent function value corresponding to each input sample point is... The amount; Indicates the first The conditional probability distribution of an input sample point under a given induced point function value; S332: Obtain the edge likelihood function according to the conditional probability distribution , the edge likelihood function and the edge likelihood function The acquisition formula is wherein: denotes the prior distribution of the inducing points; S333: marginal likelihood function Applying the variational distribution And using the Jensen inequality we get: S334: Based on S333, construct a loss function with the variational evidence lower bound ELBO as the optimization objective, and its expression is: In the formula: KL divergences, representing water depth, significant wave height, and peak period respectively, are used to constrain the distance between the variational distribution and the true prior. represents the variational lower bounds of water depth, significant wave height, and peak period, respectively; y represents the actual observed value; These represent the variational distributions of water depth, significant wave height, and peak period, respectively. These represent the prior distributions of the induction points for water depth, significant wave height, and peak period, respectively. express ; Indicates information about water depth ; Indicates the significance of wave height ; Indicates the period of peak value .