Multi-scale space-time fusion water quality prediction and anti-counterfeiting method based on dynamic graph neural network

By employing a multi-scale spatiotemporal fusion method using dynamic graph neural networks, the problems of time-varying correlation and multi-scale features in water quality prediction were solved, achieving high-precision prediction and data anti-counterfeiting, thereby improving the reliability and security of watershed water quality management.

CN120950856APending Publication Date: 2025-11-14GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510839540.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing water quality prediction models are ill-suited to the time-varying characteristics of station relationships within a watershed, cannot effectively capture the multi-scale features of water quality data, and lack data anti-counterfeiting verification mechanisms, resulting in insufficient prediction accuracy and security.

Method used

A multi-scale spatiotemporal fusion method based on dynamic graph neural networks is adopted. By combining variational mode decomposition, dynamic graph construction, graph convolution and long short-term memory networks, the spatiotemporal correlation between monitoring stations is adaptively learned. Combined with residual fusion and anti-counterfeiting verification modules, high-precision prediction and data authenticity determination are achieved.

Benefits of technology

It significantly improves the accuracy and robustness of water quality prediction, can dynamically capture time-varying correlations between stations, deeply decouples the multi-scale characteristics of water quality data, and realizes the authenticity verification of data based on high-precision prediction, providing reliable support for watershed water quality management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950856A_ABST
    Figure CN120950856A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale space-time fusion water quality prediction and anti-counterfeiting method based on a dynamic graph neural network. Comprising the following steps: 1) collecting water quality index hour data of a plurality of monitoring stations in a drainage basin; 2) decomposing the data into a plurality of intrinsic mode functions through variational mode decomposition; 3) constructing a dynamic graph neural network spatial feature extraction module, and generating a discrete dynamic graph structure; 4) constructing a multi-scale time feature extraction module, and synchronously capturing short-term fluctuation and long-term trend; 5) designing a residual fusion mechanism to integrate the spatial-temporal characteristics, and outputting a water quality prediction result through a full connection layer; and 6) calculating a path distance between the input data and a prediction result through a dynamic time warping algorithm, and comparing residual distribution by combining K-S to realize authenticity discrimination of the input data. The method can fully excavate the spatial and temporal characteristics of the basin water quality under the condition that the geographical spatial distribution of the sites is unknown, improves the prediction precision, carries out the authenticity recognition of the water quality data of an unknown source, and prevents the data from being tampered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application involves the interdisciplinary fields of information security, environmental big data and deep learning, specifically a multi-scale spatiotemporal fusion water quality prediction and anti-counterfeiting method based on dynamic graph neural networks. Background Technology

[0002] Water quality forecasting, as a core component of water resource management, is of great significance for pollution prevention and ecological governance. Traditional water quality models mainly rely on physical mechanisms or statistical learning methods, but they face significant limitations: First, watershed water quality changes exhibit complex spatiotemporal dependencies, with hydrological relationships between monitoring stations constantly evolving with dynamic processes such as rainfall and runoff. Existing graph neural network methods typically construct static graph structures based on predefined geographical topologies, making it difficult to adapt to the time-varying characteristics of station relationships in actual watersheds. Second, water quality data is highly non-stationary, with pollutant concentrations exhibiting multi-scale fluctuations under the influence of rainfall erosion and tidal effects. Single time-series models struggle to simultaneously capture short-term mutations and long-term trends. Furthermore, with the widespread adoption of IoT monitoring devices, the risks of data tampering and abnormal inputs are increasingly prominent. However, existing technologies primarily focus on improving prediction accuracy, lacking mechanisms to verify the authenticity of input data. In recent years, deep learning technology has made some progress in the field of water quality forecasting. Spatiotemporal Graph Neural Networks (STGNNs) partially solve the problem of spatial correlation modeling by fusing graph convolutions and recurrent neural networks. However, their topology often relies on prior knowledge such as geographical distance or water flow direction, making it unable to adaptively learn implicit dynamic correlations between stations. In time series processing, although models like LSTM and TCN can extract temporal features, they are insufficient in mining multi-scale characteristics: a single model struggles to balance high-frequency noise and low-frequency trends, and the hierarchical feature fusion mechanism is not perfect. Regarding data security, traditional anomaly detection methods (such as threshold alarms and statistical tests) are mostly aimed at single-point anomalies, with limited ability to identify systematic data forgery. Therefore, a new method is needed that can simultaneously meet three requirements: first, breaking through predefined topological constraints to dynamically capture time-varying correlations between stations; second, deeply decoupling the multi-scale features of water quality data to achieve collaborative modeling of fluctuations and trends; and third, establishing a data anti-counterfeiting mechanism based on high-precision prediction. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a multi-scale spatiotemporal fusion method for water quality prediction and anti-counterfeiting based on dynamic graph neural networks. This method can achieve high-precision prediction and data authenticity verification, providing reliable technical support for watershed water quality management.

[0004] The specific technical solution of this invention is as follows:

[0005] A multi-scale spatiotemporal fusion water quality prediction and anti-counterfeiting method based on dynamic graph neural networks mainly includes the following steps:

[0006] 1) Data Acquisition: Hourly water quality and hydrological data were collected from multiple monitoring stations;

[0007] 2) Multiscale temporal decomposition: The temporal data of one target pollutant in step 1) is decomposed into several intrinsic mode functions (imfs) through variational mode decomposition (VMD);

[0008] 3) Dynamic Graph Construction: Input the raw data collected in step 1), and generate a discrete dynamic graph adjacency matrix A based on the GRU gating mechanism and channel attention learning node embedding. t Completely free from predefined topological dependencies, its mathematical expression is formula (1):

[0009]

[0010] in σ represents the sigmoid function, and τ represents temperature, used to manage the interpolation process between discrete distributions and continuous classification densities. Furthermore, The connection probability between node i and node j at time t is represented by the following process:

[0011] 3-1) Edge weight calculation: Edge weight represents the association derived from the topology level and node level input. Its main calculation formulas are shown in formula (2) and formula (3):

[0012]

[0013] Where E t E represents the dynamic semantic information extracted from the input X using a 1×1 standard convolution. t =Conv(X), where E is a trainable node embedding;

[0014] 3-2) Gated Recurrent Unit (GRU): The Gated Recurrent Unit (GRU) is used as a method to capture the evolutionary representation of nodes, as shown in Equations (4), (5), (6), and (7):

[0015] r t =σ(W r E+U r E t (4),

[0016] z t =σ(W z E+U z E t (5),

[0017]

[0018] Where W and U represent trainable parameter matrices, ⊙ represents the Hadamard product, σ represents the sigmoid function, tanh represents the hyperbolic tangent function, and h t This represents the node evolution at time t. The multi-head variation based on formula (8) further enhances the stability of the learning process.

[0019]

[0020] Where LN represents LayerNorm, and These correspond to two different nonlinear mappings for the k-th head, d k Let represent the dimension of each head, and <, ·> represent the dot product operation. Dropout is applied to the edge weights to generate a more generalized trend. Finally, the final edge weights are obtained by summing the weights calculated for the k subspaces. The specific calculation is shown in formula (9):

[0021]

[0022] Where f s,4 (·) are shared trainable weights;

[0023] 3-3) Channel Attention (SENet): Channel attention adaptively determines the proportion of the three learned edge weights, calculated as shown in formula (10):

[0024]

[0025] Where F ex and F sq The symbols represent compression and activation operations, respectively, and the dot (·) indicates channel multiplication.

[0026] 3-4) Connection predictor: The connection predictor performs convolution along the channel size and outputs a scalar indicating the connection probability between node i and node j. Its calculation process is shown in formula (11):

[0027]

[0028] Where ReLU represents the activation function;

[0029] 4) Multi-scale spatiotemporal feature extraction: Multi-scale spatial features are extracted by performing graph convolution on the dynamic graph structure obtained in step 3), and multi-scale temporal features are extracted from the IMF obtained in step 2). The specific steps are as follows:

[0030] 4-1) Multi-scale spatial feature extraction: Graph Convolutional Network (GCN) is used to process dynamic graphs, and the update formula is shown in formula (12):

[0031]

[0032] Where H represents the feature matrix of the graph convolutional layer, l is the number of graph convolutional layers, and W is the weight matrix. It is the adjacency matrix A plus the identity matrix I. yes The degree matrix, where σ is the activation function;

[0033] 4-2) Multi-scale temporal feature extraction: The imfs obtained from variational mode decomposition in step 2) are used as input, and each imf is processed separately, i.e. Where k is the number of mode decompositions, which is a hyperparameter. This indicates the temporal features extracted after processing the i-th imf. The following section will explain f in detail. i (·) The process of feature mining. First, the imf i The input is fed into the TCN for shallow temporal feature extraction, and the specific process is shown in formula (13):

[0034]

[0035] Here, a four-layer TCN was set up to obtain shallow features at different scales. Then The input is then used for deep temporal feature extraction via LSTM, as shown in formulas (14), (15), (16), (17), (18), and (19).

[0036]

[0037] Therefore, IMF i The features obtained through further mining using LSTM can be represented as follows:

[0038] 5) Residual fusion and prediction output: Residual fusion is based on the residual idea and uses multi-head self-attention to fuse features. Finally, the fused features are mapped to the final prediction result through a fully connected layer. The specific process is shown in formulas (20), (21), (22), and (23):

[0039]

[0040] H = Concat(Att1, ..., Att) h W O (twenty two),

[0041] P = FC(H) (23),

[0042] in All are trainable parameters, dmodel =D model *h, where i represents the i-th attention head;

[0043] 6) Anti-counterfeiting verification: The path distance between the input data and the predicted result is calculated using Dynamic Time Warping (DTW), and the residual distribution is compared using the KS test. Let the input water quality data sequence be X = x1, x2, ..., x... T The model predicts the sequence Y = y1, y2, ..., y. T The minimum normalized path distance between them is D dtw The mathematical expression for its calculation is shown in formula (24):

[0044]

[0045] Where π represents the optimal path for aligning the two sequences, and d(·) is the Euclidean distance metric. The prediction residual ε is calculated simultaneously. t =|x t -y t The empirical distribution F of |} ε (z), and compared with the baseline normal residual distribution F ref (z) Perform the KS test, and the specific calculation process is shown in formula (25):

[0046]

[0047] Ultimately, a joint decision mechanism is used: when the path distance D... dtw >θ1 and KS statistic D ks When the value is greater than θ2, the data is determined to be non-real data.

[0048] This technical solution significantly improves water quality prediction performance through multi-module collaborative innovation: First, the VMD module is used to decouple the multi-scale characteristics of pollutant sequences, solving the problem of feature extraction from non-stationary data; the dynamic graph construction module based on GRU enhancement, combined with the channel attention mechanism, adaptively learns the implicit spatiotemporal correlations between monitoring stations, completely eliminating the limitations of predefined topologies; a TCN-LSTM cascaded time series module is designed to simultaneously capture short-term fluctuations and long-term evolution trends; an innovative multi-head self-attention residual fusion module is introduced to deeply integrate spatiotemporal features and enhance model interpretability; finally, the DTW and KS verification anti-counterfeiting modules are used to verify the authenticity of input data while outputting high-precision prediction results.

[0049] This method deeply mines the dynamic spatiotemporal characteristics of water quality, significantly improves the robustness of cross-site, multi-step prediction, and provides technical support for watershed water quality management that combines prediction accuracy and safety. Attached Figure Description

[0050] Figure 1 This is a flowchart illustrating the method in the embodiment;

[0051] Figure 2 This is an architecture diagram of the prediction model in the embodiment.

[0052] Figure 3 This is an architecture diagram of the dynamic graph construction method in the embodiment;

[0053] Figure 4 This is an architecture diagram of the multi-scale temporal feature extraction module in the embodiment; Detailed Implementation

[0054] The technical solutions of the present invention will be further described below with reference to the accompanying drawings, but this is not intended to limit the present invention.

[0055] Example:

[0056] Reference Figure 1 A multi-scale spatiotemporal fusion water quality prediction and anti-counterfeiting method based on dynamic graph neural networks includes the following steps:

[0057] 1) Data Acquisition: Hourly water quality and hydrological data were collected from multiple monitoring stations. In this example, the Lijiang River basin, located in Guilin City, Guangxi Zhuang Autonomous Region, China, was selected as the study area. The data came from the Guilin Ecological and Environmental Protection Center's multi-point continuous online monitoring stations, collecting data from six stations: Yangshuo, Jiaozhou, Guanyan, Mopanshan, Qiaotou, and Darongjiang. As shown in Table 1, the collected data covered hydrological and meteorological data as well as pollutant data. Hydrological and meteorological data included hydrological parameters (WT), turbidity (TB), pH, and precipitation (PCP), while pollutant data included dissolved oxygen (DO), ammonia nitrogen (NH3-N), total nitrogen (TN), and total phosphorus (TP). The core indicator DO was selected as the model's prediction target. The average hourly data from 2022-2024 was used as the baseline dataset, which was divided into training, validation, and test sets with proportions of 60%, 20%, and 20%, respectively. The training set was used to train the model, the validation set was used for hyperparameter search, and the test set was used to evaluate the model's performance.

[0058] Table 1. Variable Data

[0059]

[0060] 2) Multi-scale time series decomposition: The time series data of a target pollutant in step 1) is decomposed into several intrinsic mode functions (IMFs) using variational mode decomposition (VMD). In this example, the target pollutant is DO. The VMD method seeks the optimal solution of the variational model in a non-recursive manner, achieving multi-scale decomposition of the DO time series. Unlike the iterative screening and stripping used in traditional empirical mode decomposition (EMD) algorithms, VMD shifts the sequence decomposition process to a variational model. Through adaptive adjustment in the frequency domain, it ensures that each IMF closely revolves around its center frequency, effectively separating the components. This not only preserves the center frequency characteristics of each component but also provides higher resolution and accuracy, making it particularly suitable for processing complex non-stationary time series data. Based on variational principles, VMD provides richer frequency information than traditional models.

[0061] 3) Dynamic Graph Construction: Input the raw data collected in step 1), and generate a discrete dynamic graph adjacency matrix A based on the GRU gating mechanism and channel attention learning node embedding. t Completely free from predefined topological dependencies, its mathematical expression is formula (1):

[0062]

[0063] in σ represents the sigmoid function, and τ represents temperature, used to manage the interpolation process between discrete distributions and continuous classification densities. Furthermore, The connection probability between node i and node j at time t is represented by the following process:

[0064] 3-1) Edge weight calculation: Edge weight represents the association derived from the topology level and node level input. Its main calculation formulas are shown in formula (2) and formula (3):

[0065] ξ i,j = <E i E j > (2),

[0066]

[0067] Where E t E represents the dynamic semantic information extracted from the input X using a 1×1 standard convolution. t =Conv(X), where E is a trainable node embedding;

[0068] 3-2) Gated Recurrent Unit (GRU): The Gated Recurrent Unit (GRU) is used as a method to capture the evolutionary representation of nodes, as shown in Equations (4), (5), (6), and (7):

[0069] r t =σ(W r E+U r E t (4),

[0070] z t =σ(W z E+W z E t (5),

[0071]

[0072] Where W and U represent trainable parameter matrices, ⊙ represents the Hadamard product, σ represents the sigmoid function, tanh represents the hyperbolic tangent function, and h t This represents the node evolution at time t. The multi-head variation based on formula (8) further enhances the stability of the learning process.

[0073]

[0074] Where LN represents LayerNorm, and These correspond to two different nonlinear mappings for the k-th head, d k Let represent the dimension of each head, and <, ·> represent the dot product operation. Dropout is applied to the edge weights to generate a more generalized trend. Finally, the final edge weights are obtained by summing the weights calculated for the k subspaces. The specific calculation is shown in formula (9):

[0075]

[0076] Where f s,4 (·) are shared trainable weights;

[0077] 3-3) Channel Attention (SENet): Channel attention adaptively determines the proportion of the three learned edge weights, calculated as shown in formula (10):

[0078]

[0079] Where F ex and F sq The symbols represent compression and activation operations, respectively, and the dot (·) indicates channel multiplication.

[0080] 3-4) Connection predictor: The connection predictor performs convolution along the channel size and outputs a scalar indicating the connection probability between node i and node j. Its calculation process is shown in formula (11):

[0081]

[0082] Where ReLU represents the activation function. The architecture diagram of the dynamic graph construction method in this example is as follows: Figure 3 As shown, the Gumbel-Softmax technique is used to discretely sample the association probabilities to generate an adjacency matrix. The key hyperparameter temperature τ is fixed at 0.5 to control the degree of discretization: lower temperatures make the sampling results closer to a one-hot discrete distribution, but gradient propagation is difficult; higher temperatures make the output tend to be a continuous and uniform distribution, losing discreteness. The setting of τ=0.5 strikes a balance between trainability and structural discreteness, ensuring that the model can both learn the graph structure through gradient optimization and generate hard discrete edge connections to prevent the output from being too smooth.

[0083] 4) Multi-scale spatiotemporal feature extraction: Multi-scale spatial features are extracted by performing graph convolution on the dynamic graph structure obtained in step 3), and multi-scale temporal features are extracted from the IMF obtained in step 2). The specific steps are as follows:

[0084] 4-1) Multi-scale spatial feature extraction: Graph Convolutional Network (GCN) is used to process dynamic graphs, and the update formula is shown in formula (12):

[0085]

[0086] Where H represents the feature matrix of the graph convolutional layer, l is the number of graph convolutional layers, and W is the weight matrix. It is the adjacency matrix A plus the identity matrix I. yes The degree matrix, where σ is the activation function;

[0087] 4-2) Multi-scale temporal feature extraction: The imfs obtained from variational mode decomposition in step 2) are used as input, and each imf is processed separately, i.e. Where k is the number of mode decompositions, which is a hyperparameter. This indicates the temporal features extracted after processing the i-th imf. The following section will explain f in detail. i (·) The process of feature mining. First, the imf i The input is fed into the TCN for shallow temporal feature extraction, and the specific process is shown in formula (13):

[0088]

[0089] Here, a four-layer TCN was set up to obtain shallow features at different scales. Then The input is then used for deep temporal feature extraction via LSTM, as shown in formulas (14), (15), (16), (17), (18), and (19).

[0090]

[0091] Therefore, IMF i The features obtained through further mining using LSTM can be represented as follows: In this example, a multi-scale temporal feature extraction module is constructed using cascaded TCN-LSTM. The working principle of the multi-scale temporal feature extraction module is as follows: Figure 4 As shown, each IMF component undergoes a two-stage independent processing: a four-layer dilated convolutional network (TCN) with increasing dilation coefficients (1, 2, 4, 8) extracts short-term local features layer by layer; the output is then fed into an LSTM network, which is set to two layers, using a forget gate / input gate mechanism to filter key information and model long-term dependencies (such as the impact of seasonal water temperature changes on oxygen solubility). This structure overcomes the limitation of a single model in coordinating long and short periods, effectively supporting the collaborative prediction of dissolved oxygen concentration in both short and long-term evolution.

[0092] 5) Residual fusion and prediction output: Residual fusion is based on the residual idea and uses multi-head self-attention to fuse features. Finally, the fused features are mapped to the final prediction result through a fully connected layer. The specific process is shown in formulas (20), (21), (22), and (23):

[0093]

[0094] H = Concat(Att1, ..., Att) h W O (twenty two),

[0095] P = FC(H) (23),

[0096] in All are trainable parameters, d model =D model *h, where i represents the i-th attention head;

[0097] 6) Anti-counterfeiting verification: The path distance between the input data and the predicted result is calculated using Dynamic Time Warping (DTW), and the residual distribution is compared using the KS test. Let the input water quality data sequence be X = x1, x2, ..., x... T The model predicts the sequence Y = y1, y2, ..., y. T The minimum normalized path distance between them is D dtw The mathematical expression for its calculation is shown in formula (24):

[0098]

[0099] Where π represents the optimal path for aligning the two sequences, and d(·) is the Euclidean distance metric. The prediction residual ε is calculated simultaneously. t =|x t -y t The empirical distribution F of |} ε (z), and compared with the baseline normal residual distribution F ref (z) Perform the KS test, and the specific calculation process is shown in formula (25):

[0100]

[0101] Ultimately, a joint decision mechanism is used: when the path distance D... dtw >θ1 and KS statistic D ks When θ2 is greater than 0.02, the data is considered non-genuine. In this example, θ1 is taken as the 95th percentile of the DTW distance of the normal data in the training set, and θ2 is taken as the theoretical critical value of the KS test at a significance level of α = 0.05. If θ1 is set too low (e.g., 90th percentile), it will increase the sensitivity to local tampering, but may lead to natural hydrological fluctuations (such as rainstorm events) being misjudged as tampering. Conversely, if θ1 is too high (e.g., 99th percentile), it will be difficult to detect fine-grained tampering (such as replacement of 6-hour data from a single station). θ2 strictly follows hypothesis testing theory. Lowering α can reduce the missed detection of distributed tampering, but it will weaken the fault tolerance for legitimate anomalies (such as sudden pollution).

[0102] This example demonstrates a multi-scale spatiotemporal fusion water quality prediction and anti-counterfeiting method based on dynamic graph neural networks (MST-DGCN model, architecture as follows). Figure 2The model (shown in the table) was compared with benchmark models such as MLP, SVR, RNN, GRU, GCN-LSTM, GAT-LSTM, and Transformer to verify its superiority and advancement in prediction accuracy. Table 2 shows the quantitative results of RMSE, MAE, and WAPE for dissolved oxygen (DO) of these eight models at the same time step. The results indicate that MST-DGCN achieved the best prediction performance. Machine learning models such as MLP and SVR failed to effectively capture the complex relationships in time-series data, resulting in larger errors. RNN and GRU, as time-series prediction models, outperformed non-time-series models such as MLP and SVR. Furthermore, the GRU model, compared to the RNN model, effectively solved the "gradient vanishing and exploding" problem, thus capturing long-term dependencies in water quality data; therefore, GRU's prediction performance was superior to RNN. The Transformer model, containing a self-attention mechanism, could perform global modeling, thus outperforming GRU in water quality prediction. GCN and GAT, as classic graph neural networks, fused features based on the graph structure of nodes, updating node features using information from neighboring nodes. A single spatial model cannot complete the task of time series prediction. Therefore, graph neural networks are usually combined with traditional time series models to incorporate spatial features into time features. These spatiotemporal models outperform Transformers in prediction performance.

[0103] Table 2. Prediction results of different models

[0104]

[0105]

[0106] Based on the above comparisons, this method provides a systematic solution for predicting and preventing fraudulent activities related to water environment quality in complex watersheds, demonstrating its significant value in practical applications. Its core significance lies in building a data-driven support bridge for water resource and environmental management, which not only helps in the accurate identification and timely response to water quality changes but also provides a scientific basis and feasible pathways for ecological governance. Through deep fusion and modeling of multi-site, multi-scale dynamic data, this method exhibits strong generalization ability and adaptability, making it suitable for multi-point collaborative monitoring scenarios in small and medium-sized cities and even cross-regional watersheds, and possesses broad potential for widespread application.

[0107] Furthermore, this method helps to shift environmental monitoring from traditional "result presentation" to "intelligent early warning and identification," effectively reducing the governance costs and public health risks caused by water quality deterioration. It possesses high system robustness and data interpretability, and also provides key technological reserves for future applications in areas such as environmental big data governance and smart city construction.

Claims

1. A multi-scale spatiotemporal fusion water quality prediction and anti-counterfeiting method based on dynamic graph neural networks, characterized in that, Includes the following steps: 1) Data Acquisition: Hourly water quality and hydrological data were collected from multiple monitoring stations; 2) Multi-scale temporal decomposition: The temporal data of one target pollutant in step 1) is decomposed into several intrinsic mode functions (imfs) through variational mode decomposition (VMD); 3) Dynamic Graph Construction: Input the raw data collected in step 1), and generate a discrete dynamic graph adjacency matrix A based on the GRU gating mechanism and channel attention learning node embedding. t Completely free from predefined topological dependencies, its mathematical expression is formula (1): in σ represents the sigmoid function, and τ represents temperature, used to manage the interpolation process between discrete distributions and continuous classification densities. Furthermore, The connection probability between node i and node j at time t is represented by the following process: 3-1) Edge weight calculation: Edge weight represents the association derived from the topology level and node level input. Its main calculation formulas are shown in formula (2) and formula (3): Where E t E represents the dynamic semantic information extracted from the input X using a 1×1 standard convolution. t =Conv(X), where E is a trainable node embedding; 3-2) Gated Recurrent Unit (GRU): The Gated Recurrent Unit (GRU) is used as a method to capture the evolutionary representation of nodes, as shown in Equations (4), (5), (6), and (7): r t σ(W r E+U r E t ) (4), With t =σ(W z E+U z E t ) (5), Where W and U represent trainable parameter matrices, ⊙ represents the Hadamard product, σ represents the sigmoid function, tanh represents the hyperbolic tangent function, and h t This represents the node evolution at time t. The multi-head variation based on formula (8) further enhances the stability of the learning process. Where LN represents LayerNorm, and These correspond to two different nonlinear mappings for the k-th head, d k Let represent the dimension of each head, and <, ·> represent the dot product operation. Dropout is applied to the edge weights to generate a more generalized trend. Finally, the final edge weights are obtained by summing the weights calculated for the k subspaces. The specific calculation is shown in formula (9): Where f s,4 (·) are shared trainable weights; 3-3) Channel Attention (SENet): Channel attention adaptively determines the proportion of the three learned edge weights, calculated as shown in formula (10): Where F ex and F sq The symbols represent compression and activation operations, respectively, and the dot (·) indicates channel multiplication. 3-4) Connection predictor: The connection predictor performs convolution along the channel size and outputs a scalar indicating the connection probability between node i and node j. Its calculation process is shown in formula (11): Where ReLU represents the activation function; 4) Multi-scale spatiotemporal feature extraction: Multi-scale spatial features are extracted by performing graph convolution on the dynamic graph structure obtained in step 3), and multi-scale temporal features are extracted from the IMF obtained in step 2). The specific steps are as follows: 4-1) Multi-scale spatial feature extraction: Graph Convolutional Network (GCN) is used to process dynamic graphs, and the update formula is shown in formula (12): Where H represents the feature matrix of the graph convolutional layer, l is the number of graph convolutional layers, and W is the weight matrix. It is the adjacency matrix A plus the identity matrix I. yes The degree matrix, where σ is the activation function; 4-2) Multi-scale temporal feature extraction: The imfs obtained from variational mode decomposition in step 2) are used as input, and each imf is processed separately, i.e. i,k∈Z,0≤i≤k, where k is the number of mode decompositions, which is a hyperparameter. This indicates the temporal features extracted after processing the i-th imf. The following section will explain f in detail. i (·) The process of feature mining. First, the imf i The input is fed into the TCN for shallow temporal feature extraction, and the specific process is shown in formula (13): Here, a four-layer TCN was set up to obtain shallow features at different scales. Then The input is then used for deep temporal feature extraction via LSTM, as shown in formulas (14), (15), (16), (17), (18), and (19). Therefore, IMF i The features obtained through further mining using LSTM can be represented as follows: 5) Residual fusion and prediction output: Residual fusion is based on the residual idea and uses multi-head self-attention to fuse features. Finally, the fused features are mapped to the final prediction result through a fully connected layer. The specific process is shown in formulas (20), (21), (22), and (23): H=Concat(Att1,…,Att h )W O (22), P = FC(H) (23), in All are trainable parameters, d model =D model *h, where i represents the i-th attention head; 6) Anti-counterfeiting verification: The path distance between the input data and the predicted result is calculated using Dynamic Time Warping (DTW), and the residual distribution is compared using the KS test. Let the input water quality data sequence be X = x1, x2, ..., x... T The model predicts the sequence Y = y1, y2, ..., y. T The minimum normalized path distance between them is D dtw The mathematical expression for its calculation is shown in formula (24): Where π represents the optimal path for aligning the two sequences, and d(·) is the Euclidean distance metric. The prediction residual ε is calculated simultaneously. t =|x t -y t The empirical distribution F of |} ε (z), and compared with the baseline normal residual distribution F ref (z) Perform the KS test, and the specific calculation process is shown in formula (25): Ultimately, a joint decision mechanism is used: when the path distance D... dtw >θ1 and KS statistic D ks When the value is greater than θ2, the data is determined to be non-real data.

Citation Information

Cited By

  • Dissolved oxygen prediction correction method fusing graph neural network and extreme value distribution

    CN121859024A

  • A dissolved oxygen prediction correction method fusing a graph neural network and an extreme value distribution

    CN121859024B

  • Wastewater new pollutant early warning method and system, computer equipment and medium

    CN122157842A