A method for coagulant dosing prediction by fusing sparse coding and graph spatio-temporal attention

By fusing sparse coding with graph spatiotemporal attention, the problems of missing data completion and insufficient feature representation in industrial water treatment processes are solved, and high-precision coagulant dosing prediction is achieved.

CN120911711BActive Publication Date: 2026-01-27CHENGDU QIANJIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511448513.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-01-27
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing deep learning modeling methods struggle to accurately complete missing data in industrial water treatment processes and fail to effectively characterize the spatial correlation and dependence of multiple subsystems during water treatment, thus affecting the accuracy of coagulant dosing prediction.

Method used

A method combining sparse coding and graph spatiotemporal attention is adopted. A sparse data completion network is constructed by combining a Transformer encoder and decoder with a sparse self-attention mechanism and a top-k sparsification strategy. A temporal feature extraction and spatial feature aggregation module is constructed by combining a residual MLP neural network and a minimum gate unit algorithm to generate high-quality spatiotemporal feature tensors for prediction.

Benefits of technology

It improves the accuracy of missing data completion, enhances the information entropy of spatial features, dynamically focuses on key time-series nodes, improves the accuracy of coagulant addition prediction, and solves the shortcomings of traditional methods in data completion and feature representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911711B_ABST
    Figure CN120911711B_ABST
Patent Text Reader

Abstract

The application discloses a coagulant dosing prediction method combining sparse coding and graph space-time attention, and relates to the technical field of time series data modeling and prediction.The application combines the Transformer coding-decoding architecture with sparse self-attention and a top-k sparsification strategy, effectively improves the missing data completion accuracy, and provides high-quality and non-missing input data for subsequent prediction modeling; in the aspect of spatial feature expression, the graph attention network is used to capture the node correlation of the water production system multi-subsystem, the correlation weight between nodes is dynamically calculated, the neighbor features are aggregated, the information entropy of the spatial features is improved, and the expression capability of the spatial features to the multi-node coupling relationship is greatly enhanced; in the aspect of time series prediction performance, the time attention mechanism can dynamically focus on key time series nodes, further improves the utilization rate of key time series features, and improves the accuracy of coagulant dosing prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time series data modeling and prediction technology, and in particular to a method for predicting coagulant addition by fusing sparse coding and graph spatiotemporal attention. Background Technology

[0002] With economic and technological development, water treatment processes in water plants are becoming increasingly complex and their scale is continuously expanding. However, in actual production, especially when facing challenges such as frequent fluctuations in raw water quality, significant feedback delays in control systems, and harsh on-site working environments, the precise addition of coagulants plays a decisive role in the final effluent quality.

[0003] Current automated coagulant dosing technologies primarily rely on two methods: those based on process mechanism models and those based on data-driven models. The former requires in-depth analysis of the physicochemical mechanisms of each stage of the water treatment process, establishing a system of differential equations for relevant variables, and solving these equations to obtain the mass variable values ​​at each time point. The latter, however, has benefited from the development of intelligent and information technologies, particularly the widespread application of distributed control systems (DCS) which has accumulated massive amounts of process data. This data-driven approach directly utilizes historical data to construct mathematical relationship models between process variables (such as influent flow rate, turbidity, and pH) and key mass variables (such as effluent turbidity and coagulant dosage), enabling online prediction of coagulant dosage.

[0004] However, industrial processes exhibit significant online dynamic characteristics, thus fully incorporating temporal feature information into predictive modeling can significantly improve model performance. To address this need, nonlinear dynamic prediction models, represented by autoregressive methods and deep learning methods, have emerged. Among them, deep learning methods, through end-to-end learning, can automatically extract effective features, demonstrating more powerful nonlinear representation and dynamic modeling capabilities. Advanced neural network architectures, such as Long Short-Term Memory (LSTM), Gated Recurrent Units (GRU), and Minimum Gated Units (MGU), have achieved superior predictive performance in practical applications by efficiently capturing and utilizing temporal feature information through carefully designed network structures.

[0005] While existing deep learning modeling methods have improved the accuracy of industrial process predictions, significant shortcomings remain: industrial water treatment data is often incomplete due to transmission signal fluctuations, equipment failures, and sensor packet loss, making the accurate completion of such limited missing data the primary challenge in building reliable models; simultaneously, the strong coupling between multiple subsystems in the water treatment process leads to complex correlations between variables, making it difficult for traditional deep learning methods to effectively characterize these spatial correlations and dependencies, resulting in insufficient extraction of spatial feature information for prediction; furthermore, although time-series models such as MGU can capture continuity, they struggle to fully identify the dynamic differences in the importance of variables at different time points, limiting their ability to capture key time-series features and ultimately affecting the performance optimization of effluent quality prediction. Summary of the Invention

[0006] The purpose of this invention is to provide a coagulant addition prediction method that integrates sparse coding and graph spatiotemporal attention to improve the above-mentioned technical problems.

[0007] To achieve the above-mentioned objectives, the embodiments of the present invention provide the following technical solutions:

[0008] A method for predicting coagulant dosing by fusing sparse coding with graph spatiotemporal attention, comprising:

[0009] A sparse data completion network is constructed based on Transformer encoders and decoders, combined with sparse self-attention mechanism and top-k sparsification strategy;

[0010] A temporal feature extraction module is constructed based on residual MLP neural networks and minimum gate unit algorithm, combined with temporal attention mechanism; a spatial feature aggregation module is constructed based on residual MLP neural networks and graph attention mechanism.

[0011] Industrial water production data is collected by sensors and used as missing data; the missing data is filled in by a sparse data network, the time series completion matrix is ​​calculated, and complete time series data is generated.

[0012] The temporal and spatial features of the complete time series data are extracted sequentially through the spatial feature aggregation module and the temporal feature extraction module, generating a spatiotemporal feature tensor that integrates spatial neighbor association and historical temporal dependency;

[0013] Predictive values ​​for coagulant dosage are generated by using a fully connected network to predict spatiotemporal feature tensors.

[0014] Currently, industrial water treatment data collected using traditional methods is often incomplete due to transmission signal fluctuations, equipment failures, and sensor packet loss. Furthermore, the completion methods employed rely solely on local statistical patterns or simple rules, failing to consider the spatiotemporal correlations of the water treatment data (such as the dynamic dependence of a missing turbidity value at a given node on its upstream nodes), resulting in low completion accuracy. For example, forward / backward completion only copies adjacent values, which cannot handle missing values ​​caused by sudden fluctuations in water quality. This invention, however, captures the temporal dynamics and spatial correlations of missing data through multi-layer sparse self-attention, generating high-dimensional encoded features. Then, it utilizes a top-k sparsity strategy to retain only key correlations, not only reducing computational complexity but also accurately recovering missing values ​​to generate a temporal completion matrix (i.e., complete data).

[0015] Furthermore, traditional models have significant shortcomings in predicting coagulant dosage: Spatially, CNNs rely on the local receptive field of fixed-size convolutional kernels, making it difficult to overcome spatial limitations and capture long-distance coupling relationships across subsystems in the water treatment system; graph-free neural networks do not capture treatment units (such as flocculation tanks and sedimentation tanks) as independent nodes, failing to characterize the physical relationship of "upstream water quality affecting downstream coagulant demand," resulting in fragmented spatial features. Temporally, models such as MGU use fixed weight mechanisms, paying equal attention to all time points, failing to identify key temporal nodes (such as the moment of sudden turbidity change during heavy rain) and failing to quantify the differences in contribution at different time points (such as the impact of turbidity 1 hour ago being greater than 3 hours ago), thus drowning out key temporal features and making it difficult to support prediction performance optimization.

[0016] However, to address the problem of insufficient spatial correlation representation, the spatial feature aggregation module of this invention first enhances and completes the feature representation of the data through MLP to obtain a high-order latent space representation; then, it constructs a graph connectivity matrix based on the K-nearest neighbor algorithm to identify strongly correlated nodes; finally, it dynamically calculates the node correlation weights through a graph attention network, aggregates neighbor features to generate a spatial feature tensor, and overcomes the limitations of the local receptive field of CNN and the correlation fragmentation problem of graph-less models, accurately depicting the physical logic of "upstream unit water quality affecting downstream coagulant demand".

[0017] To address the limitation in identifying the dynamic importance of time series, the temporal feature extraction module of this invention integrates the Minimum Gated Unit (MGU) algorithm with a temporal attention mechanism. Based on the MGU's capture of temporal continuity, it quantifies the contribution differences at different time points through attention scores, assigning higher weights to key nodes such as sudden changes in turbidity during rainstorms. This solves the problem of key features being submerged due to the fixed weights of traditional MGU.

[0018] Therefore, the beneficial effects of the present invention are as follows:

[0019] This invention employs a full-link optimization design encompassing "sparse data completion - spatial feature aggregation - temporal feature extraction." At the data completion level, it leverages a Transformer encoder-decoder architecture combined with sparse self-attention and top-k sparsity strategies to effectively improve the accuracy of missing data completion and provide high-quality, complete input data for subsequent prediction modeling. At the spatial feature representation level, it utilizes a graph attention network to capture the node relationships within multiple subsystems of the water treatment system. By dynamically calculating the association weights between nodes and aggregating neighbor features, it increases the information entropy of spatial features, significantly enhancing their ability to express the coupling relationships between multiple nodes and strongly supporting the association learning between coagulant dosage and the state of each node. At the temporal prediction performance level, the temporal attention mechanism dynamically focuses on key temporal nodes, further improving the utilization rate of key temporal features and increasing the accuracy of coagulant dosage prediction.

[0020] Furthermore, the sparse data completion network includes an encoding module consisting of M stacked encoding layers, a decoding module consisting of M stacked decoding layers, and a fully connected layer; each encoding layer and each decoding layer includes a cascaded sparse self-attention layer, a normalization layer, and a linear transformation layer; the sparse self-attention layer includes a query vector feature transformation branch, a key vector feature transformation branch, a value vector feature transformation branch, and a top-k sparse layer; the query vector feature transformation branch and the value vector feature transformation branch each include two multilayer perceptrons; the key vector feature transformation branch includes two cascaded multilayer perceptrons and a transpose transformation layer; the top-k sparse layer includes a top-k layer, a softmax layer, and a multilayer perceptron.

[0021] In the above scheme, by constructing a network structure consisting of M encoding layers and M decoding layers stacked together, and combining sparse self-attention layers in each layer, the problem of existing data completion methods being difficult to accurately capture key spatiotemporal correlations and having excessively high computational complexity when processing long-series, multi-node data with missing data in scenarios such as industrial water treatment is solved. This achieves the goal of accurately capturing key data correlations while efficiently reducing computational complexity, improving the accuracy of missing data completion, and providing high-quality complete data for subsequent processing.

[0022] Furthermore, generating complete time-series data includes:

[0023] The high-dimensional feature vector of the missing data is extracted by the encoding module. The temporal dynamics of the missing data and the spatial correlation of different nodes are captured layer by layer by combining the sparse self-attention mechanism and the top-k sparsification strategy to generate high-dimensional encoded features.

[0024] The high-dimensional encoded features are cascaded and passed to the decoding module. The missing values ​​of the missing data are recovered through a sparse self-attention layer to generate a temporal completion matrix. Except for the input of the first decoding layer, which is the high-dimensional encoded features, the input of the other decoders is the fusion data of the high-dimensional encoded features and the previous decoding layer.

[0025] The timing completion matrix is ​​input into the fully connected layer, and the complete timing data is output.

[0026] In the aforementioned scheme, this claim effectively solves the problem that traditional missing data completion methods (such as mean imputation and simple matrix factorization) struggle to simultaneously capture the spatiotemporal correlations, computational efficiency, and completion accuracy of long-term, multi-node data through a hierarchical architecture design of "encoding-decoding-fully connected layers." Specifically, the encoding module, leveraging a combination of sparse self-attention and top-k sparsity strategies, accurately captures the temporal dynamics and spatial correlations of the data during the layer-by-layer extraction of high-dimensional features. Simultaneously, top-k sparsity filtering significantly reduces redundant computation, balancing the accuracy and efficiency of feature extraction. The decoding module, through cascading transmission and a fusion input design of "high-dimensional encoded features + previous level decoding results," achieves multi-scale feature complementarity and information propagation, avoiding the deviation in missing value recovery caused by a single input. Finally, the complete temporal data output through the fully connected layer preserves the inherent correlation structure of the original data while eliminating missing values ​​and noise interference, achieving a unified approach to missing data completion that balances accuracy, efficiency, and practicality.

[0027] Furthermore, the processing procedure of the sparse self-attention layer includes:

[0028] Retrieve the missing data and generate query vector, key vector, and value vector respectively;

[0029] The query vector and value vector are input into the query vector feature transformation branch and the value vector feature transformation branch respectively for processing. The corresponding two-layer multilayer perceptron are used to perform feature compression and feature recovery in sequence to obtain the query feature vector and the value feature vector.

[0030] The key vector is input to the key vector feature transformation branch for processing. It is then processed by two corresponding multilayer perceptrons for feature compression and feature recovery, and combined with the transpose transformation layer for dimension alignment to obtain the key feature vector.

[0031] Based on the query feature vector, value feature vector, and key feature vector, and combined with the top-k layer and softmax layer, the initial sparse self-attention value is calculated. The initial sparse self-attention value is then linearly transformed by a multilayer perceptron to align with the same dimension as the missing data, thus obtaining the sparse self-attention value.

[0032] The output of the first encoding layer is generated based on sparse self-attention values ​​and missing data.

[0033] Furthermore, the formula corresponding to the sparse self-attention layer is: ;

[0034] ;

[0035] ;

[0036] ;

[0037] in, , , These represent the query feature vector, key feature vector, and value feature vector, respectively. This represents the input data for the sparse self-attention layer. , , Both represent multilayer perceptrons. This represents the activation function. The transpose matrix representing the key eigenvectors. This represents the output data of the sparse self-attention layer. This indicates the top-k operation.

[0038] Existing technologies suffer from insufficient feature transformation, difficulty in dimensional alignment, and redundant attention calculations when processing multi-node time-series data with missing nodes, resulting in insufficient accuracy in capturing data associations. The proposed solution addresses this by designing independent feature transformation branches for the query vector, key vector, and value vector. Two multilayer perceptrons are used to perform feature compression and recovery respectively, preserving key information while reducing redundancy. The key vector branch incorporates a transpose transformation layer for precise dimensional alignment. Combined with a top-k layer to filter key associations and a softmax layer and multilayer perceptron to optimize attention calculations and align dimensions, this effectively solves the problem of insufficient capture accuracy and provides a more reliable feature foundation for subsequent encoding processes.

[0039] Furthermore, the processing procedure of the spatial feature aggregation module includes:

[0040] By using residual MLP neural networks and activation functions to enhance features of complete time-series data, the correlation between different nodes in the complete time-series data is captured, and a higher-order latent space representation is obtained.

[0041] Higher-order latent space representations are used as graph nodes; based on the higher-order latent space representations, a graph connectivity matrix is ​​constructed using the K-nearest neighbor algorithm; according to the graph connectivity matrix, it is determined whether different graph nodes are connected or strongly correlated, and the higher-order latent space representations of the connected or strongly correlated nodes are retained to determine the strongly correlated graph nodes.

[0042] Based on graph attention mechanism, exponential function and activation function, the graph attention score of each strongly correlated graph node is calculated; based on each graph attention score and the corresponding high-order latent space representation, the spatial feature tensor of industrial water treatment is generated.

[0043] In the aforementioned schemes, existing technologies struggle to accurately capture the nonlinear, long-distance relationships between multiple nodes when extracting spatial features from industrial water treatment data, and they are unable to effectively filter out irrelevant nodes, resulting in redundant spatial features and inaccurate relationship representations. This method's spatial feature aggregation module first enhances the features of the complete time-series data using a residual MLP neural network and activation functions. This deeply mines the complex relationships between nodes and obtains a high-order latent space representation, laying a high-quality feature foundation for subsequent spatial relationship analysis. Then, using the high-order latent space representation as graph nodes, a graph connectivity matrix is ​​constructed using the K-nearest neighbor algorithm to filter strongly correlated nodes, eliminating interference from irrelevant nodes and focusing on core relationships. Finally, based on a graph attention mechanism combined with exponential and activation functions, attention scores are calculated, accurately quantifying the contribution of strongly correlated nodes and aggregating features. This effectively solves the aforementioned problems, and the resulting industrial water treatment spatial feature tensor accurately preserves key spatial relationships between nodes while eliminating redundant information, significantly improving the overall model's ability to characterize spatial relationships in industrial water treatment scenarios.

[0044] Furthermore, the processing procedure of the time feature extraction module includes:

[0045] The spatial feature tensor of industrial water treatment output by the spatial feature aggregation module is obtained, and the feature is enhanced by MLP neural network and activation function to obtain the spatially enhanced feature tensor of industrial water treatment.

[0046] The temporal attention score of the spatial enhancement feature tensor of industrial water treatment is calculated using the minimum gate unit algorithm and the hyperbolic tangent activation function.

[0047] Based on the attention scores at each time and the corresponding spatial enhancement feature tensors of industrial water treatment, the temporal dependency features between the spatial enhancement feature tensors of industrial water treatment are captured, and a spatiotemporal feature tensor is generated.

[0048] In the aforementioned schemes, existing technologies struggle to accurately capture dynamic temporal dependencies between features while enhancing spatial features. Furthermore, the insufficient integration of temporal attention and gating mechanisms leads to inaccurate characterization of temporal correlations and low feature utilization. This method's temporal feature extraction module first enhances the spatial feature tensor of industrial water treatment using an MLP neural network and activation functions, providing a richer feature foundation for subsequent temporal dependency capture. Then, it utilizes the minimum gating unit algorithm (adept at handling temporal data and suppressing redundant information) combined with a hyperbolic tangent activation function to accurately calculate temporal attention scores to quantify the importance of features at different times. Finally, based on these scores and spatially enhanced features, it captures temporal dependencies, effectively solving the aforementioned problems. The resulting spatiotemporal feature tensor retains accurate spatial correlations while incorporating dynamic temporal dependencies, significantly improving the ability to characterize the spatiotemporal correlations of industrial water treatment processes.

[0049] Furthermore, the sparse data completion network, spatial feature aggregation module, and temporal feature extraction module are trained using the mean square error loss function and backpropagation algorithm.

[0050] In the above scheme, the mean square error loss function can accurately quantify the deviation between the model output and the actual water production data. The backpropagation algorithm can efficiently update the parameters of the sparse data completion network, spatial feature aggregation module and temporal feature extraction module based on this loss, thereby improving the model's data completion accuracy and spatiotemporal feature extraction capability, and providing reliable support for subsequent industrial water production related prediction tasks. Attached Figure Description

[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a flowchart of the method in an embodiment of the present invention;

[0053] Figure 2 This is a diagram showing the overall model construction corresponding to the method in the embodiments of the present invention;

[0054] Figure 3 This is a data flow diagram corresponding to the method in the embodiments of the present invention;

[0055] Figure 4 This is a diagram of the sparse data completion network structure in an embodiment of the present invention;

[0056] Figure 5 This is a structural diagram of the encoding module and decoding module in an embodiment of the present invention;

[0057] Figure 6 This is a structural diagram of the sparse self-attention layer in an embodiment of the present invention;

[0058] Figure 7 This is a structural diagram of the residual MLP neural network in an embodiment of the present invention;

[0059] Figure 8 This is a schematic diagram of the attention mechanism in an embodiment of the present invention;

[0060] Figure 9 This is a schematic diagram of the time attention mechanism in an embodiment of the present invention;

[0061] Figure 10 This is a schematic diagram of the minimum gating unit algorithm in an embodiment of the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0063] Please see Figure 1 This embodiment provides a method for predicting coagulant dosing by fusing sparse coding with graph spatiotemporal attention. Figure 1 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not impose any limitations on this.

[0064] like Figure 1 , Figure 2 and Figure 3 As shown, a method for predicting coagulant dosing by fusing sparse coding with graph spatiotemporal attention includes:

[0065] S1. Construct a sparse data completion network based on Transformer encoder and decoder, combined with sparse self-attention mechanism and top-k sparsification strategy;

[0066] like Figure 4 and Figure 5 As shown, the sparse data completion network includes an encoding module consisting of M stacked encoding layers, a decoding module consisting of M stacked decoding layers, and a fully connected layer; each encoding layer and each decoding layer includes a cascaded sparse self-attention layer, a normalization layer (LayerNorm), and a linear transformation layer (FeedForward); as... Figure 6As shown, the Sparse Self-Attention Layer (SSA) includes a query vector feature transformation branch, a key vector feature transformation branch, a value vector feature transformation branch, and a top-k sparse layer. Both the query vector feature transformation branch and the value vector feature transformation branch include two multilayer perceptrons (MLP-2 and MLP-3). The key vector feature transformation branch includes two cascaded multilayer perceptrons (MLP-2 and MLP-3) and a transpose transformation layer. The top-k sparse layer includes a top-k layer, a softmax layer, and a multilayer perceptron (MLP-3). In this embodiment, M is set to 6.

[0067] S2. Based on residual MLP neural network and minimum gate unit algorithm, combined with temporal attention mechanism, construct temporal feature extraction module; based on residual MLP neural network and graph attention mechanism, construct spatial feature aggregation module;

[0068] The spatial feature aggregation module consists of a residual MLP neural network connected in series and a graph attention layer constructed using a graph attention mechanism. The temporal feature extraction module consists of a residual MLP neural network connected in series, a temporal attention layer constructed using a temporal attention mechanism, and an MGU layer constructed using a minimum gate unit algorithm.

[0069] like Figure 7 As shown, the residual MLP neural network includes an MLP (multilayer perceptron), a summing layer, two MLPs, and a ReLU layer connected in series.

[0070] S3. Collect industrial water treatment data (such as influent flow rate, turbidity, pH value, effluent turbidity, and coagulant dosage) through sensors. Some of this data may be missing or abnormal, forming the input dataset (missing data). Missing data is imputed using a sparse data network, and a time-series completion matrix is ​​calculated to generate complete time-series data; among which, , , These represent the number of data points, the number of nodes, and the number of node features, respectively. Represents the set of real numbers. , These represent the first missing data and the last missing data, respectively.

[0071] The generation of complete time-series data includes:

[0072] S3-1. Extract high-dimensional feature vectors of missing data through the encoding module, and capture the temporal dynamics of missing data and the spatial correlation of different nodes layer by layer by combining the sparse self-attention mechanism and the top-k sparsification strategy to generate high-dimensional encoded features.

[0073] Specifically, the missing data is sequentially input into the six cascaded coding layers, the first... The output of the coding layer is the first layer. The input to the layer encoding layer, corresponding to the feature dimension Perform a mapping transformation, that is .

[0074] In each encoding layer, taking the first encoding layer as an example, the missing data is input into the sparse self-attention (SSA) layer. Through operations such as multilayer perceptron, feature transformations are performed on the query vector (Q), key vector (K), and value vector (V). Combined with a top-k sparsity strategy, the sparse self-attention score of the missing data is calculated, focusing on capturing key temporal dynamics and spatial relationships between different nodes in the missing data. This filters out the information most relevant to the completion task, initially extracting high-value features. A normalization layer normalizes the output data of the sparse self-attention layer and fuses it with the missing data, adjusting the feature distribution to a more stable state, reducing distribution differences between different samples or features, alleviating the "internal covariate shift" problem, and enabling subsequent linear transformation layers to process features more efficiently. The fused data is input into the linear transformation layer to further enhance the expressive power of the features, uncover more complex potential relationships between features, and obtain the corresponding linear features. The linear features are then fused with the fused data to stabilize the feature distribution and retain key information, obtaining the encoded information output by the first encoding layer.

[0075] Therefore, the formula for each coding layer is:

[0076] ;

[0077] in, This represents a linear transformation operation. Indicates the first The output data of the layer coding layer, This indicates a normalization operation. Indicates the first Input data for the layer coding layer, This represents the operation corresponding to the sparse self-attention layer.

[0078] The corresponding expression is:

[0079] ;

[0080] in, This represents the merged data. Represents the maximum value function. , All of these represent learnable parameters in the linear transformation layer. , Both represent the bias parameters in the linear transformation layer.

[0081] Furthermore, the processing method of each coding layer is the same as that of the first coding layer, and the corresponding feature dimensions are... Perform layer-by-layer recovery, that is The output of the last coding layer is the output of the coding module.

[0082] In this embodiment, taking the first coding layer as an example, the processing procedure of the sparse self-attention layer includes:

[0083] S3-1-1. Obtain the missing data and generate query vector, key vector, and value vector respectively;

[0084] S3-1-2. Input the query vector and value vector into the query vector feature transformation branch and the value vector feature transformation branch respectively for processing. Through the corresponding two-layer multilayer perceptron, feature compression and feature recovery are performed in sequence to obtain the query feature vector and the value feature vector.

[0085] S3-1-3. Input the key vector into the key vector feature transformation branch for processing. Perform feature compression and feature recovery sequentially through the corresponding two multilayer perceptrons, and perform dimension alignment by combining the transpose transformation layer to obtain the key feature vector.

[0086] S3-1-4. Based on the query feature vector, value feature vector, and key feature vector, and combined with the top-k layer and softmax layer, calculate the initial sparse self-attention value. Then, perform a linear transformation on the initial sparse self-attention value through a multilayer perceptron to align it with the same dimension as the missing data to obtain the sparse self-attention value.

[0087] S3-1-5. Based on sparse self-attention values ​​and missing data, generate the output of the first coding layer.

[0088] Therefore, the formula corresponding to the sparse self-attention layer is:

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] in, , , These represent the query feature vector, key feature vector, and value feature vector, respectively. This represents the input data for the sparse self-attention layer. , , Both represent multilayer perceptrons. This represents the activation function. The transpose matrix representing the key eigenvectors. This represents the output data of the sparse self-attention layer. This indicates the top-k operation.

[0094] In addition, activation function It is a normalization exponential function, mainly used to convert any real vector into a probability distribution. Its specific expression is:

[0095] ;

[0096] in, Represents the natural constant. This represents the input data for the activation function. Represents a constant. This represents the summation function. In this embodiment, The value is 256.

[0097] For learnable top-k selection operators, the sparse attention scores at each query position are used to retain the top-k operators in descending order. With 1 sparse attention score and the rest set to 0, the expression corresponding to the top-k selection operator is:

[0098] ;

[0099] in, This indicates the processing object of the top-k selection operator (the sparse attention matrix composed of sparse attention scores). Indicates the first The query position and the first Sparse attention score for each key position In the sparse attention score matrix, the first... In the row, after sorting by score from largest to smallest, the top... The set consisting of the elements.

[0100] The core of the top-k selection operator is to perform "key association filtering" on the sparse attention scores. For each query position, only the rows with the highest sparse attention scores are retained. One value is set to zero, and the rest are set to 0. In this way, on the one hand, given that the correlations between different treatment units (such as mixing tanks and flocculation tanks) in industrial water treatment data have different priorities, top-k filtering can highlight the strong correlations that are more critical to the prediction of coagulant addition (such as the strong correlation between the water quality of the mixing tank and the flocculation tank during high turbidity periods), and filter out weak correlations or noise interference; on the other hand, setting most of the unimportant attention scores to 0 can significantly reduce the number of parameters and complexity of subsequent calculations. Combined with the encoder-decoder architecture of Transformer, the sparse data completion network can ensure the accuracy of correlation capture and meet the real-time requirement of quickly completing the missing sensor data for subsequent prediction when processing long-term, multi-node industrial water treatment data.

[0101] S3-2. The high-dimensional encoded features are cascaded and passed to the decoding module. The missing values ​​of the missing data are recovered through a sparse self-attention layer to generate a temporal completion matrix. Except for the input of the first decoding layer, which is the high-dimensional encoded features, the input of the other decoders is the fusion data of the high-dimensional encoded features and the previous decoding layer. The processing of the decoding module is exactly the same as that of the encoding module.

[0102] S3-3. Input the timing completion matrix into the fully connected layer and output the complete timing data.

[0103] Existing data acquisition methods struggle to achieve accurate completion of long-term, multi-node data in industrial water treatment while maintaining computational efficiency, and they also suffer from missing values ​​in complex spatiotemporal correlations, affecting the accuracy of subsequent coagulant dosing predictions. This method employs a sparse data completion network based on a Transformer encoder-decoder architecture. The encoding module uses a 6-cascade encoding layer, combined with Sparse Self-Attention (SSA) and top-k sparsification strategies, to extract high-dimensional features layer by layer from the missing data. A top-k selection operator is used to filter key correlation attention scores, focusing on strong correlations and performing sparsified computation. The decoding module follows a similar process, using the high-dimensional features output by the encoding module to recover missing values ​​and generate a temporal completion matrix. Finally, the complete temporal data is obtained through a fully connected layer. This embodiment not only uses sparse self-attention and top-k strategies to accurately capture key spatiotemporal correlations between processing units in industrial water treatment data, effectively filtering out weak correlations and noise interference, but also significantly reduces computational complexity through sparsification, enabling the network to efficiently process long-term, multi-node water treatment data and achieve accurate and rapid completion of missing data, providing a high-quality and complete time-series data foundation for subsequent coagulant dosing prediction.

[0104] S4. The spatial features and spatial features of the complete time series data are extracted sequentially through the spatial feature aggregation module and the temporal feature extraction module to generate a spatiotemporal feature tensor that integrates spatial neighbor association and historical time series dependency.

[0105] The processing steps of the spatial feature aggregation module include:

[0106] S4-1-1. Feature enhancement of complete time series data is performed by residual MLP neural network and activation function to capture the correlation between different nodes in complete time series data and obtain a high-order latent space representation.

[0107] Specifically, according to Figure 6 The connections within will link the complete time series data. After being input into the residual MLP neural network, the data flows sequentially through a multilayer perceptron (MLP), a summation operation, two MLPs, and a ReLU activation function. Finally, it is fused with the complete time series data through residual connections to enhance feature representation capabilities. It extracts a high-order latent space representation containing complex inter-node relationships from the complete time series data. It can also overcome the limitations of simple linear models and accurately capture the nonlinear and long-distance relationships between nodes in industrial water treatment scenarios, thus completing the processing of complete time series data.

[0108] The expression corresponding to the residual MLP neural network is:

[0109] ;

[0110] ;

[0111] in, Represents matrix multiplication. , , This represents the learnable parameters in an MLP. This represents the output data of the MLP. This represents the activation function. This represents the higher-order hidden space representation.

[0112] S4-1-2. Use higher-order latent space representations as graph nodes; construct a graph connectivity matrix based on the higher-order latent space representations using the K-nearest neighbor algorithm; determine whether different graph nodes are connected or strongly correlated based on the graph connectivity matrix, retain the higher-order latent space representations of connected or strongly correlated nodes, and determine strongly correlated graph nodes; it should be noted that, as Figure 8 As shown, firstly, each higher-order latent space representation is treated as a graph node. Next, based on these higher-order latent space representations, the K-nearest neighbor algorithm is used to construct a graph connectivity matrix to determine whether different graph nodes are connected or strongly correlated. Connected or strongly correlated higher-order latent space representations are then selected and retained to identify strongly correlated graph nodes. Finally, in the graph attention mechanism, for the target graph node... It will be based on its association weight with other strongly related graph nodes. The features of these strongly correlated graph nodes are weighted and aggregated, while also considering the features of the graph nodes themselves. The final output is a target graph node that incorporates information from neighboring graph nodes, and further outputs in various dimensions are generated, thereby achieving effective capture and feature enhancement of spatial relationships between nodes. Figure 7 middle, , , Indicates the first The first, second, and third graph nodes in a high-order latent space representation Each graph node , , All of these represent nodes in a strongly correlated graph.

[0113] Therefore, the expression corresponding to the graph attention mechanism is:

[0114] ;

[0115] ;

[0116] in, This represents the graph convolution operation. In the graph connectivity matrix, the first... row, first The value of the column, i.e., the association weight, the first... The graph node and the first When the nodes of a graph are connected or strongly correlated Otherwise it is 0; , They represent the first The th high-order latent space representation in the th The first graph node, the first Each graph node Indicates intermediate variables. Indicates the first The th high-order latent space representation in the th Each graph node Represents the target graph node The set of neighboring nodes in the graph.

[0117] S4-1-3. Based on graph attention mechanism, exponential function, and activation function, calculate the graph attention score of each strongly correlated graph node; based on each graph attention score and the corresponding higher-order latent space representation, extract the spatial features of the complete time series data and generate the industrial water treatment spatial feature tensor. .

[0118] Traditional methods struggle to accurately capture the complex, nonlinear, and long-distance spatial relationships among multiple subsystems in industrial water treatment, impacting the accuracy of subsequent coagulant dosing predictions. This embodiment first utilizes a residual MLP neural network combined with activation functions to enhance the features of complete time-series data, extracting a high-order latent space representation containing complex inter-node relationships. Then, this high-order latent space representation is used as graph nodes, and a graph connectivity matrix is ​​constructed using the K-nearest neighbor algorithm to identify strongly correlated graph nodes. Finally, graph attention scores for strongly correlated graph nodes are calculated based on graph attention mechanisms, exponential functions, and activation functions, and combined with the corresponding high-order latent space representations to generate an industrial water treatment spatial feature tensor. Thus, this method effectively overcomes the limitations of simple linear models, accurately capturing the nonlinear and long-distance relationships between nodes in industrial water treatment scenarios, providing more accurate spatial feature support for subsequent coagulant dosing predictions, and improving prediction accuracy.

[0119] The processing steps of the time feature extraction module include:

[0120] S4-2-1. Obtain the industrial water treatment spatial feature tensor output by the spatial feature aggregation module, and use the MLP neural network and activation function to perform feature enhancement to obtain the industrial water treatment spatial enhancement feature tensor; the method used in S4-2-1 is the same as that in S4-1-1.

[0121] S4-2-2. Calculate the temporal attention score of the spatial enhancement feature tensor of industrial water treatment using the Minimum Gated Unit (MGU) algorithm and the hyperbolic tangent activation function; whereby the Minimum Gated Unit (MGU) algorithm at the current time... The input below is the current time. The spatial enhancement feature tensor of industrial water production and the previous time step Activation value of MGU .

[0122] S4-2-3. Based on the attention scores of each time and the corresponding spatial enhancement feature tensor of industrial water treatment, capture the temporal dependency features between the spatial enhancement feature tensors of industrial water treatment and generate spatiotemporal feature tensors.

[0123] Specifically, it should be noted that, such as Figure 9 and Figure 10 As shown, the activation value of the minimum gate unit algorithm at the previous time step is... Processed with the current industrial water treatment spatial enhancement feature tensor, through hyperbolic tangent activation function and parameter matrix ( , , The operation, combined with the Softmax function, calculates the attention score at each time point. Quantify the importance of features at different times. Calculate the attention scores at each time point. By weighted fusion with the corresponding spatial enhancement feature tensor of industrial water treatment, the feature contribution at key time points is highlighted, while capturing the dynamic dependencies between features at different times, ultimately generating a spatiotemporal feature tensor that simultaneously contains spatial correlation and temporal dynamics. .

[0124] Then, the current moment The spatial enhancement feature tensor of industrial water production and the previous time step MGU activation value Input into MGU and update to get the current time. activation value This serves both as the historical state for the next moment's temporal attention calculation and as a means to capture the temporal dependency between current and historical features. Ultimately, through multiple iterations of this kind, a spatiotemporal feature tensor that integrates spatial correlation and dynamic temporal dependency is formed, providing complete spatiotemporal information support for subsequent predictions.

[0125] Therefore, the formulas corresponding to S4-2-2 to S4-2-3 are:

[0126] ;

[0127] ;

[0128] ;

[0129] ;

[0130] in, Indicates the current time Next The output data of the hyperbolic tangent activation function corresponding to each time step. This represents the hyperbolic tangent activation function. , , All indicate the first The learnable parameter matrix of the temporal attention mechanism at each time step Indicates MGU at time Input data (spatial augmentation feature tensor of industrial water treatment). Indicates the first Bias terms of the time attention mechanism at each time step This indicates the length of the time series under consideration or the size of the relevant time window. This represents the minimum gate unit algorithm. Indicates the current time The feature vector that integrates key temporal information from multiple time steps is then used. Indicates the current time The spatial enhancement feature tensor of industrial water production.

[0131] S5. Predict the spatiotemporal feature tensor using a fully connected network to generate predicted values ​​for coagulant dosage. The corresponding expression is:

[0132] ;

[0133] in, , These represent the weight parameters and bias terms of a fully connected network, respectively.

[0134] In summary, this invention employs a full-link optimization design encompassing "sparse data completion - spatial feature aggregation - temporal feature extraction." At the data completion level, it leverages the Transformer encoder-decoder architecture combined with sparse self-attention and top-k sparsity strategies to effectively improve the accuracy of missing data completion and provide high-quality, missing-free input data for subsequent prediction modeling. At the spatial feature representation level, it utilizes a graph attention network to capture the node relationships of multiple subsystems in the water treatment system. By dynamically calculating the association weights between nodes and aggregating neighbor features, it increases the information entropy of spatial features, significantly enhancing their ability to express the coupling relationships of multiple nodes and strongly supporting the association learning between coagulant dosage and the state of each node. At the temporal prediction performance level, the temporal attention mechanism can dynamically focus on key temporal nodes, further improving the utilization rate of key temporal features and increasing the accuracy of coagulant dosage prediction.

[0135] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0136] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0137] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for predicting coagulant dosing by fusing sparse coding with graph spatiotemporal attention, characterized in that, include: A sparse data completion network is constructed based on Transformer encoders and decoders, combined with sparse self-attention mechanism and top-k sparsification strategy; A time feature extraction module is constructed based on residual MLP neural network and minimum gate unit algorithm, combined with time attention mechanism; A spatial feature aggregation module is constructed based on residual MLP neural network and graph attention mechanism; Industrial water production data is collected by sensors and used as missing data; the missing data is filled in by a sparse data network, the time series completion matrix is ​​calculated, and complete time series data is generated. The temporal and spatial features of the complete time series data are extracted sequentially through the spatial feature aggregation module and the temporal feature extraction module, generating a spatiotemporal feature tensor that integrates spatial neighbor association and historical temporal dependency; Predictive values ​​for coagulant dosage are generated by using a fully connected network to predict spatiotemporal feature tensors. The sparse data completion network includes an encoding module consisting of M stacked encoding layers, a decoding module consisting of M stacked decoding layers, and a fully connected layer; each encoding layer and each decoding layer includes a cascaded sparse self-attention layer, a normalization layer, and a linear transformation layer. The sparse self-attention layer includes a query vector feature transformation branch, a key vector feature transformation branch, a value vector feature transformation branch, and a top-k sparse layer; The top-k sparse layer includes the top-k layer, the softmax layer, and the multilayer perceptron; The generation of complete time-series data includes: The high-dimensional feature vector of the missing data is extracted by the encoding module. The temporal dynamics of the missing data and the spatial correlation of different nodes are captured layer by layer by combining the sparse self-attention mechanism and the top-k sparsification strategy to generate high-dimensional encoded features. The high-dimensional encoded features are cascaded and passed to the decoding module. The missing values ​​of the missing data are recovered through a sparse self-attention layer to generate a temporal completion matrix. Except for the input of the first decoding layer, which is the high-dimensional encoded features, the input of the other decoders is the fusion data of the high-dimensional encoded features and the previous decoding layer. The timing completion matrix is ​​input into the fully connected layer, and the complete timing data is output.

2. The method for predicting coagulant dosing by fusing sparse coding and graph spatiotemporal attention according to claim 1, characterized in that, The query vector feature transformation branch and the value vector feature transformation branch both include two layers of multilayer perceptrons; the key vector feature transformation branch includes two layers of multilayer perceptrons connected in series and a transpose transformation layer.

3. The method for predicting coagulant dosing by fusing sparse coding and graph spatiotemporal attention according to claim 2, characterized in that, The processing procedure for the sparse self-attention layer includes: Retrieve the missing data and generate query vector, key vector, and value vector respectively; The query vector and value vector are input into the query vector feature transformation branch and the value vector feature transformation branch respectively for processing. The corresponding two-layer multilayer perceptron are used to perform feature compression and feature recovery in sequence to obtain the query feature vector and the value feature vector. The key vector is input to the key vector feature transformation branch for processing. It is then processed by two corresponding multilayer perceptrons for feature compression and feature recovery, and combined with the transpose transformation layer for dimension alignment to obtain the key feature vector. Based on the query feature vector, value feature vector, and key feature vector, and combined with the top-k layer and softmax layer, the initial sparse self-attention value is calculated. The initial sparse self-attention value is then linearly transformed by a multilayer perceptron to align with the same dimension as the missing data, thus obtaining the sparse self-attention value. The output of the first encoding layer is generated based on sparse self-attention values ​​and missing data.

4. The method for predicting coagulant dosing by fusing sparse coding and graph spatiotemporal attention according to claim 1, characterized in that, The formula corresponding to the sparse self-attention layer is: ; ; ; ; in, , , These represent the query feature vector, key feature vector, and value feature vector, respectively. This represents the input data for the sparse self-attention layer. , , Both represent multilayer perceptrons. This represents the activation function. The transpose matrix representing the key eigenvectors. This represents the output data of the sparse self-attention layer. This indicates the top-k operation.

5. The method for predicting coagulant dosing by fusing sparse coding and graph spatiotemporal attention according to claim 1, characterized in that, The processing steps of the spatial feature aggregation module include: By using residual MLP neural networks and activation functions to enhance features of complete time-series data, the correlation between different nodes in the complete time-series data is captured, and a higher-order latent space representation is obtained. Higher-order latent space representations are used as graph nodes; based on the higher-order latent space representations, a graph connectivity matrix is ​​constructed using the K-nearest neighbor algorithm; according to the graph connectivity matrix, it is determined whether different graph nodes are connected or strongly correlated, and the higher-order latent space representations of the connected or strongly correlated nodes are retained to determine the strongly correlated graph nodes. Based on graph attention mechanism, exponential function and activation function, the graph attention score of each strongly correlated graph node is calculated; based on each graph attention score and the corresponding high-order latent space representation, the spatial feature tensor of industrial water treatment is generated.

6. The method for predicting coagulant dosing by fusion of sparse coding and graph spatiotemporal attention according to claim 1, characterized in that, The processing steps of the time feature extraction module include: The spatial feature tensor of industrial water treatment output by the spatial feature aggregation module is obtained, and the feature is enhanced by MLP neural network and activation function to obtain the spatially enhanced feature tensor of industrial water treatment. The temporal attention score of the spatial enhancement feature tensor of industrial water treatment is calculated using the minimum gate unit algorithm and the hyperbolic tangent activation function. Based on the attention scores at each time and the corresponding spatial enhancement feature tensors of industrial water treatment, the temporal dependency features between the spatial enhancement feature tensors of industrial water treatment are captured, and a spatiotemporal feature tensor is generated.

7. The method for predicting coagulant dosing by fusing sparse coding and graph spatiotemporal attention according to claim 1, characterized in that, The sparse data completion network, spatial feature aggregation module, and temporal feature extraction module are trained using the mean square error loss function and backpropagation algorithm.

Citation Information

Patent Citations

  • Missing data completion method based on sparse topology space-time attention mechanism

    CN118760677A

  • Memory access adaptive self-attention mechanism for transducer models

    CN119998815A