Method and system for constructing runoff data interpolation model
By constructing a runoff data interpolation model coupled to Bi-LSTM and multi-head attention module, the problem of collaborative modeling of local and global timing characteristics in the hydrological field is solved, and accurate interpolation of runoff data missing values is achieved, improving the interpolation accuracy and adaptability of the model.
Patent Information
- Application Number
- CN202510492429.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to effectively coordinate the modeling of local and global timing features in the field of hydrology, resulting in the loss of key timing dynamic information. Traditional models lack the ability to integrate space-time features in traditional models, making it difficult to accurately interpolate the missing values of runoff data.
A runoff data interpolation model coupled with space-time is constructed, combining Bi-LSTM module and multi-head attention module, through self-attention mechanism and mask self-attention mechanism, multi-level fusion of global and local timing features is achieved, and the ability to reconstruct missing values is enhanced.
The joint characterization of the spatiotemporal heterogeneity characteristics of runoff sequences is realized, the accuracy and robustness of missing value implication are improved, the limitations of traditional models are overcome, and the interpolation needs are adapted to different deletion rate scenarios.
Smart Images

Figure CN120408051A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of hydrological data processing, and more specifically, relates to a method and system for constructing a runoff data interpolation model. Background Art
[0002] In the field of hydrology, the processing of missing values in time series data is a key challenge, and its quality directly affects the accuracy of subsequent analysis, simulation, prediction and other tasks. In recent years, machine learning and deep learning have been widely used in the task of missing value interpolation in time series data.
[0003] However, machine learning relies on artificial feature construction (such as sliding window statistics, periodic features) to establish local time series mapping, but the coverage of its features is limited by prior assumptions, and it is difficult to capture long-range local and global dependencies and cross-scale hydrological dynamics. Deep learning methods (LSTM, Transformer, etc.) automatically extract multi-scale features through end-to-end learning and implicitly model time series dependencies. However, deep networks are vulnerable to gradient problems, resulting in reduced sensitivity to local mutation events, and their single capture of time series features faces the problem of weak two-way spatio-temporal feature fusion ability. Both types of methods face the bottleneck of collaborative modeling of local and global features: machine learning is prone to over-smoothing extreme events, and deep learning may lose key time series details due to architectural defects (such as convolution receptive field limitations or redundant attention calculations).
[0004] Generally speaking, the above single models are difficult to balance the collaborative modeling of local features and global long-term dependencies, which is prone to the loss of key time series dynamic information. The traditional data interpolation model Transformer faces the problem of weak two-way spatio-temporal feature fusion ability, while the simple bidirectional long short-term memory network (Bi-LSTM) interpolation model is insufficient in the collaborative capture of local time series features and global associations. These problems urgently need to be broken through by taking into account the integration of two-way spatio-temporal relationship capabilities and long sequence capture.
[0005] Therefore, how to solve the problem of the loss of key time series dynamic information in the field of hydrology is a difficult problem that urgently needs to be studied at present. Summary of the Invention
[0006] Aiming at the defects of the prior art, the purpose of this application is to provide a method and system for constructing a runoff data interpolation model, which can realize the interpolation of key time series dynamic information in the field of hydrology.
[0007] To achieve the above purpose, in the first aspect, this application provides a method for constructing a runoff data interpolation model, which is used to interpolate the missing values of time series data in the field of hydrology, and includes the following steps:
[0008] S10, based on the spatio-temporal distribution characteristics of hydrological stations in the river basin, construct a time series runoff data set;
[0009] S20, constructing a coupled spatiotemporal interpolation model architecture, including an encoder and a decoder;
[0010] The encoder consists of a Bi-LSTM module and a multi-head attention module. The Bi-LSTM module extracts local temporal features of the time series data layer by layer through a bidirectional information transmission mechanism. The multi-head attention module is used to capture the spatiotemporal relationship in the time series data. The Bi-LSTM module and the multi-head attention module work together through a hierarchical processing flow. The decoder uses a masked self-attention mechanism combined with a feature weight fusion module to dynamically correct missing values.
[0011] S30: Using the time series runoff dataset to perform model training and optimization on the interpolation model.
[0012] The present application provides the following beneficial effects: The proposed method for constructing a runoff data interpolation model deeply synergizes the Transformer's self-attention mechanism with the Bi-LSTM module, achieving multi-level fusion of global and local temporal features through structural reorganization. In the encoder, the Bi-LSTM module replaces the traditional feedforward network layer. Its bidirectional gating unit, based on the global context weights generated by the self-attention mechanism, performs forward and backward bidirectional temporal scanning of the sequence, extracting local dynamic patterns and phase-sensitivity features of adjacent time steps layer by layer. The self-attention mechanism models long-span dependencies through multi-head interaction, while the Bi-LSTM enhances the refined capture and memory update of local temporal gradients through the joint regulation of forget gates, input gates, and output gates. The two achieve feature interaction through residual connections and layer normalization: the global correlation weights output by the self-attention mechanism serve as the input to the Bi-LSTM, driving it to exploit the implicit short-term fluctuation patterns in the sequence. Simultaneously, the bidirectional temporal features output by the Bi-LSTM are recalibrated through the attention weights, enhancing the ability to reconstruct missing contextual information. This nested collaboration between the bidirectional loop structure and the self-attention mechanism not only overcomes the shortcomings of the pure attention model in insufficient analysis of local temporal dynamics, but also makes up for the gradient attenuation problem of Bi-LSTM in long-range dependency modeling, thereby achieving a joint representation of the spatiotemporal heterogeneity of runoff sequences and accurate interpolation of missing values.
[0013] As a further preferred embodiment, the multi-head attention module calculates multiple attention heads in parallel, uses multiple sets of linear projection matrices to perform global interaction on time series data, and dynamically allocates attention through Softmax normalized weights;
[0014] The Bi-LSTM module updates the hidden state time-step by time-step through the forward-backward bidirectional path timing processing, and adjusts the forward-backward spatiotemporal relationship through the gating mechanism.
[0015] As a further preference, in step S20, the hierarchical processing flow is specifically as follows:
[0016] Each encoder layer in the encoder first generates a global feature representation output by the multi-head attention module; subsequently, the Bi-LSTM module is used to perform fine-grained temporal optimization on the features; finally, feature fusion is achieved through residual connection and normalization operations.
[0017] As a further preference, the temporal runoff dataset includes a training set and a validation set;
[0018] Among them, the training set is obtained by cleaning and processing the temporal data measured at the hydrological stations; the validation set is runoff sequences with different missing rates, and this runoff sequence is generated by randomly erasing some data in the temporal data measured at the hydrological stations.
[0019] As a further preference, the validation set also includes runoff sequences with randomly generated 10% - 90% missing patterns.
[0020] As a further preference, in model training, the Adam optimizer and the MSE loss function are adopted, and an early stopping mechanism is set to prevent overfitting.
[0021] As a further preference, in model training, the feature transfer efficiency is optimized through residual connection and normalization operations.
[0022] As a further preference, in model training, the root mean square error, the mean absolute error, and the mean relative error are used as evaluation indicators to evaluate the imputation performance of the model.
[0023] As a further preference, in model optimization, the trained imputation model is deployed to an actual runoff data imputation system to perform real-time imputation on new runoff data, and based on the feedback data in actual applications, the model is further optimized and adjusted to meet the data imputation requirements of different stations and different missing rates.
[0024] In a second aspect, the present application provides a runoff data imputation model construction system, including:
[0025] A dataset construction module, configured to construct a temporal runoff dataset based on the spatio-temporal distribution characteristics of hydrological stations in a river basin;
[0026] An architecture construction module, configured to construct an imputation model architecture that couples space and time, including an encoder and a decoder;
[0027] Among them, the encoder includes a Bi-LSTM module and a multi-head attention module. The Bi-LSTM module extracts local temporal features in the forward and backward directions of the sequence layer by layer through a bidirectional information transmission mechanism. The multi-head attention module is used to capture global dependencies in the sequence. The Bi-LSTM module and the multi-head attention module complete their collaborative effect through a hierarchical processing flow; the decoder adopts a masked self-attention mechanism and combines a feature weight fusion module to dynamically correct missing values;
[0028] A model training and optimization module for training and optimizing the imputation model using the temporal runoff data set.
[0029] It can be understood that the beneficial effects of the second aspect above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings
[0030] Figure 1 is a flowchart of a method for constructing a runoff data imputation model provided by an embodiment of the present application;
[0031] Figure 2 is a model framework diagram provided by an embodiment of the present application;
[0032] Figure 3 is a structure diagram of an encoder and a decoder provided by an embodiment of the present application;
[0033] Figure 4 is a structure diagram of multi-head attention provided by a specific embodiment of the present application;
[0034] Figure 5 is a diagram of the Bi-LSTM calculation process provided by a specific embodiment of the present application;
[0035] Figure 6 are different model results of the Yichang Station data set provided by a specific embodiment of the present application;
[0036] Figure 7 are different model results of the Shashi Station data set provided by a specific embodiment of the present application;
[0037] Figure 8 are different model results of the Hankou Station data set provided by a specific embodiment of the present application;
[0038] Figure 9 are imputation results of different models under different missing rates (10%-90%) of the Yichang Station data set provided by a specific embodiment of the present application. Detailed Embodiments
[0039] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0040] It should be understood that in the description of the present application, the meaning of the term "several" is at least one, for example, one, two, etc., unless otherwise specifically defined; the meaning of the term "multiple" is two or more, unless otherwise specifically defined; the terms "first" and "second" etc. are used to distinguish different objects, rather than to describe the specific order of the objects; the term "and / or" includes any and all combinations of one or more of the related listed items.
[0041] In addition, the reference to "one embodiment" throughout this specification; the language such as "one embodiment", "one example" or the like means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of the phrase "in one embodiment;" and "in one embodiment" and similar language throughout this specification may or may not all refer to the same embodiment.
[0042] Aiming at the problem of missing hydrological time series caused by common equipment failures and human operation errors during the hydrological runoff data collection process, the present application proposes a method for constructing a runoff data interpolation model based on the time series analysis theory, provides a theoretical support for ensuring the quality of hydrological data analysis, and verifies the model accuracy based on hydrological station data and under different missing rates.
[0043] The construction idea of the runoff data imputation model proposed in this application mainly considers how to integrate the temporal dependence and spatial heterogeneity of the runoff sequence into the data repair process under the guidance of the hydrological spatio-temporal correlation mechanism. The essence of the missing hydrological time series is the manifestation of the interruption of spatio-temporal information transmission. The repair of missing values needs to consider both the temporal continuity of the sequence itself and the spatial correlation between stations in the basin. In the river network system, the spatio-temporal evolution of station runoff has significant causal transmission characteristics. The runoff fluctuations at upstream stations affect downstream stations through the confluence process, and the intensity of this influence decays with the hydrological distance and generates a time delay. Therefore, the missing values are not only restricted by the historical data of their own stations, but also implicitly affected by the comprehensive spatio-temporal characteristics of adjacent stations. Traditional imputation methods often break the spatio-temporal correlation, resulting in the repair results deviating from the true hydrological laws. In this application, a collaborative architecture of Bi-LSTM and Transformer is constructed to transform the spatial topological relationship of hydrological stations into the feature interaction channels of the model. Bi-LSTM is used to capture local bidirectional temporal patterns, and the self-attention mechanism is combined to model long-range spatial dependence, forming a dynamic imputation framework based on spatio-temporal coupling. This design is constrained by the physical mechanism of hydrological confluence, and the spatio-temporal collaborative repair of missing values is realized through a dynamic weight fusion module, providing a new technical path for improving the imputation accuracy and interpretability in complex missing scenarios.
[0044] Therefore, the design concept of the imputation model proposed in this application is to make full use of the spatio-temporal correlation features hidden in the spatial network structure of hydrological stations as the interaction interface for temporal feature learning and spatial information fusion, so that the model can not only capture the temporal continuity of the runoff sequence but also depict the spatial collaboration between stations. The construction goal is to establish a feature transfer mechanism based on the spatial proximity relationship of stations, and realize the collaborative repair of multi-station spatio-temporal features through a dynamic weight matrix. For the missing values of any station, the repair process not only depends on the historical data of the station itself but also integrates the spatio-temporal feature information of adjacent stations. This information fusion is weighted by distance decay through a spatial weight matrix to form a spatio-temporal joint feature expression with physical significance.
[0045] As Figure 1 shown, the method for constructing a runoff data imputation model provided in this application is used to impute the missing values of time series data in the hydrological field, including steps S10 to S30, which are described in detail as follows:
[0046] Step S10: Construct a time series runoff data set based on the spatio-temporal distribution characteristics of hydrological stations in the river basin.
[0047] In step S10, the time series runoff data set includes a training set and a validation set.
[0048] Among them, the training set is obtained by cleaning and processing the measured data of hydrological stations in the hydrological and water resources center. The validation set is runoff sequences with different missing rates, which are generated by randomly erasing some data in the measured data of hydrological stations.
[0049] Preferably, the validation set also includes runoff sequences with 10% - 90% randomly generated missing patterns. Subsequent model experiments are verified by constructing multi-gradient missing scenarios to target the imputation performance under different missing rates with the same missing pattern.
[0050] S20. Construct an imputation model architecture that couples space and time, namely a model that couples Transformer and Bi-LSTM (abbreviated as TBLformer imputation model). This model architecture includes an encoder and a decoder.
[0051] In step S20, the encoder includes a multi-head attention layer and a Bi-LSTM layer, which are used to jointly extract global dependency relationships and bidirectional temporal features. Among them, the multi-head attention module is used to capture the spatio-temporal relationships in the temporal data (sequence). The Bi-LSTM module extracts the local temporal features in the forward and backward directions of the sequence layer by layer through a bidirectional information transmission mechanism. The multi-head attention module and the Bi-LSTM module complete their collaborative effects through a hierarchical processing flow.
[0052] Specifically, the multi-head attention mechanism calculates multiple attention heads in parallel, and uses multiple groups of linear projection matrices (Q, K, V) to perform global interaction on the sequence, endowing the model with the ability to focus on long-distance associations between any positions. This mechanism dynamically allocates attention through Softmax-normalized weights, enabling each element in the sequence to aggregate global context information, especially suitable for inferring the association between randomly missing positions and distal data in the time series missing value imputation task. At the same time, the Bi-LSTM layer updates the hidden state step by step through the temporal processing of the forward and backward double paths. Its gating mechanism (such as the forget gate, input gate, and output gate) adjusts the spatio-temporal relationships in the forward and backward directions, enhancing the sensitivity to local temporal patterns (such as periodic changes or continuous fluctuations between adjacent time steps). The collaborative effect of the two types of modules is achieved through a hierarchical processing flow: each encoder layer first generates a global feature representation of the multi-head attention output, and then the Bi-LSTM performs fine-grained temporal optimization on the features. Finally, feature fusion is achieved through residual connection and normalization operations. The residual structure ensures the stable transmission of gradients of global context information and local temporal dynamics, avoiding the problem of information degradation in deep networks.
[0053] The decoder adopts a masked self-attention mechanism and combines a feature weight fusion module to dynamically correct missing values. The model embeds absolute position information through position encoding to ensure the effective transmission of sequence order information. Further, the encoder structure in step 2 is as Figure 3As shown, each encoder layer contains a multi-head attention module and a Bi-LSTM module. The multi-head attention mechanism realizes the dynamic association of different time steps by calculating the similarity weights of queries, keys, and values. The Bi-LSTM module captures forward and backward dependencies through bidirectional propagation, and its structure is as Figure 5 . A feature weight fusion module is designed to achieve dynamic correction of missing values. This module uses a mask matrix M to retain non-missing data, combines the Sigmoid activation function to calculate the imputation probability, and finally generates a complete sequence. This design effectively alleviates the error accumulation problem of traditional imputation methods in high missing rate scenarios.
[0054] The TBLformer imputation model provided in this embodiment structurally optimizes the problems of insufficient capture of local features and limited long-range dependence modeling in the runoff time series data imputation task by synergistically integrating the self-attention mechanism of Transformer and the bidirectional time series modeling ability of Bi-LSTM. Although the traditional Transformer relies on the self-attention mechanism to effectively capture global associations, its ability to model local dynamic patterns and phase sensitivities of adjacent time series is weak; while Bi-LSTM can extract local time series features in the forward and backward directions layer by layer through a bidirectional information transfer mechanism, but it is difficult to model long-span dependence relationships. TBLformer introduces a Bi-LSTM module in the encoder to replace the standard feed-forward network layer, enabling it to further strengthen the refined modeling of local time series features through the bidirectional gated units of Bi-LSTM on the basis of the self-attention mechanism capturing global context associations. This collaborative architecture enables the model to simultaneously fuse global attention weights and local time series gradient information, uses the long-range dependence modeling advantage of the self-attention mechanism to solve the gradient decay problem of Bi-LSTM, and at the same time uses the bidirectional feature extraction ability of Bi-LSTM to make up for the deficiency of the pure attention mechanism in parsing local dynamic patterns, so as to achieve multi-level joint representation of complex spatio-temporal heterogeneity in the runoff sequence. The complementary collaborative optimization of the two significantly improves the model's ability to reconstruct context information at missing positions, and enhances the robustness and accuracy of the imputation results.
[0055] S30. Use the time series runoff dataset to train and optimize the TBLformer imputation model.
[0056] Furthermore, in model training, the Adam optimizer and the MSE loss function can be used, and an early stopping mechanism can be set to prevent overfitting. Optimize the feature transfer efficiency through residual connection and normalization operations.
[0057] Specifically, during model optimization, the trained TBLformer interpolation model can be deployed to the actual runoff data interpolation system to perform real-time interpolation of new runoff data. Based on feedback data from actual applications, the model can be further optimized and adjusted to adapt to the data interpolation needs of different sites and different missing rates. At the same time, the model's computational efficiency and resource consumption are optimized to improve the model's applicability in large-scale data and real-time early warning systems.
[0058] The runoff data interpolation model construction method provided in this embodiment deeply synergizes the Transformer's self-attention mechanism with the Bi-LSTM module. This structural reorganization enables multi-level fusion of global and local temporal features. In the encoder, the Bi-LSTM module replaces the traditional feedforward network layer. Its bidirectional gating unit, based on the global context weights generated by the self-attention mechanism, performs forward and backward bidirectional temporal scanning of the sequence, extracting local dynamic patterns and phase-sensitivity features of adjacent time steps layer by layer. The self-attention mechanism models long-span dependencies through multi-head interaction, while the Bi-LSTM enhances the refined capture and memory update of local temporal gradients through the joint regulation of forget gates, input gates, and output gates. Both achieve feature interaction through residual connections and layer normalization. The global correlation weights output by the self-attention mechanism serve as the input to the Bi-LSTM, driving it to exploit the implicit short-term fluctuation patterns in the sequence. Simultaneously, the bidirectional temporal features output by the Bi-LSTM are recalibrated using the attention weights, enhancing the ability to reconstruct contextual information at missing locations. This nested collaboration between the bidirectional loop structure and the self-attention mechanism not only overcomes the shortcomings of the pure attention model in insufficient analysis of local temporal dynamics, but also makes up for the gradient attenuation problem of Bi-LSTM in long-range dependency modeling, thereby achieving a joint representation of the spatiotemporal heterogeneity of runoff sequences and accurate interpolation of missing values.
[0059] Based on the same inventive concept, the present application also provides a runoff data interpolation model construction system, including a data set construction module, an architecture construction module and a model training optimization module.
[0060] The dataset construction module is used to construct a time-series runoff dataset based on the spatiotemporal distribution characteristics of hydrological stations in the river basin.
[0061] The architecture construction module is used to construct an imputation model architecture that couples time and space, including an encoder and a decoder. Among them, the encoder includes a Bi-LSTM module and a multi-head attention module. The Bi-LSTM module extracts local temporal features in the forward and backward directions of the sequence layer by layer through a bidirectional information transmission mechanism. The multi-head attention module is used to capture global dependencies in the sequence. The Bi-LSTM module and the multi-head attention module complete their collaborative effects through a hierarchical processing flow; the decoder adopts a masked self-attention mechanism and combines a feature weight fusion module to dynamically correct missing values.
[0062] The model training and optimization module is used to train and optimize the imputation model using the time series runoff dataset.
[0063] Specifically, for the functions of each module provided in this embodiment, reference can be made to the detailed description in the foregoing method embodiment, and details will not be repeated here.
[0064] To more clearly illustrate the runoff data imputation model construction method provided in this application, the following is a corresponding description in combination with specific embodiments:
[0065] First, the middle reaches of the Yangtze River are selected as the research area, with geographical coordinates between 110°15′ and 117°08′ east longitude and between 27°35′ and 31°25′ north latitude. The flow observation data of 3 hydrological observation stations in the middle reaches of the Yangtze River basin are selected as the input variables for the research. The time range is from January 2008 to June 2012, and the stations are Yichang Station, Shashi Station, and Hankou Station respectively. The observed runoff data of the stations come from the measured data of hydrological stations of the Hubei Provincial Hydrology and Water Resources Center. The selected research data time range is from 2008 to 2011. After cleaning and processing the runoff data of the 3 stations, the data from 2008 / 01 to 2009 / 12 are used as the training set, a set of hyperparameters are set, and the model is trained using the training data. The data for the next 12 months, from 2010 / 01 to 2011 / 01, are used as the validation set, and the data from 2010 / 01 to 2010 / 12 are imputed for model validation and evaluation.
[0066] Step 1, based on the spatio-temporal distribution characteristics of hydrological stations in the middle reaches of the Yangtze River basin, construct a time series runoff dataset. The sequence data is the observed runoff data of stations from 2008 to 2012. To verify the effectiveness of the imputation model, a part of the data is randomly erased to construct runoff sequences with different missing rates. By artificially generating runoff sequences with 10% - 90% random missing patterns, multiple gradient missing scenarios are constructed for experimental verification, so as to target the imputation performance under different missing rates with the same missing pattern.
[0067] Step 2, construct a TBLformer imputation model architecture, including an encoder and a decoder module, as Figure 2As shown in the figure. The encoder consists of a multi-head attention layer and a Bi-LSTM layer to jointly extract global dependency relationships and bidirectional temporal features; the decoder uses a masked self-attention mechanism and combines a feature weight fusion module to dynamically correct missing values. The structures of the encoder and decoder are as shown in Figure 3 the figure. The model embeds absolute position information through positional encoding to ensure the effective transmission of sequence order information.
[0068] The input to the encoder is the processed time series. We first define the input to the encoder as P en respectively. The encoder contains N encoder layers. For the n-th layer encoder, the specific implementation process can be given by the following formula:
[0069]
[0070] The above formula can be further generalized as:
[0071]
[0072] where, represents the output of the n-th layer encoder. When n = 1, is the embedded This represents the feature information extracted after the i-th normalization layer module in the n-th layer encoder. R-Norm(·) represents the residual connection, followed by normalization. In addition, Mul-HeadAtt(·) and Bi-LSTM(·) represent the multi-head attention module and the Bi-LSTM module respectively, which will be introduced in the subsequent steps.
[0073] Step 2.1, Data Input Mapping Module. In the data preprocessing stage, the embedding layer operation is first performed: the low-dimensional representation of multivariate data is realized through feature dimensionality reduction and vectorization transformation. At the same time, the positional encoding technology is used to encode the absolute position information of the time series data, and the positional encoding vector generated by the sine and cosine functions is added element by element to the embedding vector to integrate the time series position information into the input features. Data mapping is responsible for constructing the feature space mapping of the time series. Through dimension alignment and feature normalization processing, it ensures that the spatio-temporal information of the original data is accurately captured during the feature encoding process. By learning the positional encoding, the model can better understand the relative position relationship of each element in the input sequence. The data embedding layer embeds the input time series data into the dimensional space required by the attention layer. The input layer receives a series of data points with timestamps. The calculation formula is as shown below.
[0074] For odd indices 2i + 1, there is:
[0075]
[0076] For even indices 2i, there is:
[0077]
[0078] Among them, pos represents each position in the data, i represents the dimension index, and d model represents the dimension of the model. Therefore, for different features, their positional encodings have different frequencies.
[0079] Step 2.2, Feature Encoding Module. The goal of the feature encoder is to interact with each element in the sequence through the self-attention mechanism and calculate weights based on the relationship between each element and other elements, so as to extract the deep feature representation in the sequence. In the time series data missing value imputation task, since the positions of the missing data are random, we cannot pre-determine whether there is a correlation between each element and the elements at other positions. Therefore, the feature encoder can capture higher-level abstract features by learning the context information of the sequence and the mutual relationship between elements. The multi-head attention mechanism distributes attention to the sequence data through multiple heads, which can enhance the feature extraction ability for the sequence data. The internal structure of the multi-head attention mechanism is as Figure 4 shown. Specifically, the input of the self-attention module consists of Q (Query) with dimension, K (Key), and V with dimension d V values.
[0080] Here, d K and d V represent the dimensions of the key and the value respectively, that is, the linear projection, and the calculation formula is as follows.
[0081]
[0082] In the formula, Softmax(·) represents the Softmax normalization function. In addition, the multi-head attention can be expressed as:
[0083] MultiHead(Q, K, V) = W out ·Concat(head1,…,head i )
[0084] In the formula, Q, K, V are matrices obtained through linear transformation. This self-attention can focus on all positions of the runoff data and can learn long-distance dependencies. In addition, the multi-head attention is calculated in parallel mode to reduce time consumption.
[0085] Step 2.3, Relationship capture module. The Bi-LSTM module can extract the features of runoff data layer by layer. Different from the fully connected layer in the classical Transformer, such a design can further improve the prediction performance of the proposed algorithm. This part focuses on processing the relationships between dimensions and instead hands over the temporal relationships to the Bi-LSTM module for processing, which can more efficiently capture the features and patterns of time series. Doing so can improve the overall efficiency and accuracy of the model in processing time series data.
[0086] For a single LSTM, it consists of a forget gate f t , an input gate i t and an output gate o t . f t is composed of the previous layer output h t-1 and the current input p t , with weights W f , and bias b f , and can be expressed as the following formula:
[0087]
[0088] where σ is the Sigmoid activation function. Then, the runoff sequence is input into i t , and can be expressed as:
[0089]
[0090] W i , and b i represent the weights and biases of the input gate respectively. Here, the LSTM has a state update unit, which can be given by the following formula:
[0091]
[0092] where λ is the tanh activation function, and W c , and b c are the weights and biases of the update unit respectively.
[0093] Then, the data is updated through the following formula.
[0094]
[0095] In the above formula, the circled multiplication means multiplying the corresponding elements of two matrices. Then, the updated data enters o t .
[0096]
[0097] ht = o t ⊙ λ(c t )
[0098] In the formula, h t is the final output. W o , and b o are the weight and bias of the output gate respectively. When training a single Bi-LSTM layer, first calculate the forward output and the backward output according to the above formula. Then, the output of the Bi-LSTM layer is given by the following formula:
[0099]
[0100] where the σ function represents the connection merging mode, and such a design can better extract the features of runoff time series data layer by layer.
[0101] Step 2.4, the decoder layer contains a single linear layer and is normalized. The decoder contains E decoding layers. For the E-th decoding layer, the specific implementation process can be expressed as the following formula:
[0102]
[0103] It can be further summarized as:
[0104]
[0105] In the formula, represents the output of the e-th decoding layer. When e = 1, represents the embedded represents the feature information extracted after the i-th R-Norm module in the e-th decoding layer. Among them, Linear(·) and Norm(·) represent the linear projection layer (projecting p to the target dimension) and the normalization layer respectively. Note that masking is used for the first multi-head attention mechanism.
[0106] Bi-LSTM is used as the weight fusion module in the decoder, and its calculation process is as Figure 5 shown. Each Bi-LSTM layer contains multiple LSTMs connected by bidirectional rules, and Bi-LSTM combines forward and backward propagation. Taking the Bi-LSTM module in the encoder layer as an example, the hidden layer of the bidirectional LSTM model needs to save two h t values, and the final output value obtained by combining the outputs of the forward layer and the backward layer is mathematically expressed as follows:
[0107] h t = f(w1x t + w2h t-1 )
[0108] h′ t = f(w3x t + w5h′ t-1 )
[0109] o t = g(w4h t + w6h′ t )
[0110] where, w i (i = 1, 2, …, 6) are six independent weight matrices, the weights (w1, w3) input to the forward and backward hidden layers, the weights (w2, w5) between hidden layers, and the weights (w4, w6) from the forward and backward hidden layers to the output layer. The values of the six weights are reused at each time step.
[0111] Step 3: Implement model training based on the PyTorch framework. After pre-training optimization, 3 encoder layers (i.e., N = 3) and 3 decoder layers (i.e., E = 3) are adopted. The loss function is the root mean square error (MSE), and the epoch is set to 80. The batch size for each iteration is 32, and all are optimized using the Adam optimizer to improve the convergence speed of the model. The model adopts the early stopping method with a patience value of 6, an initial learning rate of 0.001, the number of multi-head attentions is set to 8, and an early stopping mechanism is set to prevent overfitting. The feature transfer efficiency is optimized through residual connections and normalization operations.
[0112] Step 4: Perform accuracy verification on the datasets of three stations in Yichang, Shashi, and Hankou, as well as datasets with different missing rate precisions.
[0113] Figure 6 、 Figure 7 、 Figure 8The imputation results of multiple interpolation models at three stations are shown. The results indicate that there are significant differences in model performance, and a consistent optimization trend is presented among different stations. The TBLformer model demonstrates the best performance among the three stations. The statistical imputation method, Median, shows poor performance and large errors on the three datasets. Taking Yichang Station as an example, its MAE, RMSE, and MRE indicators are reduced by 83.6%, 76.0%, and 85.6% respectively compared with the imputation methods based on statistical learning, verifying the effectiveness of its spatio-temporal feature fusion architecture. Compared with Random Forest and RNN, the MAE indicator is reduced by an average of 70.05%, the RMSE indicator is reduced by an average of 75.1%, and the MRE indicator is reduced by an average of 74.7%. This indicates the limitations of these two methods in the presence of runoff time series data. For LSTM and Bi-LSTM, the MAE of TBLformer is reduced by 52.8% and 41.1% respectively compared with LSTM and Bi-LSTM. This optimization shows that TBLformer, which combines the self-attention mechanism, can capture long-range temporal dependencies more accurately and overcomes the gradient decay problem of LSTM in complex sequence modeling. The RMSE indicator further verifies its robustness. TBLformer is reduced by 51.9% and 40.0% respectively compared with LSTM and Bi-LSTM, indicating its effectiveness in suppressing extreme errors, especially in high-noise hydrological data. Through comparison, it is found that the sequence models (LSTM, Bi-LSTM) are significantly better than the non-sequence models (Random Forest, RNN). Among them, Bi-LSTM reduces the MAE by 19.8%-32.2% compared with the unidirectional LSTM through bidirectional information extraction, highlighting the enhancement effect of the bidirectional structure on temporal dependency relationships. It is worth noting that TBLformer shows better comprehensive performance compared with the standard Transformer model. The MAE of TBLformer is reduced by 17.6% compared with Transformer, the RMSE is reduced by 5.1%, and the MRE is reduced by 7.2%. This shows that it effectively captures temporal dependencies through the fusion mechanism with Bi-LSTM, making up for the deficiency of the pure self-attention mechanism of Transformer in modeling the relationships between adjacent time steps. Although the absolute reduction of RMSE and MRE is small, the consistent decline with MAE indicates that TBLformer has more advantages in the balance of error distribution.
[0114] To further explore the performance of TBLformer under different data missing rates, Yichang Station was selected, and its data missing rate was controlled within the range of 10% - 90% for experiments, and compared with other commonly used imputation methods, including Median (replacing missing values with the median in the training set), Random Forest, RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory Network), Bi-LSTM (Bidirectional Long Short-Term Memory Network), and Transformer. Tables 1 and 2 show the experimental results of the Yichang Station dataset at different missing rates from 10% to 90%. The radar chart of the imputation experimental results of each model on the Yichang Station dataset is as Figure 9 shown.
[0115] Table 1 Experimental results of different models on the Yichang Station dataset with a missing rate of 10% - 40%
[0116]
[0117]
[0118] Table 2 Experimental results of different models on the Yichang Station dataset with a missing rate of 50% - 90%
[0119]
[0120]
[0121] As the missingness rate increases from 10% to 90%, the MAE, RMSE, and MRE metrics of all models show a significant upward trend, indicating that increasing levels of missing data pose greater challenges to the imputation task. Among them, the traditional methods Median and Random Forest perform the worst across all missingness rates. The traditional Median method's MAE reaches 2.252 at a missingness rate of 90%, an increase of approximately 9.96% from a missingness rate of 10%, reflecting its limited adaptability to complex missingness patterns. In contrast, deep learning-based models generally perform better, but the degree of performance degradation with increasing missingness rate varies significantly. For example, the Transformer's MAE is 32.11% lower than that of the RNN and LSTM at low missingness rates, and 24.58% lower than that of the RNN and LSTM, respectively, demonstrating its structure's ability to capture long-range dependencies. RNNs and LSTMs, however, struggle to effectively model long-range dependencies due to the vanishing gradient problem. As the missingness ratio increases to 50%, the MAE of the TBLformer remains significantly lower than that of the RNN and LSTM, with reductions of 24.23% and 19.77%, respectively. Furthermore, the Bi-LSTM significantly outperforms the unidirectional LSTM at all missingness ratios. At a moderate missingness ratio, the MAE of the Bi-LSTM further decreases by 11.38%. At a 90% missingness ratio, the RMSE of the Bi-LSTM decreases by 8.50% and the MRE decreases by 29.08%, demonstrating that the introduction of the Bi-LSTM enables more efficient utilization of sequential information, mitigating the local bias caused by the one-way information transfer of the unidirectional LSTM.
[0122] It is worth noting that TBLformer maintains the lowest error rate under all missing rate scenarios. By integrating the self-attention mechanism of Transformer with the temporal modeling ability of Bi-LSTM, TBLformer significantly improves its robustness. At low missing rates, the MAE, RMSE, and MRE of TBLformer are respectively reduced by 13.42%, 13.56%, and 13.38% compared to Transformer, and by 17.01%, 14.98%, and 17.07% compared to Bi-LSTM, indicating its ability to synergistically capture local sequence features and global dependencies. As the missing rate increases to 50%, the three indicators of TBLformer are reduced by 23.11%, 18.98%, and 22.89% compared to Transformer, and by 9.47%, 12.61%, and 9.35% compared to Bi-LSTM respectively, highlighting the inhibitory effect of its structural improvement on error accumulation. Under extreme missing rates, the MAE of TBLformer is reduced by 24.93% and 25.91% compared to Bi-LSTM and Transformer respectively, and its error increase of 34.56% is much lower than that of Bi-LSTM (48.75%) and Transformer (57.22%). This further verifies that its stable performance advantage can still be maintained under extreme missing conditions through the collaborative optimization of the bidirectional cyclic structure and the self-attention mechanism, effectively alleviating the problem of information breakage caused by high missing rates. Figure 9 From left to right and from top to bottom are the imputation results of different models at missing rates of 10% - 90%. TBLformer has the smallest coverage area on the three indicator axes and is close to the center of the chart, indicating its optimal global performance. This result verifies the effectiveness of TBLformer in combining the self-attention mechanism of Transformer with the bidirectional temporal modeling of Bi-LSTM.
[0123] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A method for constructing a runoff data interpolation model, which is used to interpolate missing values of time series data in the hydrological field, characterized in that, It includes the following steps: S10. Based on the spatio-temporal distribution characteristics of hydrological stations in the river basin, construct a time-series runoff dataset; S20. Construct an imputation model architecture that couples space and time, including an encoder and a decoder; Among them, the encoder includes a Bi-LSTM module and a multi-head attention module. The Bi-LSTM module extracts the local time-series features in the forward and backward directions of the time-series data layer by layer through a bidirectional information transfer mechanism. The multi-head attention module is used to capture the spatio-temporal relationships in the time-series data. The Bi-LSTM module and the multi-head attention module complete their collaborative functions through a hierarchical processing flow; the decoder adopts a masked self-attention mechanism and combines a feature weight fusion module to dynamically correct missing values; S30. Use the time-series runoff dataset to train and optimize the imputation model.
2. The method for constructing a runoff data interpolation model according to claim 1, characterized in that The multi-head attention module performs global interaction on the time-series data by parallelly computing multiple attention heads, uses multiple groups of linear projection matrices, and dynamically allocates attention through Softmax-normalized weights; The Bi-LSTM module updates the hidden state step by step through the time-series processing of the forward and backward paths, and adjusts the forward and backward spatio-temporal relationships through a gating mechanism.
3. The method for constructing a runoff data interpolation model according to claim 1, wherein In step S20, the hierarchical processing flow is specifically as follows: Each encoder layer in the encoder first generates a global feature representation output by the multi-head attention module; then, the Bi-LSTM module performs fine-grained time-series optimization on the features; finally, feature fusion is achieved through residual connection and normalization operations.
4. The method for constructing a runoff data interpolation model according to claim 1, wherein The time-series runoff dataset includes a training set and a validation set; Among them, the training set is obtained by cleaning and processing the measured time-series data of hydrological stations; the validation set is runoff sequences with different missing rates, and this runoff sequence is generated by randomly erasing some data in the measured data of hydrological stations.
5. The method for constructing a runoff data interpolation model according to claim 4, wherein, The validation set also includes runoff sequences with randomly generated missing patterns of 10% - 90%.
6. The method for constructing a runoff data interpolation model according to claim 1, wherein In model training, the Adam optimizer and the MSE loss function are adopted, and an early stopping mechanism is set to prevent overfitting.
7. The method for constructing a runoff data interpolation model according to claim 1, wherein In model training, the feature transfer efficiency is optimized through residual connection and normalization operations.
8. The method for constructing a runoff data interpolation model according to claim 1, wherein In model training, the root mean square error, mean absolute error, and mean relative error are used as evaluation indicators to evaluate the imputation performance of the model.
9. The method for constructing a runoff data interpolation model according to claim 1, wherein In model optimization, the trained imputation model is deployed to an actual runoff data imputation system to perform real-time imputation on new runoff data, and according to the feedback data in actual applications, the model is further optimized and adjusted to meet the data imputation requirements of different stations and different missing rates.
10. A runoff data interpolation model construction system, characterized in that, [[ID=1,6]]It includes: A dataset construction module, which is used to construct a time-series runoff dataset based on the spatio-temporal distribution characteristics of hydrological stations in the river basin; An architecture construction module, which is used to construct an imputation model architecture that couples space and time, including an encoder and a decoder; Among them, the encoder includes a Bi-LSTM module and a multi-head attention module. The Bi-LSTM module extracts local temporal features in the forward and backward directions of temporal data layer by layer through a bidirectional information transmission mechanism. The multi-head attention module is used to capture the spatio-temporal relationships in the temporal data. The Bi-LSTM module and the multi-head attention module complete their collaborative effects through a hierarchical processing flow; the decoder adopts a masked self-attention mechanism and combines a feature weight fusion module to dynamically correct missing values; A model training and optimization module, which is used to train and optimize the imputation model by using the temporal runoff data set.
Citation Information
Cited By
Hydrometric station supervision method and device
CN120873613A
Long-sequence electrocardiosignal disease recognition system based on Transform architecture
CN120983046A
Space-time data interpolation method based on space-time decoupling correction flow framework
CN121233916A
Urban water supply prediction method based on non-stationary perception Transform-BiLSTM model
CN121543827A
Hydrological sequence missing data complementation and trend prediction system and method
CN121958786A