Water quality spatio-temporal data interpolation device and method based on multi-bias tensor decomposition

Through the combination of multi-bias tensor decomposition and differential evolution algorithm, the problem of spatiotemporal data loss in water quality is solved, fast and accurate data interpolation is achieved, data integrity and reliability are improved, and water resource management and effective maintenance of ecosystems are supported.

CN120256840APending Publication Date: 2025-07-04CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510252260.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology cannot effectively solve the problem of missing data in the spatiotemporal data of water quality, resulting in incorrect statistical analysis and suboptimal decision-making, and the inability to extract the rich knowledge hidden in the data.

Method used

The water quality spatiotemporal data interpolation device and method based on multi-bias tensor decomposition is adopted. Through the water quality spatiotemporal data reception, storage, tensor structure, multi-bias fusion and interpolation module, the differential evolution algorithm is combined to perform rapid iterative calculations to achieve accurate interpolation of missing data.

Benefits of technology

It realizes rapid and accurate interpolation of spatiotemporal data of water quality, comply with statistical laws, improves the integrity and reliability of data, and supports water resource management and effective maintenance of aquatic ecosystems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256840A_ABST
    Figure CN120256840A_ABST
Patent Text Reader

Abstract

The invention discloses a water quality spatio-temporal data interpolation device and method based on multi-bias tensor decomposition, and belongs to the field of intelligent ecology. The device is composed of a water quality spatio-temporal data receiving module, a water quality spatio-temporal data storage module, a water quality spatio-temporal data tensor construction module, a water quality spatio-temporal data multi-deviation fusion module and a water quality spatio-temporal data interpolation module. The method comprises the following steps: S1, receiving a data acquisition instruction; s2, four-tuple formatting data is carried out; s3, constructing a water quality spatio-temporal data tensor; s4, creating initial water quality deviation fusion data corresponding to a plurality of offset methods; s5, initializing parameters; and S6, constructing a loss function, establishing an updating strategy, and realizing interpolation of missing water quality spatio-temporal data. According to the method, water quality data interpolation which accords with a statistical rule and is high in accuracy can be carried out, so that the problem of rapid and accurate interpolation of missing water quality data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a water quality spatio-temporal data imputation device and method based on multi-bias tensor decomposition, belonging to the field of intelligent ecology, and particularly to water quality spatio-temporal data imputation based on multi-bias tensor decomposition. Background Art

[0002] Water is a strategic target resource for the sustainable development of human society. Due to the acceleration of global warming and the enrichment of water body nutrients, water quality deterioration has become a major environmental problem worldwide, seriously threatening human survival and ecological stability. Therefore, it is necessary to use continuously monitored water quality spatio-temporal data as the basic data resource for water resource management and aquatic ecosystem maintenance. The progress of data analysis tools, the reduction of data storage costs, and the improvement of computing performance have transformed water quality monitoring technology from manual low-frequency collection relying on professional knowledge to systematic automatic high-frequency monitoring. The water quality automatic high-frequency monitoring technology continuously monitors multiple water quality index variables of water bodies through a monitoring sensor network to obtain real-time high spatio-temporal resolution data, accurately and timely capture the change trend of water quality and analyze its change law, providing reliable information for effective water resource management and aquatic ecosystem maintenance.

[0003] The reliability of water quality spatio-temporal data is affected by various factors, mainly including sensor failures, daily maintenance, equipment aging, abnormal power-off of monitors, server paralysis, data communication interruption, and database system crashes, etc., resulting in inevitable data missing in water quality spatio-temporal data. Directly applying incomplete water quality spatio-temporal data to downstream tasks will not only lead to incorrect statistical analysis and suboptimal decisions, but also fail to extract the rich knowledge hidden in the data. On the contrary, imputing the missing data in water quality spatio-temporal data can create more opportunities for understanding the studied aquatic ecosystem. Therefore, in water resource management and aquatic ecosystem maintenance, how to efficiently impute the missing data in water quality spatio-temporal data and construct complete and reliable water quality spatio-temporal data has become a key problem that needs to be solved urgently. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above deficiencies of the prior art, and provide a water quality spatio-temporal data imputation device and method based on multi-bias tensor decomposition, aiming to use the tensor decomposition method to complete the information of water quality spatio-temporal data in multiple dimensions, and fully reflect the volatility characteristics of water quality spatio-temporal data by combining and integrating multiple biases.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A water quality spatio-temporal data imputation device based on multi-bias tensor decomposition, characterized in that the device includes:

[0007] The water quality spatio-temporal data receiving module is used to receive historical water quality spatio-temporal data and store it in the connected water quality spatio-temporal data storage module;

[0008] The water quality spatio-temporal data storage module includes a water quality historical data storage unit and a water quality interpolation value storage unit, and is used to store the received historical water quality spatio-temporal data and the interpolation values of the missing water quality data;

[0009] The water quality spatio-temporal data tensor construction module is connected to the water quality historical data storage unit; and is used to construct the stored historical water quality spatio-temporal data into a water quality spatio-temporal data tensor;

[0010] The water quality spatio-temporal data multi-bias fusion module is connected to the water quality historical data storage unit; and is used to create and store initial water quality bias fusion data by fusing multiple bias methods according to the stored historical water quality spatio-temporal data;

[0011] The water quality spatio-temporal data interpolation module is used to perform interpolation on the missing water quality spatio-temporal data by combining the constructed water quality spatio-temporal data tensor and the initial water quality bias fusion data, and store the interpolation values of the missing water quality spatio-temporal data in the water quality interpolation value storage unit;

[0012] Furthermore, the water quality spatio-temporal data interpolation module includes: a parameter initialization unit, a training unit, a hyperparameter adaptive unit, and a water quality interpolation value output unit connected in series;

[0013] The parameter initialization unit is used to initialize the relevant parameters involved in the interpolation process of the missing water quality spatio-temporal data;

[0014] The training unit is connected to the water quality spatio-temporal data tensor construction module and the water quality spatio-temporal data multi-bias fusion module; and is used to calculate the hidden features of the historical water quality spatio-temporal data and the actual biases in the water quality bias fusion data by combining the constructed water quality spatio-temporal data tensor, the initial water quality bias fusion data, and the relevant parameters initialized by the parameter initialization unit for the interpolation process of the missing water quality spatio-temporal data;

[0015] The hyperparameter adaptive unit is used to adaptively update the hyperparameters after each iterative training of the training unit;

[0016] The water quality interpolation value output unit is connected to the water quality interpolation value storage unit, and is used to calculate the interpolation values of the missing water quality spatio-temporal data according to the hidden features of the historical water quality spatio-temporal data and the actual biases in the water quality bias fusion data obtained by the training unit, and store them in the water quality interpolation value storage unit.

[0017] A water quality spatio-temporal data interpolation method based on multi-bias tensor decomposition is characterized by including the following steps:

[0018] S1: The water quality spatio-temporal data receiving module receives the interpolation instruction of the missing water quality spatio-temporal data sent by the server;

[0019] S2: The water quality spatio-temporal data storage module stores the received historical water quality spatio-temporal data in the form of quadruples;

[0020] S3: The water quality spatio-temporal data tensor construction module constructs the stored historical water quality spatio-temporal data into a water quality spatio-temporal data tensor;

[0021] S4: The water quality spatio-temporal data multi-bias fusion module creates and stores the initial water quality bias fusion data by fusing multiple bias methods according to the historical water quality spatio-temporal data;

[0022] S5: The parameter initialization unit initializes the relevant parameters involved in the interpolation training process of the water quality spatio-temporal data;

[0023] S6: Construct the objective loss function for the water quality spatio-temporal data interpolation task, perform the interpolation of the missing water quality spatio-temporal data, and store the interpolated water quality spatio-temporal data;

[0024] Among them, the representation form of the quadruple is Q = (v, s, t, p), where: v represents the water quality index variable; s represents the site where the sensor is deployed to collect data; t represents the t-th time step, and its step size h is set according to the application scenario; p represents the index variable value of the water quality index variable v at the t-th time step at the site s.

[0025] Further, the specific step S3 is as follows:

[0026] S301: Traverse the time step t ∈ [1, K], for any time step t = k corresponding quadruple Q[k] = (v, s, t = k, p) among them, use the data (v, s, p) in Q[k] to construct the slice matrix T[k], where the size of T[k] is I rows and J columns, I represents the number of water quality index variables, J represents the number of monitoring sites, 1 ≤ i ≤ I, 1 ≤ j ≤ J, and the element T[k] ij = p represents the index variable value of the i-th water quality index variable at the j-th site in the i-th time period, and K represents the number of consecutive time steps in a monitoring time period; the entire monitoring cycle contains M consecutive monitoring time periods;

[0027] S302: Construct the incomplete water quality spatio-temporal data tensor according to the chronological order of the divided time steps for the slice matrix T[k] Let the set composed of all known data in Y be represented by Λ.

[0028] Further, the specific step S4 is as follows:

[0029] S401: Create a linear deviation vector a of all water quality index variables v based on the historical water quality spatio-temporal data stored in the water quality historical data storage unit. (1) , a preprocessing deviation vector a (2) and a time series awareness deviation matrix A (3) ; where, A (3) has a size of I rows and M columns, and the element a (1) in a i (1) represents the linear deviation of the i-th water quality index variable, and the element a (2) in a i (2) represents the preprocessing deviation of the i-th water quality index variable, and the element a (3) in A im (3) represents the time series awareness deviation of the i-th water quality index variable in the m-th monitoring time period, 1 ≤ i ≤ I, 1 ≤ m ≤ M;

[0030] S402: Create a linear deviation vector b of all monitoring stations s based on the historical water quality spatio-temporal data stored in the water quality historical data storage unit. (1) , a preprocessing deviation vector b (2) and a time series awareness deviation matrix B (3) ; where, B (3) has a size of J rows and M columns, and the element b (1) in b j (1) represents the linear deviation of the j-th station, and the element b (2) in b j (2) represents the preprocessing deviation of the j-th station, and the element b (3) in B jm (3) represents the time series awareness deviation of the j-th station in the m-th monitoring time period, 1 ≤ m ≤ M, 1 ≤ j ≤ J;

[0031] S403: Create a linear deviation vector c of all time steps t based on the historical water quality spatio-temporal data stored in the water quality historical data storage unit. (1) , a preprocessing deviation vector c (2) and a time series awareness deviation matrix C (3) ; where, C (3) has a size of K rows and M columns. The c (1) in c k (1) represents the linear deviation of the k-th time step, and the c (2) in c k (2) represents the preprocessing deviation of the k-th time step, and the c (3) in C km(3) Denote the temporal perception deviation at the \(k\)-th time step in the \(m\)-th monitoring period.

[0032] Further, the step S5 includes: initializing three hidden feature matrices V, S, T; initializing three linear deviation vectors a (1) , b (1 ), c (1) ; initializing three preprocessing deviation vectors a (2) , b (2 ), c (2) ; initializing three temporal perception deviation matrices A (3) , B (3) , C (3) ; initializing the feature dimension R, the number of monitoring periods M; initializing the maximum number of training iterations L, the iteration number control variable d during the training process, and the convergence termination threshold τ; initializing the adaptive hyperparameters η, α, β in the differential evolution algorithm; initializing the control parameters NP, F, CR in the differential evolution algorithm;

[0033] Wherein: the feature dimension R determines the feature space dimension of each hidden feature matrix corresponding to the water quality spatio-temporal data tensor, the number of monitoring periods M determines the feature space dimension of the temporal perception deviation matrix, and R and M are initialized as positive integers;

[0034] The sizes of the three hidden feature matrices V, S, T are determined by each dimension value of the corresponding water quality spatio-temporal data tensor Y and the feature dimension R, that is, V is a hidden feature matrix of I rows and R columns, S is a hidden feature matrix of J rows and R columns, and T is a hidden feature matrix of K rows and R columns. For the three hidden feature matrices, they are initialized with randomly generated smaller positive numbers;

[0035] The three linear deviation vectors a (1) , b (1) , c (1) and the three preprocessing deviation vectors a (2) , b (2) , c (2) are determined by each dimension value of the corresponding water quality spatio-temporal data tensor Y, that is, a (1) , a (2) are column vectors of I rows and 1 column, b (1 ), b (2 ) are column vectors of J rows and 1 column, c (1) , c (2) are column vectors of K rows and 1 column. For the three linear deviation vectors and the three preprocessing deviation vectors, they are initialized with randomly generated smaller positive numbers;

[0036] The three temporal perception deviation matrices A (3) , B (3) , C (3)Its size is determined by the values of each dimension of the corresponding water quality spatio-temporal data tensor Y and the number of monitoring time periods M, that is, A (3) is a matrix of I rows and M columns, B (3) is a matrix of J rows and M columns, C (3) is a matrix of K rows and M columns. The three time-series perception deviation matrices are respectively initialized with randomly generated small positive numbers;

[0037] The maximum number of training iterations L is a variable that controls the upper limit of the iterative process and is initialized to a large positive integer;

[0038] The iteration number control variable d is initialized to 0;

[0039] The convergence termination threshold τ is a parameter for judging whether the iterative process has converged and is initialized with a very small positive number;

[0040] The adaptive hyperparameters in the differential evolution algorithm include the regularization parameter α corresponding to the three hidden feature matrices, the regularization parameter β corresponding to the three linear deviation vectors and the three time-series perception deviation matrices, and the learning step size η of the stochastic gradient descent update rule. Among them, η, α, and β are initialized to small positive numbers.

[0041] The control parameters in the differential evolution algorithm include the population size NP, the scaling factor F, and the crossover probability CR. NP is initialized to a positive integer, and F and CR are initialized to positive numbers.

[0042] Furthermore, the specific steps of step S6 are as follows:

[0043] S601: The training unit combines the constructed water quality spatio-temporal data tensor, the initial water quality deviation fusion data, and the relevant parameters involved in the initialized missing water quality spatio-temporal data interpolation process to construct the objective loss function of the water quality spatio-temporal data interpolation task, so as to calculate the hidden features of the historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data. The hyperparameter adaptive unit adaptively updates the hyperparameters after each iterative training to dynamically enhance the interpolation performance;

[0044] S602: The water quality interpolation value output unit calculates and stores the interpolation values of the missing water quality spatio-temporal data according to the hidden features of the obtained historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data.

[0045] Even further, the specific steps of step S601 are as follows:

[0046] S6011: Combine the constructed water quality spatio-temporal data tensor Y, the initial water quality deviation fusion data, and the relevant parameters involved in the initialized missing water quality spatio-temporal data interpolation process to construct the objective loss function ε on the known data set Λ;

[0047] Among them, the target loss function ε uses the Euclidean distance as the optimization objective, and uses Tikhonov regularization to constrain the optimization process to prevent overfitting during the optimization process. Specifically:

[0048]

[0049] represents the estimated value of the numerical value in the water quality spatio-temporal data, which is:

[0050]

[0051] S6012: To ensure the effectiveness of the update process, the random gradient descent update rule is used to iteratively optimize ε to minimize the value of ε. The formula for training iteration is as follows:

[0052]

[0053] S6013: The preprocessing deviation does not require iterative training and is obtained based on the prior information of the historical known water quality spatio-temporal data. The preprocessing deviation is calculated according to the following formula;

[0054]

[0055] Among them, μ is the mean value of all index variable values in the historical water quality spatio-temporal data. Λ(i), Λ(j), and Λ(k) are the subsets of Λ associated with i ∈ I, j ∈ J, and k ∈ K respectively. |Λ(i)|, |Λ(j)|, and |Λ(k)| are the counts of elements in the subsets. θ1, θ2, and θ3 are threshold constants that control the size of the preprocessing deviation;

[0056] S6014: To reduce the hyperparameter grid search optimization time and enhance the imputation performance, the differential evolution algorithm is used to adaptively update the regularization parameter α of the latent features, the regularization parameter β of the linear deviation and the temporal perception deviation, and the learning step size η of the random gradient descent update rule after each iterative training;

[0057] Specifically, by calculating the performance perf(d n ) of all individuals in the current population and comparing it with the current optimal performance, if it is smaller, the hyperparameters d n = [η n , α n , β n corresponding to the individual are used for update, otherwise it remains unchanged; among them, the subscript n represents the nth individual;

[0058] S6014: Repeat steps S6011 to S6014 until the maximum number of training iterations L or the absolute value of the difference between the target loss function ε value calculated after the end of this round of iterations and the target loss function ε value of the previous round is less than the convergence termination threshold τ.

[0059] The beneficial effects of the present invention are: providing a water quality spatiotemporal data interpolation device and method based on multi-bias tensor decomposition, acting on water quality spatiotemporal data, introducing multi-dimensional biases after tensor decomposition slicing for bias correction interpolation, and combining with a differential evolution algorithm to achieve fast iterative calculation, which can perform water quality data interpolation that conforms to statistical laws and has high accuracy, so as to solve the problem of fast and accurate interpolation of missing water quality data. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to make the purpose and technical solution of the present invention more clear, the present invention provides the following drawings for explanation:

[0061] Figure 1 Schematic diagram of the structure of a water quality spatiotemporal data interpolation device based on multi-bias tensor decomposition in Example 1 of the present invention;

[0062] Figure 2 Schematic diagram of the flow of a method for water quality spatiotemporal data interpolation based on multi-bias tensor decomposition in Example 1 of the present invention;

[0063] Figure 3 A distribution map of water quality monitoring stations in the monitoring waters of Dianchi Lake in Example 1 of the present invention;

[0064] Figure 4 is a comparison diagram of the convergence time of the ablation experiment in Example 1 of the present invention;

[0065] Figure 5 It is a comparison chart of interpolation accuracy of ablation experiment in Example 1 of the present invention. DETAILED DESCRIPTION

[0066] In order to make the purpose and technical solution of the present invention more clearly understood, the present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0067] Example 1: In order to predict the blue algae bloom in the waters of Dianchi Lake, there are 2 monitoring stations in the Caohai waters of Dianchi Lake, namely Broken Bridge and Caohai Center, and 8 monitoring stations in the middle of Huiwan, Luojiaying, Guanyinshan East, Guanyinshan Middle, Guanyinshan West, Baiyukou, Haikou West and Dianchi South, totaling 10 monitoring stations. These 10 monitoring stations are all national control points, and their specific locations are as follows: Figure 3As shown, the sensors are set to collect water quality data every 4 hours. Therefore, 4 hours is regarded as a time step length, with a total of 4080 time steps, and these time steps are divided into 80 monitoring time periods. The water quality data includes: water temperature (°C), pH (dimensionless), dissolved oxygen (mg / L), conductivity (μS / cm), turbidity (NTU), permanganate index (mg / L), ammonia nitrogen (mg / L), total phosphorus (mg / L), total nitrogen (mg / L), chlorophyll a (μg / L). It is known that the missing ratio of the collected water quality data dataset is 6.89%.

[0068] In order to quickly and accurately impute the missing water quality data, the present invention provides a "device for imputing water quality spatio-temporal data based on multi-bias tensor decomposition", combined with Figure 1 , including:

[0069] A water quality spatio-temporal data receiving module (110), configured to receive historical water quality spatio-temporal data and store it in a connected water quality spatio-temporal data storage module (120);

[0070] A water quality spatio-temporal data storage module (120), including a water quality historical data storage unit (121) and a water quality imputation value storage unit (122), for storing the received historical water quality spatio-temporal data and storing the imputation values of the missing water quality data;

[0071] A water quality spatio-temporal data tensor construction module (130), connected to the water quality historical data storage unit (121); for constructing the stored historical water quality spatio-temporal data into a water quality spatio-temporal data tensor;

[0072] A water quality spatio-temporal data multi-bias fusion module (140), connected to the water quality historical data storage unit (121); for fusing multiple bias methods according to the stored historical water quality spatio-temporal data to create and store initial water quality bias fusion data;

[0073] A water quality spatio-temporal data imputation module (150), for combining the constructed water quality spatio-temporal data tensor and the initial water quality bias fusion data to perform imputation of missing water quality spatio-temporal data, and storing the imputation values of the missing water quality spatio-temporal data in the water quality imputation value storage unit (122);

[0074] Furthermore, the water quality spatio-temporal data imputation module (150) includes: a parameter initialization unit (151), a training unit (152), a hyperparameter adaptive unit (153), and a water quality imputation value output unit (154) connected in series;

[0075] The parameter initialization unit (151) is used to initialize the relevant parameters involved in the imputation process of the missing water quality spatio-temporal data;

[0076] The training unit (152) is connected to the water quality spatio-temporal data tensor construction module (130) and the water quality spatio-temporal data multi-deviation fusion module (140); it is used to combine the constructed water quality spatio-temporal data tensor, the initial water quality deviation fusion data, and the relevant parameters involved in the interpolation process of the missing water quality spatio-temporal data initialized by the parameter initialization unit (151), and calculate the hidden features of the historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data.

[0077] The hyperparameter adaptive unit (153) is used to adaptively update the hyperparameters after each iterative training of the training unit (152).

[0078] The water quality interpolation value output unit (154) is connected to the water quality interpolation value storage unit (122), and is used to calculate the interpolation value of the missing water quality spatio-temporal data according to the hidden features of the historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data obtained by the training unit (152), and store it in the water quality interpolation value storage unit (122).

[0079] This device can be deployed in an existing server, or can be deployed in a separately set server dedicated to performing the task of interpolating missing water quality spatio-temporal data.

[0080] Embodiment 2: For the scenario and device of Embodiment 1, the present invention also provides a "method for interpolating water quality spatio-temporal data based on multi-bias tensor decomposition", which includes the following steps:

[0081] S1: The water quality spatio-temporal data receiving module (110) receives the interpolation instruction of the missing water quality spatio-temporal data sent by the server, and collects the historical water quality spatio-temporal data of I water quality index variables with J sites and M consecutive monitoring time periods from the server. Among them, each monitoring time period contains K consecutive time steps, and I = 10, J = 10, K = 80, M = 51.

[0082] S2: The water quality spatio-temporal data storage module (120) stores the received historical water quality spatio-temporal data in the form of a quadruple.

[0083] Among them, the quadruple representation form is Q = (v, s, t, p), where: v represents the water quality index variable; s represents the site where the sensor is deployed to collect data; t represents the t-th time step, and its step size h is set according to the application scenario; p represents the index variable value of the water quality index variable v at the t-th time step at the site s.

[0084] S3: The water quality spatio-temporal data tensor construction module (130) constructs the stored historical water quality spatio-temporal data into a water quality spatio-temporal data tensor. Specifically:

[0085] S301: Traverse the time steps \(t\in[1, K]\). For any time step \(t = k\) corresponding to the quadruple \(Q[k]=(v, s, t = k, p)\), use the data \((v, s, p)\) in \(Q[k]\) to construct the slice matrix \(T[k]\). Here, the size of \(T[k]\) is \(I\) rows and \(J\) columns. \(I\) represents the number of water quality index variables, and \(J\) represents the number of monitoring stations. \(1\leq i\leq I\), \(1\leq j\leq J\), and the element \(T[k] ij _{ij}=p\) represents the index variable value of the \(i\)-th water quality index variable at the \(j\)-th station in the \(k\)-th time period. \(K\) represents the number of consecutive time steps in a monitoring time period; the entire monitoring cycle contains \(M\) consecutive monitoring time periods;

[0086] S302: Construct an incomplete water quality spatio-temporal data tensor from the slice matrices \(T[k]\) according to the chronological order of the divided time steps Let the set composed of all known data in \(Y\) be denoted as \(\Lambda\).

[0087] S4: The water quality spatio-temporal data multi-bias fusion module (140) creates and stores initial water quality bias fusion data by fusing multiple bias methods based on historical water quality spatio-temporal data.

[0088] Specifically:

[0089] S401: Create a linear bias vector \(a\) for all water quality index variables \(v\) based on the historical water quality spatio-temporal data stored in the water quality historical data storage unit (121) (1) , a preprocessing bias vector \(a\) (2) and a time series perception bias matrix \(A\) (3) ; where the size of \(A\) (3) is \(I\) rows and \(M\) columns. The element \(a\) (1) in \(a\) i (1) represents the linear bias of the \(i\)-th water quality index variable. The element \(a\) (2) in \(a\) i (2) represents the preprocessing bias of the \(i\)-th water quality index variable. The element \(a\) (3) in \(A\) im (3) represents the time series perception bias of the \(i\)-th water quality index variable in the \(m\)-th monitoring time period, \(1\leq i\leq I\), \(1\leq m\leq M\);

[0090] S402: Create a linear bias vector \(b\) for all monitoring stations \(s\) based on the historical water quality spatio-temporal data stored in the water quality historical data storage unit (121) (1 ), a preprocessing bias vector \(b\) (2 ) and a time series perception bias matrix \(B\) (3) ; where the size of \(B\) (3) is \(J\) rows and \(M\) columns. The element \(b\) (1The element b in j (1) represents the linear deviation of the j-th station, b (2) The element b in j (2) represents the preprocessing deviation of the j-th station, B (3) The element b in jm (3) represents the time-series perception deviation of the j-th station in the m-th monitoring time period, 1 ≤ m ≤ M, 1 ≤ j ≤ J;

[0091] S403: Create the linear deviation vector c for all time steps t based on the historical water quality spatio-temporal data stored in the water quality historical data storage unit (121) (1) , the preprocessing deviation vector c (2) and the time-series perception deviation matrix C (3) ; where, C (3) has a size of K rows and M columns. c (1) In c k (1) represents the linear deviation at the k-th time step, c (2) In c k (2) represents the preprocessing deviation at the k-th time step, C (3) In c km (3) represents the time-series perception deviation at the k-th time step in the m-th monitoring time period.

[0092] S5: The parameter initialization unit (151) initializes the relevant parameters involved in the water quality spatio-temporal data interpolation training process. Including: initializing three latent feature matrices V, S, T; initializing three linear deviation vectors a (1) , b (1) , c (1) ; initializing three preprocessing deviation vectors a (2) , b (2) , c (2) ; initializing three time-series perception deviation matrices A (3) , B (3) , C (3) ; initializing the feature dimension R, the number of monitoring time periods M; initializing the maximum number of training iterations L, the iteration number control variable d during the training process, and the convergence termination threshold τ; initializing the adaptive hyperparameters η, α, β in the differential evolution algorithm; initializing the control parameters NP = 10, F = 0.5, CR = 0.5 in the differential evolution algorithm;

[0093] Among them: The feature dimension R determines the feature space dimension of each latent feature matrix corresponding to the water quality spatio-temporal data tensor, R = 5;

[0094] The sizes of the three latent feature matrices V, S, and T are determined by the values of each dimension of the corresponding water quality spatio-temporal data tensor Y and the feature dimension R. That is, V is a latent feature matrix with I rows and R columns, S is a latent feature matrix with J rows and R columns, and T is a latent feature matrix with K rows and R columns. The three latent feature matrices are respectively initialized with randomly generated small positive numbers;

[0095] The three linear deviation vectors a (1) , b (1) , c (1) and the three preprocessing deviation vectors a (2) , b (2) , c (2) are determined by the values of each dimension of the corresponding water quality spatio-temporal data tensor Y. That is, a (1) , a (2) are column vectors with I rows and 1 column, b (1 ), b (2 ) are column vectors with J rows and 1 column, c (1) , c (2) are column vectors with K rows and 1 column. The three linear deviation vectors and the three preprocessing deviation vectors are respectively initialized with randomly generated small positive numbers;

[0096] The sizes of the three time series perception deviation matrices A (3) , B (3) , C (3) are determined by the values of each dimension of the corresponding water quality spatio-temporal data tensor Y and the number of monitoring time periods M. That is, A (3) is a matrix with I rows and M columns, B (3) is a matrix with J rows and M columns, C (3) is a matrix with K rows and M columns. The three time series perception deviation matrices are respectively initialized with randomly generated small positive numbers;

[0097] The maximum number of training iterations L is a variable that controls the upper limit of the iterative process and is initialized to 1000;

[0098] The iteration number control variable d is initialized to 0;

[0099] The convergence termination threshold τ is a parameter for judging whether the iterative process has converged and is initialized to τ = 10 -6 ;

[0100] The adaptive hyperparameters in the differential evolution algorithm include the regularization parameter α corresponding to the three latent feature matrices, the regularization parameter β corresponding to the three linear deviation vectors and the three time series perception deviation matrices, and the learning step size η of the stochastic gradient descent update rule, where η, α, β are initialized to η = 0.1, α = 0.05, β = 0.01.

[0101] The control parameters in the differential evolution algorithm include the population size NP, the scaling factor F, and the crossover probability CR. NP is initialized as a positive integer, and F and CR are initialized as positive numbers.

[0102] S6: Construct the objective loss function for the water quality spatio-temporal data interpolation task, perform the interpolation of the missing water quality spatio-temporal data, and store the interpolated water quality spatio-temporal data.

[0103] Specifically:

[0104] S601: The training unit (152) combines the constructed water quality spatio-temporal data tensor, the initial water quality deviation fusion data, and the relevant parameters involved in the initialized interpolation process of the missing water quality spatio-temporal data to construct the objective loss function for the water quality spatio-temporal data interpolation task, so as to calculate the hidden features of the historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data. The hyperparameter adaptive unit (153) adaptively updates the hyperparameters after each iterative training to dynamically enhance the interpolation performance;

[0105] Furthermore, the specific steps of S601 are as follows:

[0106] S6011: Combine the constructed water quality spatio-temporal data tensor Y, the initial water quality deviation fusion data, and the relevant parameters involved in the initialized interpolation process of the missing water quality spatio-temporal data to construct the objective loss function ε on the known data set Λ;

[0107] Among them, the objective loss function ε uses the Euclidean distance as the optimization goal and uses Tikhonov regularization to constrain the optimization process to prevent overfitting during the optimization process. Specifically:

[0108]

[0109] represents the estimated value of the numerical value in the water quality spatio-temporal data, which is:

[0110]

[0111] S6012: To ensure the effectiveness of the update process, use the stochastic gradient descent update rule to iteratively optimize ε to minimize the value of ε. The formula for the training iteration is as follows:

[0112]

[0113] S6013: The preprocessing deviation does not require iterative training and is obtained according to the prior information of the historical known water quality spatio-temporal data. The preprocessing deviation is calculated according to the following formula;

[0114]

[0115] Wherein, μ is the mean value of all index variable values in the historical water quality spatio-temporal data, Λ(i), Λ(j), and Λ(k) are subsets of Λ associated with i ∈ I, j ∈ J, and k ∈ K respectively, |Λ(i)|, |Λ(j)|, and |Λ(k)| are the counts of elements in the subsets, and θ1, θ2, and θ3 are threshold constants for controlling the magnitude of the preprocessing deviation;

[0116] To reduce the hyperparameter grid search optimization time and enhance the imputation performance, the differential evolution algorithm is used to adaptively update the regularization parameter α of the latent features after each iterative training, the regularization parameter β of the linear deviation and the temporal perception deviation, and the learning step size η of the stochastic gradient descent update rule;

[0117] Specifically, by calculating the performance perf(d n ) of all individuals in the current population and comparing it with the current optimal performance, if it is smaller, the hyperparameters d n = [η n , α n , β n corresponding to the individual are used for update, otherwise it remains unchanged; wherein, The subscript n represents the nth individual;

[0118] S6014: Repeat steps S6011 to S6013 until the maximum number of training iterations L or the absolute value of the difference between the value of the target loss function ε calculated after this round of iteration and the value of the target loss function ε in the previous round is less than the convergence termination threshold τ.

[0119] S602: The water quality imputation value output unit (154) calculates the imputation value of the missing water quality spatio-temporal data based on the latent features of the obtained historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data, and stores it.

Claims

1. A water quality spatio-temporal data imputation device based on multi-bias tensor decomposition, characterized in that Including the following steps: including: A water quality spatio-temporal data receiving module (110), configured to receive historical water quality spatio-temporal data and store it in a connected water quality spatio-temporal data storage module (120); A water quality spatio-temporal data storage module (120), including a water quality historical data storage unit (121) and a water quality interpolation value storage unit (122), configured to store the received historical water quality spatio-temporal data and store the interpolation values of the missing water quality data; A water quality spatio-temporal data tensor construction module (130), connected to the water quality historical data storage unit (121); configured to construct the stored historical water quality spatio-temporal data into a water quality spatio-temporal data tensor; A water quality spatio-temporal data multi-bias fusion module (140), connected to the water quality historical data storage unit (121); configured to create and store initial water quality bias fusion data by fusing multiple bias methods according to the stored historical water quality spatio-temporal data; A water quality spatio-temporal data interpolation module (150), configured to perform interpolation on the missing water quality spatio-temporal data by combining the constructed water quality spatio-temporal data tensor and the initial water quality bias fusion data, and store the interpolation values of the missing water quality spatio-temporal data in the water quality interpolation value storage unit (122).

2. The water quality spatio-temporal data interpolation device based on multi-biased tensor decomposition according to claim 1, wherein The described water quality spatio-temporal data interpolation module (150) includes: a serially connected parameter initialization unit (151), a training unit (152), a hyperparameter adaptive unit (153), and a water quality interpolation value output unit (154); wherein, The parameter initialization unit (151) is configured to initialize the relevant parameters involved in the interpolation process of the missing water quality spatio-temporal data; The training unit (152), connected to the water quality spatio-temporal data tensor construction module (130) and the water quality spatio-temporal data multi-bias fusion module (140); configured to calculate the hidden features of the historical water quality spatio-temporal data and the actual biases in the water quality bias fusion data by combining the constructed water quality spatio-temporal data tensor, the initial water quality bias fusion data, and the relevant parameters initialized by the parameter initialization unit (151) in the interpolation process of the missing water quality spatio-temporal data; The hyperparameter adaptive unit (153) is configured to adaptively update the hyperparameters after each iterative training of the training unit (152); The water quality interpolation value output unit (154), connected to the water quality interpolation value storage unit (122), is configured to calculate the interpolation values of the missing water quality spatio-temporal data according to the hidden features of the historical water quality spatio-temporal data and the actual biases in the water quality bias fusion data obtained by the training unit (152), and store them in the water quality interpolation value storage unit (122).

3. A water quality spatio-temporal data imputation method based on multi-biased tensor decomposition, characterized in that Including the following steps: S1: Receive an interpolation instruction for the missing water quality spatio-temporal data sent by the server; S2: Store the received historical water quality spatio-temporal data in the form of quadruples; S3: Construct the stored historical water quality spatio-temporal data into a water quality spatio-temporal data tensor; S4: Create and store initial water quality bias fusion data by fusing multiple bias methods according to the historical water quality spatio-temporal data; S5: Initialize the relevant parameters involved in the water quality spatio-temporal data interpolation training process; S6: Construct the objective loss function for the water quality spatio-temporal data interpolation task, perform the interpolation of the missing water quality spatio-temporal data, and store the interpolated water quality spatio-temporal data; Among them, the form of the quadruple is expressed as Q = (v, s, t, p), where: v represents the water quality index variable; s represents the site where the sensor is deployed to collect data; t represents the t-th time step, and its step size h is set according to the application scenario; p represents the index variable value of the water quality index variable v at the t-th time step at site s.

4. The water quality spatio-temporal data interpolation method based on multi-biased tensor decomposition according to claim 3, characterized in that The specific step S3 is as follows: S301: Traverse the time steps \(t\in[1, K]\). For any time step \(t = k\) corresponding to the quadruple \(Q[k]=(v, s, t = k, p)\), use the data \((v, s, p)\) in \(Q[k]\) to construct the slice matrix \(T[k]\). Among them, the size of \(T[k]\) is \(I\) rows and \(J\) columns. \(I\) represents the number of water quality index variables, and \(J\) represents the number of monitoring stations. \(1\leq i\leq I\), \(1\leq j\leq J\). The element \(T[k] ij _{ij}=p\) represents the index variable value of the \(i\)-th water quality index variable at the \(j\)-th station in the \(k\)-th time period. \(K\) represents the number of consecutive time steps in a monitoring time period; the entire monitoring cycle contains \(M\) consecutive monitoring time periods; S302: Construct an incomplete water quality spatio-temporal data tensor from the slice matrix T[k] according to the chronological order of the divided time steps. Let the set composed of all known data in Y be denoted as Λ.

5. The water quality spatio-temporal data interpolation method based on multi-biased tensor decomposition according to claim 3, characterized in that The specific step S4 is as follows: S401: Create a linear deviation vector a of all water quality index variables v based on historical water quality spatio-temporal data (1) , preprocess the deviation vector a (2) and the time series perception deviation matrix A (3) ; where A (3) has a size of I rows and M columns, and the element a (1) in a i (1) represents the linear deviation of the i-th water quality index variable, and the element a (2) in a i (2) represents the preprocessing deviation of the i-th water quality index variable. The element a (3) in A im (3) represents the time series perception deviation of the i-th water quality index variable in the m-th monitoring time period, 1 ≤ i ≤ I, 1 ≤ m ≤ M; S402: Create the linear deviation vector b for all monitoring stations s based on historical water quality spatio-temporal data (1) , preprocess the deviation vector b (2) and the time-series aware deviation matrix B (3) ; where the size of B (3) is J rows and M columns, and the element b (1) in b j (1) represents the linear deviation of the j-th station, and the element b (2) in b j (2) represents the preprocessing deviation of the j-th station. The element b (3) in B jm (3) represents the time-series aware deviation of the j-th station in the m-th monitoring time period, where 1 ≤ m ≤ M and 1 ≤ j ≤ J; S403: Create a linear deviation vector c for all time steps t based on historical water quality spatio-temporal data (1) , preprocess the deviation vector c (2) and the time series-aware deviation matrix C (3) ; where the size of C (3) is K rows and M columns. c (1) in c k (1) represents the linear deviation at the k-th time step, and c (2) in c k (2) represents the preprocessed deviation at the k-th time step, and c (3) in C km (3) represents the time series-aware deviation at the k-th time step in the m-th monitoring time period.

6. The water quality spatio-temporal data interpolation method based on multi-bias tensor decomposition according to claim 3, wherein The step S5 includes: initializing three implicit feature matrices V, S, and T; initializing three linear deviation vectors a (1) , b (1) , c (1) ; initializing three preprocessing deviation vectors a (2) , b (2) , c (2) ; initializing three time-series perception deviation matrices A (3) , B (3) , C (3) ; initializing the feature dimension R, the number of monitoring time periods M; initializing the maximum number of training iterations L, the iteration number control variable d during the training process, and the convergence termination threshold τ; initializing the adaptive hyperparameters η, α, β in the differential evolution algorithm; initializing the control parameters NP, F, CR in the differential evolution algorithm; Among them: The feature dimension R determines the feature space dimension of each hidden feature matrix corresponding to the water quality spatio-temporal data tensor, and the number of monitoring time periods M determines the feature space dimension of the time series perception deviation matrix. R and M are initialized as positive integers; The sizes of the three hidden feature matrices V, S, and T are determined by the values of each dimension of the corresponding water quality spatio-temporal data tensor Y and the feature dimension R, that is, V is a hidden feature matrix with I rows and R columns, S is a hidden feature matrix with J rows and R columns, and T is a hidden feature matrix with K rows and R columns. The three hidden feature matrices are respectively initialized with randomly generated smaller positive numbers; Three linear deviation vectors a (1) , b (1) , c (1) and three pre - processed deviation vectors a (2) , b (2) , c (2) are determined by the values of each dimension of the corresponding water quality spatio - temporal data tensor Y, that is, a (1) , a (2) is a column vector of 1 row and I columns, b (1) , b (2) is a column vector of 1 row and J columns, c (1) , c (2) is a column vector of 1 row and K columns. The three linear deviation vectors and the three pre - processed deviation vectors are respectively initialized with randomly generated small positive numbers; Three temporal perception deviation matrices A (3) , B (3) , C (3) are determined by the values of each dimension of the corresponding water quality spatio-temporal data tensor Y and the number of monitoring time periods M, that is, A (3) is a matrix of I rows and M columns, B (3) is a matrix of J rows and M columns, C (3) is a matrix of K rows and M columns, and the three temporal perception deviation matrices are respectively initialized with randomly generated small positive numbers; The maximum number of training iterations L is a variable that controls the upper limit of the iterative process and is initialized as a large positive integer; The iteration number control variable d is initialized to 0; The convergence termination threshold τ is a parameter for judging whether the iterative process has converged and is initialized with a very small positive number; The adaptive hyperparameters in the differential evolution algorithm include the regularization parameters α corresponding to the three hidden feature matrices, the regularization parameters β corresponding to the three linear deviation vectors and the three time series perception deviation matrices, and the learning step size η of the stochastic gradient descent update rule. Among them, η, α, and β are initialized as smaller positive numbers; The control parameters in the differential evolution algorithm include the population size NP, the scaling factor F, and the crossover probability CR. NP is initialized as a positive integer, F and CR are initialized as positive numbers.

7. The water quality spatio-temporal data interpolation method based on multi-biased tensor decomposition according to claim 3, wherein The specific step S6 is as follows: S601: Combine the constructed water quality spatio-temporal data tensor, the initial water quality deviation fusion data, and the relevant parameters involved in the initialized interpolation process of the missing water quality spatio-temporal data to construct the objective loss function for the water quality spatio-temporal data interpolation task, so as to calculate the hidden features of the historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data, and adaptively update the hyperparameters after each iterative training to dynamically enhance the interpolation performance; S602: Calculate the interpolation values of the missing water quality spatio-temporal data according to the hidden features of the obtained historical water quality spatio-temporal data and the actual deviations in the water quality deviation fusion data, and store them.

8. The water quality spatio-temporal data imputation method based on multi-biased tensor decomposition according to claim 3, wherein The specific step S601 is as follows: S6011: Combine the constructed water quality spatio-temporal data tensor Y, the initial water quality deviation fusion data, and the relevant parameters involved in the initialized interpolation process of the missing water quality spatio-temporal data to construct the objective loss function ε on the known data set Λ; Among them, the objective loss function ε uses the Euclidean distance as the optimization objective and uses Tikhonov regularization to constrain the optimization process to prevent overfitting during the optimization process. Specifically: Indicates the estimated value of the numerical value in the spatio-temporal water quality data, which is: S6012: To ensure the effectiveness of the update process, the random gradient descent update rule is used to iteratively optimize ε to minimize the value of ε. The formula for training iteration is shown as follows: S6013: The preprocessing bias does not require iterative training and is obtained based on the prior information of the historical known water quality spatio-temporal data. The preprocessing bias is calculated according to the following formula: where μ is the mean of all index variable values in the historical water quality spatio-temporal data, Λ(i), Λ(j), Λ(k) are subsets of Λ associated with i ∈ I, j ∈ J, k ∈ K respectively, |Λ(i)|, |Λ(j)|, |Λ(k)| are the counts of elements in the subsets, and θ1, θ2, θ3 are threshold constants for controlling the magnitude of the preprocessing bias; S6014: To reduce the hyperparameter grid search optimization time and enhance the imputation performance, the differential evolution algorithm is used to adaptively update the regularization parameter α of the latent features, the regularization parameter β of the linear bias and the temporal perception bias, and the learning step size η of the random gradient descent update rule after each iterative training; Specifically, by calculating the performance perf(d n ) of all individuals in the current population and comparing it with the current optimal performance, if it is smaller, then the hyperparameters d n = [η n , α n , β n corresponding to the individual are used for updating, otherwise they remain unchanged; where the subscript n represents the nth individual; S6014: Repeat steps S6011 to S6014 until the maximum number of training iterations L or the absolute value of the difference between the value of the objective loss function ε calculated after the end of this round of iteration and the value of the objective loss function ε in the previous round is less than the convergence termination threshold τ.

Citation Information

Cited By

  • Hydrological data management method and system

    CN120492446A

  • River water quality prediction method and system based on big data

    CN121279555A