Construction method of multivariable time sequence missing value prediction model

By constructing a multivariate time series missing value prediction model, the mutual information, constant and seasonal period modules in the BEIN model are used to extract and fusion, the problem of sensor data missing during sewage treatment is solved, and efficient and reliable missing value interpolation is achieved.

CN120045858APending Publication Date: 2025-05-27LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510431036.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Due to the problem of missing sensor data during sewage treatment, the existing technology is difficult to effectively deal with complex patterns of high missing rates and multi-period superposition, and the long-range dependency modeling of high-dimensional variables is inefficient.

Method used

The construction method of the multivariate time series missing value prediction model is adopted. By obtaining the sensor raw data during sewage treatment, standardizing and missing marks, the mutual information module, constant module and seasonal periodic module in the initial BEIN model are used to extract the overall trend characteristics, periodic characteristics and key exogenous variable characteristics in the multivariate time series, and fusing them through residual connections to obtain the predicted missing values.

Benefits of technology

It effectively solves the problem of missing sensor data during sewage treatment, improves the accuracy, reliability and real-time nature of missing value interpolation, and can process compound cycle characteristics and strong correlation industrial timing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045858A_ABST
    Figure CN120045858A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electronic information, in particular to a method for constructing a multivariable time sequence missing value prediction model, and the method comprises the steps: carrying out the optimization training of an initial BEIN model which comprises a constant module, a seasonal period module and a mutual information module, so as to output a prediction value for interpolating missing data, according to the method, the problem of sensor data missing in the sewage treatment process is effectively solved, multi-scale component analysis is achieved through a dynamic period decomposition module aiming at composite period characteristics generally existing in industrial time series data, meanwhile, a mutual information quantification module is introduced to construct a dynamic association topology aiming at strong correlation between industrial sensor data, and the dynamic association topology is analyzed through a dynamic period decomposition module. Wastewater components highly related to target variables are screened as key exogenous variables, and a data distribution reconstruction mechanism based on physical constraints is established, so that the accuracy, reliability and real-time performance of missing value interpolation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic information technology, and in particular to a method for constructing a missing value prediction model for multivariate time series. Background Art

[0002] Driven by the global sustainable development strategy, sewage treatment plants are facing two aspects of pressure: one is to comply with more stringent pollutant discharge regulations, and the other is to achieve resource intensification and energy efficiency optimization. To address these challenges, modern sewage treatment facilities are moving towards high integration and complex process flows, which requires more refined operation control. By integrating industrial Internet of Things and automation technologies, an intelligent modeling system based on multi-source sensor data has gradually matured. Such models can real-time analyze the dynamic correlations among water quality parameters, equipment status, and energy consumption indicators, providing decision-making support for process optimization and fault warning.

[0003] However, constructing a biochemical reaction kinetics model remains a major challenge. The multiphase coupling mechanism among microbial communities, dissolved oxygen, and pollutants in the activated sludge process has strong nonlinear and time-varying characteristics, resulting in high-dimensional and low-density data of key process parameters. This data sparsity not only makes it difficult to completely characterize the state variables of the reaction process but also increases the uncertainty of detecting and identifying models. When the modeling sample space cannot cover all dynamic operating conditions, traditional mechanism models will fail due to insufficient extrapolation ability.

[0004] Existing technologies rely on simple calculations of local data when dealing with data in sewage treatment, making it difficult to capture the nonlinear variation laws of sewage flow and pollutant concentration. In the case of a high missing rate, the error increases significantly, and complex patterns with multi-period superposition cannot be processed. Although some can utilize partial spatio-temporal correlations, the modeling efficiency for long-range dependence relationships among high-dimensional variables is low. Especially when there is lag correlation across components in sensor data, the filling results are prone to temporal misalignment. In addition, in industrial scenarios, sensor data usually exhibits strong temporal correlation. Although deep learning methods show superior representation ability in the field of data augmentation, they still face challenges when dealing with interaction effects, nonlinearity, and non-standard distribution forms. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for constructing a missing value prediction model for multivariate time series.

[0006] In a first aspect, an embodiment of the present invention provides a method for constructing a missing value prediction model for multivariate time series, the method comprising:

[0007] Obtain the original sensor dataset collected during the sewage treatment process, and after standardizing and marking the missing values in the original data of the sensor original dataset, a training dataset is obtained. Each training data in the training dataset is a multivariate time series;

[0008] Calculate the mutual information of each variable in the multivariate time series based on the mutual information module in the initial BEIN model, and use the minimum redundancy maximum correlation to determine the key exogenous variables used as supplementary information;

[0009] Extract the overall trend features in the multivariate time series based on the constant module in the initial BEIN model. At the same time, apply the seasonal cycle module in the initial BEIN model to extract the periodic features in the multivariate time series;

[0010] Fuse the overall trend features, periodic features, and features of the key exogenous variables through residual connection to obtain the predicted missing values;

[0011] Iteratively optimize the initial BEIN model through a preset loss function to obtain a trained multivariate time series missing value prediction model.

[0012] Combined with the first aspect, the original data includes: one or more of influent flow rate, chemical oxygen demand, ammonia nitrogen content, pH value, external reflux flow rate, oxidation reduction potential, nitrate nitrogen content, mixed liquor suspended solid content, dissolved oxygen concentration, phosphate content, total phosphorus content, temperature, and total nitrogen content.

[0013] Combined with the first aspect, the initial BEIN model includes a time series expansion module, and the time series expansion module includes a constant module, a seasonal cycle module, and a multi-level cascaded fully connected layer.

[0014] Combined with the first aspect, each fully connected layer includes a constant module and a seasonal cycle module; the constant module is used to capture the in-cycle information in the time series; the seasonal cycle module is used to capture the between-cycle information in the time series.

[0015] Combined with the first aspect, the step of calculating the mutual information of each variable in the multivariate time series based on the mutual information module in the initial BEIN model and using the minimum redundancy maximum correlation to determine the key exogenous variables used as supplementary information includes:

[0016] For each first variable, use the K-nearest neighbor estimation method to calculate the mutual information between the first variable and each second variable;

[0017] Combined with all the mutual information, select the set of exogenous variables strongly correlated with the first variable from multiple second variables based on the minimum redundancy maximum correlation.

[0018] The steps of extracting the overall trend features in the multivariate time series based on the constant module in the initial BEIN model include:

[0019] Encode the input multivariate time series through a fully connected layer to extract high-level features;

[0020] Generate a reverse decomposition component from the high-level features through a first linear projection function, and at the same time, generate a forward prediction component from the high-level features through a second linear projection function to extract the overall trend features.

[0021] The steps of extracting the periodic features in the multivariate time series by applying the seasonal cycle module in the initial BEIN model include:

[0022] After encoding the input multivariate time series through a fully connected layer, generate decomposition coefficients and prediction coefficients through a third linear projection function and a fourth linear projection function;

[0023] Analyze the spectral energy distribution of the input multivariate time series through the fast Fourier transform FFT to construct basis functions; among them, the basis functions include backward decomposition basis functions and forward prediction basis functions;

[0024] Calculate the decomposition terms for updating the residuals by combining the decomposition coefficients and the backward decomposition basis functions, and at the same time, calculate the prediction terms for accumulating the prediction results by combining the prediction coefficients and the forward prediction basis functions;

[0025] Determine the periodic features by combining the decomposition terms and the prediction terms.

[0026] The steps of fusing the overall trend features, the periodic features, and the features of the key exogenous variables through residual connection to obtain the predicted missing values in combination with the first aspect include:

[0027] Pass the overall trend features, the periodic features, and the features of the key exogenous variables to the constant module and the seasonal cycle module through residual connection, and perform feature decomposition through a bidirectional residual structure to output the predicted missing values.

[0028] After the steps of iteratively optimizing the initial BEIN model through a preset loss function to obtain a trained multivariate time series missing value prediction model in combination with the first aspect, it further includes:

[0029] Obtain the monitoring data at a given time scale;

[0030] Perform standardization processing and missing marking on the monitoring data to determine the time data corresponding to the missing values;

[0031] In response to the time data being the current time data, input the previous time data into the multivariate time series missing value prediction model, and output the first target predicted missing value.

[0032] In combination with the first aspect, after the step of performing standardization processing and missing value marking on the monitoring data and determining the time data corresponding to the missing value, the method further includes:

[0033] In response to the time data being historical time data, input the previous time data corresponding to the time data into the multivariate time series missing value prediction model, and output the second target predicted missing value; at the same time, input the subsequent time data corresponding to the time data into the multivariate time series missing value prediction model, and output the third target predicted missing value;

[0034] Combine the second target predicted missing value and the third target predicted missing value to determine the target predicted missing value.

[0035] The embodiments of the present invention bring the following beneficial effects: The method for constructing a multivariate time series missing value prediction model provided in this application includes: obtaining the original sensor data set collected during the sewage treatment process, and performing standardization and missing value marking on the original data in the original sensor data set to obtain a training data set, where each training data in the training data set is a multivariate time series; calculating the mutual information of each variable in the multivariate time series based on the mutual information module in the initial BEIN model, and using minimum redundancy maximum correlation to determine the key exogenous variables used as supplementary information; extracting the overall trend characteristics in the multivariate time series based on the constant module in the initial BEIN model, and at the same time, applying the seasonal cycle module in the initial BEIN model to extract the periodic characteristics in the multivariate time series; fusing the overall trend characteristics, periodic characteristics, and the characteristics of the key exogenous variables through residual connection to obtain the predicted missing value; iteratively optimizing the initial BEIN model through a preset loss function to obtain the trained multivariate time series missing value prediction model. The present invention proposes that this application optimizes and trains the initial BEIN model including a constant module, a seasonal cycle module, and a mutual information module to output the predicted value for imputing missing data, effectively solving the problem of missing sensor data during the sewage treatment process. This method aims at the composite cycle characteristics commonly existing in industrial time series data, realizes multi-scale component analysis through a dynamic cycle decomposition module, and at the same time, aiming at the strong correlation between industrial sensor data, introduces a mutual information quantification module to construct a dynamic association topology, screens the wastewater components highly correlated with the target variable as the key exogenous variables, and establishes a data distribution reconstruction mechanism based on physical constraints to improve the accuracy, reliability, and real-time performance of missing value imputation.

[0036] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention are realized and attained by the structure particularly pointed out in the specification, claims and drawings.

[0037] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0038] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 Schematic flowchart of a method for constructing a missing value prediction model for a multivariate time series provided by an embodiment of the present invention;

[0040] Figure 2 Partial structural schematic diagram of an initial BEIN model in a method for constructing a missing value prediction model for a multivariate time series provided by an embodiment of the present invention;

[0041] Figure 3 Overall structure diagram of a time series expansion module in an initial BEIN model provided by an embodiment of the present invention;

[0042] Figure 4 Principle schematic diagram of the extraction of the overall trend characteristics in a multivariate time series by a constant module in an initial BEIN model provided by an embodiment of the present invention;

[0043] Figure 5 Principle schematic diagram of the extraction of the periodic characteristics in a multivariate time series by a seasonal cycle module in an initial BEIN model provided by an embodiment of the present invention;

[0044] Figure 6 Root mean square errors of missing value interpolation for test sets 1, 2, and 3 of the influent flow of a sewage treatment plant provided exemplarily;

[0045] Figure 7 Exemplarily provided with Figure 6 Average absolute percentage errors of missing value interpolation corresponding to test sets 1, 2, and 3 in

[0046] Figure 8Exemplarily provide the output schematic diagram of the constant module when the multi-variable time series missing value prediction model performs the interpolation of the missing value of the influent flow rate during the operation of the sewage treatment plant;

[0047] Figure 9 Exemplarily provide the output schematic diagram of the seasonal cycle module when the multi-variable time series missing value prediction model performs the interpolation of the missing value of the influent flow rate during the operation of the sewage treatment plant;

[0048] Figure 10 Exemplarily provide the output schematic diagram of the mutual information module when the multi-variable time series missing value prediction model performs the interpolation of the missing value of the influent flow rate during the operation of the sewage treatment plant. Detailed implementation manners

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0050] To facilitate the understanding of this embodiment, the application scenario and design concept of the embodiments of this application will be briefly introduced below.

[0051] Limitations of traditional statistical methods: mainly rely on simple calculations of local data, and it is difficult to capture the non-linear variation laws of sewage flow and pollutant concentration; in the case of a high missing rate, the error increases significantly; it is unable to effectively process complex patterns with multiple cycles superimposed.

[0052] Deficiencies of existing machine learning models: Although they can utilize some spatio-temporal correlations in the data, they are inefficient in dealing with long-range dependence relationships between high-dimensional variables; when there are lag correlations across components in sensor data, the filling results are prone to temporal misalignment.

[0053] Based on this, the embodiments of this application provide a construction method for a multi-variable time series missing value prediction model.

[0054] Embodiment 1

[0055] This application provides a construction method for a multi-variable time series missing value prediction model. As shown in Figure 1 The method includes:

[0056] S110, obtain the original sensor data set collected during the sewage treatment process, and after standardizing and missing-marking the original data in the original sensor data set, obtain a training data set, and each training data in the training data set is a multi-variable time series.

[0057] S120. Calculate the mutual information of each variable in the multivariate time series based on the mutual information module in the initial BEIN model, and use minimum redundancy maximum correlation to determine the key exogenous variables to be used as supplementary information.

[0058] S130. Extract the overall trend features in the multivariate time series based on the constant module in the initial BEIN model. Meanwhile, apply the seasonal cycle module in the initial BEIN model to extract the periodic features in the multivariate time series.

[0059] S140. Fuse the overall trend features, periodic features, and the features of the key exogenous variables through residual connection to obtain the predicted missing values.

[0060] S150. Iteratively optimize the initial BEIN model through a preset loss function to obtain a trained multivariate time series missing value prediction model.

[0061] This application optimizes and trains the initial BEIN model including a constant module, a seasonal cycle module, and a mutual information module to output predicted values for imputing missing data, effectively solving the problem of missing sensor data in the sewage treatment process. This method aims at the composite cycle features commonly existing in industrial time series data, realizes multi-scale component analysis through a dynamic cycle decomposition module. Meanwhile, aiming at the strong correlation between industrial sensor data, it introduces a mutual information quantification module to construct a dynamic association topology, screens wastewater components highly correlated with the target variable as key exogenous variables, and establishes a data distribution reconstruction mechanism based on physical constraints to improve the accuracy, reliability, and real-time performance of missing value imputation.

[0062] Combined with the first aspect, the original data in the original dataset includes one or more of the following: influent flow rate, chemical oxygen demand, ammonia nitrogen content, pH value, external reflux flow rate, oxidation-reduction potential, nitrate nitrogen content, mixed liquor suspended solid content, dissolved oxygen concentration, phosphate content, total phosphorus content, temperature, and total nitrogen content.

[0063] After collecting the sensor raw data in the sewage treatment process in step S110, perform standardization processing (such as normalization, Z-score standardization, etc.) on the raw data, and mark all missing data points. This step ensures the effectiveness and accuracy of subsequent model training.

[0064] For example, the training data after standardization and missing value marking of some detection data for three consecutive days in a sewage treatment plant is shown in Table 1 as follows:

[0065] Table 1 is an example table of each training data in the training dataset.

[0066]

[0067] Among them, NaN indicates missing data. It can be seen that missing values may occur in different variables at different time points. In response to this situation, we can fill in the missing values by constructing a prediction model for missing values in a multivariate time series to ensure the accuracy of subsequent analysis and decision-making. Combining Table 1, there are 8 variables, namely influent flow rate, chemical oxygen demand, ammonia nitrogen content, pH value, external reflux flow rate, oxidation-reduction potential, nitrate nitrogen content, and mixed liquor suspended solids content.

[0068] Combined with the first aspect, step S120 includes:

[0069] S121, for each first variable, calculate the mutual information between the first variable and each second variable using the K-nearest neighbor estimation method.

[0070] S122, combining all the mutual information, select a set of exogenous variables strongly correlated with the first variable from multiple second variables based on minimum redundancy maximum relevance.

[0071] As shown in Table 1, there are 8 variables, namely influent flow rate, chemical oxygen demand, ammonia nitrogen content, pH value, external reflux flow rate, oxidation-reduction potential, nitrate nitrogen content, and mixed liquor suspended solids content.

[0072] In this embodiment, step S120 is used to calculate mutual information and select key exogenous variables. Among them, mutual information can measure the dependence relationship between two variables and help understand the correlation between variables. Then, the minimum redundancy maximum relevance (mRMR) algorithm is used to determine which exogenous variables are the key variables for supplementary information. These key exogenous variables can provide additional information to help improve the prediction performance of the model.

[0073] Specifically, step S121 includes:

[0074] Let variable X have multiple components, and the multiple components are x 1 , x 2 ...x m . Among them, x m is the m-th component, and m is a preset quantity threshold. The entropy H(X) of variable X is defined as:

[0075] H(X) = -∑ x∈X p(x i ) log 2 p(x i ); where p(x i ) represents the probability distribution of the i-th component of variable X.

[0076] The mutual information I(X, Y) between variable X and variable Y is described as:

[0077] I(X, Y) = H(X) + H(Y) - H(X, Y).

[0078] Combining the above formula, for a pair of discrete variables (X, Y) subject to the joint distribution p(x, y), the mutual information I(X, Y) between the two variables (i.e., X and Y) is described as:

[0079]

[0080] The variable to be imputed is the first variable, and the other variables are the second variables. Combining Table 1, in this embodiment, the number of the first variables is 1, specifically "chemical oxygen demand"; the second variables are the other variables except the first variable, that is, influent flow rate, ammonia nitrogen content, pH value, external reflux flow rate, oxidation-reduction potential, nitrate nitrogen content, and mixed liquor suspended solid content. In this way, the mutual information between each second variable and the first variable can be calculated based on the above method.

[0081] Step S122 includes:

[0082] Using the minimum redundancy maximum relevance (mRMR) algorithm, iteratively select the set of exogenous variables that are most relevant to the prediction of the target variable from the candidate variables (i.e., multiple second variables):

[0083]

[0084] This equation represents minimizing the redundancy of the candidate feature X n with the selected feature set S, while ensuring that the selected feature X n (a certain second variable) has the maximum relevance to the target feature Y (i.e., the first variable), to obtain the set of exogenous variables. Among them, X j is the j-th feature selected in the feature set.

[0085] After that, step S130 inputs the initial BEIN model. The initial BEIN model includes a time series expansion module. The time series expansion module includes a constant module, a seasonal cycle module, and a multi-level cascaded fully connected layer. Combining Figure 3 as shown, each fully connected layer (FC) contains the following two main modules: a constant module for capturing the in-cycle information (overall trend feature) in the time series; a seasonal cycle module for capturing the inter-cycle information (periodic feature) in the time series.

[0086] Combining Figure 2 the l-th fully connected layer shown, the mutual information module captures the features of the key exogenous variables that are most relevant to the variable, and decomposes the periodic state of the l-th expansion term of the current input signal, denoted as Among them, is composed of the backward prediction component and the forward prediction component (not shown in the figure), the mutual information module uses to analyze the influence of relevant variables on the current layer and uses to represent the additional information purely from the l-th layer. Among them, t is the current moment, L is the forward time window, and H is the backward time window.

[0087] Among them, the constant module decomposes the state of the l-th expansion term of the input signal through a fully connected layer and a linear projection function. Combining Figure 4 as shown, the input signal of the constant module is the extended residual term of the previous fully connected layer (i.e., the l-1 layer) to output the backward decomposition component and the forward prediction component This backward decomposition component and the forward prediction component constitute the final output, which is represented as (not shown in the figure). Among them, the backward decomposition component is used to exclude the influence of historical features captured on the current fully connected layer, and the forward prediction component is used to represent the cumulative prediction purely from the l-th fully connected layer. This design ensures that each layer only learns the features of the current layer, avoids the interference of historical features, and thus improves the prediction accuracy and generalization ability of the model.

[0088] Combined with the first aspect, the extraction of the overall trend features in the multivariate time series based on the constant module in the initial BEIN model in step S130 specifically includes:[[]]

[0089] S131, encoding the input multivariate time series through a fully connected layer to extract high-level features.

[0090] S132, generating the backward decomposition component by passing the high-level features through the first linear projection function, and at the same time, generating the forward prediction component by passing the high-level features through the second linear projection function to extract the overall trend features.

[0091] Specifically, step S131 can be expressed as:[[]]

[0092]

[0093] Among them, FC() is the fully connected layer, H l is the encoded high-level feature,[[]] is the extended residual term of the l-th fully connected layer.

[0094] The decomposition operation in step S132 can be expressed as:[[]]

[0095]

[0096] Among them, Linerb is the linear projection function for inverse decomposition (i.e., the first linear projection function); the inverse decomposition component is used to exclude the influence of the historical features captured in the current layer, ensuring that each layer only learns the features of the current layer.

[0097] The prediction operation in step S132 can be expressed as:

[0098]

[0099] Among them, Linerf is the linear projection function for forward prediction (i.e., the second linear projection function); the forward prediction component represents the prediction purely from the l-th fully connected layer, and is used to accumulate the prediction results. The cumulative prediction can be expressed as:

[0100]

[0101] Among them, is the cumulative prediction of the (l - 1)-th fully connected layer.

[0102] The seasonal cycle module is another key component in the multivariate time series decomposition method, and is used to extract the periodic information in the time series.

[0103] In step S130, applying the seasonal cycle module to extract the periodic features in the multivariate time series specifically includes:

[0104] S133, after encoding the input multivariate time series through the fully connected layer, generating decomposition coefficients and prediction coefficients through the third linear projection function and the fourth linear projection function.

[0105] Specifically, encoding the input multivariate time series through the fully connected layer to extract features; subsequently, passing the encoded features through the third linear projection function (the same as the first linear projection function in this embodiment, which can be selected according to actual usage requirements and is not limited here) to generate decomposition coefficients At the same time, passing the encoded features through the fourth linear projection function (the same as the second linear projection function in this embodiment, which can be selected according to actual usage requirements and is not limited here) to generate prediction coefficients

[0106] S134, analyzing the spectral energy distribution of the input multivariate time series through the fast Fourier transform FFT to construct Fourier basis functions; among them, the basis functions include backward decomposition basis functions and forward prediction basis functions.

[0107] Specifically, first, perform Fourier transform, frequency-domain signal conversion, and smoothing on the univariate time series with N sampling points to obtain the frequency and amplitude corresponding to each sampling point among the N sampling points, where N is a preset quantity threshold.

[0108] Then, select k frequencies based on a preset algorithm to obtain a set of selected frequencies {f 1 ,..., f k}, and at the same time, combine the amplitudes corresponding to the selected frequencies to obtain a set of selected amplitudes {A 1 ,..., A k}. As an implementable method, the preset algorithm is to calculate the probability value corresponding to each frequency. After arranging the probability values in descending order, select the frequencies corresponding to the top k probability values to obtain a set containing k selected frequencies; as another implementable method, select k selected frequencies through a roulette wheel selection function, softmax function, etc.; the above methods are all implementable, and this is only an example here without limitation. f 1 is the first selected frequency, f k is the k-th selected frequency, A 1 is the amplitude corresponding to the first selected frequency, and A k is the k-th amplitude corresponding to the k-th selected frequency.

[0109] Then, construct a basis function expressed in the form of Fourier series (sin and / or cos form) based on the selected frequencies and selected amplitudes. Specifically, use the following formula:

[0110] where N 0 is a hyperparameter for controlling resonant oscillation; h(c) is the basis function, f i is the i-th selected frequency in the set of selected frequencies, is the -th selected frequency in the set of selected frequencies, A i is the amplitude corresponding to the i-th selected frequency, k is the preset quantity threshold, is the -th amplitude corresponding to the -th selected frequency, c is the output of the linear projection function in the seasonal period module, is the value obtained by rounding down

[0111] Combined with step S134, the output c of the linear projection function in the seasonal period module is the decomposition coefficient and the prediction coefficient At this time, the basis function obtained by substituting the decomposition coefficient into the above formula is the backward decomposition basis function h b ( ); substituting the prediction coefficient The basis function obtained by substituting into the above formula is the forward prediction basis function h f ( ).

[0112] In step S135, the decomposition term for updating the residual is calculated by combining the decomposition coefficients and the basis function of backward decomposition. At the same time, the prediction term for accumulating the prediction result is calculated by combining the prediction coefficients and the forward prediction basis function.

[0113] Combined with the attached Figure 5 Substitute the decomposition coefficients into the backward decomposition basis function h b ( ) to calculate the decomposition term for updating the residual

[0114] Similarly, substitute the prediction coefficients into the forward prediction basis function h f ( ) to calculate the prediction term for accumulating the prediction result

[0115] S136. Determine the periodic feature by combining the decomposition term and the prediction term.

[0116] In this embodiment, the periodic feature includes the decomposition term and the prediction term. The decomposition term and the prediction term are calculated through steps S133 - S135, thereby determining the periodic feature.

[0117] Subsequently, perform step S140 based on the cross - layer skip connection structure of the residual connection module. Transmit the decomposition residual of the previous layer to the input end of the next layer, and combine the periodic component of the seasonal cycle module to gradually strip redundant information and improve the feature learning efficiency. This design ensures that each layer only learns the features of the current layer, avoiding the interference of historical features, thereby improving the prediction accuracy and generalization ability of the model.

[0118] Combined with the first aspect, step S140 includes:

[0119] Transmit the overall trend feature, the periodic feature, and the feature of the key exogenous variable to the constant module and the seasonal cycle module through the residual connection, and perform feature decomposition through the bidirectional residual structure to output the predicted missing value.

[0120] Specifically:

[0121]

[0122]

[0123] Among them, Z t-L,t+H represents the feature of the exogenous information, represents the output of the exogenous information at the 0th layer, represents the output of the mutual information module, Represents the output of the Nth layer of external information;

[0124] x t-L,t Represents the input multivariate time series, Represents the output of the 0th layer of multivariate time series, Represents the backward component (i.e., the decomposition term) of the seasonal cycle module, Represents the backward component of the constant module, Represents the output of the Nth layer of multivariate time series;

[0125] Represents the final output, Represents the predicted value (i.e., the prediction term) of the Nth layer of multivariate time series, Represents the forward component of the constant module, Represents the forward component of the seasonal cycle module.

[0126] Where t is the current time, L is the forward time window, H is the backward time window, and N is the preset quantity threshold.

[0127] Finally, step S150 will calculate the error between the model predicted value and the true value, and the error is quantified by a preset loss function; then use the backpropagation algorithm to update the model parameters to minimize the loss function; subsequently, repeat the above steps until the model converges or reaches the preset number of iterations, so as to obtain a trained multivariate time series missing value prediction model. The preset loss function can be mean square error, mean absolute error, etc., which are not limited here.

[0128] In this embodiment, not only focuses on how to handle the data missing problem, but also particularly examines the robustness of the model prediction ability in the case of introducing measurement noise. Specifically, this study assumes that the measurement noise follows a Gaussian distribution with a mean of zero, and the noise variables are independent of each other. In this way, researchers can more realistically simulate the uncertainty factors in the actual application scenario. Next, I will explain the experimental design, evaluation metrics, and expected results in detail.

[0129] Hypothesis conditions: The measurement noise follows a Gaussian distribution with a mean of 0; the noise variables are independent of each other.

[0130] Testing method: Introduce randomly omitted input samples to simulate the data missing situation that may occur in the real world; test on the trained model to evaluate its prediction ability; use two different missing types, discrete and block missing patterns, to comprehensively test the robustness of the model.

[0131] Performance evaluation: To ensure the comprehensiveness and objectivity of the evaluation, two commonly used regression performance metrics are adopted: root mean square error (RMSE) and mean absolute percentage error (MAPE).

[0132] The root mean square error (RMSE) measures the magnitude of the difference between the predicted value and the actual value, reflecting the prediction accuracy of the model; the smaller the value, the better; the mean absolute percentage error (MAPE) represents the deviation degree of the predicted value relative to the true value in percentage form; similarly, the lower the value, the better.

[0133] Expected result model comparison: Compare the proposed model with other deep learning-based methods to verify its superiority or uniqueness.

[0134] Performance analysis: Analyze the performance of the model under different missing ratios, especially the ability to maintain high accuracy when the data is severely incomplete (such as 80% missing).

[0135] Conclusion: If the values of RMSE and MAPE are low, it indicates that the model has good generalization ability and robustness, and can make relatively accurate predictions in the presence of a large amount of missing data.

[0136] From the above steps, it can be seen that in the face of complex and changing real environments, a model with good anti-noise ability and skills in handling missing data is very important. It can not only improve the prediction accuracy, but also enhance the reliability and stability of the system. This is especially crucial for fields such as industrial process monitoring and environmental monitoring, because the data in these fields is often affected by various interferences and limitations.

[0137] Refer to the RMSE / MAPE of the BEIN model under different missing rates shown in Table 2.

[0138]

[0139] According to the data in Table 2, the RMSE / MAPE of each test set under the missing rate of 10% - 80%, it can be clearly seen that the BEIN model performs well in processing time series data with different missing rates, especially maintaining relatively low root mean square error RMSE and mean absolute percentage error MAPE values even under high missing rates. The following is a specific analysis of the results in Table 2: The RMSE / MAPE of the BEIN model under different missing rate conditions. Even in extreme cases (such as a missing rate of 80%), BEIN can still maintain relatively low errors, indicating its significant advantage in handling high-proportion missing data.

[0140] Specifically, for Test Set 1: As the missing rate decreases, the RMSE value generally shows a downward trend, but the MAPE value fluctuates at some points. For example, the MAPE slightly increases when changing from 40% to 30%. This may be because the change in data patterns affects the prediction accuracy;

[0141] For Test Set 2: Whether it is RMSE or MAPE, both metrics gradually decrease as the missing rate decreases, showing relatively stable performance;

[0142] For Test Set 3: Compared with the other two test sets, the RMSE and MAPE of Test Set 3 always remain at a relatively low and stable level, demonstrating its good adaptability to different missing rates.

[0143] Combined with the 80% missing rate in Table 2: Even at such a high missing rate, the RMSE and MAPE of the BEIN model still remain at a low level, which are 2977.58 and 0.491922, 1037.15 and 0.097359, 415.191 and 0.093141 respectively, proving the strong robustness of the model.

[0144] Combined with the other missing rates in Table 2: As the missing rate decreases, the model performance is further improved. Especially under the missing rates of 20% and 10%, the MAPE reaches the lowest points (0.081235 and 0.077713) respectively, indicating that the model can make more accurate predictions at this time.

[0145] Combined with the first aspect, after step S150, it further includes:

[0146] S160, obtaining the monitoring data under a given time scale.

[0147] S170, performing standardization processing and missing value marking on the monitoring data, and determining the time data corresponding to the missing values.

[0148] S181, in response to the time data being the current time data, inputting the previous time data corresponding to the time data into the multivariate time series missing value prediction model, and outputting the first target predicted missing value.

[0149] Combined with the first aspect, after step S170, it further includes:

[0150] S182, in response to the time data being historical time data, inputting the previous time data corresponding to the time data into the multivariate time series missing value prediction model, and outputting the second target predicted missing value; at the same time, inputting the next time data corresponding to the time data into the multivariate time series missing value prediction model, and outputting the third target predicted missing value.

[0151] S183, combining the second target predicted missing value and the third target predicted missing value to determine the target predicted missing value.

[0152] In the actual application process, the collected monitoring data is standardized and marked to determine the time data corresponding to the missing values. When dealing with the missing data of "current time", we only use the time data before the sampling point for interpolation. In this case, the future data is unavailable or not yet generated, so we can only rely on the existing historical data to fill the missing values. By inputting the time data before the sampling point into the multi-variable time series missing value prediction model and outputting the predicted value, this predicted value is used as the interpolation value for the current time.

[0153] However, there is another situation in the actual application process, that is, when the missing data is historical data, we can use the data before and after the sampling point simultaneously for more accurate estimation. Because the data of the entire time period is known at this time, more information can be obtained from two directions. First, in the way of one-way interpolation, based on the above multi-variable time series missing value prediction model, only the data before the sampling point is used for preliminary filling; then, the one-way interpolation method is applied again, but this time it is from the future (i.e., after the sampling point) to the past (i.e., the sampling point position) for reverse filling; finally, the results obtained from the two interpolations are combined, such as taking the average value or using the weighted average method, etc., to obtain a more accurate estimation.

[0154] Under the condition of a 20% missing rate, the test set is iteratively trained and the corresponding evaluation indicators are recorded such as Figure 6 , Figure 7 as shown. Combining Figure 6 the missing value interpolation results of the test sets 1, 2, and 3 for the influent flow of the sewage treatment plant provided exemplarily, where the horizontal axis is the cycle of iterative training, Figure 6 and the vertical axis in

[0155] is the first evaluation indicator: root mean square error; test set 1 is represented by a solid line, test set 2 is represented by a dashed line, and test set 3 is represented by a solid line with discrete points. These curves show the interpolation results, and it can be seen that they overlap at some points and are different at other points. Figure 7 Combining

[0156] shows the missing value interpolation errors of the test sets 1, 2, and 3 for the influent flow of the sewage treatment plant provided exemplarily. Among them, the horizontal axis is also the cycle of iterative training, and the vertical axis is the second evaluation indicator mean absolute percentage error; test set 1 is represented by a solid line, test set 2 is represented by a dashed line, and test set 3 is represented by a solid line with discrete points. The error values of the three curves are close to zero in most regions. Figures 8 - 10 Figure 8 Figure 8As shown, the constant module represents the baseline part of the time series and is used to reveal the potential trend of the data in the absence of any seasonal cycles or exogenous influences; combined with Figure 9 As shown, the seasonal cycle module captures the periodic fluctuations in the time series. As can be seen from Figure 9 , the seasonal signal shows a clear periodic fluctuation pattern, which is of great significance for timely responding to environmental changes in application scenarios with strong seasonality; combined with Figure 10 As shown, the mutual information module represents the influence of external variables on the time series. The output of the mutual information module shows a trend similar to that of the constant module, which provides another perspective for studying the rationality of the module and its impact on the performance of the benchmark module. In this way, it not only helps to identify and process missing values in the time series, but also provides valuable data support for optimizing the operation of the sewage treatment plant.

[0157] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0158] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0159] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks and other various media that can store program codes.

[0160] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0161] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for constructing a multivariate time series missing value prediction model, characterized in that: The method comprises: A sensor raw data set collected during sewage treatment is obtained, and raw data in the sensor raw data set is standardized and missing marked to obtain a training data set, wherein each training data in the training data set is a multivariate time series; The mutual information of each variable in the multivariate time series is calculated based on the mutual information module in the initial BEIN model, and the key exogenous variables used as supplementary information are determined by using minimum redundancy and maximum correlation; Extracting the overall trend features in the multivariate time series based on the constant module in the initial BEIN model, and at the same time, applying the seasonal cycle module in the initial BEIN model to extract the periodic features in the multivariate time series; The overall trend characteristics, the periodic characteristics and the characteristics of the key exogenous variables are integrated through residual connection to obtain predicted missing values; The initial BEIN model is iteratively optimized through a preset loss function to obtain a trained multivariate time series missing value prediction model.

2. The method according to claim 1, characterized in that: The raw data include: one or more of inlet flow, chemical oxygen demand, ammonia nitrogen content, pH value, external reflux flow, redox potential, nitrate nitrogen content, mixed liquor suspended solids content, dissolved oxygen concentration, phosphate content, total phosphorus content, temperature, and total nitrogen content.

3. The method according to claim 1, characterized in that The initial BEIN model includes a time series extension module, which includes a constant module, a seasonal cycle module and a multi-layer cascaded fully connected layer.

4. The method according to claim 3, characterized in that Each of the fully connected layers includes the constant module and the seasonal periodic module; the constant module is used to capture intra-periodic information in the time series; the seasonal periodic module is used to capture inter-periodic information in the time series.

5. The method according to claim 1, characterized in that The steps of calculating the mutual information of each variable in the multivariate time series based on the mutual information module in the initial BEIN model, and determining the key exogenous variables used as supplementary information by using minimum redundancy and maximum correlation include: For each first variable, using a K-nearest neighbor estimation method to calculate each mutual information between the first variable and each second variable; In combination with all of the mutual information, a set of exogenous variables that are strongly correlated with the first variable is selected from a plurality of the second variables based on minimum redundancy and maximum correlation.

6. The method according to claim 1, characterized in that The step of extracting the overall trend characteristics in the multivariate time series based on the constant module in the initial BEIN model comprises: Encoding the input multivariate time series through a fully connected layer to extract high-level features; The high-level features are used to generate reverse decomposition components through a first linear projection function, and at the same time, the high-level features are used to generate forward prediction components through a second linear projection function to extract overall trend features.

7. The method according to claim 1, characterized in that The step of applying the seasonal cycle module in the initial BEIN model to extract the periodic features in the multivariate time series comprises: After encoding the input multivariate time series through the fully connected layer, the decomposition coefficients and prediction coefficients are generated through the third linear projection function and the fourth linear projection function; Analyzing the spectrum energy distribution of the input multivariate time series by fast Fourier transform (FFT) to construct basis functions; wherein the basis functions include backward decomposition basis functions and forward prediction basis functions; Combining the decomposition coefficients and the backward decomposition basis functions to calculate a decomposition term for updating the residual, and combining the prediction coefficients and the forward prediction basis functions to calculate a prediction term for accumulating prediction results; The decomposed items and the predicted items are combined to determine periodic characteristics.

8. The method according to claim 1, characterized in that: The step of fusing the overall trend feature, the periodic feature and the feature of the key exogenous variable through residual connection to obtain the predicted missing value comprises: The overall trend characteristics, the periodic characteristics and the characteristics of the key exogenous variables are passed to the constant module and the seasonal cycle module through residual connections, and feature decomposition is performed through a bidirectional residual structure to output predicted missing values.

9. The method according to any one of claims 1 to 8, characterized in that: After the step of iteratively optimizing the initial BEIN model through a preset loss function to obtain a trained multivariate time series missing value prediction model, the method further includes: Obtain monitoring data at a given time scale; The monitoring data is standardized and marked as missing, and the time data corresponding to the missing values ​​is determined; In response to the time data being current time data, previous time data is input into the multivariate time series missing value prediction model, and a first target predicted missing value is output.

10. The method according to claim 9, characterized in that After the steps of normalizing and marking missing values ​​on the monitoring data and determining the time data corresponding to the missing values, the method further includes: In response to the time data being historical time data, the previous time data corresponding to the time data is input into the multivariate time series missing value prediction model, and a second target predicted missing value is output; at the same time, the next time data corresponding to the time data is input into the multivariate time series missing value prediction model, and a third target predicted missing value is output; The target prediction missing value is determined by combining the second target prediction missing value and the third target prediction missing value.

Citation Information

Cited By

  • Product cost prediction method and system based on multi-mode interpolation strategy, and electronic equipment

    CN121120120A

  • Solar irradiance missing value interpolation method and system

    CN122412769A