Water Environment Operation and Maintenance Monitoring Method and System Based on Data Analysis
Through the time series model of long and short-term memory network based on data analysis and the prediction of water environment treatment data, the subjectivity of human judgment and the complexity of data integration are solved, and efficient and stable operation of water environment operation and maintenance and sustainable utilization of water resources are achieved.
Patent Information
- Application Number
- CN202411592584.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-08
AI Technical Summary
In the existing water environment monitoring and management, there is subjectivity and inconsistency in human judgment, insufficient real-time response capabilities, complex data integration and high cost, which affects the accuracy and efficiency of operation and maintenance decisions.
Using a data analysis method, the front and back processing data of the water environment are time stamped, mismatched and predicted by the long and short-term memory network time series model, and the features are extracted using the LSTM layer and the fully connected layer to realize the automated processing and prediction of data.
It improves the accuracy and real-time nature of water environment operation and maintenance monitoring, reduces man-made errors, ensures the efficient and stable operation of the water treatment system, and realizes the sustainable utilization of water resources.
Smart Images

Figure CN119476719B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a water environment operation and maintenance monitoring method and system based on data analysis. Background Art
[0002] Currently, in the specific working conditions of water environment monitoring and management, the equipment involved mainly includes:
[0003] A grille for performing the front-end treatment of water, intercepting and removing solid impurities and floating objects in the water body, preventing these impurities from entering the subsequent treatment system, and ensuring the smooth and efficient water treatment process; a pumping station responsible for lifting and conveying the water body, providing power for the water flow, and ensuring the flow and circulation of water in each link of the subsequent treatment system. The subsequent treatment system includes sedimentation tanks, filters, disinfection equipment, etc.; a remote control module for realizing the remote monitoring and control of the water treatment system, often connecting each device through Internet of Things technology to provide data transmission and control interfaces; power distribution equipment for providing stable power supply for each electrical device in the water treatment system to ensure the normal operation of the system. Through the effective operation and maintenance and comprehensive management of the above equipment, the treatment and improvement effect of the water environment can be significantly improved, and the sustainable utilization of water resources can be guaranteed.
[0004] However, in the current implementation process, only the acquisition of each independent data before and after is realized, and the judgment and prediction of the effect of water environment operation and maintenance through each data are often achieved through manual means. For example: visualizing each data to facilitate the staff to intuitively but relatively subjectively judge the result of water environment operation and maintenance according to the displayed result. This kind of practice often has the following problems:
[0005] Manual judgment is highly subjective. Different personnel have different experience and knowledge levels, resulting in differences in judgment results, and the operation and maintenance decisions will be inconsistent and unstable; manual analysis and judgment take time and are difficult to respond and handle emergencies in real time, delaying the discovery and solution of problems, which may lead to more serious equipment failures and environmental pollution; the data of each device need to be manually integrated and analyzed. The amount of data is large and complex, relying on manual analysis and judgment, which requires a large amount of human resources, and the training and management costs are relatively high. In addition, it is easy to miss important information or produce errors, and the inaccurate data analysis results will affect the quality and effect of operation and maintenance decisions. Summary of the Invention
[0006] The present invention provides a water environment operation and maintenance monitoring method and system based on data analysis, which can effectively solve the problems in the background art.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] A water environment operation and maintenance monitoring method based on data analysis, including:
[0009] Collecting the pre - treatment data and post - treatment data for the water environment;
[0010] Aligning the timestamps of the pre - treatment data and the post - treatment data;
[0011] Matching the pre - treatment data at a pre - set collection moment and the post - treatment data at a post - set collection moment to obtain a data set of misaligned corresponding data;
[0012] Using a long short - term memory network time - series model and inputting the said data set;
[0013] Implementing the training, evaluation and optimization of the said time - series model;
[0014] Predicting the water environment through the optimized time - series model.
[0015] Further, aligning the timestamps of the pre - treatment data and the post - treatment data includes:
[0016] Judging whether the collection frequencies of the pre - treatment data and the post - treatment data are the same. If so, directly align the timestamps; if not, perform interpolation processing on the party with the lower collection frequency. The interpolation processing is used to fill in the missing values so that the data volumes of the pre - treatment data and the post - treatment data are the same, and then align the timestamps.
[0017] Further, the misaligned time corresponding to the misalignment is determined by the following steps:
[0018] Obtaining the pre - treatment data and the post - treatment data within a set time;
[0019] Calculating the cross - correlation between the pre - treatment data and the post - treatment data. The cross - correlation function is:
[0020]
[0021] where a t is the value of the pre - treatment data at time t, k is the number of time lag steps, and b t+k is the value of the post - treatment data after lagging k steps at time t, and are the average values of the pre - treatment data and the post - treatment data respectively, and N is the number of data points;
[0022] Using the number of lag time steps k that makes the value of the cross - correlation function the largest to calculate the misaligned time.
[0023] Furthermore, the long short-term memory network time series model includes:
[0024] An input layer that accepts the data set as input;
[0025] An LSTM layer, which consists of a number of LSTM units, and each LSTM unit contains a memory unit and three gating mechanisms to control the flow of information;
[0026] A number of fully connected layers that accept the output from the LSTM layer, and perform processing and feature transformation to map high-dimensional features to the required dimensions;
[0027] An output layer that outputs a prediction result according to the feature transformation result.
[0028] Furthermore, the number of LSTM layers is a linear function of the number of fully connected layers, and the number of fully connected layers is less than the number of LSTM layers.
[0029] Furthermore, the relationship between the number of LSTM layers and the number of fully connected layers is:
[0030] N LSTM = α × N FC ;
[0031] where N LSTM is the number of LSTM layers, N FC is the number of fully connected layers; α is a positive integer and is positively correlated with the acquisition frequency difference between the previous processed data and the subsequent processed data.
[0032] Furthermore, a peephole connection is added between the memory unit and the three gating mechanisms of the long short-term memory network time series model.
[0033] Furthermore, implementing the training, evaluation, and optimization of the time series model includes:
[0034] Converting the data set and dividing the data set into a training set, a validation set, and a test set;
[0035] Compiling the time series model and selecting a loss function, an optimizer, and evaluation metrics;
[0036] Setting the hyperparameters of the time series model, using the training set to iteratively train the time series model, and using the validation set to evaluate the time series model during the iterative training process, and adjusting the time series model according to the evaluation results of the validation set;
[0037] Using the test set to evaluate the performance of the time series model and calculating the evaluation metrics;
[0038] Optimize the time series model according to the evaluation results of the test set, and deploy the optimized time series model to the water environment monitoring system.
[0039] Further, transform the data set and divide the data set into a training set, a validation set, and a test set, including:
[0040] Define a time window, create data pairs based on the time window, and the data pairs include feature data and label data;
[0041] Divide the data set into a training set, a validation set, and a test set;
[0042] Adjust the shape of the feature data to match the input of the time series model.
[0043] A water environment operation and maintenance monitoring system based on data analysis, the system includes:
[0044] A water environment data collection module that collects pre-processing data and post-processing data for the water environment;
[0045] A timestamp alignment module that aligns the timestamps of the pre-processing data and the post-processing data;
[0046] A water environment data matching module that matches the pre-processing data at a previously set acquisition moment and the post-processing data at a subsequently set acquisition moment to obtain a data set corresponding to a time difference;
[0047] A long short-term memory network model module that uses a long short-term memory network time series model and inputs the data set;
[0048] A time series model training module that implements the training, evaluation, and optimization of the time series model;
[0049] A time series model prediction module that predicts the water environment through the optimized time series model.
[0050] Through the technical solution of the present invention, the following technical effects can be achieved:
[0051] Through the present invention, the effect of water environment operation and maintenance monitoring and prediction can be significantly improved. Through comprehensive data operation and maintenance, the current subjectivity, real-time performance, and data integration problems can be effectively solved, ensuring the efficient and stable operation of the water treatment system and realizing the sustainable utilization of water resources. Description of the Drawings
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Figure 1 It is a schematic flow diagram of a water environment operation and maintenance monitoring method based on data analysis;
[0054] Figure 2 It is a schematic structural diagram of a long short-term memory network time series model;
[0055] Figure 3 It is a schematic flow diagram of converting a data set and dividing the data set into a training set, a validation set, and a test set;
[0056] Figure 4 It is a schematic structural diagram of a water environment operation and maintenance monitoring system based on data analysis. Detailed implementation manners
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0059] Embodiment 1
[0060] As Figure 1 shown, the water environment operation and maintenance monitoring method based on data analysis includes:
[0061] A1: Collect the pre-treatment data and post-treatment data for the water environment;
[0062] A2: Align the timestamps of the pre-treatment data and the post-treatment data;
[0063] A3: Match the pre-treatment data at the previously set collection time and the post-treatment data at the subsequently set collection time to obtain a misaligned corresponding data set;
[0064] A4: Use a long short-term memory network time series model to input the data set;
[0065] A5: Implement the training, evaluation, and optimization of the time series model;
[0066] A6: Predict the water environment using the optimized time series model.
[0067] In this embodiment, the long short-term memory network can capture long-term and short-term dependencies in time series data, which is particularly important for water treatment systems because the effects of the previous treatment often take some time to be reflected in the data of subsequent treatments. By using this model, the accuracy and reliability of prediction can be improved, and the dynamic behavior of the system can be better understood and predicted. The long short-term memory network can automatically extract important features from the input data without complex manual feature engineering, reducing the complexity and error of human operations and improving the adaptability and generalization ability of the model.
[0068] In an actual scenario, there may be complex non-linear relationships between variables in the water treatment system. The time series model can effectively handle these non-linear relationships. By continuously inputting time series data, the prediction results can be updated in real time, so as to timely discover and warn of potential problems, ensure the stable operation of the system, and reduce the impact of emergencies on the system.
[0069] Since the influence of the previous treatment in the water treatment process on the subsequent treatment is not instantaneous and there is a certain time delay, in the present invention, this time delay effect can be captured through time-shifted correspondence, so as to more accurately reflect the actual operating state of the system, improve the model's understanding and prediction ability of time dynamics. Through time-shifted correspondence, the noise and error caused by data asynchronization can be reduced, and the accuracy and consistency of the input data can be ensured. In an actual system, the influence of the previous treatment effect on the subsequent treatment appears gradually, and time-shifted correspondence can better reflect this gradual influence, making the model closer to the actual operating situation and the prediction results more realistic.
[0070] Through the present invention, the effect of water environment operation and maintenance monitoring and prediction can be significantly improved, the current subjectivity, real-time performance, and data integration problems can be solved, the efficient and stable operation of the water treatment system can be ensured, and the sustainable utilization of water resources can be realized.
[0071] Furthermore, perform timestamp alignment on the previous treatment data and the subsequent treatment data, including:
[0072] Judge whether the acquisition frequencies of the previous treatment data and the subsequent treatment data are the same. If so, directly align the timestamps; if not, perform interpolation processing on the party with the lower acquisition frequency. The interpolation processing is used to fill in the missing values to make the data volumes of the previous treatment data and the subsequent treatment data the same, and then align the timestamps.
[0073] Specifically, the acquisition frequencies of the pre - process data and the post - process data are the same, and no additional processing is required; if the acquisition frequencies are different, interpolation processing needs to be performed on the party with the lower acquisition frequency. Interpolation processing is to fill in the missing values in the dataset with the lower acquisition frequency so that the data points are consistent with those of the dataset with the higher acquisition frequency. Interpolation methods such as linear interpolation, spline interpolation, and polynomial interpolation can be used for interpolation processing. By judging whether the acquisition frequencies are the same and performing interpolation processing on the party with the lower acquisition frequency, the problem of data inconsistency caused by different acquisition frequencies can be effectively solved. At the same time, interpolation processing can fill in the missing values and reduce data loss caused by inconsistent acquisition frequencies, thereby improving the integrity of the data.
[0074] As a preference of the above - mentioned embodiment, the misalignment time corresponding to the staggered time is determined by the following steps:
[0075] B1: Obtain the pre - process data and the post - process data within the set time; in the specific implementation process, the sampling time interval needs to be set, and the duration of the set time is determined according to the sampling time interval to ensure a sufficient amount of data;
[0076] B2: Calculate the cross - correlation between the pre - process data and the post - process data. The cross - correlation function is:
[0077]
[0078] where a t is the value of the pre - process data at time t, k is the number of time - lag steps, b t+k is the value of the post - process data k steps after time t, and are the average values of the pre - process data and the post - process data respectively, and N is the number of data points; among them, the number of time - lag steps k refers to the step difference between the pre - process data and the post - process data. For example, a lag of 3 steps means that the post - process data lags 3 sampling points behind the pre - process data, that is, 3 sampling time intervals;
[0079] B3: Use the number of lag time steps k that makes the value of the cross - correlation function the largest to calculate the misalignment time. The misalignment time is equal to the number of lag time steps k multiplied by the sampling time interval of each time step.
[0080] In this preferred solution, through cross - correlation analysis, the time - dependence relationship between the pre - process data and the post - process data can be accurately captured, the optimal misalignment time can be determined, and the data alignment can be made more accurate. This way of confirming the misalignment time can reduce the noise and error caused by data asynchronization and improve the quality of data alignment.
[0081] During the implementation process, when the value of the cross-correlation function reaches its maximum, the optimal time lag is calculated through the number of time steps k of lag, for the following reasons:
[0082] (1) Strongest correlation:
[0083] The maximum cross-correlation value indicates that, at the number of time steps k of lag, the correlation between the two time series is the strongest, which means that the change pattern of the data from the previous process can best explain the change of the data from the subsequent process;
[0084] (2) Capturing the actual time delay:
[0085] In the water treatment process, the effect of the previous process often takes a certain amount of time to be reflected in the data of the subsequent process. The number of time steps k of lag corresponding to the maximum cross-correlation can reflect this actual time delay as much as possible;
[0086] (3) Improving prediction accuracy:
[0087] Using the time lag corresponding to the maximum cross-correlation for data alignment can reduce the noise and errors caused by data asynchrony, improve the data quality, and thus enhance the accuracy and reliability of the input data of the prediction model.
[0088] As an optimization of the above embodiment, as Figure 2 shown, the long short-term memory network time series model includes:
[0089] An input layer that accepts a data set as input; the input layer accepts the time series data set after aligning the data from the previous process and the data from the subsequent process as input.
[0090] An LSTM layer, which consists of several LSTM units. Each LSTM unit contains a memory unit and three gating mechanisms to control the flow of information; the LSTM layer consists of several LSTM units, which are used to process time series data, capture the time dependence and sequence information in the data. The memory unit therein is used to store the long-term dependence information of the time series. The three gating mechanisms include an output gate, a forget gate, and an input gate, and they jointly control the flow of information in the memory unit. The input gate controls the proportion of new information entering the memory unit, the forget gate determines which information in the memory unit needs to be forgotten, and the output gate controls which information in the memory unit is output. Through these gating mechanisms, LSTM can effectively capture and retain the important information in the time series data, reducing the problems of gradient vanishing and explosion.
[0091] A number of fully connected layers that receive the output from the LSTM layer, process it, and perform feature transformation, mapping high-dimensional features to the desired dimension; the fully connected layers are usually placed after the LSTM layer, responsible for receiving the output of the LSTM layer and performing further processing and feature transformation. Through the processing of a number of fully connected layers, the model can better extract and transform the high-dimensional features output by the LSTM layer, thereby improving the prediction performance and accuracy of the model.
[0092] The output layer that outputs the prediction result based on the feature transformation result. Nonlinear activation functions can be applied in the fully connected layer and the output layer to introduce nonlinear characteristics, such as ReLU, Leaky ReLU, tanh, etc., to enhance the nonlinear representation ability of the model and improve the fitting ability of the model.
[0093] To make the performance of the long short-term memory network time series model in the above embodiments better, the number of LSTM layers is a linear function of the number of fully connected layers, and the number of fully connected layers is less than the number of LSTM layers.
[0094] Through this kind of optimization, the model complexity and computational efficiency can be effectively balanced. In water environment monitoring, the data of the front-end processing and the back-end processing contains a large number of time-dependent features. The LSTM layer extracts these complex features through a deeper network structure to understand the trends and patterns of water quality changes. The number of fully connected layers is less than the number of LSTM layers, reducing the computational complexity and training time of the model, improving the efficiency of real-time monitoring and prediction, and ensuring that the system can respond quickly.
[0095] During the implementation process, by setting the number of LSTM layers as a linear function of the number of fully connected layers, the model can effectively perform feature transformation and dimensionality reduction while extracting time series features. Using a linear relationship makes the network structure design intuitive, reduces the complexity of hyperparameter tuning, improves work efficiency. By setting a reasonable ratio, the number of layers can be quickly determined, reducing the time for experiments and adjustments, and quickly deploying an effective monitoring and prediction model.
[0096] As an optimization of the above embodiment, the relationship between the number of LSTM layers and the number of fully connected layers is:
[0097] N LSTM = α × N FC ;
[0098] where, N LSTM is the number of LSTM layers, N FC is the number of fully connected layers; α is a positive integer and is positively correlated with the difference in the acquisition frequencies of the front-end processing data and the back-end processing data.
[0099] When there is a large difference in the acquisition frequencies of the current-stage processed data and the post-stage processed data, the time delay and dependency relationships between the data will be more complex. In such cases, increasing the number of LSTM layers helps to better model these complex dependency relationships, thereby improving the prediction accuracy. α is positively correlated with the acquisition frequency difference, enabling the number of LSTM layers to be dynamically adjusted according to the time-dependent characteristics of the data, so that the model can maintain good performance in different monitoring scenarios, enhancing the generalization ability and adaptability.
[0100] As a preference of the above embodiment, a peephole connection is added between the memory unit of the long short-term memory network time series model and the three gating mechanisms. The peephole connection is a mechanism to enhance the long short-term memory network, enabling each gating mechanism to directly access the state of the memory unit. In the standard long short-term memory network, the calculation of each gate only depends on the current input and the hidden state at the previous moment. After introducing the peephole connection, it allows the forget gate, input gate, and output gate to directly utilize the state of the memory unit at the previous moment in addition to depending on the current input and the hidden state at the previous moment when calculating, which can better capture and utilize the long-term and short-term dependency relationships in the time series data, allowing information to flow more flexibly between the memory unit and the gating mechanisms, thereby enhancing the expressive ability of the model. It should be noted that introducing the peephole connection will increase the number of model parameters, and thus the risk of overfitting will also increase. Regularization methods can be combined for control, such as Dropout regularization, elastic net regularization, and batch normalization, etc.
[0101] Furthermore, the training, evaluation, and optimization of the time series model are implemented, including:
[0102] C1: Convert the data set and divide the data set into a training set, a validation set, and a test set;
[0103] Specifically, time series data is essentially a series of data points that change over time. To apply the time series model for prediction, the time series data can be converted into the format of a supervised learning problem, that is, a data pair of input features and target variables; after converting the data, the data set can be divided into a training set, a validation set, and a test set. For example, 70% training set, 15% validation set, 15% test set; when dividing the data set, ensure the randomness of the data to avoid biases in the division process. The obtained training set, validation set, and test set should maintain the same data distribution to ensure the reliability of the evaluation results.
[0104] C2: Compile the time series model and select a loss function, an optimizer, and evaluation metrics;
[0105] Before model training, appropriate loss functions, optimizers, and evaluation metrics can be selected to compile the model; the loss function is a function that measures the gap between the predicted values and the actual values of the model, including mean squared error, mean absolute error, cross-entropy loss, etc.; the optimizer is used to update the weight parameters of the model to minimize the loss function, including gradient descent, stochastic gradient descent, Adam optimizer, etc., and different optimizers have different update rules and hyperparameters; the evaluation metrics set when compiling the model are not used to optimize the model, but they can provide useful information during the training and evaluation processes. Mean squared error, root mean squared error, mean absolute error, etc. can be used as evaluation metrics.
[0106] C3: Set the hyperparameters of the time series model, use the training set to iteratively train the time series model, and use the validation set to evaluate the time series model during the iterative training process. Adjust the time series model according to the evaluation results of the validation set.
[0107] Based on the above embodiments, the hyperparameters of the model can be set, including learning rate, batch size, number of training epochs, etc. Then, the training set can be used to iteratively train the model, and the validation set can be used for evaluation during the training process. The specific process is to divide the training data into multiple small batches, each small batch contains a certain number of samples. Perform forward propagation on the data of each small batch. The input data propagates through the network, calculate the output of each layer until the final prediction result is obtained. Use the loss function to calculate the difference between the prediction result of each small batch and the true value. Then perform backward propagation on the data of each small batch, calculate the gradient of the loss with respect to each weight. These gradients are used to update the weights. The gradients can be calculated through the chain rule, layer by layer to calculate the gradient of the loss with respect to each parameter. The weights can be updated using the calculated gradients, and the optimizer is used to update the weights of the model. Repeat the above steps for all small batches until the entire training data set is traversed. After each training epoch, evaluate the performance of the model on the validation set, and adjust the hyperparameters according to the performance of the validation set to optimize the model performance. This helps to monitor the training progress of the model and detect overfitting. To better prevent overfitting and improve the generalization ability of the model, the model can be iteratively trained, repeating the above steps until the loss function converges or reaches the preset number of training epochs.
[0108] C4: Use the test set to evaluate the performance of the time series model and calculate the evaluation metrics.
[0109] C5: Optimize the time series model according to the evaluation results of the test set, and deploy the optimized time series model to the water environment monitoring system.
[0110] After all adjustments are completed in this step, the model is retrained. The finally trained model can be used to evaluate on the test set, calculate evaluation metrics. If the performance on the test set is not satisfactory, the model may need to be further adjusted. According to the evaluation results of the test set, the model structure, hyperparameters can be adjusted, or other modifications can be made to improve the model performance. After the adjustment, the model needs to be retrained and evaluated using the updated test set. The test set data is used to evaluate the performance of the model on unseen data, which can ensure the generalization ability of the model. At the same time, the results of the test set evaluation will provide the performance of the model in actual applications.
[0111] As a preference of this embodiment, as Figure 3 shown, convert the data set and divide the data set into a training set, a validation set and a test set, including:
[0112] D1: Define a time window, create data pairs based on the time window. The data pairs include feature data and label data;
[0113] D2: Divide the data set into a training set, a validation set and a test set;
[0114] D3: Adjust the shape of the feature data to match the input of the time series model.
[0115] This step converts time series data into a format suitable for supervised learning. A time window can be defined to determine how many past time steps of data are used to predict one or more future time steps. This step determines how much past data the model uses for prediction. A suitable window length can effectively capture the temporal dependence of the data. Based on the time window, input features and corresponding labels are generated. The input features are the data within the time window, and the labels are the prediction targets after the time window, so that the time series data can be used as the training data for the supervised learning algorithm. The data set is divided into a training set, a validation set and a test set to evaluate the performance of the model and perform the training and tuning of the model. In addition, the input data of the model needs to have a specific dimension, and the shape of the input feature data can be adjusted to meet the requirements of the model input to ensure that the data can be correctly processed by the model. Since different data sets need to have a consistent format and shape to ensure that the model can be correctly trained and evaluated, but the adjustment of the data shape may cause the data in the training set and the test set to no longer represent the original distribution, thus affecting the fairness of the model evaluation. Therefore, the shape of the data can be adjusted after dividing the data.
[0116] Embodiment 2
[0117] As Figure 4 shown, a water environment operation and maintenance monitoring system based on data analysis. The system includes:
[0118] The water environment data collection module collects the pre - processing data and post - processing data for the water environment;
[0119] The timestamp alignment module aligns the timestamps of the pre - processing data and the post - processing data;
[0120] The water environment data matching module matches the pre - processing data at a previously set acquisition moment and the post - processing data at a subsequently set acquisition moment to obtain a data set of misaligned correspondences;
[0121] The long short - term memory network model module uses the long short - term memory network time - series model and inputs the data set;
[0122] The time - series model training module implements the training, evaluation, and optimization of the time - series model;
[0123] The time - series model prediction module predicts the water environment through the optimized time - series model. The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above - mentioned embodiments. What is described in the above - mentioned embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A water environment operation and maintenance monitoring method based on data analysis, characterized in that, Including: Collecting the pre - processing data and post - processing data for the water environment; Performing timestamp alignment on the pre - processing data and the post - processing data; Matching the pre - processing data at a pre - set acquisition moment and the post - processing data at a post - set acquisition moment to obtain a data set corresponding to different times, and determining the time difference corresponding to the different times includes: Obtaining the pre - processing data and the post - processing data within a set time; Calculating the cross - correlation between the pre - processing data and the post - processing data, and the cross - correlation function is: ; where a t is the value of the previous process data at time t, is the number of time lag steps, and the number of time lag steps refers to the step difference between the previous process data and the subsequent process data. b t+k is the value of the subsequent process data after lagging by steps, and are the average values of the previous process data and the subsequent process data respectively, is the number of data points; using the number of lag time steps that maximizes the value of the cross-correlation function , calculate the misaligned time; Adopting a long short - term memory network time - series model and inputting the data set; Implementing the training, evaluation and optimization of the time - series model; Predicting the water environment through the optimized time - series model.
2. The water environment operation and maintenance monitoring method based on data analysis according to claim 1, wherein, Performing timestamp alignment on the pre - processing data and the post - processing data, including: Judging whether the acquisition frequencies of the pre - processing data and the post - processing data are the same. If so, directly align the timestamps; if not, perform interpolation processing on the party with a lower acquisition frequency. The interpolation processing is used to fill in the missing values to make the data volumes of the pre - processing data and the post - processing data consistent, and then align the timestamps.
3. The method for monitoring the operation and maintenance of water environment based on data analysis according to any one of claims 1 to 2, characterized in that The long short - term memory network time - series model includes: An input layer that accepts the data set as input; An LSTM layer composed of several LSTM units, and each LSTM unit contains a memory unit and three gating mechanisms to control the flow of information; Several fully - connected layers that accept the output from the LSTM layer, and perform processing and feature transformation to map high - dimensional features to the required dimensions; An output layer that outputs a prediction result according to the feature transformation result.
4. The water environment operation and maintenance monitoring method based on data analysis according to claim 3, wherein The number of LSTM layers is a linear function of the number of fully - connected layers, and the number of fully - connected layers is less than the number of LSTM layers.
5. The method for monitoring the operation and maintenance of water environment based on data analysis according to claim 4, characterized in that, The relationship between the number of LSTM layers and the number of fully - connected layers is: N LSTM = α × N FC ; Among them, N LSTM is the number of said LSTM layers, N FC is the number of said fully connected layers; α is a positive integer and is positively correlated with the acquisition frequency difference between the previous processed data and the subsequent processed data.
6. The method for monitoring water environment operation and maintenance based on data analysis according to claim 3, wherein Adding a peephole connection between the memory unit and the three gating mechanisms in the long short - term memory network time - series model.
7. The water environment operation and maintenance monitoring method based on data analysis according to claim 1, characterized in that Implementing the training, evaluation and optimization of the time - series model, including: Transforming the data set and dividing the data set into a training set, a validation set and a test set; Compiling the time - series model and selecting a loss function, an optimizer and evaluation metrics; Setting the hyperparameters of the time - series model, using the training set to perform iterative training on the time - series model, and using the validation set to evaluate the time - series model during the iterative training process. According to the evaluation results of the validation set, adjust the time - series model; Using the test set to evaluate the performance of the time - series model and calculating the evaluation metrics; According to the evaluation results of the test set, optimizing the time - series model and deploying the optimized time - series model to a water environment monitoring system.
8. The method for operation and maintenance monitoring of water environment based on data analysis according to claim 7, characterized in that, Transforming the data set and dividing the data set into a training set, a validation set and a test set, including: Defining a time window and creating data pairs based on the time window. The data pairs include feature data and label data; Divide the data set into a training set, a validation set, and a test set; Adjust the shape of the feature data to match the input of the time series model.
9. A water environment operation and maintenance monitoring system based on data analysis, characterized in that, Adopt the water environment operation and maintenance monitoring method based on data analysis as described in claim 1, the system comprising: A water environment data collection module for collecting pre-processing data and post-processing data for the water environment; A timestamp alignment module for aligning the timestamps of the pre-processing data and the post-processing data; A water environment data matching module for matching the pre-processing data at a previously set collection time and the post-processing data at a subsequently set collection time to obtain a data set with misaligned correspondence; A long short-term memory network model module that uses a long short-term memory network time series model and inputs the data set; A time series model training module for implementing the training, evaluation, and optimization of the time series model; A time series model prediction module for predicting the water environment through the optimized time series model.
Citation Information
Patent Citations
Equipment anomaly analysis method and device based on extreme learning machine
CN117473445A