Pump station operation characteristic prediction and safety evaluation method and device

By using the Attention-LSTM-Seq2seq prediction model and a multi-dimensional safety evaluation method, the problems of low prediction accuracy and insufficient safety evaluation of pump station unit operation were solved, achieving high-precision prediction of operating parameters and comprehensive safety status assessment, thus optimizing pump station management.

CN122490184APending Publication Date: 2026-07-31POWER CHINA KUNMING ENG CORP LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
POWER CHINA KUNMING ENG CORP LTD
Filing Date
2026-03-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing pump station unit operation prediction models have low accuracy, are difficult to reflect nonlinear characteristics, and cannot achieve multi-step accurate prediction. Safety evaluation methods cannot comprehensively analyze historical operating data and future trends, and lack a holistic and multi-dimensional prediction and evaluation system.

Method used

We employ the Attention-LSTM-Seq2seq prediction model, which optimizes the LSTM-Seq2seq model by normalizing historical data, deleting outlier data points, constructing a correlation coefficient matrix, and combining the Attention mechanism to build a time-series prediction model with multi-input unit outputs. We also conduct multi-dimensional security evaluations.

Benefits of technology

It improves the accuracy of pump station unit operating parameter prediction, enhances risk management capabilities, optimizes maintenance and operation decisions, and supports scientific data-driven management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490184A_ABST
    Figure CN122490184A_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, computer equipment, and computer-readable storage medium for predicting the operational characteristics and safety evaluation of pumping stations, relating to the field of automation technology in water conservancy and hydropower engineering. The method includes: first, standardizing the format of historical monitoring data for pumping station units, dividing the data table according to the water transfer year and unit number, identifying abnormal data and deleting invalid items, completing data interpolation, constructing a coefficient matrix through correlation analysis, eliminating redundant or weakly correlated parameters, and obtaining usable data for the prediction model; then, constructing an attention-optimized LSTM-Seq2seq time-series prediction model, selecting hyperparameters and evaluation indicators, dividing the training and test sets to determine the optimal model, forming a series of prediction models for safe operation characteristics; finally, combining real-time, historical, and predicted values ​​of the indicators, obtaining a real-time comprehensive safety evaluation value from multiple dimensions. This method improves prediction accuracy and risk management capabilities, and optimizes data-driven decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automation technology for water conservancy and hydropower projects, and in particular to a method and device for predicting the operating characteristics and safety evaluation of pumping stations. Background Technology

[0002] The accuracy of the safety monitoring parameter prediction model is the core of research on pump station unit operation status prediction methods, and the selection of the model structure has a significant impact on the prediction accuracy. Currently, prediction algorithm models generally fall into two mainstream structures: prediction models based on traditional mathematical statistics methods and prediction models based on machine learning and deep learning. Prediction models based on traditional mathematical statistics methods include: moving average (MA), autoregressive moving average (ARMA), and seasonally differentiated autoregressive moving average (SARIMA). Prediction models based on machine learning and deep learning include: random forest, support vector machine (SVM), convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), and bidirectional long short-term memory network (BiLSTM). In addition, there are hybrid models such as differential autoregressive moving average-support vector machine (ARIMA-SVM) and variational mode decomposition and support vector machine joint model (IVMD-SVM). Furthermore, based on the length of the predicted future time, they can be divided into single-time-step prediction models, multi-time-step prediction models, and variable-time-step prediction models.

[0003] The safety of pumping station units is the primary goal for the stable operation of water diversion pumping stations. Unit safety evaluation methods can be divided into single-indicator evaluation and multi-indicator comprehensive evaluation. Single-indicator evaluation can initially define the safety range and threshold of an indicator based on the statistical distribution characteristics of historical data, and evaluate the current safety status of the indicator in real time according to the defined safety level intervals. The key to multi-indicator evaluation is how to select the indicator elements participating in the comprehensive safety evaluation, how to determine the relative weight values ​​of each indicator, and how to determine the safety status to which the safety evaluation value belongs. One method is based on the "maximum membership principle," obtaining the unit health status judgment vector from the objective weight matrix of unit health assessment and the membership weight matrix, and then evaluating the unit health status based on the unit health status judgment vector. Another method considers the unit failure probability and failure consequences, multiplying the failure probability score of the evaluated unit by the failure consequence evaluation value to obtain the final relative risk value of the evaluated unit, reflecting the safety level of the unit.

[0004] Existing prediction models based on traditional mathematical statistics methods have low accuracy and are difficult to reflect the nonlinear characteristics of pump station unit operation; while current prediction models based on machine learning and deep learning algorithms have relatively simple structures, cannot achieve multi-step accurate prediction, and are difficult to meet the needs of actual engineering.

[0005] Existing prediction models do not offer a universal model structure and a universal hyperparameter tuning method for all operating parameters of pump station units, and therefore cannot meet the performance optimization requirements of accurately and quickly predicting the future trends of multiple parameters simultaneously in actual engineering projects.

[0006] Existing pump station safety evaluation methods cannot simultaneously analyze the impact on pump station unit safety from both historical operating data and future trend predictions, making it difficult to form a holistic, multi-faceted, scientific, and effective prediction and evaluation system. Summary of the Invention

[0007] The main objective of this invention is to provide a method for predicting the operating characteristics and safety evaluation of pumping stations.

[0008] Another objective of this invention is to propose a device for predicting the operating characteristics and safety evaluation of pumping stations.

[0009] The third objective of this invention is to provide an electronic device.

[0010] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0011] To achieve the above objectives, a first aspect of the present invention provides a method for predicting the operating characteristics and safety evaluation of pumping stations, comprising:

[0012] S1. Obtain historical monitoring data of pump station units, standardize the field names and data formats of historical data storage, divide data tables by water transfer year and unit number, and form abnormal data points based on the summary and organization of historical monitoring data of pump station units. S2. Based on the identified abnormal data points, invalid data is deleted, data imputation is performed, and correlation analysis is used to construct the correlation coefficient matrix between parameters. Redundant or weakly correlated attribute parameters are eliminated to obtain usable data for the prediction model after data preprocessing. S3. Construct an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. Select the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model. Set the parameter index that is strongly correlated with the output parameter in the available data as the input parameter. Construct a time series prediction model with multi-input unit output and multi-time step input and multi-time step output. S4. Based on the available data, the dataset is divided into a training set and a test set. The time series prediction model is trained and tested using the test set. The test results under different hyperparameters are compared, and the model with the best evaluation result is selected as the optimal prediction model. For different safety indicators, corresponding prediction models are obtained, and finally a series of prediction models for safe operation characteristics are formed. S5. Collect the current monitoring values ​​of each indicator participating in the calculation of the real-time safety comprehensive evaluation value, combine the historical monitoring values ​​of each indicator within a specific time period and the predicted values ​​output by the safety operation characteristic series prediction model, and form the time-series monitoring data of each indicator; evaluate the indicators from the perspectives of short-term safety, long-term safety and time-series stability, and obtain the safety comprehensive evaluation value of each indicator; obtain the real-time safety comprehensive evaluation value of the pump station unit through the pump station unit comprehensive evaluation formula.

[0013] Optionally, the process includes acquiring historical monitoring data of the pumping station units, standardizing the field names and formats for historical data storage, dividing the data tables by water transfer year and unit number, and identifying outlier data points based on the summarized historical monitoring data of the pumping station units. This also includes: Collect offline data on the operation and monitoring of pumping station units during the annual water transfer years, including technical and safety monitoring indicators. All data should be standardized and organized, and the Chinese and English field names should be consistent, as well as the accuracy of the data. Write a Python script to manage database data: using a fixed time interval as the standard, identify and filter data rows in the database that are in the "started" running state, and fill in null values ​​for data rows that are in the "stopped" running state.

[0014] Optionally, based on the identified outlier data points, invalid data is deleted, data imputation is performed, correlation analysis is used to construct a correlation coefficient matrix between parameters, redundant or weakly correlated attribute parameters are eliminated, and usable data for the prediction model after data preprocessing is obtained. This also includes: The identified invalid data outliers are processed, and the chaotic representations of missing data values ​​are standardized to null values. The identified outliers are processed by using the box plot method to detect outliers in non-normally distributed monitoring data and then deleting the detected outliers. Imputation and filling of processed data: Sequence extraction is performed based on the time interval of the source data, and then the missing values ​​are processed according to the characteristics of the missing data. No processing is performed when a large amount of time series data is missing.

[0015] Optionally, an LSTM-Seq2seq prediction model based on the Attention mechanism is constructed, selecting the hyperparameters, loss function, and training model evaluation metrics of the Attention-LSTM-Seq2seq prediction model. Parameters strongly correlated with the output parameters from the available data are set as input parameters. A time-series prediction model with multi-input unit output and multi-time-step input / multi-time-step output is constructed, further including: Building an Attention-LSTM-Seq2Seq prediction model architecture: The Seq2seq model consists of two parts: an encoder and a decoder. By introducing a two-layer LSTM network as the main body of the encoder and decoder, and adding an Attention network layer, the model can selectively focus on different parts of the input sequence. After each neural network layer in the encoder and decoder, a Dropout layer is added to reduce the complex co-adaptation relationships between neurons by randomly discarding some units in the network. Customize a loss function suitable for predicting data with volatile characteristics; Adam was chosen as the model training optimizer, a learning rate decay strategy was adopted, and early determination and model checkpoint training monitoring were set. Using the parameters that are strongly correlated with the output parameters from the available data as input parameters, and setting the input time steps and output time steps, a time series prediction model with multi-input unit output and multi-time step input and multi-time step output is constructed.

[0016] Optionally, based on the available data, the dataset is divided into a training set and a test set. The time-series prediction model is trained and tested using the test set, and the test results under different hyperparameters are compared. The model with the best evaluation performance is selected as the optimal prediction model. For different safety indicators, corresponding prediction models are obtained, ultimately forming a series of prediction models for safe operation characteristics, which also include: The minimum-maximum normalization method is used to map each feature value of the available data to the [0,1] interval and perform normalization operation. The original data is sliced ​​according to the time series, with a length equal to the sum of the model's input and output steps. Training and test samples suitable for the model's input are constructed. After slicing the data, samples with missing values ​​are removed to ensure that each sample contains complete historical information for prediction. Divide the dataset after normalization and data slicing into training and testing sets; The model is trained multiple times using the training set, and the model parameters are continuously updated through optimization algorithms. Cross-validation is used to further confirm the model's performance and stability. The test set is input into the trained model to generate prediction results. The model's performance under different hyperparameter configurations is evaluated by calculating the model's evaluation index on the test set. For safety indicators such as temperature, vibration, sway, pressure pulsation, and noise, model training and hyperparameter adjustment are carried out according to the requirements of different indicators to obtain the corresponding optimal prediction model, and finally a series of prediction models for safe operation characteristics are obtained.

[0017] Optionally, the current monitoring values ​​of each indicator participating in the calculation of the real-time comprehensive safety evaluation value are collected, and combined with the historical monitoring values ​​of each indicator within a specific time period and the predicted values ​​output by the obtained series prediction model of safe operation characteristics, to form the time-series monitoring data of each indicator; the indicators are evaluated from the perspectives of short-term safety, long-term safety, and time-series stability to obtain the comprehensive safety evaluation value of each indicator; the real-time comprehensive safety evaluation value of the pump station unit is obtained through the comprehensive evaluation formula of the pump station unit, and also includes: By training the optimal prediction model that meets the requirements, the corresponding multi-step prediction values ​​are output in real time through the input of historical monitoring data; the current monitoring index value, the historical monitoring value within a specific time period, and the prediction value within a specific time period are combined into a set of time-series monitoring data that can be used for historical data analysis, real-time anomaly monitoring, and predictive maintenance. The safety of the indicators is comprehensively evaluated from three aspects: short-term safety, long-term safety, and time-series stationarity. Short-term safety is measured by the safety score obtained from the monitoring data value at the current moment; long-term safety is measured by the safety score obtained from the average value of the total monitoring data series; and time-series stationarity is measured by the safety score obtained from the volatility value of the total monitoring data series. Based on the real-time comprehensive safety evaluation value of multiple safety evaluation indicators, the real-time comprehensive safety evaluation value of the pump station unit is obtained through the comprehensive evaluation formula of the pump station unit. The evaluation criteria for safety status levels are formulated based on industry standards, past case studies, and the professional experience of experts. The safety status characteristics of each level are described in detail, along with the corresponding comprehensive safety evaluation value range for the pump station unit. Based on the obtained real-time comprehensive safety evaluation value, the real-time safety status of the unit can be determined.

[0018] To achieve the above objectives, a second aspect of the present invention provides a device for predicting the operating characteristics and safety evaluation of pumping stations, comprising: The monitoring module is used to acquire historical monitoring data of pump station units, standardize the field names and data formats of historical data storage, divide data tables by water transfer year and unit number, and form abnormal data points based on the summary and organization of historical monitoring data of pump station units. The identification module is used to delete invalid data based on the identified abnormal data points, perform data imputation, construct the correlation coefficient matrix between parameters using correlation analysis, eliminate redundant or weakly correlated attribute parameters, and obtain usable data for the prediction model after data preprocessing. The module is used to build an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. It selects the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model, sets the parameters that are strongly correlated with the output parameters in the available data as input parameters, and builds a time series prediction model with multi-input unit output and multi-time step input and multi-time step output. The prediction module is used to divide the available data into training and testing sets; the time series prediction model is trained and tested using the testing set, and the test results under different hyperparameters are compared. The model with the best evaluation result is selected as the optimal prediction model; corresponding prediction models are obtained for different safety indicators, and finally a series of prediction models for safe operation characteristics are formed. The evaluation module is used to collect the current monitoring values ​​of each indicator participating in the calculation of the real-time comprehensive safety evaluation value, combine the historical monitoring values ​​of each indicator within a specific time period with the predicted values ​​output by the series prediction model of the safety operation characteristics, and form the time-series monitoring data of each indicator; evaluate the indicators from the perspectives of short-term safety, long-term safety, and time-series stability to obtain the comprehensive safety evaluation value of each indicator; and obtain the real-time comprehensive safety evaluation value of the pump station unit through the comprehensive evaluation formula of the pump station unit.

[0019] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0020] To achieve the above objectives, a third aspect of this application provides an electronic device, including a processor and a memory; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, for implementing the pump station operation characteristic prediction and safety evaluation method as described in the first aspect embodiment.

[0021] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the pump station operation characteristic prediction and safety evaluation method as described in the first aspect embodiment.

[0022] The embodiments of the present invention have the following beneficial effects: Improve prediction accuracy: By using the LSTM-Seq2seq model combined with the attention mechanism, key operating parameters of pump station units, such as temperature and vibration, can be predicted more accurately, thereby anticipating potential faults and anomalies in a timely manner and reducing unexpected downtime and maintenance costs.

[0023] Enhancing risk management capabilities: The invention's multi-indicator, multi-dimensional safety evaluation method integrates historical and real-time data to provide a comprehensive safety status assessment. This approach not only increases the understanding of short-term and long-term risks but also improves the overall control over the operational safety of pumping stations.

[0024] Optimize maintenance and operational decisions: By accurately predicting future conditions, operations teams can better plan preventative maintenance, optimize resource allocation, and improve operational efficiency.

[0025] Enhancing data-driven decision-making: This invention supports data-driven decision-making, making pump station management more scientific and precise, and reducing experience-based uncertainty and potential operational errors. Attached Figure Description

[0026] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a method for predicting the operating characteristics and evaluating the safety of a pumping station, provided in an embodiment of the present invention; Figure 2 A flowchart of data preprocessing, prediction model construction and training is provided in an embodiment of the present invention; Figure 3 A flowchart of a safety rating method for pump station units provided in an embodiment of the present invention; Figure 4 This is a structural diagram of a pump station operation characteristic prediction and safety evaluation device provided in an embodiment of the present invention; Detailed Implementation It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] The following describes, with reference to the accompanying drawings, a method and apparatus for predicting the operating characteristics and evaluating the safety of pumping stations according to embodiments of the present invention.

[0029] Example 1 This invention provides a method for predicting the operating characteristics and evaluating the safety of pumping stations. Figure 1This is a flowchart illustrating a method for predicting the operating characteristics and evaluating the safety of a pumping station according to an embodiment of the present invention. Figure 2 A flowchart of data preprocessing, prediction model construction and training is provided in an embodiment of the present invention; Figure 3 A flowchart illustrating a safety rating method for pump station units provided in an embodiment of the present invention. Figure 1-3 As shown, the method includes the following steps: S1. Obtain historical monitoring data of pump station units, standardize the field names and data formats of historical data storage, divide the data tables by water transfer year and unit number, and form abnormal data points based on the summarized and organized historical monitoring data of pump station units.

[0030] In order to achieve standardized collection and processing of pump station unit operation data and lay a solid data foundation for subsequent unit analysis and judgment, this application carries out data collection, sorting, storage and anomaly identification and processing work and builds an original database.

[0031] In this embodiment, historical monitoring data of the pumping station units are first comprehensively acquired, and various offline data of the pumping station units' operation monitoring during the annual water transfer years are collected simultaneously. The formats of the offline data include mainstream data formats such as xlsx and csv files exported from the sensors, without format limitations, to ensure the integrity and comprehensiveness of data collection.

[0032] In this embodiment of the application, the collected pump station unit operation monitoring data mainly includes two categories of monitoring indicators. The first category is the technical monitoring indicators of pump station operation. This type of technical monitoring indicator includes multiple core data during the operation of the pump station, specifically including basic information such as the monitoring and recording time and operating status, operating condition indicators such as the water level of the intake pool, the water level of the outlet pool, the head, the instantaneous flow rate, and the blade angle, as well as operating stability indicators such as the cumulative flow rate, active power, reactive power, excitation current, excitation voltage, active energy, reactive energy, and operating efficiency, comprehensively reflecting the operating conditions and technical operating status of the pump station unit.

[0033] In this application embodiment, the second category is pump station operating characteristic indicators. These operating characteristic indicators are core indicators that can directly reflect the safe operating status of the unit. Specifically, they include the monitored and recorded stator temperature, thrust bearing temperature, upper and lower guide bearing temperature, upper and lower cylinder temperature, as well as the X and Y direction vibration of the upper and lower frames, the X and Y direction vibration of the impeller housing, and the X and Y direction swing of the main shaft. They also include axial displacement, blade pulsation pressure, pump noise, and other related data. Through these operating characteristic indicators, the safe operating status of the pump station unit can be accurately grasped, providing data support for the assessment of safe operation of the unit.

[0034] In this embodiment, a systematic and standardized process is carried out on all collected pump station unit data. The primary task is to standardize the data fields, unifying the Chinese and English field names for all data fields. The rules for these field names and corresponding data units are designed to balance the standardization of water conservancy industry terminology with ease of understanding for pump station operation and management personnel, thus forming a standardized field naming convention. For example, monitoring time (record_time, yyyy / MM / dd) The data includes HH:mm:ss, running state, in-water level (m), out-water level (m), head (m), active power (kW), excitation current (A), excitation voltage (kV), instantaneous flow (m³ / s), cumulative flow (ten thousand m³), ​​and unit efficiency (%). This application also standardizes the accuracy of all data and performs format normalization on all data.

[0035] In this embodiment of the application, all data tables that have been standardized and organized are scientifically classified and categorized according to the water diversion year and unit number. All standardized data tables are uniformly stored in an SQL database. Based on the summarized and organized historical monitoring data of the pump station units, a clear, uniform, and easily accessible original database of the pump station units is built, providing a complete foundational database for subsequent data processing, data analysis, and data retrieval.

[0036] In this embodiment of the application, after the original database is built, abnormal data identification and processing is carried out on all the data in the database to accurately identify various abnormal data points such as missing values ​​and outliers in the database, so as to ensure the validity and accuracy of the data in the database.

[0037] In this embodiment, a dedicated Python script is written to automate the management and anomaly processing of the completed SQL database. The Python script uses a fixed time interval detection time as the data filtering standard to specifically identify and filter data rows in the database that are in the "started" running state. Each of these "started" data rows is checked to accurately remove data rows containing missing values, outliers, or other abnormal data points. At the same time, for all data rows in the database that are in the "stopped" running state, the corresponding monitoring data is filled with null according to data specifications, thus completing the standardized anomaly processing of all data.

[0038] In this embodiment of the application, a standardized raw database is built by collecting, organizing, storing and identifying anomalies in the historical monitoring data and offline data of the pump station units in a standardized manner, which lays the foundation for obtaining usable data for the prediction model after data preprocessing.

[0039] S2. Based on the identified abnormal data points, invalid data is deleted, data imputation is performed, and correlation analysis is used to construct the correlation coefficient matrix between parameters. Redundant or weakly correlated attribute parameters are eliminated to obtain usable data for the prediction model after data preprocessing.

[0040] In order to improve the data quality of pump station units and eliminate invalid interference parameters, and to provide high-quality usable data for prediction models, this application carries out abnormal data processing, data interpolation and attribute parameter screening.

[0041] In this embodiment, the identified invalid data anomalies are first standardized. Various chaotic representations of missing values ​​in the database are standardized, including the number 0, formula error identifier #DIV / 0!, numerical error identifier #VALUE!, no matching value identifier #N / A, and null value identifier Null. This application standardizes all the chaotic representations of the above-mentioned missing values ​​into null values, thereby standardizing the missing value identifier and laying a unified data foundation for subsequent data imputation and filling.

[0042] In this application embodiment, targeted deletion processing is carried out on the identified outliers. This application uses the box plot method to accurately detect outliers in the non-normally distributed monitoring data of the pump station units, and directly deletes the outliers identified by the box plot method, effectively eliminating extreme abnormal data points in the data. At the same time, considering that the monitoring data under unstable operating conditions of the unit has no effective reference value, in order to reduce the interference of such data on the subsequent training of prediction models, this application uniformly deletes the unstable operating condition data collected within a certain period after the unit starts up and before the unit shuts down, further ensuring the effectiveness and rationality of the data source.

[0043] In this embodiment, after invalid data deletion and outlier removal, a refined imputation and filling operation is performed on the data to maximize data utilization and ensure data integrity. This application first extracts time-series data based on predetermined time intervals from the source data, then classifies the missing data according to different characteristics. For different types of missing data, corresponding imputation processing methods are matched. The imputation processing methods selected in this application include linear interpolation, pchip interpolation, cubic spline interpolation, and the minimum distance method based on the sliding window concept. Furthermore, for cases with a large number of missing time-series data, this application adopts an unprocessed approach, using differentiated imputation strategies to optimize data filling. The final result is a standardized database containing several long-term consecutive gaps. When performing dataset slicing before building the network model, this application directly deletes slices containing missing values. These slices are not used as training or testing sets for the model, thus avoiding the impact of missing data on model training from the source.

[0044] In this application embodiment, differentiated processing methods are adopted according to the characteristics of different missing data. Specifically, for data rows containing only a few missing parameter values ​​and all other parameters being normal, this application uses the valid data of the corresponding parameters in the adjacent time steps above and below the data row to perform linear interpolation on the missing parameters. For cases where the entire data row is missing but the data in the time steps above and below the missing data row are normal, or where two consecutive time step data rows are missing, this application uses linear interpolation to fill in the data for each parameter. For cases where the entire data row is missing and up to six consecutive time step data rows are missing, this application uses third-order Hermitian interpolation (i.e., pchip interpolation) or third-order spline interpolation (i.e., cubic_spline interpolation) to interpolate each parameter. Furthermore, this application requires that when constructing the corresponding interpolation function… To improve the accuracy of interpolation, the interpolation points are set to known data points contained in the number of missing data rows before and after the missing data rows. For cases where the entire data row is missing and there are no more than 12 consecutive missing data rows, this application considers that traditional interpolation methods will produce large interpolation errors in such scenarios. Therefore, for such data rows, the minimum distance method based on the sliding window idea is used for interpolation. The effective monitoring data with the smallest distance to the missing data is accurately found from the historical operation data of the pump station unit, and the data is completely filled into the missing data row. For a large amount of data that is missing more than 12 consecutive time steps, this application determines that all kinds of interpolation methods will produce large errors in this case. Such interpolated data will have an adverse effect on the training effect of the subsequent prediction model. Therefore, no interpolation processing is performed on such a large amount of consecutively missing data.

[0045] In this embodiment, the minimum distance method based on the sliding window concept employs specific steps. First, the sample data to be matched is normalized to effectively ignore the interference from different dimensions and numerical ranges of the n missing monitoring parameters. Then, the number of consecutive missing data steps, t, is determined. Before and after the time node corresponding to the missing data, t rows of valid data are selected, forming a total of 3n time steps, with the data at time steps t~2t in the middle being in a missing state. Next, based on the core idea of ​​the sliding window, the remaining valid data is processed in 3t time steps. The sample is sliced, and all data with missing values ​​in the slices are deleted to build a complete effective sample set. Then, in the 3t×n dimensional feature space, the Euclidean distance between the interpolation target sequence and the corresponding parameters and corresponding time nodes of each sample in the sample set is calculated accurately. The optimal interpolation target sample that is closest to the interpolation target sequence is selected. Finally, the selected optimal interpolation target sample is denormalized, and the effective data rows in the middle t~2t of the sample are filled into the missing positions of the interpolation target sequence to complete the interpolation operation for long time-series missing data.

[0046] In this embodiment of the application, after completing the above-mentioned data preprocessing operations, the application further uses correlation analysis to construct a correlation coefficient matrix between parameters of the processed complete data. Through quantitative analysis of the correlation coefficient matrix, redundant attribute parameters and weakly correlated attribute parameters in the data are accurately identified and eliminated, and invalid feature interference is removed. Finally, high-quality and effective data that has completed all data preprocessing steps and can be directly used for predictive model training is obtained.

[0047] In this embodiment, by standardizing abnormal data, performing differential interpolation and filling, and filtering redundant parameters based on the correlation coefficient matrix, the quality of pump station unit data is improved, resulting in high-quality and effective data that can be directly used for predictive model training, thus laying the foundation for subsequent predictive model construction.

[0048] S3. Construct an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. Select the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model. Set the parameters that are strongly correlated with the output parameters from the available data as input parameters. Construct a time series prediction model with multi-input unit output and multi-time step input and multi-time step output.

[0049] To achieve high-precision time-series prediction of pump station unit operation data and solve the gradient vanishing problem of traditional LSTM models, this application constructs an LSTM-Seq2seq prediction model optimized based on the Attention mechanism and completes the relevant parameter configuration.

[0050] In this embodiment, the core architecture of the Attention-LSTM-Seq2Seq prediction model is constructed. The Seq2seq model used in this application consists of two main parts: an encoder and a decoder. The encoder part is responsible for receiving and processing the input time-series data sequence, while the decoder part accurately generates the corresponding target prediction sequence based on the feature information output by the encoder. In this embodiment, a two-layer LSTM network is used as the core network of the encoder and decoder. By utilizing the long short-term memory characteristic of the LSTM network, the temporal dependencies of the pump station unit's time-series operation data can be effectively processed, and the temporal correlation features in the data can be fully explored.

[0051] In this embodiment, to address the vanishing gradient defect and the inability of the model to selectively focus on key input information that are common in traditional LSTM models, this application adds an Attention network layer during the decoding stage. This Attention mechanism calculates the relative importance weights between the current state of the decoder and the state of the encoder at each time step, allowing the prediction model to selectively focus on information content at different positions in the input sequence and concentrate on the input data part that is most relevant to the current output prediction. This effectively solves the vanishing gradient defect of LSTM models and allows the prediction model to focus on more critical information, significantly improving the prediction accuracy of the model. This is also the core optimization point of this application for the traditional LSTM-Seq2seq model.

[0052] In this embodiment, to further improve the generalization ability of the model and avoid overfitting, a Dropout layer is added to the back end of each neural network layer in the encoder and decoder. The Dropout layer can effectively reduce the complex co-adaptation relationship between neurons by randomly discarding some neuron units in the network during model training, weaken the model's excessive dependence on local features, and thus significantly enhance the generalization ability and robustness of the Attention-LSTM-Seq2Seq prediction model constructed in this application.

[0053] In this application embodiment, considering that the operation monitoring data of the pump station unit itself has certain random fluctuation characteristics, conventional loss functions cannot accurately reflect the fitting effect of such data fluctuation. Based on this, this application defines a combined loss function suitable for predicting data with fluctuation characteristics. This application combines the basic mean square error (MSE) with the fluctuation error that can accurately reflect the fitting of the data fluctuation characteristics to form the core loss function for model training. This customized loss function has dual characteristics: it can effectively measure the magnitude of the numerical deviation between the model's predicted value and the actual value, and it can accurately capture the consistency between the fluctuation of the model's predicted data and the fluctuation of the actual operating data, ensuring the model's adaptability and fitting effect on fluctuating time series data.

[0054] The formula for calculating the mean square error is:

[0055] in, For the first One predicted value, For the first A true value, This represents the number of samples in this batch.

[0056] The formula for calculating volatility error is:

[0057] in, , This represents the standard deviation function, which is applied to both the difference and the true value.

[0058] Based on the two error values ​​mentioned above, the final loss function formula is determined as follows:

[0059] in, The weighting coefficient for volatility error, ranging from The weighting ratio can be flexibly adjusted according to the fluctuation characteristics of the actual pump station operation data.

[0060] In this application embodiment, the hyperparameters and model evaluation metrics for the Attention-LSTM-Seq2Seq prediction model are selected. The appropriate selection of hyperparameters has a direct and crucial impact on the final training effect of the model. The model hyperparameters selected in this application include, but are not limited to, the learning rate, batch size, number of LSTM layers, number of hidden units, random dropout rate of the Dropout layer, and the PATIENCE value of the early stopping strategy. To comprehensively and accurately evaluate the training effect and prediction accuracy of the model, this application selects multi-dimensional model evaluation metrics, specifically including mean squared error (MSE), mean absolute error (MAE), mean absolute percentage error (MAPE), and... Coefficient of determination.

[0061] In this embodiment of the application, the formula for calculating the mean absolute error is:

[0062] The formula for calculating the mean absolute percentage error is:

[0063] The formula for calculating the coefficient of determination is:

[0064] in, It is the mean of the true values. The total number of samples is used as an evaluation metric to comprehensively assess the predictive performance of the model.

[0065] In this embodiment, the configuration related to model training is selected. Adam is selected as the core optimizer for model training. To ensure model training efficiency and convergence, a learning rate decay strategy is adopted to dynamically adjust the learning rate of model training. In addition, a dual training monitoring mechanism of early stopping and model checkpoints is set up during model training. Early stopping avoids overfitting caused by overtraining, while model checkpoints save the optimal model parameters in real time during training to maximize the training quality of the model.

[0066] In this embodiment, based on the available data for the prediction model obtained after the previous data preprocessing, the input and output parameters of the model are configured and the final time series prediction model is constructed. Specifically, this application selects various parameter indicators that are strongly correlated with the output parameters from the available data as the input parameters of the model. Combined with the time series characteristics of the pump station unit operation data, the number of input time steps and output time steps of the model are reasonably set. Finally, an Attention-LSTM-Seq2seq time series prediction model with multi-input unit output and multi-time step input and multi-time step output is constructed. Through the above complete model construction and configuration process, this application enables the constructed prediction model to fully adapt to the time series operation data characteristics of the pump station unit and achieve high-precision time series data prediction.

[0067] In this embodiment, by building an LSTM-Seq2seq architecture with an Attention mechanism and Dropout layer, a custom fluctuation adaptation loss function, and configuring multi-dimensional evaluation indicators, a high-precision prediction model adapted to the time series data of pump station units is constructed, laying the foundation for subsequent model testing.

[0068] S4. Based on the available data, the dataset is divided into a training set and a test set. The time series prediction model is trained and tested using the test set. The test results under different hyperparameters are compared, and the model with the best evaluation result is selected as the optimal prediction model. For different safety indicators, corresponding prediction models are obtained, and finally, a series of prediction models for safe operation characteristics are formed.

[0069] In order to complete model training and validation and optimal model selection, and to meet the prediction requirements of various safety indicators of pump station units, this application processes the available data, conducts model training and evaluation, and constructs a series of prediction models for safe operation characteristics.

[0070] In this embodiment, the available data for the prediction model is first normalized. This application uses the minimum-maximum normalization method as the core approach for data normalization, mapping each feature value in the preprocessed available data to the [0,1] interval, thus completing the normalization of the entire dataset. This eliminates differences in the dimensions and numerical ranges of different feature parameters, improving the training efficiency and convergence effect of the model. The minimum-maximum normalization formula used in this application is x_norm=(x-min(x)) / (max(x)-min(x)), where min(x) is the minimum value of the corresponding feature value, max(x) is the maximum value of the corresponding feature value, x is the original feature value, and x_norm is the feature value after normalization.

[0071] In this embodiment, after data normalization, the normalized dataset is sliced ​​to obtain a sample set suitable for model training and testing. The normalized raw data is sliced ​​sequentially according to the time series, with the slice length set to the sum of the input and output steps of the prediction model. This slicing method accurately constructs training and testing samples suitable for the Attention-LSTM-Seq2seq prediction model built in this application. After slicing, all sliced ​​samples are verified, and samples with missing values ​​are uniformly deleted to ensure that each sample contains complete historical time series information, effectively supporting the model's accurate prediction.

[0072] In this embodiment, the complete dataset after normalization and data slicing is divided into a training set and a test set. These two datasets serve different model application functions. The training set is primarily used for iterative model training, continuously adjusting various model parameters to minimize the model's loss function. Typically, 80% of the sample data is selected as the training set. The test set is mainly used for performance evaluation of the trained model. By calculating the model's prediction error on test data not used in training, the model's generalization ability and actual prediction effect are verified. Typically, 20% of the sample data is selected as the test set. This division ratio balances the model's training effectiveness and validation validity.

[0073] In this embodiment, based on multiple preset hyperparameter configurations, the constructed Attention-LSTM-Seq2seq prediction model is trained in multiple rounds using a partitioned training set. During training, the model's parameters are continuously updated iteratively using a selected optimization algorithm. Simultaneously, cross-validation is employed to further confirm the model's actual performance and operational stability, effectively avoiding overfitting and underfitting issues. After training, a partitioned test set is input into the trained prediction model, which generates corresponding prediction results. Subsequently, the model's prediction performance under different hyperparameter configurations is quantitatively evaluated by calculating various preset evaluation metrics on the test set. The model's test performance under different hyperparameter combinations is compared one by one, and finally, the hyperparameter combination with the best overall performance is selected to determine the optimal prediction model.

[0074] In this embodiment, considering that the preset evaluation index is the mean error between the model's multi-step prediction results and the actual values, relying solely on quantitative indicators cannot fully reflect the model's prediction accuracy. Therefore, this application also supplements the evaluation of the model from the perspective of observing multiple instances. Specifically, it draws trend comparison charts between the model's predicted values ​​and the actual values ​​to intuitively observe the overall prediction fitting effect of the model. At the same time, it selects multiple typical test samples for targeted analysis to thoroughly evaluate the model's prediction performance on various typical samples. This allows for a comprehensive understanding of the model's prediction capability and operational stability under different operating scenarios, forming a dual evaluation system that combines quantitative evaluation of evaluation indicators with visual observation of instances.

[0075] In this application embodiment, specialized model training and optimization are carried out for the core safety indicators of the pump station unit, specifically including temperature indicators, vibration indicators, sway indicators, pressure pulsation indicators, and noise indicators. Based on the data characteristics of each type of safety indicator and the actual prediction requirements, this application conducts independent model training and hyperparameter tuning for each type of indicator, matching the corresponding optimal prediction model for each type of safety indicator, and finally integrating them to form a series of prediction models for the safe operation characteristics of the pump station unit that are suitable for full-dimensional safety monitoring.

[0076] In this embodiment of the application, by normalizing, slicing and dividing the available data, and combining multiple rounds of model training and multi-dimensional evaluation to select the optimal hyperparameters, a series of predictive models for the safe operation characteristics of pump station units that are adapted to multiple safety indicators are obtained, laying the foundation for subsequent comprehensive evaluation.

[0077] S5. Collect the current monitoring values ​​of each indicator participating in the calculation of the real-time safety comprehensive evaluation value, combine the historical monitoring values ​​of each indicator within a specific time period and the predicted values ​​output by the safety operation characteristic series prediction model, and form the time-series monitoring data of each indicator; evaluate the indicators from the perspectives of short-term safety, long-term safety and time-series stability, and obtain the safety comprehensive evaluation value of each indicator; obtain the real-time safety comprehensive evaluation value of the pump station unit through the pump station unit comprehensive evaluation formula.

[0078] In order to achieve accurate quantitative assessment of the operating status of pump station units and determine the real-time safety level, this application constructs a multi-dimensional safety evaluation system and calculates the real-time comprehensive safety evaluation value of the unit.

[0079] In this application embodiment, the core indicator factors participating in the calculation of the real-time safety comprehensive evaluation value of the pump station unit are first determined. These indicator factors are key parameters that can directly reflect the safe operating status of the pump station unit, and may specifically include safety indicators such as temperature, vibration, swing, pressure pulsation, and noise. This application carries out subsequent safety evaluation work based on the selected evaluation indicator factors.

[0080] In this embodiment, based on the optimal prediction model that meets the requirements and has been trained in the early stage, real-time prediction of indicator data is carried out. Specifically, by inputting real-time collected historical monitoring data into the optimal prediction model, the model outputs multi-step prediction values ​​of the corresponding indicators. On this basis, this application integrates the monitoring indicator value of each indicator at the current moment, the historical monitoring value within a specific time period, and the prediction value within a specific time period to form a complete set of time-series monitoring data. This time-series monitoring data can simultaneously support multiple application needs such as historical data analysis, real-time anomaly monitoring, and predictive maintenance.

[0081] In this embodiment of the application, the calculation rules for the time steps of time series monitoring data are clarified, and the specific calculation formula is as follows:

[0082] Where T is the length of the monitoring data window, which is the total number of time series steps; The selected historical monitoring period window length based on requirements; The formula allows for flexible adjustment of the time coverage of time-series monitoring data to suit different evaluation scenarios, based on the selected future forecast window length according to requirements.

[0083] In this embodiment, the integrated time-series monitoring data is comprehensively evaluated from three dimensions: short-term security, long-term security, and time-series stability, thereby achieving a comprehensive assessment of the indicator's operational status. Specifically, short-term security is measured using a security score derived from the current monitoring data value, directly reflecting the indicator's current operational security status; long-term security is measured using a security score derived from the average value of the entire monitoring data sequence, reflecting the overall operational security level of the indicator over a period of time; and time-series stability is measured using a security score derived from the volatility value of the entire monitoring data sequence, characterizing the stability of the indicator's operational status.

[0084] In this embodiment of the application, a method for calculating the volatility of the monitoring data sequence is specified, and the specific calculation formula is as follows:

[0085] Where V is the volatility of the monitoring data window, which is the sample standard deviation; The monitored values ​​within this window; This formula represents the average value within a window, and it can accurately quantify the degree of fluctuation in indicator data, providing data support for the evaluation of time series stationarity.

[0086] In this application embodiment, the application clarifies the method for determining the safety scores of short-term safety, long-term safety, and time-series stability. Specifically, a safety score correspondence table can be pre-formulated based on relevant standards for the safe operation of pump station units and the experience of industry experts. In the actual evaluation process, the corresponding safety score can be directly read from the safety score correspondence table based on the monitoring data value, sequence average value, and sequence volatility of the indicators, ensuring the objectivity and professionalism of the safety scores.

[0087] In this application embodiment, a method for calculating the comprehensive safety evaluation value of a certain indicator is specified, and the specific calculation formula is as follows:

[0088] in, This represents the real-time comprehensive security evaluation value for the i-th indicator. The safety score obtained for the short-term safety of indicator i; The safety score obtained for indicator i over a long period of time; The safety score is obtained by considering the time-series stationarity of index i. The weights are respectively for short-term security, long-term security, and time-series stationarity.

[0089] In this application embodiment, the method for determining the above-mentioned weights is clearly defined. The weighting method can be obtained by fuzzy hierarchical analysis (FAHP). The specific steps are as follows: First, the three evaluation elements of short-term security, long-term security, and time series stationarity need to be compared in pairs, and a scale of 1 to 9 is used to represent the importance of one element relative to another. Then, the fuzzy judgment matrix is ​​processed by fuzzy mathematics theory and synthesis operation methods, such as fuzzy multiplication and weighted average, to finally obtain the relative weight of each evaluation element, so as to ensure the scientific and reasonable weight allocation.

[0090] In this embodiment, based on the real-time comprehensive safety evaluation value of multiple safety evaluation indicators, the real-time comprehensive safety evaluation value of the pump station unit is calculated using the comprehensive evaluation formula for the pump station unit. The specific calculation formula is as follows:

[0091] Where K is the real-time comprehensive safety evaluation value of the pump station unit; n is the number of indicators participating in the safety rating of the pump station; This represents the real-time comprehensive security evaluation value for the i-th indicator. The weight is calculated for the i-th indicator. It can also be obtained by fuzzy hierarchical analysis (FAHP), thereby ensuring the accuracy of the calculation of the unit's comprehensive evaluation value.

[0092] In this embodiment, an evaluation standard for the safety status level of pump station units is formulated based on relevant safety operation standards in the pump station industry, research results of past unit operation safety cases, and the professional experience of industry experts. This evaluation standard describes in detail the safety status characteristics of the unit corresponding to each safety level, and clarifies the range of comprehensive safety evaluation values ​​for pump station units corresponding to each safety level. In practical applications, the calculated real-time comprehensive safety evaluation value of the unit can be matched with this evaluation standard to quickly obtain the real-time safety status of the unit, providing an intuitive and accurate decision-making basis for the safe operation management of pump station units.

[0093] In this embodiment of the application, a time-series monitoring sequence that integrates historical, real-time and predictive data is constructed to conduct safety evaluation of indicators from multiple dimensions and calculate a comprehensive evaluation value by combining weights, so as to achieve accurate determination of the real-time safety status of pump station units.

[0094] Example 2 This invention provides a device for predicting the operating characteristics and evaluating the safety of pumping stations, such as... Figure 4 As shown, the device includes: Monitoring module 100 is used to acquire historical monitoring data of pump station units, standardize the field names and field data formats of historical data storage, divide data tables by water transfer year and unit number, and form abnormal data points based on the summary and organization of historical monitoring data of pump station units. The identification module 200 is used to delete invalid data based on the identified abnormal data points, perform data interpolation, construct a correlation coefficient matrix between parameters using correlation analysis, eliminate redundant or weakly correlated attribute parameters, and obtain usable data for the prediction model after data preprocessing. Module 300 is used to build an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. It selects the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model, sets the parameter index that is strongly correlated with the output parameter in the available data as the input parameter, and builds a time series prediction model with multi-input unit output and multi-time step input and multi-time step output. The prediction module 400 is used to divide the available data into training and testing sets; after training the time series prediction model using the testing set, it is tested and the test results under different hyperparameters are compared. The model with the best evaluation result is selected as the optimal prediction model; for different safety indicators, corresponding prediction models are obtained, and finally a series of prediction models for safe operation characteristics are formed. The evaluation module 500 is used to collect the current monitoring values ​​of each indicator participating in the calculation of the real-time comprehensive safety evaluation value, and combine the historical monitoring values ​​of each indicator within a specific time period with the predicted values ​​output by the safety operation characteristic series prediction model to form the time-series monitoring data of each indicator; evaluate the indicators from the perspectives of short-term safety, long-term safety, and time-series stability to obtain the comprehensive safety evaluation value of each indicator; and obtain the real-time comprehensive safety evaluation value of the pump station unit through the comprehensive evaluation formula of the pump station unit.

[0095] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0096] Example 3 To implement the methods of the above embodiments, the present invention also provides an electronic device, which includes a memory and a processor; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the various steps of the methods described above.

[0097] Example 4 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.

[0098] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0099] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0100] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0101] The monitoring module is used to acquire historical monitoring data of pump station units, standardize the field names and data formats of historical data storage, divide data tables by water transfer year and unit number, and form abnormal data points based on the summary and organization of historical monitoring data of pump station units. The identification module is used for; The module is used to build an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. It selects the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model, sets the parameters that are strongly correlated with the output parameters in the available data as input parameters, and builds a time series prediction model with multi-input unit output and multi-time step input and multi-time step output. The prediction module is used for; The evaluation module is used for...

Claims

1. A method for predicting the operating characteristics and evaluating the safety of a pumping station, characterized in that, include: S1. Obtain historical monitoring data of pump station units, standardize the field names and data formats of historical data storage, divide data tables by water transfer year and unit number, and form abnormal data points based on the summary and organization of historical monitoring data of pump station units. S2. Based on the identified abnormal data points, invalid data is deleted, data imputation is performed, and correlation analysis is used to construct the correlation coefficient matrix between parameters. Redundant or weakly correlated attribute parameters are eliminated to obtain usable data for the prediction model after data preprocessing. S3. Construct an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. Select the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model. Set the parameter index that is strongly correlated with the output parameter in the available data as the input parameter. Construct a time series prediction model with multi-input unit output and multi-time step input and multi-time step output. S4. Based on the available data, the dataset is divided into a training set and a test set. The time series prediction model is trained and tested using the test set. The test results under different hyperparameters are compared, and the model with the best evaluation result is selected as the optimal prediction model. For different safety indicators, corresponding prediction models are obtained, and finally a series of prediction models for safe operation characteristics are formed. S5. Collect the current monitoring values ​​of each indicator participating in the calculation of the real-time safety comprehensive evaluation value, combine the historical monitoring values ​​of each indicator within a specific time period and the predicted values ​​output by the safety operation characteristic series prediction model, and form the time-series monitoring data of each indicator; evaluate the indicators from the perspectives of short-term safety, long-term safety and time-series stability, and obtain the safety comprehensive evaluation value of each indicator; obtain the real-time safety comprehensive evaluation value of the pump station unit through the pump station unit comprehensive evaluation formula.

2. The method according to claim 1, characterized in that, The process of acquiring historical monitoring data of pumping station units, standardizing the field names and formats for historical data storage, dividing data tables by water transfer year and unit number, and forming abnormal data points based on the summarized and organized historical monitoring data of pumping station units also includes: Collect offline data on the operation and monitoring of pumping station units during the annual water transfer years, including technical and safety monitoring indicators. All data should be standardized and organized, and the Chinese and English field names should be consistent, as well as the accuracy of the data. Write a Python script to manage database data: using a fixed time interval as the standard, identify and filter data rows in the database that are in the "started" running state, and fill in null values ​​for data rows that are in the "stopped" running state.

3. The method according to claim 1, characterized in that, The process of deleting invalid data based on identified abnormal data points, performing data imputation, constructing a correlation coefficient matrix between parameters using correlation analysis, eliminating redundant or weakly correlated attribute parameters, and obtaining usable data for the preprocessed prediction model also includes: The identified invalid data outliers are processed, and the chaotic representations of missing data values ​​are standardized to null values. The identified outliers are processed by using the box plot method to detect outliers in non-normally distributed monitoring data and then deleting the detected outliers. Imputation and filling of processed data: Sequence extraction is performed based on the time interval of the source data, and then the missing values ​​are processed according to the characteristics of the missing data. No processing is performed when a large amount of time series data is missing.

4. The method according to claim 1, characterized in that, The construction of an optimized LSTM-Seq2seq prediction model based on the Attention mechanism, including selecting hyperparameters, loss functions, and training model evaluation metrics of the Attention-LSTM-Seq2seq prediction model, setting parameters strongly correlated with output parameters from available data as input parameters, and constructing a time-series prediction model with multi-input unit output and multi-time-step input and multi-time-step output, further includes: Building an Attention-LSTM-Seq2Seq prediction model architecture: The Seq2seq model consists of two parts: an encoder and a decoder. By introducing a two-layer LSTM network as the main body of the encoder and decoder, and adding an Attention network layer, the model can selectively focus on different parts of the input sequence. After each neural network layer in the encoder and decoder, a Dropout layer is added to reduce the complex co-adaptation relationships between neurons by randomly discarding some units in the network. Customize a loss function suitable for predicting data with volatile characteristics; Adam was chosen as the model training optimizer, a learning rate decay strategy was adopted, and early determination and model checkpoint training monitoring were set. Using the parameters that are strongly correlated with the output parameters from the available data as input parameters, and setting the input time steps and output time steps, a time series prediction model with multi-input unit output and multi-time step input and multi-time step output is constructed.

5. The method according to claim 1, characterized in that, Based on the available data, the dataset is divided into a training set and a test set. The time series prediction model is trained and tested using the test set. The test results under different hyperparameters are compared, and the model with the best evaluation result is selected as the optimal prediction model. For different safety indicators, corresponding prediction models are obtained, ultimately forming a series of prediction models for safe operation characteristics, which also include: The minimum-maximum normalization method is used to map each feature value of the available data to the [0,1] interval and perform normalization operation. The original data is sliced ​​according to the time series, with a length equal to the sum of the model's input and output steps. Training and test samples suitable for the model's input are constructed. After slicing the data, samples with missing values ​​are removed to ensure that each sample contains complete historical information for prediction. Divide the dataset after normalization and data slicing into training and testing sets; The model is trained multiple times using the training set, and the model parameters are continuously updated through optimization algorithms. Cross-validation is used to further confirm the model's performance and stability. The test set is input into the trained model to generate prediction results. The model's performance under different hyperparameter configurations is evaluated by calculating the model's evaluation index on the test set. For safety indicators such as temperature, vibration, sway, pressure pulsation, and noise, model training and hyperparameter adjustment are carried out according to the requirements of different indicators to obtain the corresponding optimal prediction model, and finally a series of prediction models for safe operation characteristics are obtained.

6. The method according to claim 1, characterized in that, The current monitoring values ​​of each indicator involved in the calculation of the real-time comprehensive safety evaluation value are collected, combined with the historical monitoring values ​​of each indicator within a specific time period and the predicted values ​​output by the safety operation characteristic series prediction model, to form the time-series monitoring data of each indicator; the indicators are evaluated from the perspectives of short-term safety, long-term safety and time-series stability to obtain the comprehensive safety evaluation value of each indicator; The real-time comprehensive safety evaluation value of the pump station unit is obtained through the comprehensive evaluation formula, which also includes: By training the optimal prediction model that meets the requirements, the corresponding multi-step prediction values ​​are output in real time through the input of historical monitoring data; the current monitoring index value, the historical monitoring value within a specific time period, and the prediction value within a specific time period are combined into a set of time-series monitoring data that can be used for historical data analysis, real-time anomaly monitoring, and predictive maintenance. The safety of the indicators is comprehensively evaluated from three aspects: short-term safety, long-term safety, and time-series stationarity. Short-term safety is measured by the safety score obtained from the monitoring data value at the current moment; long-term safety is measured by the safety score obtained from the average value of the total monitoring data series; and time-series stationarity is measured by the safety score obtained from the volatility value of the total monitoring data series. Based on the real-time comprehensive safety evaluation value of multiple safety evaluation indicators, the real-time comprehensive safety evaluation value of the pump station unit is obtained through the comprehensive evaluation formula of the pump station unit. The evaluation criteria for safety status levels are formulated based on industry standards, past case studies, and the professional experience of experts. The safety status characteristics of each level are described in detail, along with the corresponding comprehensive safety evaluation value range for the pump station unit. Based on the obtained real-time comprehensive safety evaluation value, the real-time safety status of the unit can be determined.

7. A device for predicting the operating characteristics and evaluating the safety of a pumping station, characterized in that, include: The monitoring module is used to acquire historical monitoring data of pump station units, standardize the field names and data formats of historical data storage, divide data tables by water transfer year and unit number, and form abnormal data points based on the summary and organization of historical monitoring data of pump station units. The identification module is used to delete invalid data based on the identified abnormal data points, perform data imputation, construct the correlation coefficient matrix between parameters using correlation analysis, eliminate redundant or weakly correlated attribute parameters, and obtain usable data for the prediction model after data preprocessing. The module is used to build an optimized LSTM-Seq2seq prediction model based on the Attention mechanism. It selects the hyperparameters, loss function, and training model evaluation index of the Attention-LSTM-Seq2seq prediction model, sets the parameters that are strongly correlated with the output parameters in the available data as input parameters, and builds a time series prediction model with multi-input unit output and multi-time step input and multi-time step output. The prediction module is used to divide the available data into training and testing sets; the time series prediction model is trained and tested using the testing set, and the test results under different hyperparameters are compared. The model with the best evaluation result is selected as the optimal prediction model; corresponding prediction models are obtained for different safety indicators, and finally a series of prediction models for safe operation characteristics are formed. The evaluation module is used to collect the current monitoring values ​​of each indicator participating in the calculation of the real-time comprehensive safety evaluation value, combine the historical monitoring values ​​of each indicator within a specific time period with the predicted values ​​output by the series prediction model of the safety operation characteristics, and form the time-series monitoring data of each indicator; evaluate the indicators from the perspectives of short-term safety, long-term safety, and time-series stability to obtain the comprehensive safety evaluation value of each indicator; and obtain the real-time comprehensive safety evaluation value of the pump station unit through the comprehensive evaluation formula of the pump station unit.

8. An electronic device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.