Water outlet pressure prediction method and device

By combining the prediction methods of LSTM and random forests, the problem that the prior art is difficult to accurately predict the water effluent pressure of the water plant is solved, and the accurate prediction of the water effluent pressure value is achieved, taking into account the timing and nonlinear characteristics.

CN120067898AInactive Publication Date: 2025-05-30SHENZHEN DANGKANG TECH CO LTD

Patent Information

Application Number
CN202510551247.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict water outflow pressure in water plants, traditional methods cannot process nonlinear data, while artificial intelligence methods are prone to overfitting and cannot capture time features.

Method used

Using a prediction method combining LSTM and random forest, by obtaining the data set to be predicted within the preset time range, first use the trained LSTM prediction model to predict, obtain the first LSTM prediction result, and then fuse it with the data set to be predicted to form the first fusion data set, and finally input it into the trained random forest prediction model to obtain the water pressure prediction value.

Benefits of technology

Accurate prediction of the water outlet pressure value is achieved, taking into account timing and nonlinear characteristics, avoiding overfitting, and reducing the complexity of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067898A_ABST
    Figure CN120067898A_ABST
Patent Text Reader

Abstract

The invention provides a water outlet pressure prediction method and device.The water outlet pressure prediction method comprises the steps that a to-be-predicted data set is obtained, and the to-be-predicted data set comprises data within a preset time range and corresponding to important feature columns obtained in advance; obtaining a first LSTM prediction result according to a pre-trained LSTM prediction model and the to-be-predicted data set; fusing the first LSTM prediction result and the prediction data set to obtain a first fused data set; and according to a pre-trained random forest prediction model and the fusion data set, obtaining a water outlet pressure prediction value. By adopting the water outlet pressure prediction method provided by the invention, the water outlet pressure of the water plant can be accurately predicted, and the water plant is helped to dynamically balance the water outlet amount and the water outlet pressure in real time so as to ensure the stable operation of the water plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water plant management, and in particular, to a method and device for predicting outlet water pressure. Background Art

[0002] The main production task of a water plant is to provide high-quality water for various industries. Its stable operation requires real-time dynamic balancing of the water output and the outlet water pressure. Therefore, in related technologies, the outlet water pressure is predicted to facilitate the dynamic balancing of the water output and the outlet water pressure.

[0003] Currently, traditional methods or artificial intelligence methods are mainly used to predict the outlet water pressure. Among them, the traditional methods are mainly based on time series and cannot accurately predict non-linear complex data, but one of the main characteristics of the outlet water pressure is non-linearity; the artificial intelligence methods can better fit non-linear data, but are prone to overfitting and may also fail to capture time characteristics, but another main characteristic of the outlet water pressure is time series; thus, both the traditional methods and the artificial intelligence methods are difficult to accurately predict the outlet water pressure. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art or related technologies.

[0005] In view of this, one or more embodiments of this specification provide a method and device for predicting outlet water pressure, which can accurately predict the outlet water pressure value to help maintain the stable operation of the water plant.

[0006] According to a first aspect of one or more embodiments of this specification, a method for predicting outlet water pressure is proposed. The method includes: Obtain a dataset to be predicted, where the dataset to be predicted includes data corresponding to pre-acquired important feature columns within a preset time range; According to a pre-trained LSTM prediction model and the dataset to be predicted, obtain a first LSTM prediction result; Fuse the first LSTM prediction result and the dataset to be predicted to obtain a first fused dataset; According to a pre-trained random forest prediction model and the first fused dataset, obtain the predicted value of the outlet water pressure.

[0007] In some alternative embodiments, the method for obtaining the pre-acquired important feature columns includes: Obtain the historical data of the operation of the water plant; Preprocess the historical data to obtain an initial dataset. The initial dataset is divided into several feature subsets based on time values. The data corresponding to the same time value in the historical data belong to the same feature subset. Each feature subset includes a time value and an outlet water pressure value; Filter out a second data set from the initial data set based on a predetermined time interval unit. The second data set is sorted in chronological order and includes subsets X 1 , X 2 ,... X n , where n is a natural number; Form a third data set including subsets Y 1 , Y 2 ,... Y n-1 based on the second data set. The subset Y i in the third data set includes the subset X i in the second data set, and the water outlet pressure values in the subset X i+1 . Among them, the water outlet pressure value in the subset X i+1 is used as the water outlet pressure prediction value in Y i , 1 ≤ i < n, and i is a natural number; Construct a feature screening model, and use the third data set to train the feature screening model to obtain the importance of each feature in the third data set, and obtain a feature table by sorting in descending order of importance; Select features with importance greater than or equal to the preset threshold from the feature table based on the preset threshold to form an important feature column.

[0008] In some alternative embodiments, the training method of the pre-trained random forest prediction model includes: Select data corresponding to the important feature column and the water outlet pressure prediction value from the third data set to form a fourth data set; Use a pre-trained LSTM prediction model to predict the fourth data set to obtain a second LSTM prediction result; Fuse the second LSTM prediction result and the fourth data set to obtain a second fusion data set; Construct a random forest prediction model, and use the second fusion data set to train and optimize the random forest prediction model to obtain a trained random forest prediction model, where the water outlet pressure prediction value in the second fusion data set is the target value, and other data are input values.

[0009] In some alternative embodiments, the construction of the random forest prediction model includes: Construct an initial random forest prediction model, set the number of sub-decision trees to 120, and limit the maximum depth of each sub-decision tree to 15; Specify that the number of features considered at each split is the square root of the total number of features in the important feature column; Set the random seed random_state to 42.

[0010] In some alternative embodiments, the training method of the pre-trained LSTM prediction model includes: Select data corresponding to the important feature columns and the predicted water outlet pressure values from the third data set to form a fifth data set; Construct an LSTM prediction model, and use the fifth data set to train and evaluate the LSTM prediction model to obtain a trained LSTM prediction model, where the predicted water outlet pressure value in the fifth data set is the target value, and other data are input values.

[0011] In some alternative embodiments, the construction of the LSTM prediction model includes: Construct an input layer and an output layer, where the number of nodes in the input layer is the same as the total number of features of the important feature columns, and the number of nodes in the output layer is 1; Construct a dropout layer, where the dropout rate of the dropout layer is 0.2; Construct a fully connected layer; Set the random seed random_state to 42.

[0012] According to the second aspect of one or more embodiments of this specification, a water outlet pressure prediction device is proposed, and the device includes: A data acquisition module for acquiring a data set to be predicted, where the data set to be predicted includes data corresponding to the pre-acquired important feature columns within a preset time range; A first prediction module for obtaining a first LSTM prediction result according to the pre-trained LSTM prediction model and the data set to be predicted; A data fusion module for fusing the first LSTM prediction result and the data set to be predicted to obtain a first fusion data set; A second prediction module for obtaining the predicted water outlet pressure value according to the pre-trained random forest prediction model and the first fusion data set.

[0013] According to the third aspect of one or more embodiments of this specification, an electronic device is proposed, including a memory, a processor, and computer instructions stored on the memory, and the processor executes the computer instructions to implement the method as described in the first aspect.

[0014] According to the fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is proposed, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in the first aspect are implemented.

[0015] According to a fifth aspect of one or more embodiments of the present specification, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the method described in the first aspect.

[0016] As can be seen from the above technical solutions, in one or more embodiments of the present specification, data within a preset time range and corresponding to a pre-acquired important feature column is obtained as data to be predicted. Then, the trained LSTM prediction model is first used to predict the data to be predicted, and a first LSTM prediction result is obtained; then, the first LSTM prediction result is concatenated and fused with the data to be predicted to obtain a first fusion data set; finally, the first fusion data set is input into the trained random forest prediction model to obtain the final predicted value of the water outlet pressure. In the technical solution provided in the present application, after the data to be predicted is obtained, the LSTM prediction model is first used to obtain the first LSTM prediction result, and the first LSTM prediction result is fused with the data to be predicted to obtain the first fusion data set, and then the first fusion data set is input into the random forest prediction model to obtain the predicted value of the water outlet pressure. By combining the time series modeling ability of LSTM and the regression ability of random forest, the time series and non-linearity of the water outlet pressure value can be taken into account, and accurate prediction of the water outlet pressure value is achieved; moreover, the important feature column is determined in advance from several features of the water plant, and then the data corresponding to the important feature column is used as the data set to be predicted, which can avoid excessive feature quantities from increasing the complexity of the prediction model and reducing the prediction accuracy.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 is a flowchart of a method for predicting the water outlet pressure provided in Embodiment 1; Figure 2 is a flowchart of a method for obtaining a pre-acquired important feature column provided in Embodiment 1; Figure 3 is a comparison diagram of the actual water outlet pressure and the predicted values of three prediction methods provided in Embodiment 1; Figure 4 is a schematic structural diagram of an electronic device provided in Embodiment 2; Figure 5 is a block diagram of a device for predicting the water outlet pressure provided in Embodiment 2. Reference Numerals: 402, processor; 404, internal bus; 406, network interface; 408, memory; 410, non-volatile memory; 1, data acquisition module; 2, first prediction module; 3, data fusion module; 4, second prediction module. Detailed Implementation Manner

[0018] Here, exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The specific manners described in the following exemplary embodiments do not represent all the solutions consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims. It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0019] In the related technologies of outlet water pressure prediction, both traditional methods and artificial intelligence methods have defects and deficiencies: For traditional methods, it is impossible to accurately predict non-linear complex data. For example, ARIMA (Autoregressive Integrated Moving Average Model) is a statistical model for time series analysis and prediction. Such methods fully consider the time series and periodic fluctuations of the outlet water pressure, but have weak capabilities in data regression and capturing the relevance of various features; For artificial intelligence methods, although they can better fit non-linear data, they cannot fully consider the time series of the water outlet pressure and are also prone to overfitting. For example, LightGBM (Light Gradient Boosting Machine) is an efficient machine learning algorithm based on gradient boosting decision trees (GBDT, Gradient Boosting Decision Tree), which has the advantages of fast training speed, high efficiency, and high prediction accuracy. However, it is prone to overfitting, sensitive to noise data, and unable to capture time features. Another example is SVM (Support Vector Machine), which is a generalized linear classifier for binary classification of data in a supervised learning manner. It processes non-linear relationships through kernel tricks and is suitable for high-dimensional data, but it cannot distinguish noise and cannot capture time features. Another example is LSTM (Long Short-Term Memory), which is a special recurrent neural network (RNN, Recurrent Neural Network). Although it takes into account both the time series and non-linearity of the data, has a short training time and high prediction accuracy, it is also prone to overfitting, sensitive to noise, and has low interpretability.

[0020] Therefore, providing a prediction method that can take into account both the time series and non-linearity of the water outlet pressure is of great significance for helping the water plant operate stably.

[0021] Next, the following embodiments are used to further illustrate one or more embodiments of this specification: Embodiment 1: Figure 1 It is a flowchart of a water outlet pressure prediction method provided by an exemplary embodiment of this application. As Figure 1 shown, the method may include the following steps: Step 101, obtain a dataset to be predicted, where the dataset to be predicted includes data within a preset time range and corresponding to important feature columns obtained in advance.

[0022] In this embodiment, the preset time range is set to four predetermined time interval units. The data to be predicted is: the feature value at the current time point, the feature value at the time point one predetermined time interval unit before the current time point, the feature value at the time point two predetermined time interval units before the current time point, and the feature value at the time point three predetermined time interval units before the current time point. Four feature subsets are formed based on the time values, and all the feature subsets constitute the data set to be predicted, where each of the above feature values corresponds to all the features in the important feature column. For example, assuming that the current time point is exactly twelve o'clock and the predetermined time interval unit is ten minutes, the data to be predicted is: the feature values corresponding to all the features in the important feature column at 11:30, 11:40, 11:50, and 12:00. The data at each time point forms a feature subset, and the four feature subsets constitute the data set to be predicted.

[0023] Specifically, the data during the operation of the water plant will be uploaded to the digital water plant database. In this embodiment, the required data to be predicted can be extracted from this digital water plant database.

[0024] The outlet water pressure is affected by many factors such as time. However, among the predicted feature sets, it is not the case that the more feature factors there are, the higher the prediction accuracy. Too many feature quantities will increase the complexity of the prediction model and reduce the prediction accuracy. To avoid too many feature quantities, in this embodiment, the important feature column is obtained in advance, and all the features in the important feature column are used as the feature set for prediction, so as to reduce the complexity of the prediction model and improve the prediction accuracy.

[0025] In one embodiment, as Figure 2 shown, the method for obtaining the pre-obtained important feature column can be but is not limited to the following methods and steps: Step 201: Obtain the historical data of the water plant operation.

[0026] When performing this step 201, all the data in the digital water plant database can be obtained as the historical data.

[0027] Step 202: Preprocess the historical data to obtain an initial data set. The initial data set is divided into several feature subsets based on the time value. The data corresponding to the same time value in the historical data belongs to the same feature subset, and each feature subset includes the time value and the outlet water pressure value.

[0028] When performing this step 202, it can be but is not limited to the following methods and steps: Perform a pivoting operation on the historical data obtained in step 201; Merge all the data corresponding to the same historical time into one row record; Extract the hour and minute of each said historical time as a time value, and incorporate the time value into the corresponding row record to obtain a time series data set.

[0029] In this step, each row record is a feature subset, and all row records with time values constitute a time series data set, and the time series data set is an initial data set divided based on time values.

[0030] Step 203, based on a predetermined time interval unit, screen out a second data set from the initial data set. The second data set is sorted in chronological order and includes subsets X 1 , X 2 ,... X n , where n is a natural number.

[0031] When performing this step 203, the following methods and steps can be adopted but are not limited to: Take the first historical data as the starting data, and screen out the feature subsets at intervals of 10 minutes in the initial data set obtained in step 202; Obtain a second data set {X 1 , X 2 ,... X n} sorted in chronological order.

[0032] It should be noted that in this step, the predetermined time interval unit is 10 minutes. In other embodiments, it can be changed according to actual needs, such as being set to 5 minutes, 20 minutes, 30 minutes, etc. The specific value of n can be jointly determined by the number of historical data obtained and the predetermined time interval unit, or it can be a preset maximum value, such as 1000.

[0033] Step 204, form a third data set including subsets Y 1 , Y 2 ,... Y n-1 based on the second data set. The subset Y i in the third data set includes the subset X i in the second data set, and the outlet water pressure value in the subset X i+1 . Among them, the outlet water pressure value in the subset X i+1 is used as the predicted outlet water pressure value in Y i , where 1 ≤ i < n and i is a natural number.

[0034] When performing this step 204, the following methods and steps can be adopted but are not limited to: Extract the outlet water pressure value of each subset X i+1 in the second data set obtained in step 203; Add the effluent pressure value to the previous subset X i as the predicted effluent pressure value to form a new subset Y i ; Obtain the third data set {Y 1 、Y 2 、...Y n-1}.

[0035] In this step, since there is no corresponding data for the next 10 minutes in the subset X of the second data set n , the third data set has only n - 1 subsets.

[0036] Step 205: Construct a feature screening model and use the third data set to train the feature screening model to obtain the importance of each feature in the third data set, and obtain a feature table by sorting the importance in descending order.

[0037] As a robust and powerful intelligent classification algorithm, the Random Forest (RF) algorithm has the ability to measure variable importance and can analyze complex and interacting features, and is widely used in the selection of high-dimensional data features.

[0038] When performing this step 205, in order to avoid the adverse impact of artificially subjective feature selection on the prediction accuracy of the effluent pressure, a feature screening model is constructed according to the Random Forest (RF) algorithm, and then the third data set is used to train the feature screening model. After training, the importance of each feature in each third data set can be obtained according to the feature_importances_ function of the feature screening model, and then a feature table is obtained by sorting the importance in descending order. Specifically, the number of trees in the feature screening model is 100, the random seed random_state is 42, and the feature table can be a csv file; a schematic example of the feature table is shown in Table 1 below. The first column is the feature name, the second column is the importance score, and the third column is the number of non-empty rows of the feature.

[0039] Table 1:

[0040] Step 206: Based on a preset threshold, select features with importance greater than or equal to the preset threshold from the feature table to form an important feature column.

[0041] Specifically, when performing this step 206, the following methods and steps can be adopted but are not limited to: Select features with importance greater than or equal to 0.01 (preset threshold) from the feature table; Determine whether the number of selected features is not less than 10 and not greater than 15; If so, use the selected features as the important feature column; If it is less than 10, reduce the preset threshold, and then select features from the feature table based on the reduced threshold; If it is greater than 15, increase the preset threshold, and then select features from the feature table based on the increased threshold.

[0042] Among them, increasing the preset threshold can be adding 0.005 to the current preset threshold or adding 0.01 to the current preset threshold; reducing the preset threshold can be subtracting 0.001 from the current preset threshold or subtracting 0.005 from the current preset threshold. In other embodiments, other rules for increasing and decreasing can also be used to adjust the preset threshold.

[0043] Step 103, according to the pre-trained LSTM prediction model and the to-be-predicted data set, obtain the first LSTM prediction result.

[0044] The advantages of the long short-term memory (LSTM) neural network algorithm are mainly reflected in its ability to process long sequence data and good learning ability. In one or more embodiments of this specification, an LSTM prediction model is constructed based on the long short-term memory (LSTM) neural network algorithm, and the trained LSTM prediction model is used to predict the obtained to-be-predicted data, which can fully consider the timing and periodic fluctuations of the water outlet pressure.

[0045] In this embodiment, a pre-trained LSTM prediction model is established in advance, that is, the pre-trained LSTM prediction model. The to-be-predicted data set is input into the pre-trained LSTM prediction model. After receiving the input data, the LSTM prediction model will make a prediction based on this data and output the first LSTM prediction result.

[0046] Specifically, the LSTM prediction model itself is a model with time characteristics. The first LSTM prediction result output by the LSTM prediction model learns the dynamic characteristics of the data in terms of timing. Therefore, the first LSTM prediction result can be used as a new time value to supplement the to-be-predicted data set.

[0047] In a feasible embodiment, the training method of the pre-trained LSTM prediction model can adopt, but is not limited to, the following methods and steps: Select the data corresponding to the important feature column and the water outlet pressure prediction value from the third data set to form the fifth data set; Build an LSTM prediction model, and use the fifth dataset to train and evaluate the LSTM prediction model to obtain a trained LSTM prediction model. Among them, the predicted value of the outlet pressure in the fifth dataset is the target value, and other data are input values.

[0048] Specifically, in this embodiment, building an LSTM prediction model includes: building an input layer and an output layer. Among them, the number of nodes in the input layer is the same as the total number of features in the important feature columns, and the number of nodes in the output layer is 1, which is used to output the obtained LSTM prediction result; building a dropout layer, where the dropout rate of the dropout layer is 0.2, and the purpose is to randomly ignore some neurons to prevent overfitting; building a fully connected layer for feature transformation and non-linear combination; setting the random seed random_state to 42 for easy model tuning. Extract the feature values corresponding to all features in the important feature columns determined in the previous steps and the predicted value of the outlet pressure from the third dataset to form the fifth dataset; then take 25% of the fifth dataset as the fifth test dataset, and the other 75% as the fifth training dataset; further use the fifth training dataset to train the built LSTM prediction model, learn the data patterns and features, capture the rules and relationships in the data, and obtain the trained LSTM prediction model; use the fifth test dataset to evaluate the trained LSTM prediction model, which is used to evaluate the generalization ability of the trained LSTM prediction model on unseen data, can effectively prevent overfitting, and optimize the LSTM prediction model according to the evaluation situation to finally obtain the trained LSTM prediction model.

[0049] It should be noted that in other examples, the fifth dataset can also be formed first and then the LSTM prediction model is built; this specification does not limit the execution order of the two; the data ratio for training the LSTM prediction model can also be adjusted according to actual needs; the ratio of the fifth training dataset to the fifth test dataset can also be 7:3, 8:2, or other ratios.

[0050] Preferably, in this example, before training the LSTM prediction model, that is, during initialization, a parameter tuning strategy such as grid search or random search can also be used to determine the best parameter combination and use this combination as the initial parameters of the LSTM prediction model.

[0051] Preferably, in other embodiments, before training the LSTM prediction model, the fifth dataset can also be preprocessed, such as normalization.

[0052] Step 105: Fuse the first LSTM prediction result and the dataset to be predicted to obtain the first fused dataset.

[0053] When performing this step 105, the LSTM prediction result obtained in step 103 is concatenated with other data in the dataset to be predicted to obtain a first fusion dataset.

[0054] Step 107: Obtain the predicted water outlet pressure value according to the pre-trained random forest prediction model and the first fusion dataset.

[0055] In one embodiment, the training method of the pre-trained random forest prediction model can adopt, but is not limited to, the following methods and steps: Select the data corresponding to the important feature columns and the predicted water outlet pressure value from the third dataset to form a fourth dataset; Use the pre-trained LSTM prediction model to predict the fourth dataset to obtain a second LSTM prediction result; Fuse the second LSTM prediction result and the fourth dataset to obtain a second fusion dataset; Construct a random forest prediction model, and use the second fusion dataset to train and optimize the random forest prediction model to obtain a trained random forest prediction model, where the predicted water outlet pressure value in the second fusion dataset is the target value and other data are input values.

[0056] Specifically, in this embodiment, constructing a random forest prediction model includes: constructing an initial random forest prediction model, setting the number of sub-decision trees to 120, and limiting the maximum depth of each sub-decision tree to 15; specifying that the number of features considered at each split is the square root of the total number of features in the important feature columns; setting the random seed random_state to 42; Take 25% of the second fusion dataset as the second fusion test dataset, and the other 75% as the second fusion dataset; When training the random forest prediction model, use a random sampling method with replacement to extract samples from the second fusion test dataset to generate multiple sub-datasets; independently train sub-decision trees on each sub-dataset, and according to the set feature selection rule (that is, the number of features at each split is the square root of the total number of features in the important feature columns), use feature information and GINI coefficient value to select the optimal splitting feature.

[0057] When optimizing the random forest prediction model, use the out-of-bag error rate for verification and analysis, continuously adjust and optimize the model, and finally obtain a trained random forest prediction model.

[0058] It should be noted that in other instances, the fourth data set can also be formed first, and then the random forest prediction model can be constructed; this specification does not limit the execution order of the two; the data ratio for training the random forest prediction model can also be adjusted according to actual needs; the ratio of the fourth training data set to the fourth test data set can also be 7:3, 8:2, or other ratios.

[0059] Preferably, in other embodiments, before performing step 107 and using the random forest prediction model for prediction, the first fusion data set can also be preprocessed, such as normalization processing.

[0060] As Figure 3 shown, in this specification, the single LSTM model (yellow line), random forest algorithm (red line), LSTM-RF hybrid model algorithm (the method provided in this specification, black line) and the actual water outlet pressure (blue line) are also compared under the same data samples, and the results are as shown in Table 2 below. Compared with the traditional LSTM algorithm and random forest algorithm, the root mean square error of the prediction results of the LSTM-RF hybrid model is reduced by more than 10%.

[0061] Table 2:

[0062] As can be seen from the above technical solutions, in one or more embodiments of this specification, data within a preset time range and corresponding to the pre-acquired important feature columns are obtained as the data to be predicted. Then, the trained LSTM prediction model is first used to predict the data to be predicted to obtain the first LSTM prediction result; then, the first LSTM prediction result is spliced and fused with the data to be predicted to obtain the first fusion data set; finally, the first fusion data set is input into the trained random forest prediction model to obtain the final predicted value of the water outlet pressure. In the technical solution provided in this application, after obtaining the data to be predicted, the LSTM prediction model is first used to obtain the first LSTM prediction result, and the first LSTM prediction result is fused with the data to be predicted to obtain the first fusion data set. Then, the first fusion data set is input into the random forest prediction model to obtain the predicted value of the water outlet pressure. Combining the time series modeling ability of LSTM and the regression ability of random forest can take into account the time series and non-linearity of the water outlet pressure value, and achieve accurate prediction of the water outlet pressure value; moreover, by pre-determining the important feature columns from several features of the water plant and then using the data corresponding to the important feature columns as the data set to be predicted, it is possible to avoid excessive feature quantities from increasing the complexity of the prediction model and reducing the prediction accuracy.

[0063] Embodiment 2: Corresponding to the foregoing embodiment of a method for predicting water outlet pressure, this application also provides an embodiment of a device for predicting water outlet pressure. An embodiment of the water outlet pressure prediction device in this specification can be applied to an electronic device. Figure 4 It is a schematic structural diagram of an electronic device in an exemplary embodiment. Please refer to Figure 4 , at the hardware level, the electronic device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, there may also be other hardware required for other services. The processor 402 reads the corresponding computer instructions from the non-volatile memory 410 into the memory and then runs, and the water outlet pressure prediction device at the logical level. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 402 reads the corresponding computer instructions from the non-volatile memory 410 into the memory 408 and then runs. Of course, in addition to the software implementation method, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logic device. Please refer to Figure 5 , Figure 5 It is a block diagram of a water outlet pressure prediction device in an exemplary embodiment. As Figure 5 shown, the device may include: A data acquisition module 1, configured to acquire a dataset to be predicted, where the dataset to be predicted includes data corresponding to important feature columns within a preset time range and obtained in advance; A first prediction module 2, configured to obtain a first LSTM prediction result according to a pre-trained LSTM prediction model and the dataset to be predicted; A data fusion module 3, configured to fuse the first LSTM prediction result and the prediction dataset to obtain a first fusion dataset; A second prediction module 4, configured to obtain a water outlet pressure prediction value according to a pre-trained random forest prediction model and the first fusion dataset.

[0064] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method for details, and will not be elaborated here.

[0065] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of at least one embodiment of the present disclosure. A person of ordinary skill in the art can understand and implement it without creative work.

[0066] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices. In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory. The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium. The computer-readable medium includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves. It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element. The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0067] The terms used in one or more embodiments of this specification are for the purpose of describing particular embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the" and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "upon" or "in response to determining". The above are only the preferred embodiments of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.

Claims

1. A method for predicting water outlet pressure, characterized in that: The method comprises: Acquire a data set to be predicted, where the data set to be predicted includes data within a preset time range and corresponding to the important feature columns acquired in advance; Obtaining a first LSTM prediction result according to a pre-trained LSTM prediction model and the data set to be predicted; Fusing the first LSTM prediction result and the data set to be predicted to obtain a first fused data set; According to the pre-trained random forest prediction model and the first fusion data set, a predicted value of the water outlet pressure is obtained.

2. The method for predicting water outlet pressure according to claim 1, characterized in that: The method for obtaining the pre-acquired important feature columns includes: Obtain historical data on water plant operations; Preprocessing the historical data to obtain an initial data set, wherein the initial data set is divided into a plurality of subsets based on the time value, wherein data corresponding to the same time value in the historical data belongs to the same subset, and each subset includes the time value and the outlet water pressure value; Based on a predetermined time interval unit, a second data set is selected from the initial data set, and the second data set is sorted in chronological order, including subsets X1, X2, ...X n , where n is a natural number; Based on the second data set, a subset Y1, Y2, ... Y n-1 A third data set, a subset Y of the third data set i Include the subset X in the second dataset i , and the subset X i+1 The outlet pressure value in the subset X i+1 The outlet water pressure value in is taken as Y i The predicted value of the outlet water pressure in , 1≤i<n, i is a natural number; Constructing a feature screening model, and using the third data set to train the feature screening model, obtaining the importance of each feature in the third data set, and arranging the features in descending order of importance to obtain a feature table; Based on a preset threshold, features whose importance is greater than or equal to the preset threshold are selected from the feature table to form an important feature column.

3. The method for predicting water outlet pressure according to claim 2, characterized in that: The training method of the pre-trained random forest prediction model includes: Selecting data corresponding to the important feature column and the predicted value of the outlet water pressure from the third data set to form a fourth data set; Using a pre-trained LSTM prediction model to predict the fourth data set to obtain a second LSTM prediction result; Fusing the second LSTM prediction result and the fourth data set to obtain a second fused data set; A random forest prediction model is constructed, and the random forest prediction model is trained and optimized using the second fused data set to obtain a trained random forest prediction model, wherein the outlet water pressure prediction value in the second fused data set is the target value, and other data are input values.

4. The method for predicting outlet water pressure according to claim 3, characterized in that: The construction of the random forest prediction model includes: Construct an initial random forest prediction model, set the number of child decision trees to 120, and limit the maximum depth of each child decision tree to 15; The number of features considered at each split is specified to be the square root of the total number of features in the important feature column; Set the random seed random_state to 42.

5. The method for predicting outlet water pressure according to any one of claims 2 to 4, characterized in that: The training method of the pre-trained LSTM prediction model includes: Selecting data corresponding to the important feature column and the predicted value of the outlet water pressure from the third data set to form a fifth data set; Construct an LSTM prediction model, and use the fifth data set to train and evaluate the LSTM prediction model to obtain a trained LSTM prediction model, wherein the outlet water pressure prediction value in the fifth data set is the target value, and other data are input values.

6. The method for predicting outlet water pressure according to claim 5, characterized in that: The construction of the LSTM prediction model includes: Constructing an input layer and an output layer, wherein the number of nodes in the input layer is the same as the total number of features in the important feature column, and the number of nodes in the output layer is 1; Constructing a discard layer, wherein the discard rate of the discard layer is 0.2; Build a fully connected layer; Set the random seed random_state to 42.

7. A water outlet pressure prediction device, characterized in that: The device comprises: A data acquisition module, used to acquire a data set to be predicted, wherein the data set to be predicted includes data within a preset time range and corresponding to the important feature columns acquired in advance; A first prediction module, used for obtaining a first LSTM prediction result according to a pre-trained LSTM prediction model and the data set to be predicted; A data fusion module, used for fusing the first LSTM prediction result and the prediction data set to obtain a first fused data set; The second prediction module is used to obtain a predicted value of the outlet water pressure according to a pre-trained random forest prediction model and the first fusion data set.

8. An electronic device comprising a memory, a processor and computer instructions stored in the memory, characterized in that: The processor executes the computer instructions to implement the steps of the method according to any one of claims 1-6.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Credit risk prediction method and device, electronic equipment and storage medium

    CN112750029A

  • Dam deformation monitoring data anomaly detection method and system based on LSTM-RF

    CN117609889A

  • Urban real-time water supply flow prediction method and system and water supply scheduling method and system

    CN119168306A

Cited By

  • Outlet water pressure prediction method based on random forest and long short-term memory network

    CN121093177A