A method, system, device and medium for daily runoff prediction in a reservoir basin

Through the fusion of isolated forest detection outliers, DD-ACMD decomposition and TCN-IFELM models, combined with IPSA optimization, the shortcomings of the existing runoff prediction model in nonlinear and non-stationary data are solved, and a higher accuracy of daily runoff prediction is achieved.

CN119917948BActive Publication Date: 2025-07-11POWERCHINA HUADONG ENG CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510416459.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing runoff prediction model is difficult to capture dynamic changes in hydrological processes when processing nonlinear and non-stationary runoff data, resulting in insufficient prediction accuracy and performance is more unstable in extreme weather conditions.

Method used

Outlier detection is used in isolated forests, combined with data-driven adaptive chirped modal decomposition (DD-ACMD) to decompose runoff data, build a fusion model of time convolution network (TCN) and inverse-free limit learning machine (IFELM), and optimize model parameters through improved PID search algorithm (IPSA) to improve prediction performance.

Benefits of technology

It improves the accuracy and robustness of daily runoff prediction, can effectively capture the nonlinear relationship of runoff data, and enhances the adaptability and prediction ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917948B_ABST
    Figure CN119917948B_ABST
Patent Text Reader

Abstract

The present invention provides a daily runoff prediction method, system, device and medium for a reservoir basin. The method includes the following steps: S1. Obtain the daily runoff data within a certain basin, use the Isolation Forest to detect outliers in the daily runoff data, and perform data cleaning; S2. Use data-driven adaptive chirp mode decomposition to decompose the cleaned daily runoff data to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set; S3. Construct a fusion model of a time convolutional network and an inverse-free extreme learning machine and train it; S4. Introduce the half-uniform initialization and variable spiral update strategy into the PID search algorithm to obtain the IPSA algorithm; use the IPSA algorithm to optimize the TCN-IFELM model; S5. Use the test set to obtain the prediction results of each component and add them up to obtain the predicted value of the daily runoff at a future time. The present invention can effectively capture the time series characteristics and nonlinear relationships of runoff data, thereby improving the prediction accuracy of daily runoff.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hydrological data prediction, and particularly relates to a daily runoff prediction method, system, device and medium for a reservoir basin. Background Art

[0002] Runoff refers to the water flow that, after precipitation, except for evaporation, infiltration, and plant interception, flows into rivers through the surface and underground. Runoff prediction is a key task in hydrology and water resources management, aiming to predict the runoff in the future for a certain period based on previous hydrological and meteorological elements. However, due to the influence of factors such as human activities and climate change on the runoff process, it is a highly complex non-linear process. Especially in the context of frequent extreme weather, the runoff sequence shows greater volatility.

[0003] Traditional runoff prediction models include physical cause analysis methods and mathematical statistics methods. Physical cause analysis methods require a large amount of measured data and fine parameter adjustment, and it is difficult to adapt to complex and changing environmental conditions. Mathematical statistics methods have limited modeling capabilities for non-linear and non-stationary time series, and it is difficult to capture the dynamic changes of the hydrological process.

[0004] Currently, with the progress of technology, runoff prediction models increasingly tend to use big data and machine learning technologies to improve the accuracy and efficiency of prediction. For example, the rapid development of deep learning models in time series processing methods provides the possibility for accurate runoff prediction. At the same time, the combined prediction model based on "decomposition - prediction - reconstruction" shows good practicability, can better process non-linear and non-stationary runoff data, decomposes the original hydrological time series through time-frequency signal decomposition technology, and can effectively extract the effective information in the hydrological time series, thereby improving the prediction accuracy of the model. Summary of the Invention

[0005] The first object of the present invention is to provide a daily runoff prediction method for a reservoir basin, hoping to effectively capture the non-linear relationship of runoff data, thereby improving the prediction accuracy of daily runoff.

[0006] To achieve the above object, the following technical solutions are adopted in the present invention:

[0007] A daily runoff prediction method for a reservoir basin includes the following steps:

[0008] S1. Pre-acquire the daily runoff data of the monitoring stations in the reservoir basin, use the Isolation Forest to detect outliers in the daily runoff data, and perform data cleaning;

[0009] S2. Decompose the daily runoff data processed in step S1 using data-driven adaptive chirp mode decomposition (DD-ACMD) to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set;

[0010] S3. Build a TCN-IFELM fusion model by fusing a time convolutional network (TCN) and an inverse-free extreme learning machine (IFELM), and input the obtained training set into the TCN-IFELM model for training;

[0011] S4. Introduce a half-uniform initialization and variable helix update strategy into the PID search algorithm (PSA) to obtain an improved PID search algorithm IPSA; use the IPSA algorithm to optimize parameters such as the dropout rate, the number of residual blocks, and the number of neurons in the hidden layer of the TCN-IFELM model to improve the prediction performance of the TCN-IFELM model;

[0012] S5. Input the test set into the optimized fusion model for component prediction, and add the prediction results of each component to obtain the predicted value of the daily runoff at a future moment.

[0013] While adopting the above technical solutions, the present invention can also adopt or combine the following technical solutions:

[0014] As a preferred technical solution of the present invention: In step S1, the implementation process of using an isolation forest to detect outliers in the daily runoff data is as follows:

[0015] S11. Construct an isolation forest ;

[0016] ;

[0017] wherein, represents a binary tree; w is the number of binary trees;

[0018] S12. Randomly select V data points from the daily runoff data Q to form a subspace X , randomly select one data feature of X , and randomly select a cutting point within the sample interval corresponding to this feature, and construct a hyperplane from this cutting point to divide X into two different branches under the current node;

[0019] S13. Recursively execute step S12 for the child nodes under the two different branches until the data cannot be further divided;

[0020] S14. Calculate and determine the daily runoff data set VThe data points v Degree of abnormality :

[0021] ;

[0022] Among them, Is the node depth of the data point v In the binary tree, Is the data point v In w The average path length of the tree, Is from Q The average path length of the tree established by the points;

[0023] S15. Cut out the outlier and abnormal data points to determine the outliers.

[0024] As a preferred technical solution of the present invention: The specific process of step S2 is as follows:

[0025] S21. Differentiate the daily runoff time series signal to enhance the high-frequency mode, and its derivative can be expressed as:

[0026] ;

[0027] Among them, Represents the original daily runoff time series signal; Is the derivative of the signal ; K Represents the total number of modes in the signal; Is the i Amplitude of the th mode; The i Phase function of the th mode; Is the i Instantaneous frequency of the th mode; Is Derivative of; t Represents the time variable;

[0028] S22. Estimate the initial instantaneous frequency of the highest frequency mode by normalizing the derivative of the signal. The specific process is as follows:

[0029] ;

[0030] Among them, Is the frequency modulation signal of the th mode after normalization; i The The i Phase function of the th mode; Represents the instantaneous frequency of the highest frequency mode; d Represents differentiation, referring to a very small change;dt Represents a tiny time interval; Is a differential operator, representing the derivative with respect to time t ;

[0031] S23. Introduce a time-varying low-pass filter to remove the noise above the instantaneous frequency of the current signal. The process is as follows:

[0032] S231. Rewrite the frequency-modulated signal as:

[0033] ;

[0034] Where, Represents the frequency-modulated signal; Represents a time interval; Represents An infinitesimal change in; Represents the instantaneous frequency of the signal at point ; Represents the infinitesimal change in the instantaneous frequency Over a time interval ; Represents the initial phase;

[0035] S232. Calculate the analytic signal of the frequency-modulated signal:

[0036] ;

[0037] Where, Is the Analytic signal of; Represents the frequency-modulated signal; Represents the initial phase; Represents The Hilbert transform of; j Represents the imaginary part of the signal;

[0038] S233. Define the demodulation operator And the modulation operator :

[0039] ;

[0040] ;

[0041] Where, Is the demodulation frequency; Represents the carrier frequency;

[0042] S234. Multiply the analytic signal by the demodulation operator to obtain the demodulated signal:

[0043] ;

[0044] Among them, represents the demodulated signal;

[0045] S235. Remove the noise higher than the carrier frequency through low-pass filtering, and then multiply by the modulation operator to obtain a new analytical signal:

[0046] ;

[0047] Among them, represents the signal after modulation; is the demodulated signal after low-pass filtering;

[0048] S24. After each mode extraction, use a time-varying low-pass filter to remove the noise above the mode to reduce the impact on subsequent mode extraction, and sequentially extract the signal modes from high frequency to low frequency through a recursive framework; after each mode is extracted, update the remaining signal and repeat the above process until the stop condition is met; DD-ACMD can decompose the daily runoff time series signal into independent components for subsequent analysis and prediction.

[0049] As a preferred technical solution of the present invention: The specific process of step S3 is as follows:

[0050] S31. Capture the dependencies in the daily runoff time series through TCN, and the specific process of its dilated convolution is as follows:

[0051] ;

[0052] Among them, is the input sequence; f is the filter; d is the dilation factor; z is the convolution kernel position; is the convolution kernel size; represents the elements of the previous layer;

[0053] S32. Adopt a residual structure to add the input and the output of the convolutional network, and the calculation formula is:

[0054] ;

[0055] Among them, is the activation function;

[0056] S33. Define the hidden node addition strategy of IFELM:

[0057] ;

[0058] Among them, lis the number of currently hidden nodes; and are l the input weights and biases when there are the newly added input weights, and T represents the transpose operation;

[0059] S34. Update the output weights according to the following rules:

[0060] ;

[0061] ;

[0062] where is the output weight matrix after adding a new hidden node; Y represents the training output matrix; represents the current l pseudoinverse matrix of is the pseudoinverse matrix after adding a new hidden node; and are two blocks of c used to update the pseudoinverse matrix; H represents the vector after adding a new hidden node; I represents the hidden layer output matrix; is the identity matrix; c represents the transpose of the vector

[0063] S35. Construct the fusion model of TCN-IFELM; use TCN to extract the deep features between runoff data, convert the input high-dimensional data into low-dimensional feature representations, and IFELM sequentially takes the features extracted by TCN as inputs and the corresponding runoff as the target output; by gradually increasing the hidden nodes and updating the weights, IFELM avoids the matrix inversion operation in the traditional extreme learning machine, thereby improving the computational efficiency.

[0064] As a preferred technical solution of the present invention: The specific process of step S4 is as follows:

[0065] S41. Set the population size, maximum number of iterations, and the interval range of the parameters to be optimized of the IPSA algorithm;

[0066] S42. Improve the population initialization part of the original PSA algorithm using the half-uniform initialization strategy. The specific process is as follows:

[0067] ;

[0068] where represents the number of iterations, and are the initial values of the individuals respectively, , ; up and low represent the upper and lower limits of the variable; n is the population size; is a random number between 0 and 1; the half-uniform initialization strategy not only retains the randomness of the population but also avoids the concentrated distribution of the initialized individuals, which is beneficial to improving the diversity of the population;

[0069] S43. Define the objective function of the algorithm as the root mean square error between the predicted value and the actual value of the daily runoff, and calculate the fitness value of the population through the objective function;

[0070] S44. Calculate the system deviation: the deviation of the population at the iteration number is defined as:

[0071] ;

[0072] where represents the best individual value in the population; represents the individual value at the iteration;

[0073] S45. The adjustment process of PID: the output value of PID adjustment at the iteration number is defined as:

[0074] ;

[0075] where , and are random number vectors between 0 and 1; , and are the adjustment coefficients of proportional, integral and differential respectively;

[0076] S46. Population update: To prevent the algorithm from falling into local optimum, the variable spiral update strategy is adopted to improve the population update process of the PSA algorithm. The specific process is as follows:

[0077] ;

[0078] where is the value after individual perturbation; represents the average value of the individuals with the top 10% fitness, is the dynamic weight; b is the parameter for adjusting the spiral shape; A random number uniformly distributed in [-1, 1]; D is the distance between the individual and ; The improved population update formula is:

[0079] ;

[0080] where, is a random number between 0 and 1; represents the output value of PID regulation when the iteration number is ; represents the value of the individual at the next iteration; The variable helix update strategy increases the exploration of the local area while maintaining the global search ability of the algorithm, thereby improving the search accuracy and convergence speed of the algorithm;

[0081] S47. Optimize parameters such as the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model through the IPSA algorithm, so that the root mean square error between the predicted daily runoff value and the actual value continuously decreases;

[0082] S48. Determine whether the algorithm termination condition is reached through the given maximum number of iterations, and finally output the optimal parameters of the TCN-IFELM model within the maximum number of iterations.

[0083] The second object of the present invention is to provide a daily runoff prediction system for a reservoir basin, including the following modules:

[0084] A daily runoff data acquisition module, which is used to pre-acquire the daily runoff data of the monitoring stations in the reservoir basin, detect outliers in the daily runoff data using the isolation forest, and perform data cleaning;

[0085] A data processing module, which is used to decompose the daily runoff data processed by the daily runoff data acquisition module using data-driven adaptive chirp mode decomposition to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set;

[0086] A TCN-IFELM fusion model construction and training module, which is used to construct a TCN-IFELM fusion model that fuses a time convolutional network (TCN) and an inverse-free extreme learning machine (IFELM), and input the training set obtained by the data processing module into the TCN-IFELM model for training;

[0087] TCN-IFELM fusion model optimization module, which is used to introduce half-uniform initialization and variable helix update strategy into the PID search algorithm to obtain an improved PID search algorithm IPSA; use the IPSA algorithm to optimize the dropout rate, the number of residual blocks, and the number of neurons in the hidden layer of the TCN-IFELM model constructed by the TCN-IFELM fusion model construction and training module;

[0088] Prediction module, which is used to input the test sets of each component obtained by the data processing module into the TCN-IFELM fusion model optimized by the TCN-IFELM fusion model optimization module for component prediction, and add the prediction results of each component obtained to get the predicted value of the daily runoff at a certain future moment.

[0089] The third object of the present invention is to provide an electronic device, which includes a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory complete communication with each other through the communication bus.

[0090] Memory, which is used to store computer programs;

[0091] Processor, which is used to execute the computer program stored on the memory to implement the steps of the reservoir basin daily runoff prediction method as described above.

[0092] Another object of the present invention is to provide a non-transitory readable storage medium, in which a program is stored. When the program is executed by a processor, it implements the steps of the reservoir basin daily runoff prediction method as described above.

[0093] The present invention provides a method, system, device and medium for predicting daily runoff in a reservoir basin. The method includes the following steps: S1. Obtain the daily runoff data within a certain basin, detect outliers in the daily runoff data using the Isolation Forest, and perform data cleaning; S2. Decompose the cleaned daily runoff data using data-driven adaptive chirp mode decomposition (DD-ACMD) to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set; S3. Construct a fusion model of a temporal convolutional network (TCN) and an inverse-free extreme learning machine (IFELM) (TCN-IFELM), and input the obtained training set into the TCN-IFELM model for training; S4. Introduce a half-uniform initialization and variable spiral update strategy into the PID search algorithm (PSA) to obtain an improved PID search algorithm (IPSA); use the IPSA algorithm to optimize parameters such as the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model; S5. Obtain the prediction results of each component using the test set and add them up to obtain the predicted value of the daily runoff at a future time.

[0094] Compared with the prior art, the present invention has the following beneficial effects:

[0095] 1). The present invention proposes to decompose the daily runoff data using data-driven adaptive chirp mode decomposition (DD-ACMD). Considering that runoff data usually contains signals of multiple time scales and frequencies, DD-ACMD can decompose these complex signals into independent components, which is convenient for subsequent analysis and prediction; it not only has high adaptability and can adapt to different signal characteristics, but also has good noise robustness.

[0096] 2). The present invention constructs a fusion model of TCN-IFELM. TCN can extract the temporal features of long-term dependencies in runoff data; IFELM avoids the matrix inversion operation in the traditional ELM by gradually increasing the hidden nodes and updating the weights, thereby improving the calculation efficiency; the TCN-IFELM model combines the advantages of both, enhancing the generalization ability of the model.

[0097] 3). The present invention proposes an IPSA algorithm that combines multiple improvement strategies. Using the half-uniform initialization strategy not only retains the randomness of the population but also avoids the centralized distribution of initialized individuals, which is beneficial to improving the diversity of the population; at the same time, the variable spiral update strategy is used to improve the population update process of the PSA algorithm. While maintaining the global search ability of the algorithm, it increases the exploration of local areas, thereby improving the search accuracy and convergence speed of the algorithm. Using IPSA to optimize parameters such as the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model optimizes the structure of the model and further improves the prediction performance of the model.

[0098] 4) The present invention can effectively capture the time series characteristics and non-linear relationships of runoff data, thereby improving the prediction accuracy of daily runoff. Description of the Drawings

[0099] Figure 1 It is a flowchart of the daily runoff prediction method for the reservoir basin provided by the present invention.

[0100] Figure 2 It is a structural diagram of the TCN model.

[0101] Figure 3 It is a structural diagram of the IFELM model.

[0102] Figure 4 It is a radar chart for comparing the NSE indicators of each model.

[0103] Figure 5 It is a box plot for comparing the prediction errors of each model.

[0104] Figure 6 It is a line comparison chart of the predicted values and measured values of the TCN-IFELM fusion model provided by the present invention. Detailed Embodiment

[0105] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0106] As Figure 1 shown, a daily runoff prediction method for a reservoir basin specifically includes the following steps:

[0107] (1) Pre-acquire the daily runoff data of the reservoir basin monitoring stations, and use the Isolation Forest to detect outliers in the daily runoff data and perform data cleaning;

[0108] The implementation process of using the Isolation Forest to detect outliers in the daily runoff data is as follows:

[0109] (11) Construct an Isolation Forest ;

[0110] (1)

[0111] Among them, represents a binary tree; w is the number of binary trees;

[0112] (12) Randomly select V data points from the daily runoff data Q to form a subspace X , randomly select X a data feature, and randomly select a cutting point within the sample interval corresponding to this feature, and construct a hyperplane from this cutting point to divideX Split into two different branches under the current node;

[0113] (13) Recursively execute step (12) for the child nodes under the two different branches until the data cannot be further split;

[0114] (14) Calculate and determine the abnormality degree of the data points V in v the daily runoff dataset :

[0115] (2)

[0116] Among them, is the node depth of the data point v in the binary tree, is the average path length of the path of the data point v on w the tree, is the average path length of the tree established by Q points;

[0117] (15) Cut out the outlier and abnormal data points to determine the outliers.

[0118] (2) Use data-driven adaptive chirp mode decomposition (DD-ACMD) to decompose the daily runoff data in step (1) to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set;

[0119] (21) Differentiate the daily runoff time series signal to enhance the high-frequency mode. Its derivative can be expressed as:

[0120] (3)

[0121] Among them, represents the original daily runoff time series signal; is the derivative of the signal ; K represents the total number of modes in the signal; is the i th amplitude of the mode; the i th phase function of the mode; is the i th instantaneous frequency of the mode; is 's derivative; t represents the time variable;

[0122] (22) Estimate the initial instantaneous frequency of the highest frequency mode by normalizing the derivative of the signal; the specific process is as follows:

[0123] (4)

[0124] wherein, is the frequency modulation signal of the i th mode after normalization; represents the instantaneous frequency of the highest frequency mode; d represents differentiation, referring to a very small change amount; dt represents a very small time interval; is a differential operator, representing the derivative with respect to time t ;

[0125] (23) Introduce a time-varying low-pass filter to remove the noise above the estimated instantaneous frequency of the current signal; the process is as follows:

[0126] (231) Rewrite the frequency modulation signal as:

[0127] (5)

[0128] wherein, represents the frequency modulation signal; represents a time interval; represents an infinitesimal change amount of; represents the instantaneous frequency of the signal at point ; represents the infinitesimal change amount of the instantaneous frequency over a time interval ; represents the initial phase;

[0129] (232) Calculate the analytic signal of the frequency modulation signal:

[0130] (6)

[0131] wherein, is the analytic signal of; represents the initial phase; represents the Hilbert transform of; j represents the imaginary part of the signal;

[0132] (233) Define the demodulation operator and the modulation operator :

[0133] (7)

[0134] (8)

[0135] wherein, is the demodulation frequency; represents the carrier frequency;

[0136] (234) Multiply the parsed signal by the demodulation operator to obtain the demodulated signal:

[0137] (9)

[0138] where, represents the demodulated signal;

[0139] (235) Remove the noise higher than the carrier frequency through low-pass filtering, and then multiply by the modulation operator to obtain a new parsed signal:

[0140] (10)

[0141] where, represents the modulated signal; is the demodulated signal after low-pass filtering;

[0142] (24) After each mode extraction, use a time-varying low-pass filter to remove the noise above the mode to reduce the impact on subsequent mode extraction, and sequentially extract the signal modes from high frequency to low frequency through a recursive framework; after each mode is extracted, update the remaining signal and repeat the above process until the stop condition is met; DD-ACMD can decompose the daily runoff time series signal into independent components for subsequent analysis and prediction.

[0143] (3) Construct a fusion model of a time convolutional network (TCN) and an inverse-free extreme learning machine (IFELM) (TCN-IFELM), and input the obtained training set into the TCN-IFELM model for training;

[0144] (31) Capture the dependencies in the daily runoff time series through TCN, and the specific process of its dilated convolution is as follows:

[0145] (11)

[0146] where, is the input sequence; f is the filter; d is the dilation factor; z is the convolution kernel position; is the convolution kernel size; represents the elements of the previous layer;

[0147] (32) Use a residual structure to add the input and the output of the convolutional network for addition operation, and the calculation formula is:

[0148] (12)

[0149] Among them, is the activation function; the structure of the TCN is as Figure 2 shown.

[0150] (33) Define the hidden node addition strategy of IFELM

[0151] (13)

[0152] Among them, l is the number of current hidden nodes; and are l the input weights and biases when there are hidden nodes, is the newly added input weight, T is the newly added bias;

[0153] (34) Update the output weights through the following rules:

[0154] (14)

[0155] (15)

[0156] Among them, is the output weight matrix after adding a new hidden node; Y represents the training output matrix; represents the current l pseudo-inverse matrix of hidden nodes; and are two blocks of c for updating the pseudo-inverse matrix; H represents the vector after adding a new hidden node; I represents the hidden layer output matrix; is the identity matrix; c represents the transpose of the vector

[0157] (35) Construct the fusion model of TCN-IFELM; use TCN to extract the deep features between runoff data, convert the input high-dimensional data into low-dimensional feature representations, and IFELM sequentially takes the features extracted by TCN as inputs and the corresponding runoff as the target output; by gradually increasing the hidden nodes and updating the weights, IFELM avoids the matrix inversion operation in the traditional extreme learning machine, thereby improving the computational efficiency; the structure of IFELM is as Figure 3 shown.

[0158] (4) Introduce the half-uniform initialization and variable helix update strategy into the PID search algorithm (PSA) to obtain the improved PID search algorithm (IPSA); use the IPSA algorithm to optimize the parameters such as the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model to improve the prediction performance of the TCN-IFELM model; as Figure 1 shown;

[0159] (41) Set the population size, the maximum number of iterations, and the interval range of the parameters to be optimized of the IPSA algorithm;

[0160] (42) Improve the population initialization part of the original PSA algorithm using the half-uniform initialization strategy; the specific process is as follows:

[0161] (16)

[0162] where, represents the number of iterations, and are the initial values of the individual respectively, , ; up and low represent the upper and lower limits of the variable; n is the population size; is a random number between 0 and 1; the half-uniform initialization strategy not only retains the randomness of the population but also avoids the concentrated distribution of the initialized individuals, which is beneficial to improving the diversity of the population;

[0163] (43) Define the objective function of the algorithm as the root mean square error between the predicted value and the actual value of the daily runoff, and calculate the fitness value of the population through the objective function;

[0164] (44) Calculate the system deviation; the deviation of the population at the iteration number is defined as:

[0165] (17)

[0166] where, represents the best individual value in the population; represents the individual value at the iteration;

[0167] (45) The adjustment process of PID; the output value of PID adjustment at the iteration number is defined as:

[0168] (18)

[0169] where, , and is a random number vector between 0 and 1; , and are the adjustment coefficients of proportional, integral, and derivative respectively;

[0170] (46) Population update; To prevent the algorithm from falling into local optimum, a variable helix update strategy is adopted to improve the population update process of the PSA algorithm. The specific process is as follows:

[0171] (19)

[0172] where is the value after individual perturbation; represents the average value of the individuals with the top 10% fitness, is the dynamic weight; b is the parameter for adjusting the helix shape; is a random number uniformly distributed in [-1, 1]; D is the distance between the individual and ; The improved population update formula is:

[0173] (20)

[0174] where is a random number between 0 and 1; represents the output value of PID regulation when the iteration number is ; represents the value of the individual at the next iteration; The variable helix update strategy increases the exploration of local areas while maintaining the global search ability of the algorithm, thereby improving the search accuracy and convergence speed of the algorithm;

[0175] (47) Optimize the parameters such as the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model through the IPSA algorithm, so that the root mean square error between the predicted value and the actual value of the daily runoff continuously decreases;

[0176] (48) Judge whether the algorithm termination condition is reached through the given maximum number of iterations, and finally output the optimal parameters of the TCN-IFELM model within the maximum number of iterations.

[0177] (5) Input the test set into the optimized fusion model for component prediction, and add the prediction results of each component to obtain the predicted value of the daily runoff at the future moment.

[0178] To verify the accuracy of the daily runoff prediction model (DD-ACMD-IPSA-TCN-IFELM) provided by the present invention in predicting daily runoff time series and prove the effectiveness and significance of the model in the field of hydrological prediction, the daily runoff data of the Wujiangdu Hydrological Observation Station located on the main stream of the Wujiang River in Zunyi County, Guizhou Province was taken as the research object, and the daily runoff (m 3 / s) data from May 1, 2000 to April 29, 2020 of the Wujiangdu Hydrological Observation Station were selected, with a total of 7,304 data points; an algorithm program was written in MATLAB language to construct the daily runoff prediction model (DD-ACMD-IPSA-TCN-IFELM) provided by the present invention, and comparative experiments were carried out with six other prediction models, namely the error backpropagation algorithm (BP), extreme learning machine (ELM), inverse-free extreme learning machine (IFELM), temporal convolutional network (TCN), TCN-IFELM, and DD-ACMD-TCN-IFELM. Through comparative analysis, it was found that the DD-ACMD-IPSA-TCN-IFELM model provided by the present invention has higher accuracy than the other six models and has certain superiority in daily runoff prediction. Table 1 shows the evaluation indexes of various different models, including root mean square error (RMSE), mean absolute error (MAE), correlation coefficient (R), symmetric mean absolute percentage error (SMAPE), and Nash efficiency coefficient (NSE).

[0179] Table 1 Statistical table of performance indexes between the model provided by the present invention and the control group models

[0180] model RMSE MAE R SMAPE NSE Back Propagation (BP) algorithm 191.3771 115.5188 0.9202 30.4079 0.8407 Extreme Learning Machine (ELM) 180.8141 102.6418 0.9264 26.4456 0.8578 Inverse-Free Extreme Learning Machine (IFELM) 177.7680 102.1859 0.9288 26.3197 0.8626 Temporal Convolutional Network (TCN) 171.7469 102.0259 0.9352 26.2213 0.8717 TCN-IFELM 169.0253 99.5835 0.9366 26.0195 0.8757 DD-ACMD-TCN-IFELM 109.8021 72.6003 0.9741 21.3179 0.9476 DD-ACMD-IPSA-TCN-IFELM 98.3114 59.7773 0.9790 20.2754 0.9580

[0181] It can be seen from Table 1 that the DD-ACMD-PSA-TCN-IFELM model is superior to other models in the control group in terms of various indexes. Its RMSE, MAE, and SMAPE are the smallest, which are 98.3114, 59.7773, and 20.2754 respectively, lower than those of the BP, ELM, IFELM, TCN, TCN-IFELM, and DD-ACMD-TCN-IFELM models. This indicates that the DD-ACMD-PSA-TCN-IFELM model has significant advantages in prediction accuracy. The constructed TCN-IFELM model combines the respective advantages of TCN and IFELM, and its prediction effect is better than those of the BP, ELM, IFELM, and TCN models. In addition, as shown by the NSE indexes of each model Figure 4 The NSE index of the model proposed in the present invention reaches 0.9580, which is better than other comparative models. It should be noted that after preprocessing the original data through the DD-ACMD decomposition technology, the prediction effect of the DD-ACMD-TCN-IFELM model has been greatly improved. In Figure 4It is visually reflected that its NSE index reaches 0.9476, verifying the effectiveness of the DD-ACMD decomposition technology adopted in the present invention.

[0182] From Figure 5 It can be seen that the TCN-IFELM fusion model provided by the present invention exhibits high robustness in runoff prediction and has the fewest outliers. From the comparison between the prediction results and the measured values, the DD-ACMD-IPSA-TCN-IFELM model (i.e., the TCN-IFELM fusion model) provided by the present invention demonstrates excellent fitting ability. As Figure 6 shown, the model is highly consistent with the measured values during the flat period of runoff volume. When the runoff volume surges, although there are certain deviations, the overall predicted values are basically consistent with the actual values, fully proving the effectiveness of the model.

[0183] The present invention also provides a daily runoff prediction system for a reservoir basin, including the following modules:

[0184] Daily runoff data acquisition module, which is used to pre-acquire the daily runoff data of the monitoring stations in the reservoir basin, detect outliers in the daily runoff data using Isolation Forest, and perform data cleaning;

[0185] Data processing module, which is used to decompose the daily runoff data processed by the daily runoff data acquisition module using data-driven adaptive chirp mode decomposition to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set;

[0186] TCN-IFELM fusion model construction and training module, which is used to construct a temporal convolutional network and an inverse-free extreme learning machine to obtain a TCN-IFELM fusion model, and input the training set obtained by the data processing module into the TCN-IFELM model for training;

[0187] TCN-IFELM fusion model optimization module, which is used to introduce half-uniform initialization and variable helix update strategy into the PID search algorithm to obtain an improved PID search algorithm IPSA; use the IPSA algorithm to optimize the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model constructed by the TCN-IFELM fusion model construction and training module;

[0188] A prediction module, which is used to input the test sets of each component obtained by the data processing module into the TCN-IFELM fusion model optimized by the TCN-IFELM fusion model optimization module for component prediction, and add the prediction results of each component obtained to get the predicted value of the daily runoff at a certain future moment.

[0189] The present invention also provides an electronic device, which includes a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0190] The memory is used to store a computer program.

[0191] The processor is used to execute the computer program stored on the memory to implement the steps of the daily runoff prediction method for the reservoir basin as described above.

[0192] The present invention also provides a non-transitory readable storage medium, in which a program is stored. When the program is executed by a processor, it implements the steps of the daily runoff prediction method for the reservoir basin as described above.

[0193] So far, the technical solution of the present invention has been described in combination with the specific experimental process shown in the drawings. However, the protection scope of the present invention is not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A daily runoff prediction method for a reservoir basin, characterized in that It includes the following steps: S1. Pre-acquire the daily runoff data of the reservoir basin monitoring stations, detect outliers in the daily runoff data using the Isolation Forest, and perform data cleaning; S2. Use data-driven adaptive chirp mode decomposition to decompose the daily runoff data processed in step S1 to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set; S3. Build a TCN-IFELM fusion model that fuses a time convolutional network and an inverse-free extreme learning machine, and input the obtained training set into the TCN-IFELM model for training; S4. Introduce a half-uniform initialization and variable helix update strategy into the PID search algorithm to obtain an improved PID search algorithm IPSA; use the IPSA algorithm to optimize the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model to improve the prediction performance of the TCN-IFELM model; S5. Input the test set into the optimized fusion model for component prediction, and add the prediction results of each component to obtain the predicted value of the daily runoff at a certain future moment; The specific process of step S4 is as follows: S41. Set the population size, the maximum number of iterations, and the interval range of the parameters to be optimized of the IPSA algorithm; S42. Use the half-uniform initialization strategy to improve the population initialization part of the original PSA algorithm. The specific process is as follows: Among them, represents the number of iterations, and are the initial values of individuals respectively, where p = 1, 2, …, n / 2, q = n / 2 + 1, …, n; up and low represent the upper and lower limits of the variables; n is the population size; r1 is a random number between 0 and 1; the half-uniform initialization strategy not only retains the randomness of the population but also avoids the concentrated distribution of the initialized individuals, which is beneficial to improving the diversity of the population; S43. Define the objective function of the algorithm as the root mean square error between the predicted value and the actual value of the daily runoff, and calculate the fitness value of the population through the objective function; S44. Calculate the system deviation: The deviation of the population at the iteration number is defined as: when Among them, represents the best individual value in the population; represents at the individual value during the S45. Adjustment process of PID: Number of iterations Output value of PID adjustment at is defined as: where r2, r3, and r4 are vectors of random numbers between 0 and 1; K p , K i and K d are the adjustment coefficients for proportional, integral, and derivative, respectively; S46. Population update: To prevent the algorithm from falling into a local optimum, use a variable helix update strategy to improve the population update process of the PSA algorithm. The specific process is as follows: Among them, is the value after individual perturbation; represents the average value of the top 10% individuals in terms of fitness, ω is the dynamic weight; b is the parameter for adjusting the spiral shape; υ is a random number uniformly distributed in [-1, 1]; D is the distance between the individual and ; the improved population update formula is: where η is a random number between 0 and 1; denotes the output value of PID regulation when the iteration number is ; denotes the value of the individual at the next iteration; The variable helix update strategy increases the exploration of the local area while maintaining the global search ability of the algorithm, thereby improving the search accuracy and convergence speed of the algorithm; S47. Optimize the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model through the IPSA algorithm, so that the root mean square error between the predicted value and the actual value of the daily runoff is continuously reduced; S48. Determine whether the algorithm termination condition is reached by the given maximum number of iterations, and finally output the optimal parameters of the TCN-IFELM model within the maximum number of iterations.

2. The method according to claim 1, wherein In step S1, the implementation process of detecting outliers in the daily runoff data using the Isolation Forest is as follows: S11. Construct an Isolation Forest Π; Π = {π1, π2, …, π w} where, π represents a binary tree; w is the number of binary trees; S12. Randomly extract Q data points from the daily runoff data V to form a subspace X, randomly select a data feature of X, and randomly select a cutting point within the sample interval corresponding to this feature, and construct a hyperplane from this cutting point to divide X into two different branches under the current node; S13. Recursively execute step S12 for the child nodes under the two different branches until the data cannot be further divided; S14. Calculate and discriminate the outlier degree δ of the data point v in the daily runoff data set V: δ(v,Q) = 2 -E[L(v)] / C(Q) where, L(v) is the node depth of the data point v in the binary tree, E[L(v)] is the average path length of the data point v on w trees, and C(Q) is the average path length of the tree established by Q points; S15. Cut out the outlier and abnormal data points to determine the outliers.

3. The method according to claim 1, wherein The specific process of step S2 is as follows: S21. Differentiate the daily runoff time series signal to enhance the high-frequency pattern, and its derivative can be expressed as: Among them, s(t) represents the original daily runoff time series signal; s′(t) is the derivative of the signal s(t); K represents the total number of modes in the signal; a i (t) is the amplitude of the i-th mode; φ i (t) is the phase function of the i-th mode; φ′ i (t) is the instantaneous frequency of the i-th mode; a i ′(t) is the derivative of a i (t); t represents the time variable; S22. Estimate the initial instantaneous frequency of the highest frequency pattern by normalizing the derivative of the signal. The specific process is as follows: where, g i (t) is the frequency-modulated signal of the i-th mode after normalization; φ i (t) is the phase function of the i-th mode; f i (t) represents the instantaneous frequency of the highest frequency mode; d represents differentiation, referring to a very small change; dt represents a tiny time interval; is the differential operator, representing the derivative with respect to time t; S23. Introduce a time-varying low-pass filter to remove the noise above the estimated instantaneous frequency of the current signal. The process is as follows: S231. Rewrite the frequency modulation signal as: where p(t) represents the frequency modulation signal; τ represents a time interval; dτ represents an infinitesimal change in τ; m(τ) represents the instantaneous frequency of the signal at point τ; m(τ)dτ represents the small change in the instantaneous frequency m(τ) over a time interval τ; θ0 represents the initial phase; S232. Calculate the analytic signal of the frequency modulation signal: where y(t) is the analytic signal of p(t); p(t) represents the frequency modulation signal; θ0 represents the initial phase; H[p(t)] represents the Hilbert transform of p(t); j represents the imaginary part of the signal; S233. Define the demodulation operator Φ - (t) and the modulation operator Φ + (t): where m d (τ) is the demodulation frequency; m c represents the carrier frequency; S234. Multiply the analytic signal by the demodulation operator to obtain the demodulated signal: y d (t) = y(t)Φ - (t) Among them, y d (t) represents the demodulated signal; S235. Remove the noise above the carrier frequency through low-pass filtering, and then multiply by the modulation operator to obtain a new analytic signal: Among them, represents the modulated signal; is the demodulated signal after low-pass filtering; S24. After each mode extraction, use a time-varying low-pass filter to remove the noise above the mode to reduce the impact on subsequent mode extraction. Extract the signal modes from high frequency to low frequency in sequence through a recursive framework; after each mode is extracted, update the remaining signal and repeat the above process until the stop condition is met; DD-ACMD can decompose the daily runoff time series signal into independent components for subsequent analysis and prediction.

4. The method according to claim 1, wherein The specific process of step S3 is as follows: S31. Capture the dependencies in the daily runoff time series through TCN. The specific process of its dilated convolution is as follows: where χ is the input sequence; f is the filter; d is the dilation factor; z is the convolution kernel position; λ is the convolution kernel size; s - dz represents the elements of the previous layer; S32. Use a residual structure to perform an addition operation on the input χ and the output F(χ) of the convolution network. The calculation formula is: O = Activation(χ + F(χ)) where Activation(·) is the activation function; S33. Define the hidden node increment strategy of IFELM: where l is the number of current hidden nodes; A l and b l are the input weights and biases when there are l hidden nodes, α is the newly added input weight, and b l+1 is the newly added bias; T represents the transpose operation; S34. Update the output weights according to the following rules: W l+1 = YB l+1 B l+1 = [B l+1,1 B l+1,2 ​ Among them, W l+1 is the output weight matrix after adding a new hidden node; Y represents the training output matrix; B l represents the pseudoinverse matrix of the current l hidden nodes; B l+1 is the pseudoinverse matrix after adding a new hidden node; B l+1,1 and B l+1,2 are two blocks of B l+1 for updating the pseudoinverse matrix; c represents the vector after adding a new hidden node; H represents the hidden layer output matrix; I is the identity matrix; c T represents the transpose of the vector c; S35. Construct a fusion model of TCN-IFELM; use TCN to extract the deep features between runoff data, convert the input high-dimensional data into a low-dimensional feature representation, and IFELM sequentially takes the features extracted by TCN as inputs and the corresponding runoff as the target output; IFELM avoids the matrix inversion operation in the traditional extreme learning machine by gradually increasing the hidden nodes and updating the weights, thereby improving the calculation efficiency.

5. A daily runoff prediction system for a reservoir basin, characterized in that, It includes the following modules: A daily runoff data acquisition module, which is used to pre-acquire the daily runoff data of the reservoir basin monitoring stations, detect outliers in the daily runoff data using Isolation Forest, and perform data cleaning; Data processing module, which is used to decompose the daily runoff data processed by the daily runoff data acquisition module using data-driven adaptive chirp mode decomposition to obtain multiple components; construct an input matrix for each component and divide it into a training set and a test set; TCN-IFELM fusion model construction and training module, which is used to construct a TCN-IFELM fusion model that combines a temporal convolutional network and an inverse-free extreme learning machine, and input the training set obtained by the data processing module into the TCN-IFELM model for training; TCN-IFELM fusion model optimization module, which is used to introduce a half-uniform initialization and variable helix update strategy into the PID search algorithm to obtain an improved PID search algorithm IPSA; use the IPSA algorithm to optimize the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model constructed by the TCN-IFELM fusion model construction and training module; Prediction module, which is used to input the test sets of each component obtained by the data processing module into the TCN-IFELM fusion model optimized by the TCN-IFELM fusion model optimization module for component prediction, and add the prediction results of each component obtained to obtain the predicted value of the daily runoff at a future moment; The running process of the TCN-IFELM fusion model optimization module is as follows: S41. Set the population size, maximum number of iterations, and the range of the parameters to be optimized of the IPSA algorithm; S42. Improve the population initialization part of the original PSA algorithm using the half-uniform initialization strategy. The specific process is as follows: Among them, represents the number of iterations, and are the initial values of the individuals respectively, where p = 1, 2, …, n / 2, q = n / 2 + 1, …, n; up and low represent the upper and lower limits of the variables; n is the population size; r1 is a random number between 0 and 1; the half-uniform initialization strategy not only retains the randomness of the population but also avoids the concentrated distribution of the initialized individuals, which is beneficial to improving the diversity of the population. S43. Define the objective function of the algorithm as the root mean square error between the predicted value and the actual value of the daily runoff, and calculate the fitness value of the population through the objective function; S44. Calculate the system deviation: The deviation of the population at the iteration number is defined as: when Among them, represents the best individual value in the population; represents at the individual value during the S45. Adjustment process of PID: Iteration count Output value of PID adjustment at is defined as: where r2, r3, and r4 are random number vectors between 0 and 1; K p , K i and K d are the adjustment coefficients for proportional, integral, and differential, respectively; S46. Population update: To prevent the algorithm from falling into a local optimum, a variable helix update strategy is used to improve the population update process of the PSA algorithm. The specific process is as follows: Among them, is the value after individual perturbation; represents the average value of the individuals with the top 10% fitness, ω is the dynamic weight; b is the parameter for adjusting the spiral shape; υ is a random number uniformly distributed in [-1, 1]; D is the distance between the individual and ; the improved population update formula is: where η is a random number between 0 and 1; represents the output value of PID regulation when the iteration number is ; represents the value of the individual at the next iteration; The variable helix update strategy increases the exploration of the local area while maintaining the global search ability of the algorithm, thereby improving the search accuracy and convergence speed of the algorithm; S47. Optimize the dropout rate, the number of residual blocks, and the number of hidden layer neurons of the TCN-IFELM model through the IPSA algorithm, so that the root mean square error between the predicted value and the actual value of the daily runoff continues to decrease; S48. Determine whether the algorithm termination condition is reached by the given maximum number of iterations, and finally output the optimal parameters of the TCN-IFELM model within the maximum number of iterations.

6. An electronic device, which includes a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory complete communication with each other through the communication bus. It is characterized in that Memory, which is used to store computer programs; Processor, which is used to execute the computer programs stored on the memory to implement the steps of the reservoir basin daily runoff prediction method described in any one of claims 1-4.

7. A non-transitory readable storage medium, characterized in that, The non-transitory readable storage medium stores a program, which, when executed by a processor, implements the steps of the reservoir basin daily runoff prediction method described in any one of claims 1-4.

Citation Information

Patent Citations

  • System load prediction method integrating isolated forest and long-term and short-term memory network

    CN111738520A

  • PID parameter optimization method based on improved sparrow algorithm

    CN114397807A