A Customer-Side Flexible Load Forecasting System Based on HHO-LSTM
The customer-side flexible load prediction system with parameter optimization and front-end separation architecture of LSTM through Harry Eagle Optimization Algorithm (HHO) solves the local optimization and parameter complexity of the existing model, improves prediction accuracy and simplifies the process, and is suitable for flexible load scheduling of the power grid.
Patent Information
- Application Number
- CN202211009527.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-22
AI Technical Summary
The existing customer-side flexible load prediction model based on LSTM and group optimization algorithms has the problem of easily falling into local optimization, complex parameter settings and lack of intelligent web operating systems, which is difficult to meet the needs of power grid decision-making and intelligent operation.
The Harry Eagle Optimization Algorithm (HHO) is used to optimize the parameters of the long and short-term memory network (LSTM), and the web system is developed in combination with the front-end and back-end separation architecture, including data analysis and preprocessing, model parameter optimization, LSTM prediction and customer information management modules, and the HHO algorithm is used to find the optimal parameters for prediction.
It achieves the accuracy of customer-side flexible power load prediction and process simplification, reduces labor costs, and is suitable for scheduling decisions of flexible loads.
Smart Images

Figure CN115409251B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a customer-side flexible load prediction system based on HHO-LSTM, belonging to the technical field of electric load prediction. Background Art
[0002] As a supplement to power generation scheduling, customer-side flexible load scheduling can cut peaks and fill valleys, balance the fluctuations of intermittent energy, and provide auxiliary services. It is a means to enrich the regulation methods of power grid dispatching operation. Therefore, the improvement and optimization of the methods and means of flexible load prediction need to be emphasized. In the existing prediction models based on LSTM and swarm optimization algorithms, there are problems such as being easily trapped in local optima and having complex parameter settings. There is also a lack of research on relevant intelligent web operating systems, which is not conducive to meeting the needs of power grid decision-making and intelligent operation. As a swarm optimization algorithm, the HHO optimization algorithm can replace classical algorithms such as the particle swarm algorithm for model parameter optimization in machine learning or deep learning. At present, there is no research on applying it to the parameter optimization of LSTM. Summary of the Invention
[0003] Technical Problem: The present invention now proposes to use the Harris Hawks Optimization (HHO) algorithm to optimize the parameters of the Long Short-Term Memory network (LSTM), and then use the found optimal parameters and the corresponding model to predict the customer-side flexible load data. And a web system is developed based on the proposed combined algorithm using a front-end and back-end separated architecture. This system realizes the improvement of the accuracy of customer-side flexible power load prediction and the simplification of the process, and is used for the scheduling decision of flexible loads.
[0004] Technical Solution: To achieve the above objectives, the present invention adopts the following technical solutions:
[0005] A new short-term prediction system for customer-side flexible power loads based on web development technology, deep learning (LSTM) model, and HHO algorithm mainly includes the following four modules:
[0006] (1) Data analysis and preprocessing module;
[0007] (2) Model parameter optimization module based on HHO;
[0008] (3) Prediction module based on the Long Short-Term Memory network (LSTM) algorithm;
[0009] (4) Management module for customer information and data sets.
[0010] Furthermore, the data analysis and preprocessing module includes the following steps:
[0011] (1-1) Data set upload and preview
[0012] The dataset includes a weather dataset and an electricity load dataset. Both are at one-hour intervals. The weather dataset includes five eigenvalue features: time, temperature, dew point temperature, humidity, wind speed, and air pressure. The electricity load dataset includes two eigenvalue features: time and load.
[0013] The dataset can be uploaded by uploading files of types csv, txt, xls, and xlsx through the front-end interface. The front-end page is written based on Vue.js, built with the vue scaffolding, and uses the Element-UI framework to implement the page layout and user interaction functions. The dataset is uploaded to the back-end interface through the el-upload component of the framework. The back-end is built using the classic Python web framework Django. The back-end stores the incoming files and user information in the Mongodb database in the data storage layer, with the user ID as the key and the dataset as the value, storing all the datasets uploaded by users. After successful upload, it is convenient to send requests from the front-end to the back-end, and the back-end reads the database to perform a series of subsequent steps.
[0014] Data preview includes tabular display of data and display of various visualization pictures. The tabular display of data uses el-table to display data in a format similar to an Excel table, and corresponding data can be queried by setting the time range. The display of data visualization pictures facilitates the intuitive feature analysis of the dataset. For the load dataset, the back-end can be requested according to the specified time range. The back-end uses various mathematical analysis tools in Python for analysis and drawing. After drawing the line chart for the specified time range, it can be displayed to the user on the front-end. Through this kind of chart, load distribution characteristics can be seen, such as seasonal electricity consumption characteristics and holiday electricity consumption characteristics. The weather dataset is also visualized in the same way.
[0015] (1-2) Dataset preprocessing
[0016] The dataset preprocessing is specific to processing each column. In the system page, the corresponding column in the dataset can be selected using a drop-down box, and a series of classic dataset preprocessing operations can be performed. By sending a request from the front-end, after the back-end receives it, various data processing toolkits in Python are used to execute specific business processing. The processed dataset is updated to the database, and after returning a successful response signal, the front-end is automatically refreshed to display the processed database.
[0017] The specific processing operations include: outlier detection, missing value filling, discrete feature extraction, and normalization processing.
[0018] Among them, two optional methods are implemented for outlier handling: box plot and 3-Sigma criterion. The box plot regards the outliers beyond the quartile boundaries as outliers in the dataset. The 3-Sigma criterion is applicable to datasets that follow a normal distribution. That is, within the 3-Sigma range (μ–3σ, μ+3σ), 99.73% of the data is normal data, where σ represents the standard deviation and μ represents the mean, and x = μ is the axis of symmetry of the graph. The detected outliers are set to null values and filled together with the missing values for missing value filling.
[0019] Missing value filling includes filling with data from the previous hour or the next hour, filling with the mean value of data from the previous and next two hours, or filling with data from the corresponding moment of the previous day, filling with data from the corresponding moment of the next day, or filling with the mean value of data from the corresponding moments of the previous and next two days. Both load and weather data have the characteristics of time series. Therefore, using the above-mentioned methods of filling data based on adjacent time data is more reasonable and effective compared to other traditional data preprocessing methods.
[0020] Discrete feature extraction includes feature extraction of whether it is a holiday. By judging the time corresponding to the data, the sample data of holidays is marked as 1, and non-holidays are marked as 0, and merged into the dataset as a new column; there is also feature extraction of seasons. Spring, summer, autumn, and winter are respectively marked as 0, 1, 2, and 3 and merged into the dataset as a new column. The air-conditioning load in the flexible load on the customer side is a key and important component, and the use of air conditioners is greatly affected by holidays and seasons. Therefore, it is necessary to extract these two factors as new features.
[0021] There are two optional methods for normalization: maximum-minimum normalization and standard normalization. The data after standard normalization conforms to the standard normal distribution, that is, the standard normal distribution with a mean of 0 and a standard deviation of 1. The formula for the standard deviation is as shown in Equation (1). Considering all sample data, it is less affected by outliers. The maximum-minimum normalization method is as shown in Equation (2). This method can normalize the data between 0 and 1. However, if the outliers are extremely large, the difference between the maximum and minimum values will be relatively large. For example, once there is an outlier (a particularly large value) in the data, after normalization, this outlier will be extremely close to 1, while other values will be extremely close to 0. Therefore, this method is generally not selected.
[0022]
[0023] In the formula: μ is the mean of all samples, and δ is the standard deviation of all sample data.
[0024]
[0025] In the formula: x is the original sample data, x′ is the data after normalization processing, x max and xmin They are the sample maximum value and the minimum value respectively.
[0026] Furthermore, the HHO-based model parameter optimization module includes the following steps:
[0027] First, select the load data set in the drop-down list, and then merge it with the corresponding weather data set to generate the total data set. The sliding window method is used for the data set to construct a data set suitable for the LSTM input format. One hour is used as the sliding window step size, and the sliding window size is used as one of the optimization parameters.
[0028] Use pytorch to build an LSTM model with two hidden layers and one fully connected layer. Select the number of neurons Num_Units, the number of training iterations Epoch, and the above-mentioned data sliding window size Window_Size as the three target parameters [Num_Units, Window_Size, Epoch] that need to be optimized by HHO.
[0029] Initialize the parameters of the HHO algorithm through the front-end interface to set the size of the population and the upper and lower bounds of the eagle group positions. Then, an error evaluation method can be selected at the radio box of the evaluation criteria on the interface as the fitness function. There are four optional error evaluation methods: mean absolute error (MAE), root mean squared error (RMSE), mean absolute percentage error (MAPE), and symmetric mean absolute percentage error (SMAPE). The expressions are as follows:
[0030]
[0031]
[0032]
[0033]
[0034] In the formula, N represents the total number of samples, and y iThey respectively represent the predicted value and the actual value of the i-th sample data. The smaller the above four indicators, the better the prediction accuracy of the model. Both MAE and RMSE are absolute indicators. MAE reflects the average of the absolute errors. Compared with MAE, RMSE uses the square term to amplify the gap between larger errors and smaller errors, making RMSE more sensitive to data with larger prediction deviations. MAPE and SMAPE are relative indicators. MAPE uses the actual value as the denominator and reflects the average of the percentage errors. SMAPE is a modified version of MAPE, which solves the problems that MAPE cannot be calculated when the actual value is 0 and that MAPE punishes negative errors more severely than positive errors. The value range of SMAPE is [0, 200], and the smaller the SMAPE, the better the prediction performance.
[0035] Therefore, it is used as the fitness function to search for the parameters that minimize its value during the optimization process. After the iteration in the Python backend is completed according to the HHO iteration process, the three optimal parameters found, the global optimal fitness of each round, the model corresponding to the optimal parameters, and the total iteration time plus the start timestamp are saved as a parameter optimization record. And the convergence curve of the fitness is visually displayed in the form of a line chart on the front end.
[0036] Compared with the traditional PSO, the HHO algorithm does not require initializing the upper and lower bounds of the velocity and other hyperparameters, and the manual configuration is more concise. Coupled with the development of the web system, the entire process is more user-friendly and convenient, reducing the labor cost of the customer-side flexible load prediction. And currently, no other research on applying HHO to LSTM parameter optimization has been seen except for this patent.
[0037] Furthermore, the prediction module based on the Long Short-Term Memory (LSTM) algorithm includes the following steps:
[0038] This module aims to predict the newly added dataset of the user using the optimal parameters found by the HHO algorithm described in claim 3 and the saved model. For the page settings, only the dataset and the parameter optimization record saved in claim 3 need to be selected, and then the saved model and optimal parameters can be loaded to perform the sliding window processing of the dataset and the result prediction. The prediction result is saved as a record with a timestamp in the database, and clicking on the result details can return it to the front end for display in the form of a line chart. The numerical results can also be directly displayed in the form of a table and can be saved and downloaded locally in csv format.
[0039] Furthermore, the management module for customer information and datasets includes the following steps:
[0040] This module includes user management and data management. The former is mainly for the personal account management of users, and the latter is for the management of the data sets corresponding to users. After a user registers in the system, personal information will be saved to the database, and information verification will be performed during login. As described in claim 2, a user can upload multiple data sets, which are stored in the database in the form of key-value pairs like <user, data set>. The data preview of the system only shows the list of data sets owned by the user. Data set management mainly includes the addition, deletion, modification, and query of data sets, as well as the deletion of the parameter optimization results described in claim 3 and the deletion of the prediction results described in claim 4.
[0041] Advantageous effects: The present invention aims at the problems of inaccurate parameter optimization of LSTM, complex parameter settings of the optimization algorithm, and cumbersome data processing and prediction processes. It proposes to use the Harris Hawk Optimization (HHO) algorithm to optimize the parameters of the Long Short-Term Memory network (LSTM), and then use the obtained optimal parameters and the corresponding model to predict the flexible load data on the customer side. Moreover, a web system is developed based on the proposed combined algorithm using a front-end and back-end separated architecture. This system improves the accuracy of predicting the flexible power load on the customer side and simplifies the process, and is used for the scheduling decision of flexible loads. Description of the Drawings
[0042] Figure 1 It is a functional module diagram of the system of the present invention.
[0043] Figure 2 It is an overall architecture diagram of the system of the present invention.
[0044] Figure 3 It is an overall flow chart of the system of the present invention. Detailed Embodiments
[0045] The present invention will be further described below with reference to the drawings.
[0046] As Figure 3 shown, the present invention relates to a new short-term prediction system for flexible power loads on the customer side based on web development technology, deep learning (LSTM) model, and HHO algorithm. Each module will be described below.
[0047] (1): Data analysis and preprocessing module.
[0048] (1-1) Data set upload and preview
[0049] The data sets include weather data sets and power load data sets, both with a one-hour time step. The weather data set includes five characteristic values: time, temperature, dew point temperature, humidity, wind speed, and air pressure. The power load data set includes two characteristic values: time and load.
[0050] The uploaded dataset can be a csv, txt, xls, or xlsx file uploaded through the front-end interface. The front-end page is written based on Vue.js, built with the vue scaffolding, and uses the Element-UI framework to implement the page layout and user interaction functions. The dataset is uploaded to the back-end interface through the el-upload component of the framework. The back-end is built using the classic Python web framework Django. The back-end stores the incoming file and user information in the Mongodb database in the data storage layer, with the user ID as the key and the dataset as the value, storing all the datasets uploaded by users. After successful upload, it is convenient to send requests from the front-end to the back-end, and the back-end reads the database to perform a series of subsequent steps.
[0051] Data preview includes tabular display of data and various visual images. The tabular data is displayed using el-table to show data in a format similar to an Excel table, and corresponding data can be queried by setting the time range. The visual image display of data facilitates the intuitive analysis of the characteristics of the dataset. For the load dataset, requests can be sent to the back-end according to the specified time range. The back-end uses various mathematical analysis tools in Python for analysis and drawing. After drawing the line chart for the specified time range, it can be displayed to the user on the front-end. Through this chart, load distribution characteristics such as seasonal electricity consumption characteristics and holiday electricity consumption characteristics can be seen. The weather dataset is also visually displayed in the same way.
[0052] (1-2) Dataset preprocessing
[0053] Dataset preprocessing specifically processes each column. In the system page, the corresponding column in the dataset can be selected using a dropdown box, and a series of classic dataset preprocessing operations can be performed. By sending requests from the front-end, after the back-end receives them, various data processing toolkits in Python are used to execute specific business processing. The processed dataset is updated to the database, and after returning a successful response signal, the front-end is automatically refreshed to display the processed database.
[0054] The specific processing operations include: outlier detection, missing value filling, discrete feature extraction, and normalization processing.
[0055] Among them, two optional methods are implemented for outlier processing: box plot and 3-Sigma criterion. The box plot regards the outliers beyond the quartile boundaries as outliers in the dataset. The 3-Sigma criterion is applicable to datasets that follow a normal distribution, that is, 99.73% of the data is normal within the 3-Sigma range (μ–3σ, μ+3σ), where σ represents the standard deviation, μ represents the mean, and x = μ is the axis of symmetry of the graph. The detected outliers are set to null values and filled together with the missing values.
[0056] Missing value filling includes filling with data from the previous hour or the next hour, filling with the average value of data from the previous and next two hours, or filling with data at the corresponding time of the previous day, filling with data at the corresponding time of the next day, or filling with the average value of data at the corresponding time of the previous and next two days. Both load and weather data have the characteristics of time series, so using the above-mentioned methods of filling based on data from similar times is more reasonable and effective than other traditional data preprocessing methods.
[0057] Discrete feature extraction includes extracting the feature of whether it is a holiday. By judging the time corresponding to the data, the sample data of holidays are marked as 1, and non-holidays are marked as 0, and merged into the dataset as a new column; there is also the feature extraction of seasons, where spring, summer, autumn, and winter are respectively marked as 0, 1, 2, and 3 and merged into the dataset as a new column. The air-conditioning load in the flexible load on the customer side is a key and important component, and the use of air conditioners is greatly affected by holidays and seasons, so it is necessary to extract these two factors as new features.
[0058] There are two optional methods for normalization: maximum-minimum normalization and standard normalization. The data after standard normalization conforms to the standard normal distribution, that is, the standard normal distribution with a mean of 0 and a standard deviation of 1. The formula for the standard deviation is as shown in Equation (1). Considering all sample data, it is less affected by outliers. The maximum-minimum normalization method is as shown in Equation (2). This method can normalize the data between 0 and 1, but if the outlier is extremely large, the difference between the maximum and minimum values is relatively large. For example, once there is an outlier (a particularly large value) in the data, after normalization, this outlier will be particularly close to 1, while other values will be particularly close to 0, so this method is generally not selected.
[0059]
[0060] In the formula: μ is the mean of all samples, and δ is the standard deviation of all sample data.
[0061]
[0062] In the formula: x is the original sample data, x′ is the data after normalization processing, x max and x min are the sample maximum and minimum values respectively.
[0063] (2): Model parameter optimization module based on HHO.
[0064] First, select the load dataset in the drop-down list, and then merge it with the corresponding weather dataset to generate the total dataset. The sliding window method is used for the dataset to construct a dataset suitable for the input format of LSTM, with one hour as the sliding window step size, and the sliding window size as one of the optimization parameters.
[0065] Build an LSTM model with two hidden layers and one fully connected layer using PyTorch. Select the number of neurons Num_Units, the number of training iterations Epoch, and the data sliding window size Window_Size described above as the three target parameters [Num_Units, Window_Size, Epoch] to be optimized using HHO.
[0066] Initialize the parameters of the HHO algorithm through the front-end interface to set the size of the population and the upper and lower bounds of the eagle population positions. Then, you can select an error evaluation method as the fitness function at the evaluation criterion radio box on the interface. There are four optional error evaluation methods: mean absolute error (MAE), root mean squared error (RMSE), mean absolute percentage error (MAPE), and symmetric mean absolute percentage error (SMAPE). The expressions are as follows:
[0067]
[0068]
[0069]
[0070]
[0071] In the formula, N represents the total number of samples, and y i represent the predicted value and the actual value of the i-th sample data, respectively. The smaller the above four indicators, the better the prediction accuracy of the model. Both MAE and RMSE are absolute indicators. MAE reflects the average of the absolute errors. Compared with MAE, RMSE uses the square term to amplify the gap between larger and smaller errors, making RMSE more sensitive to data with larger prediction deviations. MAPE and SMAPE are relative indicators. MAPE uses the actual value as the denominator and reflects the average of the percentage errors. SMAPE is a modified version of MAPE, which solves the problem that MAPE cannot be calculated when the actual value is 0 and the problem that MAPE punishes negative errors more severely than positive errors. The value range of SMAPE is [0, 200], and the smaller the SMAPE, the better the prediction performance.
[0072] Therefore, it is used as a fitness function to search for the parameters that minimize its value during the optimization process. After the iteration is completed on the Python backend according to the HHO iteration process, the three optimal parameters found, the global optimal fitness of each round, the model corresponding to the optimal parameters, and the total iteration time plus the start timestamp are saved as a parameter optimization record. And the convergence curve of the fitness is visually displayed in the form of a line chart on the front end.
[0073] Compared with the traditional PSO, the HHO algorithm does not require initializing the velocity upper and lower bounds and other hyperparameters, and the manual configuration is more concise. Coupled with the development of the web system, the entire process is more user-friendly and convenient, reducing the labor cost of flexible load prediction on the customer side. And currently, no other research on applying HHO to LSTM parameter optimization has been seen except for this patent.
[0074] (III): Prediction module based on the Long Short-Term Memory (LSTM) algorithm.
[0075] This module includes user management and data management. The former is mainly the personal account management of users, and the latter is the management of the data sets corresponding to users. After a user registers in the system, personal information will be saved to the database, and information verification will be performed during login. As described in claim 2, a user can upload multiple data sets, which are stored in the database in the form of key-value pairs <user, data set>. The data preview of the system only shows the list of data sets owned by the user. Data set management mainly includes adding, deleting, modifying, and querying data sets, as well as deleting the parameter optimization results described in claim 3 and the prediction results described in claim 4.
[0076] (IV): Management module for customer information and data sets.
[0077] This module includes user management and data management. The former is mainly the personal account management of users, and the latter is the management of the data sets corresponding to users. After a user registers in the system, personal information will be saved to the database, and information verification will be performed during login. As described in claim 2, a user can upload multiple data sets, which are stored in the database in the form of key-value pairs <user, data set>. The data preview of the system only shows the list of data sets owned by the user. Data set management mainly includes adding, deleting, modifying, and querying data sets, as well as deleting the parameter optimization results described in claim 3 and the prediction results described in claim 4.
Claims
1. A customer-side flexible load prediction system based on HHO-LSTM, characterized in that, It includes the following modules: (1) Data analysis and preprocessing module; (2) Model parameter optimization module based on HHO; (3) Prediction module based on LSTM algorithm; (4) Management module for customer information and dataset; The (2) model parameter optimization module based on HHO specifically includes: First, select the load dataset in the drop-down list, and then merge it with the corresponding weather dataset to generate the total dataset; Use the sliding window method for the dataset to construct a dataset suitable for the LSTM input format, with one hour as the sliding window step size, and the sliding window size as one of the optimization parameters; Use pytorch to build an LSTM model with two hidden layers and one fully connected layer, and select the number of neurons Num_Units, the number of training iterations Epoch, and the above-mentioned data sliding window size Window_Size as the three target parameters [Num_Units, Window_Size, Epoch] to be optimized using HHO; Initialize the parameters of the HHO algorithm through the front-end interface to set the size of the population and the upper and lower bounds of the eagle group positions, and then you can select an error evaluation method as the fitness function at the radio box of the evaluation criteria on the interface; there are four optional error evaluation metrics: Mean Absolute Error MAE, Root Mean Square Error RMSE, Mean Absolute Percentage Error MAPE, and Symmetric Mean Absolute Percentage Error SMAPE, and the expressions are as follows: where N represents the total number of samples, and y i represent the predicted value and the actual value of the i-th sample data respectively; the smaller the above four indicators, the better the prediction accuracy of the model; both MAE and RMSE are absolute indicators; MAE reflects the average of the absolute errors; compared with MAE, RMSE uses the square term to amplify the gap between larger and smaller errors, making RMSE more sensitive to data with larger prediction deviations; MAPE and SMAPE are relative indicators; MAPE uses the actual value as the denominator and reflects the average of the percentage errors; SMAPE is a modified version of MAPE, which solves the problems that MAPE cannot be calculated when the actual value is 0 and MAPE punishes negative errors more severely than positive errors; the value range of SMAPE is [0, 200], and the smaller the SMAPE, the better the prediction performance; Therefore, use it as the fitness function to search for the parameters that minimize its value during the optimization process; after the iteration is completed in the python backend according to the HHO iteration process, save the three optimal parameters found, the global optimal fitness of each round, the model corresponding to the optimal parameters, and the total iteration time plus the start timestamp as a parameter optimization record; and the convergence curve of the fitness is visually displayed in the form of a line chart on the front end.
2. The client-side flexible load prediction system based on HHO-LSTM according to claim 1, wherein The (1) data analysis and preprocessing module specifically includes: (1-1) Dataset upload and preview The dataset includes a weather dataset and an electric load dataset; both are at one-hour intervals; the weather dataset includes five characteristic values: time, temperature, dew point temperature, humidity, wind speed, and air pressure; the electric load dataset includes two characteristic values: time and load; Upload the dataset by uploading files of csv, txt, xls, xlsx types through the front-end interface. The front-end page is written based on Vue.js, built using the vue scaffolding, and uses the Element-UI framework to implement the page layout and user interaction functions. The dataset is uploaded to the backend interface through the el-upload component of the framework. The backend is built using the classic python web framework Django. The backend stores the incoming files and user information in the Mongodb database in the data storage layer, with the user ID as the key and the dataset as the value, storing all the datasets uploaded by users; after the upload is successful, it is convenient to send requests from the front end to the backend, and the backend reads the database to perform a series of subsequent steps; Data preview includes tabular display of data and display of various visual images; tabular data is displayed using el-table to show data in a format similar to an Excel table, and corresponding data can be queried by setting a time range; visual image display of data facilitates intuitive feature analysis of the dataset. For the load dataset, the backend is requested according to the defined time range. The backend uses various mathematical analysis tools in Python for analysis and drawing. After drawing a line chart for the specified time range, it can be displayed to the user on the front end, and the load distribution characteristics can be seen from this chart; the weather dataset is also visually displayed in the same way. (1-2) Dataset preprocessing Dataset preprocessing is specifically carried out for each column. In the system page, the corresponding column in the dataset is selected using a dropdown box, and a series of classic dataset preprocessing operations are performed. By sending a request from the front end, after the backend receives it, various data processing toolkits in Python are used to execute specific business processing, and the processed dataset is updated to the database. After returning a successful response signal, the front end is automatically refreshed to display the processed database. Specific processing operations include: outlier detection, missing value filling, discrete feature extraction, and normalization processing. Among them, two optional methods are implemented for outlier processing: box plot and 3-Sigma criterion; the box plot regards the outliers beyond the quartile boundaries as outliers in the dataset, and the 3-Sigma criterion is applicable to datasets that follow a normal distribution, that is, 99.73% of the data within the 3-Sigma range (μ–3σ, μ+3σ) is normal data, where σ represents the standard deviation and μ represents the mean, and x = μ is the axis of symmetry of the graph; the detected outliers are set to null values and filled together with the missing values. Missing value filling includes filling with data from the previous hour or the next hour, filling with the mean of the data in the previous and next two hours, or filling with data at the corresponding time of the previous day, filling with data at the corresponding time of the next day, or filling with the mean of the data at the corresponding time of the previous and next two days; both load and weather data have the characteristics of time series. Discrete feature extraction includes feature extraction of whether it is a holiday. The time corresponding to the data is judged, and the sample data of holidays is marked as 1, and non-holidays are marked as 0, which is merged into the dataset as a new column; there is also feature extraction of seasons, where spring, summer, autumn, and winter are respectively marked as 0, 1, 2, and 3 and merged into the dataset as a new column. There are two optional methods for normalization processing: maximum-minimum normalization and standard normalization; the data processed by standard normalization conforms to the standard normal distribution, that is, the standard normal distribution with a mean of 0 and a standard deviation of 1; the formula for the standard deviation is as shown in Equation (1). Considering all sample data, it is less affected by outliers; the maximum-minimum normalization method is as shown in Equation (2); this method can normalize the data between 0 and 1, but if the outlier is particularly large, then the difference between the maximum and minimum values is relatively large. Once there is an outlier in the data, this outlier will be very close to 1 after normalization, while other values will be very close to 0. Where: μ is the mean of all samples, and δ is the standard deviation of all sample data; where: x is the original sample data, x ′ is the data after normalization processing, x max and x min are the maximum and minimum values of the sample respectively.
3. The client-side flexible load prediction system based on HHO-LSTM according to claim 2, wherein The prediction module based on the LSTM algorithm in (3) specifically includes: This module aims to predict the newly added dataset of the user using the optimal parameters found by the HHO algorithm and the saved model. For page settings, only the dataset and the saved parameter optimization record need to be selected, and then the saved model and optimal parameters can be loaded to perform sliding window processing on the dataset and result prediction; the predicted results are saved as a record with a timestamp in the database, and clicking on the result details will return a line chart to the front end for display. The numerical results can also be directly displayed in a table form and saved in csv format for downloading to the local.
4. The client-side flexible load prediction system based on HHO-LSTM according to claim 3, characterized in that, The management module for customer information and datasets in (4) specifically includes: This module includes user management and data management. The former is mainly for the personal account management of users, and the latter is for the management of the datasets corresponding to users; after a user registers in the system, personal information will be saved in the database, and information verification will be performed during login. A user can upload multiple datasets, which are stored in the database in the form of key-value pairs like <user, dataset>. The data preview of the system only shows the list of datasets owned by the user; dataset management mainly includes adding, deleting, modifying, and querying datasets, as well as deleting the parameter optimization results and the prediction results mentioned above.