Enterprise energy use prediction and scheduling algorithm based on big data analysis

Through the Internet of Things, multi-source heterogeneous data is collected and combined with machine learning and intelligent optimization algorithms to generate optimal scheduling solutions, the problems of insufficient data utilization and low prediction accuracy in enterprise energy prediction and scheduling are solved, and efficient energy management and decision-making support are achieved.

CN120258379APending Publication Date: 2025-07-04NANJING YUSHAN INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510291496.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology has problems such as insufficient data utilization, poor real-time performance, low prediction accuracy, insufficient scheduling optimization and lack of visual support in enterprise energy prediction and scheduling, resulting in low energy management efficiency.

Method used

Multi-source heterogeneous data is collected through the Internet of Things, combined with machine learning algorithms for prediction, and genetic algorithms or particle swarm algorithms to generate the optimal scheduling scheme, and display it through a graphical interface, and dynamic update of the model is achieved by combining the system feedback module to achieve closed-loop feedback and continuous optimization.

Benefits of technology

It improves prediction accuracy and scheduling efficiency, ensures long-term accuracy and adaptability of the model, provides intuitive decision-making support, and improves energy management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258379A_ABST
    Figure CN120258379A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise energy use prediction and scheduling algorithm based on big data analysis, and the algorithm comprises the steps: S1, forming a multi-source heterogeneous big data set, and recording constraint conditions; s2, carrying out cleaning and fusion; s3, extracting a plurality of key features, and training an energy use prediction model through a machine learning algorithm; s4, inputting a future production plan and environmental parameters, and predicting an energy demand result; s5, generating an initial scheduling scheme according to the energy demand result; verifying by using constraint conditions to obtain an optimal scheduling scheme; s6, displaying through a graphical interface; and S7, dynamically updating the energy use prediction model according to the feedback data. A data set of multi-source heterogeneous data is obtained through the Internet of Things, and the energy use condition is comprehensively reflected; a machine learning algorithm is used for prediction, an optimal scheduling scheme is generated in combination with an intelligent optimization algorithm, and precision and efficiency are ensured; and a system feedback module is also used for feeding back, the model is dynamically updated, and closed-loop feedback and continuous optimization are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enterprise energy consumption prediction, and particularly relates to an algorithm for enterprise energy usage prediction and scheduling based on big data analysis. Background Art

[0002] With the continuous growth of global energy demand, the energy consumption of enterprises accounts for an important proportion in their operating costs. In order to reduce energy costs and improve energy utilization efficiency, enterprises need to effectively predict and schedule energy usage.

[0003] Currently, when enterprises predict and schedule energy usage, they often rely on systems such as resource management to monitor the usage of computing resources, adjust the computing resource allocation strategy according to the resource usage and task characteristics, then train the BP neural network model or other neural network models based on historical data, and finally use the trained model for prediction and optimization.

[0004] However, the above existing solutions have at least the following drawbacks: 1) Insufficient data utilization and poor real-time performance: Although traditional energy prediction and scheduling methods use neural networks for training, they mainly rely on simple statistical analysis of historical energy consumption data, lacking the comprehensive utilization of multi-source heterogeneous big data, unable to accurately reflect complex and changeable energy usage patterns, and it is also difficult to achieve rapid processing and response to real-time data, unable to timely adjust the prediction and scheduling plan, resulting in a large lag. 2) Low prediction accuracy and still insufficient scheduling optimization: Due to the lack of improvement in data analysis algorithms, the prediction accuracy of energy demand is relatively low, unable to meet the requirements of enterprise refined management; at the same time, there is a lack of intelligent optimization means for the content of scheduling prediction, unable to fully consider various constraints and optimization objectives, resulting in unreasonable energy allocation and difficulty in achieving the minimization of energy costs. 3) Lack of visualization and decision support: There is a lack of an intuitive visualization interface, unable to provide effective decision support for managers, affecting the energy management efficiency. Summary of the Invention

[0005] In view of the above three problems, the object of the present invention is to propose an algorithm for enterprise energy use prediction and scheduling based on big data analysis. By collecting various energy data through Internet of Things technology, combining multi-source heterogeneous data such as historical energy consumption, equipment operation parameters, environmental factors and production plans, a complete data set is constructed to comprehensively reflect the enterprise energy use situation and realize the comprehensive utilization of multi-source heterogeneous data. An advanced machine learning algorithm is also used for energy demand prediction, and an intelligent optimization algorithm combining genetic algorithm or particle swarm algorithm is used to generate an optimal scheduling plan to ensure prediction accuracy and scheduling efficiency. At the same time, the actual operation data is fed back into the model through the system feedback module, and incremental learning or online learning technology is used to dynamically update the model to realize the continuous optimization and enhanced adaptability of the model, ensure the accuracy of long-term operation and achieve closed-loop feedback. In addition, it is intuitively displayed through a graphical interface to facilitate decision-makers to analyze and check.

[0006] It is achieved through the following technical solutions: An algorithm for enterprise energy use prediction and scheduling based on big data analysis, comprising the following steps: S1. In multiple data acquisition modules, multiple basic parameters of the energy used within the enterprise are collected through the Internet of Things to form a multi-source heterogeneous big data set. The multiple basic parameters include historical energy consumption data, equipment operation parameters, environmental factor data and production plan data; at the same time, the constraint conditions for using energy are recorded. S2. In the data processing module, the big data set in step S1 is cleaned and fused, including: using the data cleaning unit to perform missing value processing, outlier removal, normalization processing and unified time stamp in sequence; using the data fusion unit to fuse the historical energy consumption data, equipment operation parameters, environmental factor data and production plan data that have completed data cleaning in the big data set to form a specific data set. S3. In the prediction model construction module, multiple key features are extracted from the specific data set in step S2 using data mining technology, and then the multiple key features are used as inputs to train an energy use prediction model through a machine learning algorithm based on time series. S4. The energy use prediction module, based on the energy use prediction model completed in training in step S3, inputs future production plans and environmental parameters to predict the future energy demand results of the enterprise. S5. In the scheduling optimization module, first, according to the energy demand results predicted in step S4, an initial scheduling plan is generated using an intelligent optimization algorithm, where the intelligent optimization algorithm is a genetic algorithm or a particle swarm algorithm; then, the initial scheduling plan is verified using the constraint conditions in step S1, and the parts of the initial scheduling plan that do not meet the constraint conditions are adjusted to obtain an optimal scheduling plan. S6. In the visualization and decision support module, the optimal scheduling plan in step S5 and the energy demand results in step S4 are displayed through a graphical interface. S7. After running according to the optimal scheduling plan displayed in step S6, the system feedback module transmits the data during actual operation as feedback data to the energy usage prediction model, calculates the error index between the feedback data and the predicted energy demand results, and the error index is used to represent the model accuracy to the enterprise; then, based on the feedback data, the energy usage prediction model is retrained, and incremental learning or online learning technology is adopted to dynamically update the energy usage prediction model.

[0007] For a large amount of data from different sources, comprehensive optimization is carried out from data collection, processing, prediction to scheduling, greatly improving the accuracy and effectiveness of prediction; at the same time, a feedback mechanism is also set up, which can dynamically update the prediction model, so as to effectively improve the model accuracy in the long term and perform real-time intelligent optimization; in addition, the results are visually displayed, intuitively presented to decision-makers, and convenient for decision-makers to analyze and check.

[0008] Preferably, in step S1, the Internet of Things at least includes multiple flow meters, multiple pressure sensors, multiple intelligent water meters, multiple power monitors, a communication module, a central management platform and a gateway; each flow meter is used to obtain flow data, each pressure sensor is used to obtain pressure data, each intelligent water meter is used to obtain water consumption, and each power monitor is used to obtain power consumption data; the communication module receives the corresponding data of each flow meter, each pressure sensor, each intelligent water meter and each power monitor through the gateway and stores them in the central management platform, and the central management platform is used to transmit the data stored by itself to the data processing module; among them, the flow data, pressure data, water consumption and power consumption data all belong to the device operation parameters. Transmitting data through the Internet of Things composed of a variety of monitoring devices can effectively obtain the data related to each device, so as to quickly transmit the corresponding data.

[0009] Preferably, the constraint conditions in step S1 all include internal constraint conditions and external constraint conditions. The internal constraint conditions include resource constraints, time constraints and quality constraints, and the external constraint conditions include grid power consumption restrictions and environmental restrictions. The internal constraint conditions are determined by the corresponding parameters of each device itself and the needs of the enterprise, and the external constraint conditions are the special specifications and conditions of the external working hours of the enterprise. Only by meeting these two categories of constraint conditions can each device operate stably.

[0010] Preferably, when dealing with missing values in step S2, data quality inspection is first carried out. In the case where missing values are detected, if the number of missing values is less than the set threshold, the mean method or the median method is used for supplementation; if the number of missing values is not less than the set threshold, the K-nearest neighbor algorithm is used to predict each missing value for supplementation. Using different methods to supplement missing values according to the number of missing values can improve the accuracy of the data, and thus increase the accuracy in subsequent prediction.

[0011] Preferably, when removing outliers in step S2, the interquartile range method is used for detection, and any data less than Q1 - 1.5×IQR or greater than Q3 + 1.5×IQR is regarded as an outlier and removed. Detecting and removing outliers using the interquartile range method effectively reduces the impact of outliers on subsequent analysis and prediction, improves the quality of the data. At the same time, the interquartile range method is simple and intuitive to calculate and is suitable for large-scale data processing.

[0012] Preferably, when performing standardization processing in step S2, the deviation of each data is determined based on the Z-Score standardization formula, and then each data is controlled within the same scale size based on each deviation; among them, the Z-Score standardization formula is: Z = (X - μ) / σ, where X is the original data, μ is the mean, and σ is the standard deviation. Controlling each data within the same scale size eliminates the dimensional difference between data of different sizes, facilitating data calculation and modeling.

[0013] Preferably, the data fusion in step S2 includes attribute-level fusion and record-level fusion; among them, attribute-level fusion includes merging the data corresponding to each same basic parameter collected by different data acquisition modules; record-level fusion uses the key identification method to fuse multiple data with the same key identification collected by different data acquisition modules to form an energy entity record. Attribute-level fusion and record-level fusion ensure that data from different data sources can be effectively integrated, reducing data redundancy and inconsistency problems.

[0014] Preferably, when using data mining techniques to extract multiple key features in step S3, first use feature engineering to extract each potential feature from a specific dataset; then, for each linearly correlated potential feature, use correlation analysis or principal component analysis or mutual information method to screen out each corresponding key feature related to energy use. For each non-linearly correlated potential feature, use the feature importance evaluation method based on the tree model or the attention mechanism in deep learning to identify each corresponding key feature. Key features can be screened through correlation analysis, principal component analysis, and mutual information method, ensuring the effectiveness and representativeness of the features. For non-linearly correlated potential features, methods such as the attention mechanism in deep learning are used to further improve the accuracy and comprehensiveness of feature extraction.

[0015] Preferably, the machine learning algorithm based on time series in step S3 is an ARIMA (Autoregressive Integrated Moving Average) model or an LSTM (Long Short-Term Memory) network model. Both different machine algorithm models can capture complex energy usage patterns over time.

[0016] Preferably, the energy demand results in step S4 include multiple characteristic parameters, and the multiple characteristic parameters at least include energy cost parameters, energy efficiency parameters, time parameters, storage capacity parameters, and power parameters; when obtaining the initial scheduling plan in step S5, each characteristic parameter in the energy demand results is first encoded to generate multiple types of chromosomes; then, a fitness function is defined, and the fitness of each chromosome under each type is calculated through the objective function. Each chromosome is updated through selection operations, crossover operations, and mutation operations, and the chromosome with the highest fitness under each type is respectively used as the optimal chromosome. Multiple optimal chromosomes form the optimal scheduling plan. Generating the initial scheduling plan through a genetic algorithm or a particle swarm optimization algorithm and adjusting it according to the constraint conditions ensures the optimality and feasibility of the scheduling plan. The design of the fitness function improves the solution efficiency and result quality. When considering multiple parameters, multiple parameters such as energy cost, efficiency, time, storage capacity, and power are comprehensively considered, ensuring the comprehensiveness and rationality of the scheduling plan.

[0017] The beneficial effects of the present invention compared with the prior art are as follows: The technical solution of the present invention collects various energy data through the Internet of Things technology, combines multi-source heterogeneous data such as historical energy consumption, equipment operation parameters, environmental factors, and production plans to construct a complete data set, comprehensively reflects the enterprise's energy usage situation, and realizes the comprehensive utilization of multi-source heterogeneous data; also uses advanced machine learning algorithms for energy demand prediction, combines intelligent optimization algorithms such as genetic algorithms or particle swarm optimization algorithms to generate the optimal scheduling plan, ensuring prediction accuracy and scheduling efficiency; at the same time, the actual operation data is fed back into the model through the system feedback module, and incremental learning or online learning technology is used to dynamically update the model, realizing the continuous optimization and adaptability of the model, ensuring the accuracy of long-term operation, and realizing closed-loop feedback and continuous optimization; in addition, it is intuitively displayed through a graphical interface, facilitating decision-makers to analyze and check. Description of the Drawings

[0018] Figure 1 It is a flowchart of an algorithm for enterprise energy usage prediction and scheduling based on big data analysis; Figure 2 It is a schematic framework diagram of a system corresponding to an algorithm for enterprise energy usage prediction and scheduling based on big data analysis. Detailed Embodiments

[0019] Next, the appendix in the embodiments of the present invention will be combinedFigure 1 and 2 , a detailed description of the technical solutions in the embodiments of the present invention will be given.

[0020] As Figure 1 shown, it is a flowchart of an algorithm for enterprise energy use prediction and scheduling based on big data analysis; as Figure 2 shown, it is a schematic framework diagram of a system corresponding to an algorithm for enterprise energy use prediction and scheduling based on big data analysis; combining Figure 1 and Figure 2 shown, when performing prediction and scheduling, a dataset of multi-source heterogeneous data is obtained through the Internet of Things, comprehensively reflecting the enterprise's energy use situation; machine learning algorithms are also used for prediction, and combined with intelligent optimization algorithms to generate an optimal scheduling plan to ensure accuracy and efficiency; at the same time, a graphical interface is used for intuitive display; finally, a system feedback module is used for feedback to dynamically update the model and achieve closed-loop feedback and continuous optimization. The method specifically includes the following steps: S1. In multiple data acquisition modules, multiple basic parameters of the energy used within the enterprise are collected through the Internet of Things to form a big data set of multi-source heterogeneity. The multiple basic parameters include historical energy consumption data, equipment operation parameters, environmental factor data, and production plan data; at the same time, the constraint conditions for using energy are recorded. The historical energy consumption data is the energy data consumed by each device within the collected time period; the equipment operation parameters include data such as equipment power, operation duration, and the number of equipment; the environmental factor data includes temperature, humidity, and light intensity, etc., and these data mainly affect the operation status of climate control devices such as air conditioners; the production plan data includes expected output or the number of equipment expected to be used, etc. These basic parameters are the main factors affecting the enterprise's energy consumption, and using them as a big data set can ensure the accuracy of subsequent predictions.

[0021] In this embodiment, when the enterprise involves multiple devices using water and electricity, in step S1, the Internet of Things at least includes multiple flow meters, multiple pressure sensors, multiple intelligent water meters, multiple power monitors, a communication module, a central management platform, and a gateway. Each flow meter is used to obtain flow data, each pressure sensor is used to obtain pressure data, each intelligent water meter is used to obtain water consumption, and each power monitor is used to obtain power consumption data; the communication module receives the corresponding data of each flow meter, each pressure sensor, each intelligent water meter, and each power monitor through the gateway and stores them in the central management platform, and the central management platform is used to transmit the data stored by itself to the data processing module. Among them, the flow data, pressure data, water consumption, and power consumption data all belong to the equipment operation parameters. Transmitting data through the Internet of Things composed of multiple monitoring devices can effectively obtain the data related to each device, and thus quickly transmit the corresponding data.

[0022] In this embodiment, the constraint conditions in step S1 include both internal constraint conditions and external constraint conditions. The internal constraint conditions include resource constraints, time constraints, and quality constraints. For example, the total amount of resources available within an enterprise, the construction period requirements, and the load rate requirements for equipment, etc.; the external constraint conditions include power grid power consumption restrictions and environmental restrictions. For example, the power grid standards stipulated by the state and the temperature of environmental restrictions, etc. The internal constraint conditions are determined by the corresponding parameters of each device itself and the requirements of the enterprise, while the external constraint conditions are the special specifications and conditions outside the enterprise. Only by meeting these two categories of constraint conditions can each device operate stably.

[0023] S2. The data processing module includes a data cleaning unit and a data fusion unit. The data processing module is used to clean and fuse the large dataset in step S1, including: using the data cleaning unit to perform missing value processing, outlier removal, standardization processing, and unified timestamp in sequence; using the data fusion unit to fuse the historical energy consumption data, equipment operation parameters, environmental factor data, and production plan data that have completed data cleaning in the large dataset to form a specific dataset.

[0024] In this embodiment, when performing missing value processing in step S2, first perform data quality inspection. In the case where missing values are detected, if the number of missing values is less than the threshold set in advance by humans, it means that the number of missing values is small and has little impact on the overall distribution. The mean method or median method can be used for supplementation; if the number of missing values is not less than the set threshold, at this time the number of missing values is large. Therefore, the K-Nearest Neighbors algorithm (also known as the KNN algorithm) can be used to predict each missing value for supplementation, so that each part with missing values can be supplemented more accurately. K in the K-Nearest Neighbors algorithm is a parameter that can be freely defined, indicating the number of nearest neighbors to refer to during prediction. By giving a sample, find its K closest neighbors, and predict the value of this sample based on the values of these neighbors, thereby completing the supplementation of missing values. Using different methods to supplement missing values according to the number of missing values can improve the accuracy of the data, and further increase the accuracy during subsequent prediction.

[0025] In this embodiment, when removing outliers in step S2, the IQR interquartile range method is used for detection. Any data that is less than Q1 - 1.5×IQR or greater than Q3 + 1.5×IQR is regarded as an outlier and removed. Q1 refers to the first quartile, that is, the position where 25% of the data points are below it, and Q3 is the third quartile, that is, the position where 75% of the data is below it. Using the interquartile range method to detect and remove outliers effectively reduces the impact of outliers on subsequent analysis and prediction, improves the quality of the data, and at the same time, the interquartile range method is simple and intuitive to calculate and is suitable for large-scale data processing.

[0026] In this embodiment, when performing the standardization process in step S2, the deviation of each data is determined based on the Z-Score standardization formula, that is, the Z-score standardization formula, where Z is the standardized value; then each data is controlled within the same scale size based on each deviation; among them, the Z-Score standardization formula is: Z = (X - μ) / σ, where X is the original data, μ is the mean, and σ is the standard deviation. Controlling each data within the same scale size eliminates the dimensional difference between data of different sizes and facilitates data calculation and modeling.

[0027] In this embodiment, the data fusion in step S2 includes attribute-level fusion and record-level fusion; among them, attribute-level fusion includes merging the data corresponding to each same basic parameter collected by different data acquisition modules; record-level fusion uses the key identification method to perform record fusion on multiple data with the same key identification collected by different data acquisition modules to form an energy entity record. The attribute-level and record-level fusion methods ensure that data from different data sources can be effectively integrated, reducing data redundancy and inconsistency problems.

[0028] It should be noted that multiple data acquisition modules are used to collect a large number of devices in the enterprise. Each device will be collected by at least 2 data acquisition modules, which is equivalent to at least two departments in the enterprise for equipment monitoring to improve accuracy. For these devices collected by at least 2 data acquisition modules, when performing attribute-level fusion and record-level fusion, in attribute-level fusion, for example, the attributes such as equipment operation time and power in the equipment energy consumption data recorded by different departments within the enterprise are integrated, and the integration method can be the weighted average method to ensure the integrity and accuracy of the equipment energy consumption information. In record-level fusion, based on key identifiers such as equipment ID or timestamp, the records of the same equipment from different data sources can be merged to form corresponding entity records.

[0029] During the fusion process, it is also necessary to solve the data conflict problem. According to factors such as different accuracies and timeliness of different data acquisition modules, a conflict resolution strategy is formulated to ensure the quality of the fused data. For example, when there are differences in the energy consumption data of the same equipment between two data acquisition modules, if data acquisition module A has higher accuracy than B and the acquisition time of A is closer to the current moment, then the data of data acquisition module A is preferentially adopted.

[0030] S3. The prediction model construction module includes a feature extraction unit and a model training unit. The feature extraction unit is used to extract multiple key features from the specific data set in step S2 using data mining techniques, and then transmit the multiple key features to the model training unit as model inputs to train the energy usage prediction model through a machine learning algorithm based on time series.

[0031] In this embodiment, when extracting multiple key features using data mining techniques in step S3, each potential feature is first extracted from a specific dataset using feature engineering. Then, for each linearly correlated potential feature, each corresponding key feature related to energy use is screened out using correlation analysis, principal component analysis, or mutual information method. For each non-linearly correlated potential feature, each corresponding key feature is identified using the feature importance evaluation method based on a tree model or the attention mechanism in deep learning. Through correlation analysis, principal component analysis, or mutual information method, key features can be screened out, ensuring the effectiveness and representativeness of the features. For non-linearly correlated potential features, the attention mechanism in deep learning or the feature importance evaluation method based on a tree model is used to further improve the accuracy and comprehensiveness of feature extraction.

[0032] In this embodiment, the machine learning algorithm based on time series in step S3 is the ARIMA autoregressive integrated moving average model or the LSTM long short-term memory network model. Both different machine algorithm models can capture complex energy use patterns over time.

[0033] Taking the LSTM long short-term memory network model as an example, a network structure including an input layer, multiple LSTM hidden layers, and an output layer is constructed in sequence. The input layer is used to input each key feature. Each LSTM hidden layer consists of multiple LSTM units, and these units can learn the time series features and long-term dependencies in the data, so as to perform calculations based on each key feature and finally output at the output layer. During the training process, set the hyperparameters of the model. For example, the initial learning rate is set to 0.01 - 0.001, the number of iterations is set to 100 - 500 times, and the number of neurons in the hidden layer is 64 or 128. The specific values here can be set and adjusted by relevant staff according to the actual scenario. Then, use the backpropagation algorithm to calculate the loss function, such as the mean square error loss function , where yi is the true energy consumption value, y ^ i is the predicted value, n is the number of samples, and update the model parameters according to the opposite direction of the gradient, that is, the direction in which the loss function drops fastest, and continuously adjust the model to minimize the loss function, so that the model can better fit the historical energy consumption data and learn the energy consumption patterns and rules.

[0034] When performing model training, test according to the division of 70% training set and 30% test set, and use the test set to evaluate the trained model. Evaluation metrics can use root mean square error RMSE, mean absolute error MAE, mean absolute percentage error MAPE, etc. The RMSE calculation formula is , reflecting the average magnitude of the deviation between the predicted value and the true value; the MAE calculation formula is , measuring the average absolute value of the error between the predicted value and the true value; the MAPE calculation formula is , representing the relative percentage of the prediction error. When evaluating these evaluation metrics, if the model performance does not meet the expectations, the model can be optimized by adjusting the model structure or reselecting hyperparameters or increasing the amount of training data, etc.; adjusting the model structure, such as increasing or decreasing the number of LSTM hidden layers, adjusting the number of neurons in the hidden layer; reselecting hyperparameters, such as adjusting the initial value of the learning rate and the number of optimization iterations.

[0035] S4. The energy usage prediction module, based on the energy usage prediction model completed in training in step S3, inputs the future production plan and future environmental parameters, and predicts the future energy demand results of the enterprise.

[0036] In this embodiment, the energy demand results in step S4 include multiple characteristic parameters, and the multiple characteristic parameters at least include energy cost parameters, energy efficiency parameters, time parameters, storage capacity parameters, and power parameters. When considering multiple parameters, multiple parameters such as energy cost, efficiency, time, storage capacity, and power are comprehensively considered, ensuring the comprehensiveness and rationality of the scheduling plan.

[0037] S5. The scheduling optimization module includes a constraint condition setting unit, an optimization target setting unit, and a scheduling algorithm unit. First, in the optimization target setting unit, the staff sets corresponding data targets, such as target energy cost, target energy utilization efficiency, and target duration, etc.; then, according to the energy demand results predicted in step S4, in the scheduling algorithm unit, an intelligent optimization algorithm is used to generate an initial scheduling plan, where the intelligent optimization algorithm is a genetic algorithm or a particle swarm algorithm; then, the constraint conditions in step S1 stored in the constraint condition setting unit and the possible externally input constraint conditions are used to verify the initial scheduling plan, and the parts of the initial scheduling plan that do not meet the constraint conditions are adjusted to obtain the optimal scheduling plan.

[0038] When obtaining the initial scheduling plan, first encode each characteristic parameter in the energy demand result to generate multiple types of chromosomes; then, define the fitness function and calculate the fitness of the chromosomes under each type through the objective function. For example, the objective function is to minimize the energy cost or maximize the energy efficiency, and then perform subsequent operations through the fitness function. For example, perform selection operations, crossover operations, and mutation operations in sequence to update each chromosome, and take the chromosome with the highest fitness under each type as the optimal chromosome respectively. Multiple optimal chromosomes form the optimal scheduling plan. Generating the initial scheduling plan through genetic algorithms or particle swarm algorithms and adjusting it according to the constraint conditions ensures the optimality and feasibility of the scheduling plan, and the design of the fitness function improves the solution efficiency and result quality. Among them, the selection operation can be roulette wheel selection, where the probability of each chromosome being selected is proportional to its fitness, or tournament selection, where the chromosome with the highest fitness is selected from several randomly selected individuals; the crossover operation can be single-point crossover or multi-point crossover, both of which can exchange some genes of two different chromosomes; the mutation operation can be randomly changing the value of a certain gene on some chromosomes.

[0039] In addition, for the optimal scheduling plan, there is also decision-level fusion, which fuses the decision results obtained from multiple data sources through different models or algorithms. For example, in energy scheduling decisions, the energy demand results predicted based on machine learning and the energy demand results obtained from production plan analysis are integrated to obtain a more accurate basis for energy scheduling decisions.

[0040] S6. In the visualization and decision support module, the optimal scheduling plan in step S5 and the energy demand result in step S4 can be respectively displayed through different graphical interfaces. For example, display the bar chart of the number of devices to be run in a specific time period accounting for the total number of devices, and compare and display the predicted operating power of each different device to be run, etc.

[0041] S7. After running according to the optimal scheduling plan shown in step S6, the system feedback module transmits the data during actual operation as feedback data to the energy usage prediction model, calculates the error index between the feedback data and the predicted energy demand result, and the error index is used to represent the model accuracy to the enterprise; then retrain the energy usage prediction model based on the feedback data, using incremental learning or online learning techniques to dynamically update the energy usage prediction model. Incremental learning is a machine learning method suitable for application scenarios that need to be continuously updated. The online learning technique can be the stochastic gradient descent method to update the data in small batches, making the model pay more attention to the actual data.

[0042] In summary, the present invention collects various energy data through Internet of Things technology, combines multi-source heterogeneous data such as historical energy consumption, equipment operation parameters, environmental factors, and production plans to construct a complete data set, comprehensively reflects the enterprise's energy usage situation, and realizes the comprehensive utilization of multi-source heterogeneous data; it also uses advanced machine learning algorithms for energy demand prediction, combines intelligent optimization algorithms such as genetic algorithms or particle swarm algorithms to generate an optimal scheduling plan to ensure prediction accuracy and scheduling efficiency; at the same time, through the system feedback module, the actual operation data is fed back into the model, and incremental learning or online learning technology is used to dynamically update the model, realizing the continuous optimization and adaptability of the model, ensuring the accuracy of long-term operation, and realizing closed-loop feedback and continuous optimization; in addition, it is intuitively displayed through a graphical interface, facilitating decision-makers to analyze and check, and has significant progressiveness.

[0043] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. An algorithm for enterprise energy usage prediction and scheduling based on big data analysis, characterized in that, It includes the following steps: S1. In multiple data acquisition modules, multiple basic parameters of the energy used within the enterprise are collected through the Internet of Things to form a multi-source heterogeneous big data set. The multiple basic parameters include historical energy consumption data, equipment operation parameters, environmental factor data, and production plan data. At the same time, the constraint conditions for using energy are recorded; S2. In the data processing module, the big data set in step S1 is cleaned and fused, including: using the data cleaning unit to perform missing value processing, outlier removal, normalization processing, and unified timestamp in sequence; using the data fusion unit to fuse the historical energy consumption data, equipment operation parameters, environmental factor data, and production plan data that have completed data cleaning in the big data set to form a specific data set; S3. In the prediction model construction module, multiple key features are extracted from the specific data set in step S2 using data mining techniques, and then the multiple key features are used as inputs to train an energy usage prediction model through a machine learning algorithm based on time series; S4. The energy usage prediction module, based on the energy usage prediction model completed in training in step S3, inputs future production plans and environmental parameters to predict the future energy demand results of the enterprise; S5. In the scheduling optimization module, first, according to the energy demand results predicted in step S4, an initial scheduling plan is generated using an intelligent optimization algorithm, where the intelligent optimization algorithm is a genetic algorithm or a particle swarm algorithm; then, the constraint conditions in step S1 are used to verify the initial scheduling plan, and the parts of the initial scheduling plan that do not meet the constraint conditions are adjusted to obtain an optimal scheduling plan; S6. In the visualization and decision support module, the optimal scheduling plan in step S5 and the energy demand results in step S4 are displayed through a graphical interface; S7. After running according to the optimal scheduling plan displayed in step S6, the system feedback module transmits the data during actual operation as feedback data to the energy usage prediction model, calculates the error index between the feedback data and the predicted energy demand results, and the error index is used to represent the model accuracy to the enterprise; then, based on the feedback data, the energy usage prediction model is retrained, and incremental learning or online learning techniques are used to dynamically update the energy usage prediction model.

2. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, In step S1, the Internet of Things at least includes multiple flow meters, multiple pressure sensors, multiple intelligent water meters, multiple power monitors, a communication module, a central management platform, and a gateway; each flow meter is used to obtain flow data, each pressure sensor is used to obtain pressure data, each intelligent water meter is used to obtain water consumption, and each power monitor is used to obtain power consumption data; the communication module receives the corresponding data of each flow meter, each pressure sensor, each intelligent water meter, and each power monitor through the gateway and stores them in the central management platform, and the central management platform is used to transmit the data stored by itself to the data processing module; among them, the flow data, pressure data, water consumption, and power consumption data all belong to the equipment operation parameters.

3. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, The constraint conditions in step S1 all include internal constraint conditions and external constraint conditions. The internal constraint conditions include resource constraints, time constraints, and quality constraints, and the external constraint conditions include power grid power consumption restrictions and environmental restrictions.

4. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, When dealing with missing values in step S2, first perform data quality inspection. In the case where missing values are detected, if the number of missing values is less than the set threshold, the mean method or median method is used for supplementation; if the number of missing values is not less than the set threshold, the K-nearest neighbor algorithm is used to predict each missing value for supplementation.

5. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, When removing outliers in step S2, the interquartile range method is used for detection, and any data that is less than Q1 - 1.5×IQR or greater than Q3 + 1.5×IQR is regarded as an outlier and removed.

6. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, When performing normalization processing in step S2, the deviation of each data is determined based on the Z-Score normalization formula, and then each data is controlled at the same scale size based on each deviation; among them, the Z-Score normalization formula is: Z = (X - μ) / σ, where X is the original data, μ is the mean, and σ is the standard deviation.

7. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, The data fusion in step S2 includes attribute-level fusion and record-level fusion; Among them, attribute-level fusion includes merging the data corresponding to each same basic parameter collected in different data acquisition modules; record-level fusion uses the key identification method to perform record fusion on multiple data with the same key identification collected in different data acquisition modules to form an energy entity record.

8. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, When using data mining technology to extract multiple key features in step S3, first use feature engineering to extract each potential feature from a specific dataset; then, for each potentially linearly correlated feature, use correlation analysis or principal component analysis or mutual information method to screen out each corresponding key feature related to energy use, and for each potentially non-linearly correlated feature, use the feature importance evaluation method based on the tree model or the attention mechanism in deep learning to identify each corresponding key feature.

9. An algorithm for enterprise energy use prediction and scheduling based on big data analysis according to claim 1, characterized in that, The machine learning algorithm based on time series in step S3 is the ARIMA autoregressive integrated moving average model or the LSTM long short-term memory network model.

10. An algorithm for enterprise energy usage prediction and scheduling based on big data analysis according to claim 1, characterized in that, The energy demand results in step S4 include multiple characteristic parameters, and the multiple characteristic parameters at least include energy cost parameters, energy efficiency parameters, time parameters, storage capacity parameters, and power parameters; When obtaining the initial scheduling plan in step S5, first encode each characteristic parameter in the energy demand results to generate multiple types of chromosomes; Then, define the fitness function, calculate the fitness of the chromosomes under each type through the objective function, update each chromosome through selection operation, crossover operation, and mutation operation, and take the chromosome with the highest fitness under each type as the optimal chromosome respectively. The multiple optimal chromosomes form the optimal scheduling plan.

Citation Information

Patent Citations

  • Multi-source power dispatching method and system based on big data

    CN117767446A

  • Multi-source integrated energy system energy-saving optimization method based on data mining

    CN119250307A

  • New energy intelligent distribution and prediction system based on big data optimization

    CN119448426A

  • Energy demand prediction system based on big data analysis

    CN119558472A

  • System and method facilitating forecasting, optimization and visualization of energy data for an industry

    WO2013102932A2

Cited By

  • Multi-level energy management system based on multi-dimensional data

    CN120494450A