Method for constructing power load accurate prediction model by using big data
Through the construction method of the large data power load accurate prediction model, the deep learning model is used to process complex power load data, which solves the problems of large errors and insufficient computing capabilities of traditional prediction methods, and realizes high-precision load prediction and optimization management of power systems.
Patent Information
- Application Number
- CN202510203445.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Traditional power load prediction methods are difficult to capture complex nonlinear factors and user behavior diversity, resulting in huge prediction errors and limited computing power, making it difficult to process massive multi-source data, resulting in insufficient accuracy and timeliness.
The power load accurate prediction model construction method of big data is adopted, real-time data is collected through IoT devices, and deep learning models are used to combine LSTM and CNN to process time series data and extract hidden modes, and feature engineering and model construction are carried out to achieve efficient data storage and processing.
It realizes high-precision power load prediction, optimizes power resource allocation, reduces operating costs, improves power supply reliability, enhances user satisfaction, and supports the stable operation of the power system and energy conservation and emission reduction.
Smart Images

Figure CN120031335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power load prediction, and in particular to a method for constructing a precise power load prediction model using big data. Background Art
[0002] In today's society, the stable operation of the power system plays a vital role in the normal operation of economic development and social life. With the acceleration of industrialization and urbanization, the demand for electricity has shown explosive growth, and the fluctuation of power load has become increasingly complex. Traditional power load forecasting methods can no longer meet the needs of refined management of modern power systems.
[0003] On the one hand, early power load forecasting mostly relied on simple statistical analysis of historical data, such as linear regression-based methods, which only considered the roughly linear trend of power consumption over time. However, the actual power load is affected by the interaction of many factors and cannot be summarized by a simple linear relationship. The changes in temperature and humidity brought about by the change of seasons will greatly change the power consumption patterns of residents and enterprises. In the hot summer, refrigeration equipment is fully turned on, and in the cold winter, heating appliances are frequently used. These nonlinear factors make the traditional linear prediction model have huge errors.
[0004] Furthermore, previous prediction methods lacked consideration of the diversity of user behavior. For users in different industries, industrial production has off-season and peak season, and its electricity load fluctuates greatly with orders and production process adjustments; commercial users' electricity demand soars during holidays and promotional activities; residential users have different living habits, such as nine-to-five office workers and those who work from home, and their peak electricity consumption hours are completely different. Traditional methods cannot accurately capture this complex electricity consumption characteristic on the user side, resulting in prediction results that deviate from reality.
[0005] In addition, at the data collection level, the past technical means were scattered and inefficient. Smart meters have not yet been popularized, and they mostly rely on manual meter reading. The data update frequency is low, making it difficult to obtain real-time electricity consumption data. In addition, the data between different departments and systems are isolated, and meteorological data, user information and power operation data cannot be effectively integrated, which cannot provide a comprehensive and accurate data source for load forecasting.
[0006] From the perspective of technical architecture, traditional prediction models have limited computing power and are difficult to process massive, multi-source data. Emerging technologies such as deep learning have been slow to penetrate the power sector. Faced with the flood of data in the big data era, there is a lack of efficient data storage, processing and model building mechanisms, and it is impossible to deeply mine the value of data, which greatly reduces the accuracy and timeliness of power load forecasting, seriously restricting the optimization and scheduling of power systems, energy saving and loss reduction, and the improvement of power supply reliability. A new accurate prediction solution based on big data is urgently needed to break the deadlock. Summary of the invention
[0007] The present invention proposes a method for constructing an accurate prediction model of power load using big data to solve the problems mentioned in the above-mentioned prior art.
[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: a method for constructing an accurate prediction model of power load using big data, comprising the following steps: Data collection steps: Use IoT devices to collect real-time operation data at each node of the power system at intervals of 15 minutes, use cleaning algorithms to remove erroneous data, and then transmit the processed data to the data storage link through a secure encrypted channel; suppose the data transmission packet loss rate formula is , the packet loss rate is required to be controlled below 0.1%, where LP represents the number of lost packets, and TP represents the total number of packets; Data storage steps: Distributed storage architecture is used to classify and store the collected data. The Ceph storage system uses the CRUSH algorithm for data distribution. The retrieval efficiency improvement formula is: , it is expected that the value is not less than 80%, where ORT represents the original retrieval time and NRT represents the new retrieval time. Feature engineering steps: retrieve information from the stored data, and filter out features that are strongly correlated with power load by calculating the Pearson correlation coefficient. Suppose the feature correlation strength formula is , select features with an absolute value greater than 0.5 so that the data value range is between [0,1], and after processing, transfer the data to the model building stage; Model construction steps: Build a prediction model based on deep learning, a hybrid model that combines the long short-term memory network LSTM with the convolutional neural network CNN, where LSTM is used to process time series data and CNN is used to extract hidden patterns in the feature space. During the training process, the model convergence speed formula is set to , it is expected to reach convergence within 80% of the training rounds, where TE represents the total training rounds of TotalEpochs and CE represents the current training round of CurrentEpoch; the model is trained using the training set until the root mean square error RMSE index of the model on the test set is controlled within 0.1. The RMSE calculation formula is ,in is the true value, is the predicted value, n is the number of samples; Model evaluation and update steps: Use the collected data to evaluate the model, use the mean absolute error (MAE) and mean absolute percentage error (MAPE) indicators to measure the performance of the model, and also use the mean absolute percentage error (MAPE) indicator to measure the performance of the model. The MAE calculation formula is: , the MAPE calculation formula is: ,in is the true value, is the predicted value, n is the number of samples; Forecast application steps: deploy the constructed accurate forecast model to the power dispatching center to predict the power load in different periods in the future; present the forecast results to the dispatching personnel in the form of visual charts, and the dispatching personnel will formulate power generation plans and power allocation plans based on the forecast results.
[0009] Furthermore, it also includes: Data fusion verification step: After the data collection step is completed, the data obtained from different data sources are fused. For data with the same attributes, the weighted average method is used for integration. For the current data of the same area collected from different smart meters, weights are assigned according to the meter accuracy. The weight calculation formula is as follows: ,in The accuracy of the th meter is measured, and the data consistency verification algorithm is used at the same time.
[0010] Furthermore, it also includes: Feature dimension reduction optimization step: After selecting features in the feature engineering step, the principal component analysis (PCA) algorithm is used to reduce the dimension of the features. The variance contribution rate threshold of the retained principal component is set to 90%. The formula for the retention rate of the principal component is: ,in is the variance of the th principal component, k is the number of retained principal components, and n is the number of original features. Then the random forest algorithm is used to screen the reduced-dimensional features again to remove redundant features.
[0011] Furthermore, it also includes: Model adaptive adjustment step: During the model building step, if the model shows signs of overfitting after a certain number of training rounds, including a continuous increase in the validation set loss value, the learning rate is reduced and the L2 regularization coefficient is increased. The adjustment coefficient formula is set as The learning rate adjustment range is expected to be between 50%-80%, and the L2 regularization coefficient increase range is expected to be between 20%-50%, where OV represents the original value of OriginalValue and NV represents the new value of NewValue; on the contrary, if it is underfitting, increase the learning rate so that the model can adapt to different data distributions.
[0012] Furthermore, in the model evaluation and update step, a user feedback collection mechanism is constructed to quantify the feedback information into weights and integrate it into the model performance evaluation index system. The feedback weight calculation formula is: , where SS represents the SeverityScore, IS represents the ImportanceScore, and TS represents the TotalScore. The weight is used to adjust the proportion of the model evaluation indicators.
[0013] Furthermore, in the prediction application step, when presenting the prediction results to the dispatcher, interactive operations are supported and a comparison function is provided.
[0014] Further, including: Data acquisition and transmission subsystem: responsible for executing data acquisition steps and data storage steps; Model training management subsystem: undertakes the responsibilities of feature engineering step, model building step, and model evaluation and update step in claim 1, and builds, optimizes and maintains the prediction model; Prediction result application subsystem: corresponds to the prediction application step in claim 1, providing the prediction results to the power dispatching personnel in a visualized form for formulating power dispatching strategies.
[0015] Furthermore, the data storage between the data storage center in the data acquisition and transmission subsystem and the model training management subsystem adopts a high-speed optical fiber channel with a bandwidth of no less than 10Gbps. The transmission delay reduction formula is: , it is expected that the value is not less than 60%, where OD represents the original delay and ND represents the new delay.
[0016] Furthermore, when the model training management subsystem encounters insufficient computing resources during model construction, it automatically connects to the cloud computing platform and optimizes resource allocation. The cost reduction rate formula is set as , requiring the value to reach or exceed 20%, where OC represents the original cost and NC represents the new cost.
[0017] Furthermore, the prediction result application subsystem uses virtual reality VR technology when presenting the prediction results, so that the power dispatching personnel can intuitively feel the trend of power load changes. Suppose the decision efficiency improvement formula is , it is expected that the value is not less than 30%, where ODT represents the original decision time and NDT represents the new decision time.
[0018] Compared with the prior art, the present invention has the following beneficial effects: First, for the operation of the power system, high-precision load forecasting helps to formulate accurate power generation plans, avoid energy waste caused by excess power generation or power outage risks caused by insufficient power generation, optimize power resource allocation, and reduce operating costs. Accurate forecasting allows dispatchers to arrange unit start-up and shutdown and adjust power generation output in advance to ensure the balance of power supply and demand and improve system stability.
[0019] Secondly, the economic benefits of power companies are significantly improved. By accurately grasping the load changes and reasonably arranging the equipment maintenance period, the impact of power outages on users can be reduced, the service quality can be improved, and user satisfaction can be enhanced, thereby consolidating market share. At the same time, unnecessary standby power generation capacity can be reduced, investment costs can be reduced, and capital utilization can be improved.
[0020] Furthermore, from the perspective of energy conservation and emission reduction, accurate prediction can avoid excessive energy production and transmission losses, and help achieve the goals of carbon peak and carbon neutrality. Based on the prediction results, the power grid flow distribution can be optimized, line losses can be reduced, and energy utilization efficiency can be improved.
[0021] Finally, it brings users a high-quality electricity experience. Stable power supply reduces the risk of electrical equipment being damaged due to voltage fluctuations, ensures the continuity of corporate production, and protects residents from power outages, creating a good electricity environment for social production and life. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A schematic block diagram of a method for constructing an accurate prediction model of power load using big data proposed by the present invention; Figure 2 This is a schematic block diagram of a system for building an accurate power load prediction model using big data proposed by the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0025] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, and it can be the internal connection of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0026] Reference Figure 1-2 A method for constructing an accurate prediction model of power load using big data, characterized in that it includes the following steps: Data collection steps: Use smart meters, sensors and other IoT devices to collect real-time operating data such as voltage, current, and power factor at high frequency at intervals of 15 minutes at each node in the power system. At the same time, establish connections with the power company's customer information system and the meteorological department's database to obtain multi-source data such as user type, electricity usage habits, weather conditions, temperature, humidity, wind speed, etc. Smart meters use high-precision chips with a voltage measurement accuracy of ±0.5%, a current accuracy of ±0.2%, and a power factor accuracy of ±0.01. Use a rule-based cleaning algorithm to remove duplicate, erroneous, and obviously abnormal data. For example, remove data with voltage values that exceed the normal range of ±15%, and then transmit the processed data to the data storage link through a secure encrypted channel. Suppose the data transmission packet loss rate formula is , requiring the packet loss rate to be controlled below 0.1% to ensure complete data transmission, where LP stands for LostPackets (the number of lost packets) and TP stands for TotalPackets (the total number of packets).
[0027] Data storage steps: In order to efficiently store and manage these massive amounts of data, we use a distributed storage architecture. Among them, the Ceph distributed file system is a good choice. Through its unique design, Ceph can establish indexes based on multiple dimensions such as data type, region, and time. For example, for data types, it can be subdivided into structured data (such as table data in the database), semi-structured data (such as data in XML and JSON formats), and unstructured data (such as pictures, videos, documents, etc.). For the geographical dimension, it can be divided according to the geographical location where the data is generated, which is particularly important in data management for multinational companies or those with multiple branches. The time dimension can be indexed according to the time sequence of data generation, which facilitates the management and query of data in different time periods. Through this multi-dimensional index establishment, the collected data is classified and stored, which can ensure that the system can achieve fast retrieval within 2 seconds when processing data query requests. This fast retrieval capability is crucial for business scenarios that require instant data acquisition, such as real-time data analysis in financial transactions and user behavior analysis on e-commerce platforms. The Ceph storage system uses the CRUSH algorithm for data distribution. In the actual data storage process, the system will comprehensively consider factors such as storage device performance and network topology to intelligently allocate data. Specifically, for storage device performance, the system monitors the read and write speed, storage capacity, I / O load and other indicators of each storage node, and prioritizes data allocation to nodes with better performance and lower load. For network topology, the system considers factors such as network latency and bandwidth between nodes to ensure data transmission efficiency during storage and retrieval. In order to measure the improvement in retrieval efficiency, the retrieval efficiency improvement formula is , the value is expected to be no less than 80%, where ORT represents OriginalRetrievalTime and NRT represents NewRetrievalTime; at the same time, in order to ensure the integrity and reliability of the data, we will also use data backup strategies. In actual operation, a combination of regular full backup and real-time incremental backup is adopted. Regular full backup means that all data is fully backed up at fixed time intervals (such as weekly and monthly) to ensure that the complete data state can be restored to a certain point in time when a major data disaster occurs. Real-time incremental backup records and backs up data changes in real time (such as adding, modifying, and deleting operations). This can reduce the amount of backup data and ensure that the latest state of the data is protected in a timely manner. This double-insurance backup strategy can maximize the security and availability of data during storage.
[0028] Feature engineering steps: retrieve the required information from the stored data, use statistical methods and professional knowledge in the power field, and calculate the Pearson correlation coefficient to screen out features that are strongly correlated with power load, such as temperature during a specific period of time, historical peak power consumption of different user types, etc. Suppose the feature correlation strength formula is , select features with absolute values greater than 0.5, normalize the selected features so that the data value range is between [0,1], and transmit the data to the model building stage after processing.
[0029] Model construction steps: Build a prediction model based on deep learning, specifically a hybrid model that combines the long short-term memory network (LSTM) with the convolutional neural network (CNN). LSTM is good at processing time series data, and CNN is used to extract hidden patterns in the feature space. Set the hyperparameters of the model, such as setting the number of LSTM units to 128, the learning rate to 0.001, and using the adaptive moment estimation (Adam) optimization algorithm to divide the historical load data into training set, validation set, and test set at a ratio of 80%, 10%, and 10%. During the training process, the model convergence speed formula is set to , it is expected to reach convergence within 80% of the training rounds, where TE represents TotalEpochs (total training rounds) and CE represents CurrentEpoch (current training round); the model is trained using the training set, and the model is adjusted and optimized on the validation set until the root mean square error (RMSE) of the model on the test set is controlled within 0.1, thereby obtaining an accurate prediction model. The RMSE calculation formula is ,in is the true value, is the predicted value, and n is the number of samples.
[0030] Model evaluation and update steps: Regularly, such as weekly, use the latest collected data to evaluate the model. In addition to the root mean square error (RMSE), the mean absolute error (MAE), mean absolute percentage error (MAPE) and other indicators are used to measure the performance of the model. The MAE calculation formula is , the MAPE calculation formula is: When the performance of the model declines, for example, when the RMSE rises by more than 10%, the model update mechanism is triggered, recent data is recollected, feature engineering and model building steps are repeated, and the model is optimized to ensure that the accuracy of the prediction can continue to improve.
[0031] Prediction application steps: Prediction application plays a vital role in the operation and management of power systems. First, the carefully constructed accurate prediction model needs to be deployed to the power dispatching center. The construction process of this prediction model is a complex and rigorous task, involving a large amount of data collection and analysis. The data sources include historical operation data of various links in the power system, such as the power generation of power stations, the load of transmission lines, the voltage and current data of substations, etc., and the impact of external factors such as weather conditions, seasonal changes, holidays, etc. on power demand will also be comprehensively considered. Through advanced data mining technology and machine learning algorithms, these data are deeply analyzed and trained to build a high-precision prediction model. After this prediction model is successfully deployed to the power dispatching center, it will have powerful functions. Its core capability is to be able to receive the current operation data of the power system in real time. This means that it will establish a close connection with various monitoring equipment and data acquisition systems in the power system, and obtain the latest data information every minute and every second, including but not limited to the real-time power generation of each power station, the real-time load of each transmission line, and the real-time voltage and current data of each substation. Based on these real-time received data, the prediction model can then predict the power load at different time periods in the future. For example, it can accurately predict the change of power load every hour in the next 24 hours, and the power demand trend in each period in the next 48 hours. These forecast results are very valuable decision-making basis for power dispatchers. In order to make these forecast results more intuitive, easier to understand and use, they will be presented to dispatchers in the form of visual charts. There are various types of these visual charts, including line charts and bar charts. The line chart can clearly show the change trend of power load over time. Dispatchers can judge whether power demand is rising or falling, and the magnitude of the change by observing the direction of the line. The bar chart can be used to compare the power load situation in different time periods, such as the power demand comparison of the same time period on different dates or different time periods on the same day, to help dispatchers quickly find the peak and trough periods of power load. After receiving these intuitive forecast results, dispatchers will formulate reasonable power generation plans and power allocation plans based on these results. When formulating power generation plans, they will comprehensively consider factors such as the power generation capacity, fuel reserves, and equipment maintenance plans of each power station. For example, if it is predicted that the demand for electricity will increase significantly in a certain period of time in the future, the dispatcher will notify the thermal power station in advance to increase the coal reserve to ensure that there is enough power output during the high-demand period; for the hydroelectric power station, the operation time of the turbine will be reasonably arranged according to the water level of the reservoir and the water inflow forecast. In terms of power allocation, the dispatcher will reasonably adjust the power transmission direction and transmission volume of the transmission line according to the distribution of power load.For example, when it is predicted that the power demand in a certain area will increase sharply, the power distribution of nearby transmission lines will be adjusted in time to transmit more power from areas with more abundant power supply to this area to avoid insufficient power supply. At the same time, it is also necessary to prevent the situation of excess power, because excess power may lead to unnecessary operation of power generation equipment, increase energy waste and operating costs. Through such a complete set of prediction application steps, the stable operation of the power system can be effectively guaranteed, ensuring the reasonable allocation and efficient use of power resources while meeting the power needs of users.
[0032] The present invention also includes a data fusion verification step: after the data collection step is completed, the data obtained from different data sources are fused, and the data with the same attributes are integrated using a weighted average method. For example, for the current data of the same area collected from different smart meters, weights are assigned according to the meter accuracy. The weight calculation formula is: ,in To ensure the accuracy of each electricity meter, a data consistency verification algorithm is used. At the same time, a data consistency verification algorithm is used to check whether there are logical contradictions in the fused data. If it is found that the temperature in a certain period of time is inconsistent with the usual climate laws of that period, the data source is traced back in time to re-collect data to ensure data accuracy and provide a reliable foundation for subsequent model training.
[0033] The present invention also includes a feature dimension reduction optimization step: after the feature engineering step selects the features, the principal component analysis (PCA) algorithm is used to reduce the dimension of the features, converting the original high-dimensional and complex feature space into a low-dimensional and concise feature representation, and setting the variance contribution rate threshold of the principal component to be retained to 90%, while reducing the computational complexity and retaining the information valuable to the power load forecast to the greatest extent. Assume that the principal component retention rate formula is ,in is the variance of the th principal component, k is the number of retained principal components, and n is the number of original features. Then the random forest algorithm is used to screen the reduced-dimensional features again to remove redundant features and further improve the quality of the model input.
[0034] The present invention also includes a model adaptive adjustment step: in the process of training the model in the model building step, after each certain training round (such as 10 epochs), the hyperparameters are dynamically adjusted according to the performance of the current model on the validation set. If the model shows signs of overfitting, such as the validation set loss value continues to rise, the learning rate is reduced and the L2 regularization coefficient is increased. The adjustment coefficient formula is set to be The learning rate adjustment range is expected to be between 50%-80%, and the L2 regularization coefficient increase range is expected to be between 20%-50%, where OV represents OriginalValue and NV represents NewValue. On the contrary, if it is underfitting, the learning rate should be appropriately increased so that the model can adapt to different data distributions and continuously improve the prediction performance.
[0035] In the present invention, in the model evaluation and update step, a user feedback collection mechanism is constructed to collect feedback on the prediction results from power dispatchers and power users, such as whether the prediction accuracy meets the actual needs, whether there is a large deviation in the load prediction results for a certain period of time that is not captured by the model, and the feedback information is quantified into weights and integrated into the model performance evaluation index system to make the model update more suitable for actual application scenarios. Suppose the feedback weight calculation formula is , where SS stands for Severity Score, IS stands for Importance Score, and TS stands for Total Score. The weight is used to adjust the proportion of model evaluation indicators to make model optimization more targeted.
[0036] In the present invention, in the prediction application step, when presenting the prediction results to the dispatcher, a rich interactive operation function is provided. The dispatcher can view the detailed load forecast data of different regions and different user types by clicking and dragging the mouse. Specifically, when viewing the data of different regions, the dispatcher can use the mouse to click on a specific geographical area on the visual interface, such as various subdistricts in the city, different industrial parks, commercial areas or residential areas, and the system will immediately present the detailed power load forecast data of the area. These data include the peak and valley values of power demand in different time periods in the future, as well as the changing trend of power load. At the same time, through the mouse drag operation, the dispatcher can flexibly switch between different areas to view, which is convenient for a comprehensive and detailed analysis of the entire power supply area. For viewing the load forecast data of different user types, the system can distinguish different types such as industrial users, commercial users, and residential users. When the dispatcher clicks the corresponding user type option, the system will display the power usage forecast of this type of user in different time periods in the future. For example, for industrial users, it can display the power demand forecast of enterprises of different sizes, such as large factories and small workshops, in detail, including the power consumption during the peak and trough periods of production; for commercial users, it will present the power load forecast data of places such as shopping malls, office buildings, and hotels, taking into account the impact of business hours, holidays and other factors on power demand; for residential users, it will predict the power usage in different time periods based on the daily living habits of residents, seasonal changes and other factors. In addition, the system also provides a powerful comparison function. This function can take a snapshot of the current forecast results with the same period in history and the same period last week, and compare them. When comparing with the same period in history, the system will automatically obtain the power load data of the same time period in the past few years and present it in the form of intuitive charts. For example, by showing the current predicted power load curve and the power load curve of the same date in the past few years through a line chart, the dispatcher can clearly see the changing trend of power demand and determine whether there is any abnormality. When comparing with the same period last week, the system will quickly obtain the power load data of the same period last week and compare it with the current forecast results. This comparison can help dispatchers understand the changes in power demand in the short term, especially when considering some cyclical factors (such as the difference in power demand between working days and weekends), so as to better formulate corresponding power generation plans and power allocation plans. Through these rich interactive operations and comparison functions, dispatchers can make more accurate decisions, ensure the stable operation of the power system, reasonably allocate power resources, and avoid insufficient or excessive power supply.
[0037] In the present invention, the data acquisition and transmission subsystem is responsible for executing the data acquisition steps and data storage steps to ensure data acquisition, transmission and storage; Model training management subsystem: responsible for feature engineering steps, model building steps, and model evaluation and update steps, and building, optimizing, and maintaining prediction models; Prediction result application subsystem: corresponding to the prediction application steps, the prediction results are provided to the power dispatchers in a visualized form for formulating power dispatch strategies.
[0038] In the present invention, the data storage center in the data acquisition and transmission subsystem and the data storage center in the data acquisition and transmission subsystem and the data transmission between the model training management subsystem adopt a high-speed optical fiber channel with a bandwidth of not less than 10Gbps, which reduces data snapshots, reduces data transmission delays, and improves model training efficiency. Assume that the transmission delay reduction formula is , it is expected that the value is not less than 60%, where OD represents OriginalDelay and ND represents NewDelay.
[0039] In the present invention, when the model training management subsystem encounters insufficient computing resources during model construction, it automatically accesses the cloud computing platform, such as Alibaba Cloud and Tencent Cloud, and uses elastic computing resources to ensure smooth model construction and optimization. At the same time, it optimizes resource allocation and reduces cloud computing costs by more than 20%. Assume that the cost reduction rate formula is , requiring the value to reach or exceed 20%, where OC stands for Original Cost and NC stands for New Cost.
[0040] In the present invention, when presenting the prediction results, the prediction result application subsystem adopts virtual reality (VR) or augmented reality (AR) technology to provide an immersive load forecast display for power dispatchers. In the specific implementation process, when power dispatchers need to view the load forecast results, they are no longer limited to traditional two-dimensional charts and data tables. Instead, they enter a virtual power system operation scene by wearing VR equipment. In this virtual scene, the architecture of the entire power system is presented in a three-dimensional form, including various components such as power stations, substations, and transmission lines. The dispatcher seems to be in the real power system operation site and can observe the details of the power system in 360 degrees. Each power station has an accurate model presentation in the virtual scene, and its appearance and scale are consistent with the actual power station. The current operating status of the power station is indicated by different colors and dynamic effects. For example, a power station with sufficient power generation may be displayed as bright blue, while a power station close to full load operation may be warned in red. The substation displays its voltage conversion and power distribution functions in the virtual scene. Dispatchers can view its detailed data information, such as the current voltage level, power input and output, by approaching the substation. Transmission lines in the virtual scene are crisscrossed like real power grids. The brightness and color of the lines are changed to reflect the flow of electricity. For example, lines with heavy power loads may be displayed with thicker lines and brighter colors, while lines with light loads are relatively thin and dim.
[0041] For the display of load forecast results, in the virtual scene, the brightness and dynamic effects of different areas are used to indicate the changing trend of future power load. For example, when it is predicted that the power demand in a certain area will increase significantly in the future, the area will gradually become brighter in the virtual scene, and some dynamic arrows or ripples may appear to indicate the rising trend of power load. Dispatchers can view the power load forecasts in different areas by walking or flying in the virtual scene. This immersive experience allows them to more intuitively feel the spatiotemporal distribution and changing laws of power load. If augmented reality (AR) technology is used, dispatchers can wear AR glasses to see the virtual model and forecast data of the power system superimposed on the real scene in the real environment. For example, on the large screen of the power dispatching center, the virtual power system operation scene and load forecast results can be projected directly on the screen through AR technology. When the dispatcher views the actual geographical map or power system layout diagram, he can see the virtual forecast data and model at the same time. This combination of virtual and real allows dispatchers to enjoy the convenience brought by advanced technology in a familiar working environment, feel the trend of power load changes more intuitively, and make more accurate and timely power dispatching decisions. Assume that the decision efficiency improvement formula is , it is expected that the value is not less than 30%, where ODT represents OriginalDecisionTime and NDT represents NewDecisionTime.
[0042] The above are only preferred specific implementation modes of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and inventive concepts of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for constructing an accurate prediction model of power load using big data, characterized in that: The following steps are involved: Data collection steps: Use IoT devices to collect real-time operation data at each node of the power system, use cleaning algorithms to remove erroneous data, and then transmit the processed data to the data storage link through a secure encrypted channel; suppose the data transmission packet loss rate formula is , where LP represents the number of lost packets LostPackets, and TP represents the total number of TotalPackets; Data storage steps: Distributed storage architecture is used to classify and store the collected data. The Ceph storage system uses the CRUSH algorithm for data distribution. The retrieval efficiency improvement formula is: , where ORT represents the original retrieval time, and NRT represents the new retrieval time; Feature engineering steps: retrieve information from the stored data, and filter out features that are strongly correlated with power load by calculating the Pearson correlation coefficient. Suppose the feature correlation strength formula is , select features with an absolute value greater than 0.5 so that the data value range is between [0,1], and after processing, transfer the data to the model building stage; Model construction steps: Build a prediction model based on deep learning, a hybrid model that combines the long short-term memory network LSTM with the convolutional neural network CNN. During the training process, the model convergence speed formula is set as , where TE represents the total training rounds of TotalEpochs, and CE represents the current training round of CurrentEpoch; the model is trained using the training set, and the RMSE calculation formula is: ,in is the true value, is the predicted value, n is the number of samples; Model evaluation and update steps: Use the collected data to evaluate the model, use the mean absolute error (MAE) and mean absolute percentage error (MAPE) indicators to measure the performance of the model, and also use the mean absolute percentage error (MAPE) indicator to measure the performance of the model. The MAE calculation formula is: , the MAPE calculation formula is: ,in is the true value, is the predicted value, n is the number of samples; Prediction application steps: deploy the constructed accurate prediction model to the power dispatching center; present the prediction results to the dispatching personnel in the form of visual charts.
2. The method for constructing an accurate prediction model of power load using big data according to claim 1, characterized in that: Also includes: Data fusion verification step: After the data collection step is completed, the data obtained from different data sources are fused. For data with the same attributes, the weighted average method is used for integration. For the current data of the same area collected from different smart meters, weights are assigned according to the meter accuracy. The weight calculation formula is as follows: ,in The accuracy of the th meter is measured, and the data consistency verification algorithm is used at the same time.
3. The method for constructing an accurate prediction model of power load using big data according to claim 1, characterized in that: Also includes: Feature dimension reduction optimization step: After selecting features in the feature engineering step, the principal component analysis (PCA) algorithm is used to reduce the dimension of the features. The variance contribution rate threshold of the retained principal component is set to 90%. The formula for the retention rate of the principal component is: ,in is the variance of the th principal component, k is the number of retained principal components, and n is the number of original features. Then the random forest algorithm is used to screen the reduced-dimensional features again to remove redundant features.
4. The method for constructing an accurate prediction model of power load using big data according to claim 1, characterized in that: Also includes: Model adaptive adjustment step: During the model building step, if the model shows signs of overfitting after a certain number of training rounds, including a continuous increase in the validation set loss value, the learning rate is reduced and the L2 regularization coefficient is increased. The adjustment coefficient formula is set as The learning rate adjustment range is expected to be between 50%-80%, and the L2 regularization coefficient increase range is expected to be between 20%-50%, where OV represents the original value of OriginalValue and NV represents the new value of NewValue; on the contrary, if it is underfitting, increase the learning rate so that the model can adapt to different data distributions.
5. The method for constructing an accurate prediction model of power load using big data according to claim 1, characterized in that: In the model evaluation and update step, a user feedback collection mechanism is constructed, feedback information is quantified into weights, and integrated into the model performance evaluation index system. The feedback weight calculation formula is set as , where SS represents the SeverityScore, IS represents the ImportanceScore, and TS represents the TotalScore. The weight is used to adjust the proportion of the model evaluation indicators.
6. The method for constructing an accurate prediction model of power load using big data according to claim 1, characterized in that: In the prediction application step, when presenting the prediction results to the dispatcher, interactive operations are supported and a comparison function is provided.
7. A system for implementing the method for constructing an accurate prediction model of power load using big data as described in any one of claims 1 to 6, characterized in that: include: Data acquisition and transmission subsystem: responsible for executing the functions of the data acquisition step and the data storage step described in claim 1; Model training management subsystem: undertakes the responsibilities of the feature engineering step, model building step, and model evaluation and update step described in claim 1, and builds, optimizes, and maintains the prediction model; Prediction result application subsystem: corresponds to the prediction application step described in claim 1, providing the prediction results to the power dispatching personnel in a visualized form for formulating power dispatching strategies.
8. The system for constructing an accurate prediction model of power load using big data according to claim 7, characterized in that: The data transmission between the data storage center in the data acquisition and transmission subsystem and the model training management subsystem adopts a high-speed optical fiber channel with a bandwidth of no less than 10Gbps. The transmission delay reduction formula is: , it is expected that the value is not less than 60%, where OD represents the original delay and ND represents the new delay.
9. The system for constructing an accurate prediction model of power load using big data according to claim 7, characterized in that: When the model training management subsystem encounters insufficient computing resources when building the model, it automatically connects to the cloud computing platform and optimizes resource allocation. The cost reduction rate formula is set as , requiring the value to reach or exceed 20%, where OC represents the original cost and NC represents the new cost.
10. The system for constructing an accurate prediction model of power load using big data according to claim 7, characterized in that: When presenting the prediction results, the prediction result application subsystem uses virtual reality VR technology to enable power dispatchers to intuitively feel the trend of power load changes. The decision-making efficiency improvement formula is set as , it is expected that the value is not less than 30%, where ODT represents the original decision time and NDT represents the new decision time.
Citation Information
Patent Citations
Power load scheduling method and system based on big data
CN116646933A
Power load prediction method and device based on machine learning, and storage medium
CN117787075A
Power dispatching method and system based on power load prediction
CN119029836A
Charging station load prediction method and system considering multiple influence factors
CN119029838A
Cited By
Green data center computing power demand prediction and energy consumption control method, system and device
CN120508401A
Active power distribution network dynamic load prediction model based on multi-scale feature fusion
CN121146210A
Power load prediction agent construction method based on large model
CN121282850A
A Method for Constructing an Intelligent Agent for Power Load Forecasting Based on a Large Model
CN121282850B