Power load prediction method and system based on big data
By deploying intelligent sensing nodes at various levels of the power grid and constructing a distributed collaborative prediction network, combined with a multi-timescale and multi-spatial dimension partitioning mechanism, the problem of decreased accuracy and insufficient risk identification in traditional power load forecasting methods when facing complex load changes is solved, achieving efficient power load forecasting and adaptive optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional power load forecasting methods struggle to effectively couple the propagation of local anomalies with global trend evolution when faced with increased load volatility, randomness, and spatiotemporal correlations brought about by energy structure transformation and renewable energy integration. Furthermore, the lack of distributed coordination mechanisms leads to decreased forecast accuracy and insufficient risk identification.
A power load forecasting system based on big data is constructed by deploying intelligent sensing nodes at various levels of the power grid, building a distributed collaborative forecasting network, setting up a multi-time scale and multi-spatial dimension partitioning mechanism, configuring initial model parameters, and using a risk diffusion network for adaptive optimization and early warning of error propagation and amplification.
It improves the model's adaptability to different scenarios, integrates the real-time data acquisition capabilities of various levels of the power grid, reduces data transmission pressure and latency, realizes real-time identification and adaptive optimization of high-risk nodes, and improves the response speed and reliability of the prediction system.
Smart Images

Figure CN121791129A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power planning technology, and specifically to a power load forecasting method and system based on big data. Background Technology
[0002] Electricity load forecasting is a crucial link in power grid dispatching, operation, and planning management, and its accuracy directly affects the safety, stability, and economic operation of the power grid. With the transformation of the energy structure and the large-scale integration of renewable energy, the volatility, randomness, and spatiotemporal correlation of power grid loads are increasing, posing a severe challenge to traditional forecasting methods. Traditional methods mainly use structured historical load data, which is insufficient for correlation mining of multidimensional environmental data (such as meteorological, geographical, and user behavior data), making it difficult to comprehensively capture the complex driving factors of load changes. Furthermore, fixed-structure forecasting models often cannot adapt to the differentiated characteristics of different regions, seasons, or power grid levels, leading to decreased forecasting accuracy in spatially heterogeneous and temporally non-stationary scenarios. Existing methods are mostly centralized forecasting methods, lacking distributed collaborative mechanisms based on power grid topology and operating status. This makes it difficult to achieve effective coupling analysis of local anomaly propagation and global trend evolution. Most methods focus on point or interval forecasting, failing to establish a visualization and quantitative analysis network for forecast error propagation paths, thus hindering early identification and warning of regional and cascading load fluctuation risks. To address these issues, this paper proposes a big data-based power load forecasting method and system. Summary of the Invention
[0003] The purpose of this invention is to provide a power load forecasting method and system based on big data to address the shortcomings of the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: A power load forecasting method based on big data includes the following steps: Step S1: Obtain historical load data and associated historical environmental records, and build a predictive model component library based on the historical load data and historical environmental records; Step S2: Deploy smart sensing nodes at all levels of the power grid to collect multi-source operation data in real time, and construct a distributed collaborative prediction network based on the multi-source operation data and the spatial distribution of smart sensing nodes. Step S3: Set up a multi-timescale and multi-spatial dimension partitioning mechanism to classify the distributed collaborative prediction network into different monitoring areas, and configure initial model parameters for each node based on historical load data; Step S4: Call up multi-source operational data to label the load trend of the monitoring area, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. Based on the propagation and amplification process of prediction error in the risk diffusion network, adaptively optimize the model parameters of the prediction nodes and provide early warning.
[0005] Furthermore, the historical load data covers load-related data at all levels of the power grid; The historical environmental records include meteorological environmental data, social activity data, and power grid operation environment data.
[0006] Furthermore, the process of building a predictive model component library based on data features and prediction objectives includes: Data features and prediction targets are extracted from historical load data, including trend features, periodic features, and abrupt change features. The correlation strength between various historical environmental records and historical load data is analyzed to screen out key influencing features. The forecasting objectives are set as multi-dimensional forecasting requirements, including short-term load forecasting, medium-term load forecasting, and long-term load forecasting. Each forecasting objective corresponds to different forecasting accuracy requirements and time granularity. Based on the data characteristics and prediction objectives, various suitable prediction model modules are selected, including time series prediction components, environmental correlation prediction components, and hybrid integrated prediction components.
[0007] Furthermore, the process of deploying intelligent sensing nodes at various levels of the power grid to collect multi-source operational data in real time includes: A data collection period of k is set with an annual data collection cycle, where k is a natural number greater than 5. Data collection equipment is deployed at key monitoring points at each level of the power grid to collect multi-source operation data during the corresponding data collection period. The multi-source operational data includes real-time load data, real-time environmental data, equipment operating status data, and dynamic data of the power grid topology.
[0008] Furthermore, the process of constructing a distributed collaborative prediction network based on multi-source operational data and the spatial location distribution of intelligent sensing nodes includes: Establish a three-dimensional spatial coordinate system and map all intelligent sensing nodes onto this three-dimensional spatial coordinate system to form a node spatial distribution map; The neighbor association algorithm is used to determine the neighboring nodes of the predicted node. That is, the predicted nodes that are within a preset distance of kilometers and belong to the same or adjacent power supply areas are neighboring nodes. Each predicted node establishes a direct data interaction link only with its neighboring nodes, forming a distributed mesh topology.
[0009] Furthermore, the process of setting up a multi-timescale and multi-spatial dimension partitioning mechanism to classify the distributed collaborative prediction network into different monitoring areas includes: The multi-timescale division mechanism: combining the actual application needs of power load forecasting, it divides the forecasting timescale into three timescales: short-term timescale, medium-term timescale, and long-term timescale. Set a corresponding prediction time point for each time scale; for short-term time scales, each time granularity corresponds to one prediction time point. The multi-dimensional spatial division mechanism employs a two-layer spatial division strategy, namely basic spatial division and dynamic optimization division, to achieve precise definition of the monitoring area. If the load density of two adjacent basic monitoring areas is lower than the preset threshold and the load characteristic similarity is higher than the preset value, then the two monitoring areas will be merged into one monitoring area. Based on the monitoring area determined by the multi-spatial dimension division mechanism, all prediction nodes in the distributed collaborative prediction network are classified into the corresponding monitoring areas according to their spatial location and power grid level.
[0010] Furthermore, the process of configuring initial model parameters for each prediction node based on historical load data includes: For each monitoring area corresponding to a prediction node, the statistical characteristics of its historical load data are extracted. At the same time, the sensitivity of the load influencing factors in the monitoring area is analyzed, that is, the contribution of various environmental factors to load changes. The contribution is represented by the standardized regression coefficient in multiple regression analysis. Based on the prediction node's level, the type of monitored object, and the corresponding prediction time scale, select the appropriate model component from the prediction model component library; The historical load data corresponding to the prediction node is divided into training set and validation set in a 7:3 ratio. All parameter combinations are traversed within the preset parameter range. The prediction error is calculated through the validation set, and the parameter combination with the smallest prediction error is selected as the initial model parameters.
[0011] Furthermore, the process of calling multi-source operational data to label the load trends in the monitored area, and then matching and generating a risk diffusion network in the prediction model component library based on the load trend labels includes: First, a short-term load change rate sequence is obtained based on multi-source operation data. Then, a load trend curve is fitted by combining real-time load data and historical load data to obtain the slope of the trend curve. By combining the load change rate sequence with the slope of the trend curve, the load trend is divided into rapid growth trend, slow growth trend, rapid decline trend, slow decline trend, stable trend, and fluctuating trend. For different load trends, the corresponding prediction model components are matched. The nodes of the risk diffusion network include prediction nodes, error monitoring nodes, and regional collaboration nodes. Taking the monitoring area as a unit, the matched prediction model components are bound to the corresponding prediction nodes, the prediction nodes are connected to the error monitoring nodes, and each prediction node is connected to the regional collaboration nodes according to the error transmission coefficient to form a local risk diffusion network. The local risk diffusion networks of all monitoring areas are interconnected to form a global risk diffusion network.
[0012] Furthermore, based on the propagation and amplification process of prediction errors in the risk diffusion network, the process of adaptively optimizing model parameters and providing early warnings for prediction nodes includes: The prediction error of each error monitoring node is monitored in real time. If the prediction error of a prediction node exceeds the error threshold, the error propagation path tracing is triggered. Through the risk diffusion network, the neighboring prediction nodes of the prediction node and the downstream nodes affected by its error are found to determine the propagation range and main propagation path of the error. The optimization is verified using multi-source running data at the next prediction time point. If the prediction error is lower than the threshold, the optimization is considered effective. If the prediction error is still higher than the threshold, the above optimization process is repeated until the prediction error meets the requirements. The warning levels are divided according to the magnitude of the prediction error, the range of error propagation, and the scale of the impact load. When the risk diffusion network predicts that the error will reach the triggering conditions of the corresponding warning level, the warning is triggered in advance and the warning information is sent. Then, corresponding response measures are formulated for the warning information of different warning levels.
[0013] A power load forecasting system based on big data includes a power grid data acquisition module, a power grid status analysis module, and a load early warning analysis module; The power grid data acquisition module is used to deploy intelligent sensing nodes at all levels of the power grid, collect multi-source operation data in real time, acquire historical load data and related historical environmental records, and build a predictive model component library based on historical load data and historical environmental records. The power grid status analysis module is used to construct a distributed collaborative prediction network based on multi-source operation data and the spatial location distribution of intelligent sensing nodes. It sets up a multi-time scale and multi-spatial dimension division mechanism to classify the distributed collaborative prediction network into different monitoring areas and configure initial model parameters for each node according to historical load data. The load early warning analysis module is used to label the load trend of the monitoring area based on multi-source operation data, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. Based on the propagation and amplification process of prediction error in the risk diffusion network, the module performs adaptive optimization of model parameters and early warning for prediction nodes.
[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention constructs a predictive model component library, which flexibly matches the optimal model components based on data features and prediction targets. Combined with multi-source operational data, it enhances the model's adaptability to different scenarios. At the same time, the distributed collaborative prediction network built based on intelligent sensing nodes effectively integrates the real-time data acquisition capabilities of various levels of the power grid. Through information interaction and collaborative computing between nodes, it reduces the data transmission pressure and latency of centralized processing, and improves the system response speed and operational reliability.
[0015] 2. This invention models the propagation and amplification process of prediction errors through a risk diffusion network, which can identify high-risk nodes and regions in real time and drive adaptive adjustment of model parameters. This mechanism enables the prediction system to have online learning and self-optimization capabilities, and to a certain extent, it can respond to sudden load fluctuations or environmental anomalies in a timely manner. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0017] Figure 1 This is a flowchart of a power load forecasting method based on big data as described in this invention.
[0018] Figure 2 This is a system block diagram of a power load forecasting system based on big data as described in this invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Please see Figure 1 As shown, a power load forecasting method based on big data includes the following steps: Step S1: Obtain historical load data and associated historical environmental records, and build a predictive model component library based on the historical load data and historical environmental records; Step S2: Deploy smart sensing nodes at all levels of the power grid to collect multi-source operation data in real time, and construct a distributed collaborative prediction network based on the multi-source operation data and the spatial distribution of smart sensing nodes. Step S3: Set up a multi-timescale and multi-spatial dimension partitioning mechanism to classify the distributed collaborative prediction network into different monitoring areas, and configure initial model parameters for each node based on historical load data; Step S4: Call up multi-source operational data to label the load trend of the monitoring area, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. Based on the propagation and amplification process of prediction error in the risk diffusion network, adaptively optimize the model parameters of the prediction nodes and provide early warning.
[0021] Furthermore, step S1 is implemented through the following process: Step S101: Obtain historical load data and associated historical environmental records. The specific process includes: Historical load data and related historical environmental records are obtained through data sources such as the Internet and power grid databases. The historical load data covers load-related data at all levels of the power grid, specifically including residential electricity load data, commercial electricity load data, and industrial electricity load data on the user side; transformer load rate data, three-phase current imbalance data, and voltage deviation data on the distribution transformer side; and feeder current data, power factor data, and line loss rate data on the feeder side. The historical environmental records include meteorological environmental data, social activity data, and power grid operation environment data. The meteorological environmental data includes historical temperature change data, historical humidity change data, historical light intensity data, historical precipitation data, historical wind speed data, etc. The social activity data includes historical holiday schedule data, historical records of major events, historical regional population flow data, etc. The power grid operation environment data includes historical power grid topology adjustment data, historical equipment maintenance records, historical electricity price adjustment data, etc.
[0022] Step S102: Construct a prediction model component library based on data features and prediction objectives. The specific process includes: Data features and prediction targets are extracted from historical load data, including trend features (such as long-term growth trends and seasonal fluctuation trends), periodic features (such as daily, weekly, and monthly cycles), and abrupt change features (such as load abrupt changes during holidays and extreme weather). The correlation strength between various historical environmental records and historical load data is analyzed to screen out key influencing features. The forecasting objectives are set to multi-dimensional forecasting needs, including short-term load forecasting (forecast duration 15 minutes - 24 hours), medium-term load forecasting (forecast duration 1-7 days), and long-term load forecasting (forecast duration 1-30 days). Each forecasting objective corresponds to different forecasting accuracy requirements and time granularity. Short-term forecasting focuses on minute-level and hour-level accuracy, medium-term forecasting focuses on daily-level accuracy, and long-term forecasting focuses on weekly-level and monthly-level accuracy. Based on the data characteristics and prediction objectives, various suitable prediction model modules are selected, including time series prediction components, environment-related prediction components, and hybrid integrated prediction components. The time-series prediction component includes LSTM neural network modules, which are suitable for capturing the time-series variation patterns of load data; the environmental correlation prediction component includes multiple linear regression modules, random forest regression modules, etc., which are suitable for quantifying the mapping relationship between environmental factors and load; the hybrid ensemble prediction component includes Boosting ensemble modules, etc. Each model component has preset basic parameter ranges and adaptive feature type labels, which are convenient for subsequent calls and combinations according to actual scenarios.
[0023] Furthermore, step S2 is implemented through the following process: Step S201: Deploy intelligent sensing nodes at various levels of the power grid to collect multi-source operational data in real time. The specific process includes: The data collection period is set to k data collection periods with an annual data collection cycle. It should be noted that the time span of each data collection period is not exactly the same. That is, there are data collection periods with a length of one week and data collection periods with a length of one month. Here, k is a natural number greater than 5. This setting method can take into account both short-term load fluctuation patterns and long-term load change trends, and ensure that the data covers load characteristics under different time dimensions. Data acquisition equipment is deployed at key monitoring points at each level of the power grid (user side, distribution transformer side, and feeder side). Specifically, this includes deploying smart meters and load monitoring terminals at user-side distribution areas, distribution transformer monitoring terminals and voltage and current sensors at the distribution transformer side, and feeder monitoring devices and power sensors at the feeder side. Simultaneously, j types of environmental sensing sensors are configured, including temperature sensors, humidity sensors, light sensors, and precipitation sensors, and deployed at key locations such as regional meteorological monitoring stations, densely populated areas, and industrial clusters, where j is a natural number greater than 0. At the start of each data acquisition period, all data acquisition devices and environmental sensing sensors are reset to clear residual data from the previous acquisition period and then acquire multi-source operational data for the corresponding data acquisition period. The multi-source operation data includes real-time load data, real-time environmental data, equipment operation status data, and power grid topology dynamic data. The data types contained in the real-time load data correspond one-to-one with those contained in the historical load data in step S1, ensuring the continuity and consistency of the data.
[0024] Step S202: Construct a distributed collaborative prediction network based on multi-source operational data and the spatial location distribution of intelligent sensing nodes. The specific process includes: A three-dimensional spatial coordinate system is established, with the x and y axes corresponding to geographical latitude and longitude coordinates, and the z axis corresponding to the power grid level (level 1 for the user side, level 2 for the distribution transformer side, and level 3 for the feeder side). All intelligent sensing nodes are mapped onto this three-dimensional spatial coordinate system based on their GPS positioning information and level information to form a node spatial distribution map. Each prediction node corresponds one-to-one with a smart sensing node and has the functions of data preprocessing, local feature extraction, preliminary prediction and data interaction between nodes. Based on the spatial distribution of smart sensing nodes and the topology of the power grid, the neighbor association algorithm is used to determine the neighboring nodes of the prediction node. That is, prediction nodes that are within a preset kilometer range (the preset kilometer range is generally taken as 5 kilometers) and belong to the same power supply area or adjacent power supply areas are neighboring nodes. Each prediction node establishes a direct data interaction link only with its neighboring nodes, forming a distributed mesh topology.
[0025] Furthermore, step S3 is implemented through the following process: Step S301: Set up a multi-timescale and multi-spatial dimension partitioning mechanism to classify the distributed collaborative prediction network into different monitoring areas. The specific process includes: The multi-timescale division mechanism, based on the practical application needs of power load forecasting, divides the forecasting timescale into three segments: short-term, medium-term, and long-term. The division criteria and characteristics of each timescale are as follows: Short-term time scale: The prediction duration is 15 minutes to 24 hours, and the time granularity is 15 minutes / segment. It is suitable for scenarios such as real-time power grid dispatch and load control. At this scale, the load is greatly affected by real-time environmental changes and users' immediate electricity consumption behavior, and the fluctuation frequency and amplitude are high. Medium-term time scale: The forecast duration is 1-7 days, and the time granularity is 1 hour / segment. It is suitable for scenarios such as power grid operation and maintenance planning and electricity market transaction bidding. At this scale, the load is significantly affected by weather change trends, differences between weekdays and holidays, and regional economic activity patterns. Long-term time scale: The forecast duration is 1-30 days, and the time granularity is 24 hours. It is suitable for scenarios such as power grid planning and power source construction planning. At this scale, the load is greatly affected by seasonal changes, industry production cycles, and macroeconomic policies, showing obvious trends and cycles. For each time scale, a corresponding prediction point is set. For the short-term time scale, each time granularity corresponds to one prediction point, that is, 96 prediction points are set every day; for the medium-term time scale, each time granularity corresponds to one prediction point, that is, 24 prediction points are set every day; for the long-term time scale, each time granularity corresponds to one prediction point, that is, one prediction point is set every day. The multi-dimensional spatial division mechanism employs a two-layer spatial division strategy, namely basic spatial division and dynamic optimization division, to achieve precise definition of the monitoring area. The basic spatial division is based on both administrative regions and power grid supply areas, dividing the entire power supply range into several monitoring areas. The administrative region division is based on the city, district (county), and township (street) levels to ensure that the monitoring areas are consistent with the social management boundaries. The power grid supply area division is based on the power supply range of substations and the power supply range of feeders to ensure that the monitoring areas are consistent with the power grid operation boundaries. The dynamic optimization division is based on the historical load fluctuation intensity and load density distribution, and the monitoring area is dynamically adjusted. The load fluctuation coefficient (the ratio of the difference between the maximum and minimum load values in a certain period of a monitoring area to the average value) is used as the judgment index. If the load fluctuation coefficient of a monitoring area is greater than the preset threshold (the preset threshold is 0.8) for two consecutive data collection periods, the area is divided into multiple smaller monitoring areas to improve the prediction accuracy. If the load density (average load per unit area) of two adjacent basic monitoring areas is lower than the preset threshold (the preset threshold is 0.5MW / km²) and the load characteristic similarity (calculated using cosine similarity, with a threshold of 0.9) is higher than the preset value, then the two monitoring areas will be merged into one monitoring area to reduce the calculation cost. Based on the monitoring areas determined by the multi-spatial dimension division mechanism, all prediction nodes in the distributed collaborative prediction network are classified into corresponding monitoring areas according to their spatial location and power grid level: First, extract the GPS positioning information and hierarchical identifier of each prediction node, determine the administrative region and power grid supply area where the node is located, and initially match it to the corresponding basic monitoring area; Then, based on the results of dynamic optimization, the preliminary matching results are adjusted to ensure that each predicted node belongs to only one monitoring area. Finally, a monitoring area-predicted node mapping table is established to record information such as the number of predicted nodes, node identifiers, and node hierarchical distribution in each monitoring area. At the same time, the boundary range and the predicted nodes contained in each monitoring area are marked in a three-dimensional spatial coordinate system to form a visualized area-node distribution map.
[0026] Step S302: Configure initial model parameters for each prediction node based on historical load data. The specific process includes: For each monitoring area corresponding to a prediction node (such as a user-side transformer area, a distribution transformer unit, or a feeder section), extract the statistical characteristics of its historical load data, including mean, variance, peak value, valley value, peak value occurrence period, valley value occurrence period, trend slope, and periodic characteristic parameters. Simultaneously, the sensitivity of load-influencing factors in the monitoring area is analyzed, that is, the contribution of various environmental factors to load changes. The contribution is represented by the standardized regression coefficient in multiple regression analysis. Based on the prediction node's level, the type of monitored object, and the corresponding prediction time scale, select suitable basic model components from the prediction model component library; for example, for short-term prediction of the user-side resident load prediction node, prioritize the LSTM neural network module (capturing time series features) and the random forest regression module (quantifying the impact of environmental factors). The historical load data corresponding to the prediction node is divided into training and validation sets in a 7:3 ratio. All parameter combinations are iterated within a preset parameter range (preset in the model component library). The model is trained using the training set, and the prediction error (using root mean square error RMSE as the evaluation metric) is calculated using the validation set. The parameter combination with the smallest prediction error is selected as the initial model parameters. For example, the initial parameters of the LSTM neural network module include the number of hidden layer neurons (preset range 32-256), learning rate (preset range 0.001-0.01), number of iterations (preset range 100-500), and batch size (preset range 8-64), etc.
[0027] Because the prediction nodes within the same monitoring area have load correlations (such as the complementary electricity consumption periods of industrial and residential loads within the same area), it is necessary to perform regional collaborative calibration on the initial parameters of each prediction node. This involves obtaining the fitting degree between the prediction results corresponding to the initial parameters of all prediction nodes within the monitoring area and the historical data of the total regional load. If the fitting degree is lower than a preset threshold (the preset threshold is 0.85), the initial parameters of the relevant prediction nodes are adjusted (the adjustment range does not exceed 20% of the initial parameters) until the fitting degree between the prediction results of the total regional load and the historical data meets the requirements, ensuring that the prediction results of each node within the area have synergy and consistency.
[0028] Furthermore, step S4 is implemented through the following process: Step S401: Call multi-source operational data to label the load trend of the monitoring area, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. The specific process includes: At the start of each prediction point, each prediction node in the distributed collaborative prediction network calls the multi-source operational data collected in real time by the corresponding intelligent sensing node through the data input module; Then, load trends are labeled for each monitoring area based on multi-source operational data: First, real-time load data for 1 hour, 3 hours, and 6 hours before the current forecast time are selected based on multi-source operation data. The load change rate of adjacent time periods is calculated (change rate = (load of the next time period - load of the previous time period) / load of the previous time period × 100%) to obtain the short-term load change rate sequence. The sliding window method (window size is 24 historical time points) is used to fit the load trend curve by combining real-time load data and historical load data. The slope of the trend curve is obtained. A positive slope indicates that the load is increasing, a negative slope indicates that the load is decreasing, and an absolute value of the slope less than 0.01 indicates that the load is stable. Combining the load change rate sequence with the slope of the trend curve, the load trend is divided into the following categories: rapid growth trend (the load change rate is greater than 5% for any two consecutive periods and the slope of the trend curve is greater than 0.05), slow growth trend (the load change rate is between 1% and 5% and the slope of the trend curve is between 0.01 and 0.05), rapid decline trend (the load change rate is less than -5% for any two consecutive periods and the slope of the trend curve is less than -0.05), slow decline trend (the load change rate is between -5% and 1% and the slope of the trend curve is between -0.05 and 0.01), stable trend (the absolute value of the load change rate is less than 1% and the absolute value of the slope of the trend curve is less than 0.01), and fluctuating trend (the load change rate alternates between positive and negative and the absolute value is greater than 3%). The credibility of trend labeling is evaluated using the trend consistency coefficient (the similarity between real-time trends and historical trends of the same period). A trend consistency coefficient greater than 0.8 indicates high credibility, a trend consistency coefficient between 0.6 and 0.8 indicates medium credibility, and a trend consistency coefficient less than 0.6 indicates low credibility. For different load trends, corresponding prediction model components are matched, such as a fast growth trend matching LSTM neural network module + gradient boosting tree module (to enhance the ability to capture sudden growth trends), and a risk diffusion network is established. The nodes of the risk diffusion network include prediction nodes, error monitoring nodes, and regional coordination nodes. The prediction nodes are the prediction nodes in the distributed collaborative prediction network, which are responsible for outputting local load prediction values. The error monitoring nodes correspond one-to-one with the prediction nodes and are responsible for calculating the prediction error of the prediction nodes (prediction error = |predicted load value - actual load value| / actual load value × 100%). The regional coordination nodes correspond to the regional coordination centers of the monitoring areas and are responsible for summarizing the error information of all prediction nodes in the area and analyzing the error propagation path. Taking the monitoring area as a unit, the matched prediction model components are bound to the corresponding prediction nodes, the prediction nodes are connected to the error monitoring nodes, and each prediction node is connected to the regional collaborative nodes according to the error transmission coefficient to form a local risk diffusion network. The local risk diffusion networks of all monitoring areas are interconnected (the regional collaborative nodes of adjacent monitoring areas establish connections) to form a global risk diffusion network.
[0029] Step S402: Based on the propagation and amplification process of prediction errors in the risk diffusion network, adaptive optimization of model parameters and early warning are performed on the prediction nodes. The specific process includes: The prediction error of each error monitoring node is monitored in real time. If the prediction error of a prediction node exceeds the error threshold (5% for short-term prediction error, 8% for medium-term prediction error, and 10% for long-term prediction error), the error propagation path tracking is triggered. Through the risk diffusion network, the neighboring prediction nodes of the prediction node and the downstream nodes affected by its error are found to determine the propagation range and main propagation path of the error. The optimization is verified using multi-source running data at the next prediction time point. If the prediction error is lower than the threshold, the optimization is considered effective. If the prediction error is still higher than the threshold, the above optimization process is repeated until the prediction error meets the requirements. The number of iterations does not exceed 5 (to avoid overfitting). Based on error propagation prediction using risk diffusion networks, a multi-level early warning mechanism is established: Warning level classification: Based on the magnitude of the forecast error, the range of error propagation, and the scale of the affected load, the warning level is divided into three levels: Level 1 warning (severe), Level 2 warning (relatively severe), and Level 3 warning (general). Level 1 Warning Conditions: The total regional error exceeds 15%, the error propagation range covers more than 50% of the forecast nodes, and the affected load exceeds 30% of the total regional load; Level 2 Warning Conditions: The total regional error is between 10% and 15%, the error propagation range covers 30% to 50% of the forecast nodes, and the affected load is between 15% and 30%; Level 3 Warning Conditions: The total regional error is between the threshold and 15%, the error propagation range covers less than 30% of the forecast nodes, and the affected load is less than 15%. When the risk diffusion network predicts that the error will reach the triggering condition of the corresponding warning level, it will trigger the warning 1-3 prediction time points in advance (1 time point for short-term warnings, 2 time points for medium-term warnings, and 3 time points for long-term warnings) and send the warning information. The early warning information includes the warning level, affected area, estimated error, scale of impact on load, and recommended measures. Develop corresponding response measures for different warning levels: Level 1 warning: activate the emergency load control plan, the dispatch center adjusts the power grid operation mode in real time, guides high-energy-consuming users to use electricity during off-peak hours, and arranges maintenance personnel for on-site inspections; Level 2 warning: optimize the power output distribution in the region, adjust the charging and discharging strategies of energy storage equipment, and strengthen the monitoring of key prediction nodes; Level 3 warning: only adjust the model parameters of the corresponding prediction nodes, closely monitor the error change trend, and no additional dispatching operations are required.
[0030] Please see Figure 2 As shown, a power load forecasting system based on big data includes a power grid data acquisition module, a power grid status analysis module, and a load early warning analysis module. The power grid data acquisition module is used to deploy intelligent sensing nodes at all levels of the power grid, collect multi-source operation data in real time, acquire historical load data and related historical environmental records, and build a predictive model component library based on historical load data and historical environmental records. The power grid status analysis module is used to construct a distributed collaborative prediction network based on multi-source operation data and the spatial location distribution of intelligent sensing nodes. It sets up a multi-time scale and multi-spatial dimension division mechanism to classify the distributed collaborative prediction network into different monitoring areas and configure initial model parameters for each node according to historical load data. The load early warning analysis module is used to label the load trend of the monitoring area based on multi-source operation data, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. Based on the propagation and amplification process of prediction error in the risk diffusion network, the module performs adaptive optimization of model parameters and early warning for prediction nodes.
[0031] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A power load forecasting method based on big data, characterized in that, Includes the following steps: Step S1: Obtain historical load data and associated historical environmental records, and build a predictive model component library based on the historical load data and historical environmental records; Step S2: Deploy smart sensing nodes at all levels of the power grid to collect multi-source operation data in real time, and construct a distributed collaborative prediction network based on the multi-source operation data and the spatial distribution of smart sensing nodes. Step S3: Set up a multi-timescale and multi-spatial dimension partitioning mechanism to classify the distributed collaborative prediction network into different monitoring areas, and configure initial model parameters for each node based on historical load data; Step S4: Call multi-source operational data to label the load trend of the monitoring area, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. Based on the propagation and amplification process of prediction error in the risk diffusion network, adaptively optimize the model parameters of the prediction nodes and provide early warning.
2. The power load forecasting method based on big data according to claim 1, characterized in that, The historical load data covers load-related data at all levels of the power grid; The historical environmental records include meteorological environmental data, social activity data, and power grid operation environment data.
3. The power load forecasting method based on big data according to claim 2, characterized in that, The process of building a predictive model component library based on data features and prediction objectives includes: Data features and prediction targets are extracted from historical load data. Data features include trend features, periodic features, and abrupt change features. The correlation strength between various historical environmental records and historical load data is analyzed to screen out key influencing features. The forecasting objectives are set as multi-dimensional forecasting requirements, including short-term load forecasting, medium-term load forecasting, and long-term load forecasting. Each forecasting objective corresponds to different forecasting accuracy requirements and time granularity. Based on the data characteristics and prediction objectives, various suitable prediction model modules are selected, including time series prediction components, environmental correlation prediction components, and hybrid integrated prediction components.
4. The power load forecasting method based on big data according to claim 3, characterized in that, The process of deploying intelligent sensing nodes at various levels of the power grid and collecting multi-source operational data in real time includes: A data collection period of k is set with an annual data collection cycle, where k is a natural number greater than 5. Data collection equipment is deployed at key monitoring points at each level of the power grid to collect multi-source operational data during the corresponding data collection period. The multi-source operational data includes real-time load data, real-time environmental data, equipment operating status data, and power grid topology dynamic data.
5. The power load forecasting method based on big data according to claim 4, characterized in that, The process of constructing a distributed collaborative prediction network based on multi-source operational data and the spatial location distribution of intelligent sensing nodes includes: Establish a three-dimensional spatial coordinate system and map all intelligent sensing nodes onto this three-dimensional spatial coordinate system to form a node spatial distribution map; The adjacent nodes of the prediction node are determined. Prediction nodes that are within a preset distance of kilometers and belong to the same or adjacent power supply areas are considered to be adjacent nodes. Each prediction node establishes a direct data interaction link only with its adjacent nodes, forming a distributed mesh topology.
6. The power load forecasting method based on big data according to claim 5, characterized in that, The process of setting up a multi-timescale and multi-spatial dimension partitioning mechanism to classify distributed collaborative prediction networks into different monitoring areas includes: The multi-timescale division mechanism combines the actual application needs of power load forecasting, divides the forecasting timescale into three timescales: short-term, medium-term, and long-term, and sets corresponding forecasting time points for each timescale. The multi-spatial dimension partitioning mechanism adopts a two-layer spatial partitioning strategy. If the load density of two adjacent basic monitoring areas is lower than a preset threshold and the load feature similarity is higher than a preset value, the two monitoring areas are merged into one monitoring area; otherwise, the monitoring areas are split. Based on the monitoring area determined by the multi-spatial dimension division mechanism, all prediction nodes in the distributed collaborative prediction network are classified into the corresponding monitoring areas according to their spatial location and power grid level.
7. The power load forecasting method based on big data according to claim 6, characterized in that, The process of configuring initial model parameters for each forecast node based on historical load data includes: Based on the prediction node's level, the type of monitored object, and the corresponding prediction time scale, suitable model components are selected from the prediction model component library. The historical load data corresponding to the prediction node is divided into a training set and a validation set in a 7:3 ratio. All parameter combinations are traversed within the preset parameter range. The prediction error is calculated through the validation set, and the parameter combination with the smallest prediction error is selected as the initial model parameters.
8. The power load forecasting method based on big data according to claim 7, characterized in that, The process of calling multi-source operational data, labeling load trends in the monitored area, and then matching and generating a risk diffusion network in the prediction model component library based on the load trend labels includes: Short-term load change rate sequences are obtained from multi-source operation data, and load trend curves are fitted by combining real-time load data and historical load data to obtain the slope of the trend curve. By combining the load change rate sequence with the slope of the trend curve, the load trend is divided into rapid growth trend, slow growth trend, rapid decline trend, slow decline trend, stable trend, and fluctuating trend. For different load trends, the corresponding prediction model components are matched. The nodes of the risk diffusion network include prediction nodes, error monitoring nodes, and regional collaboration nodes. Taking the monitoring area as a unit, the matched prediction model components are bound to the corresponding prediction nodes, the prediction nodes are connected to the error monitoring nodes, and each prediction node is connected to the regional collaboration nodes according to the error transmission coefficient to form a local risk diffusion network. The local risk diffusion networks of all monitoring areas are interconnected to form a global risk diffusion network.
9. The power load forecasting method based on big data according to claim 8, characterized in that, Based on the propagation and amplification process of prediction errors in risk diffusion networks, the process of adaptively optimizing model parameters and providing early warnings for prediction nodes includes: The prediction error of each error monitoring node is monitored in real time. If the prediction error of a prediction node exceeds the error threshold, the error propagation path tracing is triggered. Through the risk diffusion network, the neighboring prediction nodes of the prediction node and the downstream nodes affected by its error are found to determine the propagation range and main propagation path of the error. The optimization is verified using multi-source running data at the next prediction time point. If the prediction error is lower than the threshold, the optimization is considered effective. If the prediction error is still higher than the threshold, the above optimization process is repeated until the prediction error meets the requirements. The warning levels are divided according to the magnitude of the prediction error, the range of error propagation, and the scale of the impact load. When the risk diffusion network predicts that the error will reach the triggering conditions of the corresponding warning level, the warning is triggered in advance and the warning information is sent. Then, corresponding response measures are formulated for the warning information of different warning levels.
10. A big data-based power load forecasting system, used to implement the big data-based power load forecasting method according to any one of claims 1-9, characterized in that, It includes a power grid data acquisition module, a power grid status analysis module, and a load early warning analysis module; The power grid data acquisition module is used to deploy intelligent sensing nodes at all levels of the power grid, collect multi-source operation data in real time, acquire historical load data and related historical environmental records, and build a predictive model component library based on historical load data and historical environmental records. The power grid status analysis module is used to construct a distributed collaborative prediction network based on multi-source operation data and the spatial location distribution of intelligent sensing nodes. It sets up a multi-time scale and multi-spatial dimension division mechanism to classify the distributed collaborative prediction network into different monitoring areas and configure initial model parameters for each node according to historical load data. The load early warning analysis module is used to label the load trend of the monitoring area based on multi-source operation data, and then match and generate a risk diffusion network in the prediction model component library based on the load trend label. Based on the propagation and amplification process of prediction error in the risk diffusion network, the module performs adaptive optimization of model parameters and early warning for prediction nodes.