A power grid load prediction and dynamic scheduling optimization method based on big data
By employing multi-dimensional data acquisition and processing, attention-enhanced prediction models, and multi-objective scheduling optimization, the system addresses the issues of data systematization and insufficient prediction models in power grid load forecasting. This achieves greater accuracy in power grid load forecasting and greater rationality in scheduling schemes, ensuring the stability and security of power grid operation and supporting full-process data traceability.
Patent Information
- Application Number
- CN202511357593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Existing technologies for power grid load forecasting suffer from several problems, including a lack of systematic and comprehensive multi-source data processing capabilities, insufficient depth of feature fusion in forecasting models and a lack of time-specific targeting, and the absence of a complete optimization system encompassing forecasting, scheduling, adjustment, and tracing.
Raw data from the user side, equipment side, and environment side are collected using multi-dimensional sensing devices. Through processing of missing value classification and imputation, outlier removal by the 3σ criterion, and Z-Score standardization, an attention-enhanced load forecasting model is constructed. The key influencing factors are screened by combining LSTM network and random forest, and an attention mechanism is introduced for weighted fusion to construct a multi-objective scheduling function of "supply and demand balance first, energy consumption optimization second". An improved whale optimization algorithm is used to solve the function, and real-time closed-loop adjustment is achieved through high-frequency load monitoring. The data is archived to a time-series database to support traceability.
It achieves consistency and availability of multi-source data quality, improves the accuracy of load forecasting and the rationality of scheduling schemes, ensures the stability and security of power grid operation, and supports full-process data traceability.
Smart Images

Figure CN120855327B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power grid load prediction, in particular to a power grid load prediction and dynamic scheduling optimization method based on big data. BACKGROUND
[0002] With the increase of new energy grid connection ratio and the diversification of user-side power load (such as electric vehicle charging, industrial and commercial flexible load, etc.), the volatility and uncertainty of power grid load have significantly increased. Precise load prediction and dynamic scheduling optimization have become the core demand to ensure the safe and stable operation of power grid and improve energy utilization efficiency. In the current power grid operation, the accuracy of load prediction results directly affects the power output allocation and equipment operation and maintenance plan formulation, and the rationality of scheduling scheme is related to the physical safety of equipment (such as transformer load rate and line current control) and the economic efficiency of system operation (such as energy consumption optimization), so it is urgent to build a whole-process technical scheme covering "data processing-load prediction-scheduling optimization-closed loop adjustment-data tracing".
[0003] For example, Chinese patent CN202411002209.1 discloses a power load prediction method and system based on big data driving. The method obtains regional power station data, divides regional power consumption types, and evaluates regional power complexity. Then, based on the preset complexity threshold, time series load analysis is performed to generate high complexity regional power peak load data and low complexity regional power base load data. The core is to improve the prediction reliability by combining data classification and complexity evaluation with historical real-time data. Chinese patent CN202510413310.4 discloses a power grid load prediction improvement method and device based on AI intelligent algorithm. The method constructs an AI intelligent algorithm library, clusters the power grid load by industry, selects algorithms for load prediction by industry, and then constructs a temperature correction model to combine predicted weather data for real-time correction of industry load prediction data. The focus is on using industry classification and environmental factors to optimize AI prediction effect.
[0004] Although the above technical solutions have corresponding design advantages in data classification and algorithm adaptation of power load prediction, the above technical solutions still have the following technical defects: firstly, the multi-source data processing lacks systematicness and full-dimensional coverage capability: Chinese patent CN202411002209.1 only classifies and evaluates the complexity of regional power station data, does not include user-side power consumption parameters, equipment-side operation state and other key data, and does not establish a classification filling strategy and an abnormal value standardization elimination mechanism when data is missing; Chinese patent CN202510413310.4 introduces temperature and other environmental factors for correction, but still does not form a full-dimensional data processing flow covering the user side, the equipment side and the environment side, and it is difficult to guarantee the quality consistency and availability of multi-source data; secondly, the feature fusion depth of the prediction model is insufficient and the time period is not targeted: Chinese patent CN202411002209.1 does not realize the deep fusion of time series load features and multi-dimensional key influence factor features, and does not design a differentiated feature weight distribution strategy for peak power consumption period; Chinese patent CN202510413310.4 only optimizes the prediction by industry clustering and temperature correction, and does not effectively integrate time series features and equipment operation, user behavior and other influence factor features, which cannot accurately capture the strong fluctuation characteristics of peak period load; thirdly, a "prediction-scheduling-adjustment-tracing" full-chain optimization system is not constructed: the above two patents only focus on the load prediction link and do not extend to the design of dynamic scheduling scheme of power grid, and lack multi-objective scheduling optimization logic with equipment physical safety as the priority constraint; at the same time, there is no real-time closed-loop adjustment mechanism after the prediction deviation occurs, and there is no design for archiving and tracing of key data in the whole process, which cannot meet the full-scene operation needs of power grid from load prediction to scheduling execution, deviation correction and data tracing. In view of this, we propose a power grid load prediction and dynamic scheduling optimization method based on big data. SUMMARY
[0005] The purpose of the present application is to provide a power grid load prediction and dynamic scheduling optimization method based on big data, to solve the problems of lack of systematicness and full-dimensional coverage capability in multi-source data processing, insufficient feature fusion depth of the prediction model and lack of time period targeting, and not constructing a "prediction-scheduling-adjustment-tracing" full-chain optimization system in the background art.
[0006] To solve the above technical problems, the present application provides a power grid load prediction and dynamic scheduling optimization method based on big data, comprising the following steps:
[0007] S100, power grid multi-source data preprocessing: multi-dimensional sensing equipment is used to collect original data of user side, equipment side and environment side, and the data is processed through a combination process of "missing value classification filling + 3σ criterion abnormal value elimination + Z-Score standardization";
[0008] S200, attention-enhanced load forecasting model construction: based on the standardized data set of S100, the time series load characteristics are extracted through the LSTM network, and the key influence factor characteristics are selected through the random forest; the attention mechanism is introduced to weight and fuse the two types of characteristics, and the feature weight of the power consumption peak period is highlighted; the Adam optimizer is used to train the model, the mean absolute percentage error and the root mean square error are used to verify the model precision, and the load prediction result is output;
[0009] S300, hierarchical constraint improved whale optimization scheduling scheme generation: based on the load prediction result of S200 and the real-time operation parameters of the power grid, a multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization auxiliary" is constructed; and an improved whale optimization algorithm is used to solve, the algorithm introduces an adaptive weight factor to adjust the search strategy, and the priority of meeting the physical safety limit of the equipment is judged through hierarchical constraint, and a scheduling scheme is generated;
[0010] S400, real-time closed-loop adjustment: high-frequency load monitoring is used to calculate the deviation rate; when the deviation exceeds the reasonable range, the combination of "incremental data training + scheduling parameter fine-tuning" is used to make the deviation fall within the reasonable range;
[0011] S500, data archiving and tracing: using a time series database to classify and archive the key data of S100-S400; the data retention period meets the power grid equipment operation and maintenance tracing requirements, supports time, region, and equipment number retrieval, and ensures data traceability.
[0012] As a further improvement of the technical solution, in the S100, the process of collecting user-side, device-side, and environment-side original data using multi-dimensional sensing devices includes the following steps:
[0013] S110.1, selection and matching of collection equipment: for user-side original data collection, multi-dimensional sensing devices suitable for different power consumption scenarios are selected to collect real-time active power, reactive power, and cumulative power consumption; for device-side original data collection, sensing devices matching the types of transformers, transmission lines, and switch devices are selected to collect transformer oil temperature, line current, line voltage, and switch opening and closing state; for environment-side original data collection, sensing devices suitable for outdoor power grid scenarios are selected to collect air temperature, relative humidity, and wind speed;
[0014] S110.2, collection process data integrity monitoring: during the collection process, the data stream transmission state of the sensing device is monitored in real time, and when it is detected that a single device has not uploaded data for M consecutive times or the missing data field ratio exceeds the preset threshold, it is determined that there is a collection anomaly, and a device self-checking program is triggered;
[0015] S110.3, Abnormal emergency treatment: if the self-checking procedure still cannot recover normal collection, enable the standby sensing device in the same area to continue collection, and record the number of abnormal devices, abnormal occurrence time and abnormal type;
[0016] S110.4, Raw data identification and temporary storage: the normally collected raw data are identified according to "collection object type-device unique number-collection time", and temporarily stored in the nearest edge storage node. Data integrity is verified by using data check code during temporary storage.
[0017] As a further improvement of the technical solution, in the S100, the "missing value classification filling + 3σ criterion outlier removal + Z-Score standardization" combined process includes the following steps:
[0018] S120.1, Missing value classification filling: the raw data temporarily stored in S110.4 are classified according to user side, device side and environment side. For missing values in each type of data, linear interpolation method is used to fill in the missing values for short-term continuous missing, and mean value of the same type of data in the same region in the same period is used to fill in the missing values for long-term continuous missing.
[0019] S120.2, 3σ criterion outlier removal: for the data filled in S120.1, the mean value and standard deviation of each type of data are calculated according to user side, device side and environment side respectively. The values deviating from the mean value by more than 3 times the standard deviation are determined as outliers and removed.
[0020] S120.3, Z-Score standardization: for the data after removing outliers in S120.2, Z-Score standardization processing is performed according to the type of data, so as to eliminate the magnitude difference of different dimension data and make the processed data dimensionless.
[0021] As a further improvement of the technical solution, in the S200, the process of extracting time series load features by LSTM network and screening key influence factor features by random forest includes the following steps:
[0022] S210.1, LSTM network feature extraction: the standardized data set of S100 is divided into continuous input units according to time sequence , as time index, input LSTM network; the network controls new information inflow through input gate , forget gate clears redundant information, output gate selects output information, generates hidden layer state , finally outputs time series load feature vector ;
[0023] S210.2, Random Forest Feature Screening: Combining environmental features and equipment operation features into a feature set of influencing factors. With load data Input a random forest model and calculate the node splitting gain for each feature. Determine the contribution level, and select features that meet preset conditions to form a feature set of key influencing factors. ;
[0024] S210.3 Feature Dimension Alignment: Aligning Time-Series Load Feature Vectors Based on Timestamps Key Influencing Factors Feature Set By unifying the time scales and matching the dimensions of the two sets through dimensional expansion or compression, a set of features to be fused is formed. .
[0025] As a further improvement to this technical solution, in step S200, the process of introducing an attention mechanism to weightedly fuse the two types of features includes the following steps:
[0026] S220.1 Feature Correlation Analysis: Calculate the set of features to be fused Various characteristics and actual load changes correlation coefficient To form a correlation matrix ;
[0027] S220.2 Attention Weight Allocation: Based on the Relevance Matrix During peak electricity consumption periods Feature assignment weights For off-peak hours Feature assignment weights ,in To form a dynamic weight matrix ;
[0028] S220.3, Feature-weighted fusion: This involves using matrix operations to combine the dynamic weight matrix... With the feature set to be fused Fusion to generate a comprehensive feature matrix .
[0029] As a further improvement to this technical solution, in step S200, the process of training the model using the Adam optimizer and verifying the model accuracy through mean absolute percentage error and root mean square error includes the following steps:
[0030] S230.1, Model Training Configuration: Integrating the Feature Matrix Divided into training set and verification set Set the first moment coefficients of the Adam optimizer second moment coefficient , initializing model parameters ;
[0031] S230.2, iterative training: input the training set into the model, update the model parameters iteratively through the optimizer ; , calculate the loss value after each round of iteration using the validation set ; , stop training when the loss value does not decrease for a preset number of consecutive rounds, and save the optimal parameters ;
[0032] S230.3, accuracy verification: input the test set into the optimal model, calculate the predicted load ; , and the mean absolute percentage error and root mean square error of the actual load , output the load prediction result that meets the error requirement .
[0033] As a further improvement of the technical solution, in the S300, the process of constructing the multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization as auxiliary" includes the following steps:
[0034] S310.1, scheduling parameter definition: define the load prediction result output by S200 as the baseline load demand , the real-time operation parameters of the power grid include the output of each power source node , the transmission line loss coefficient , and the device energy consumption characteristic parameters ;
[0035] S310.2, supply-demand balance target construction: take the matching degree of the regional total power supply and the baseline load demand as the core target, and the matching degree is represented by the supply-demand deviation coefficient , and the deviation threshold is set based on the deviation degree calculation ; ; ; , when the deviation is less than the threshold, it is determined that the basic supply-demand balance is met ;
[0036] S310.3, energy consumption optimization target construction: take the minimum of the grid comprehensive energy consumption as the auxiliary target , establish a mapping relationship based on , , , and the scheduling duration , and quantify the correlation between output and energy consumption through energy consumption characteristic parameters
[0037] S310.4 Multi-objective integration: Introducing priority weights , and The supply and demand balance objective and the energy consumption optimization objective are weighted and integrated to form a multi-objective scheduling function. .
[0038] As a further improvement to this technical solution, in step S300, the process of using the improved whale optimization algorithm to solve the problem and generating a scheduling scheme through hierarchical constraints includes the following steps:
[0039] S320.1 Algorithm Parameter Initialization: Set the population size for the improved whale optimization algorithm. Maximum number of iterations Initialize population individuals (Each individual corresponds to a set of scheduling parameter combinations);
[0040] S320.2 Adaptive Weight Factor Adjustment: Introducing a weight factor that dynamically changes with the iteration process. Early iteration hour To enhance global search capabilities, larger values are selected in later iterations. hour To improve the accuracy of local optimization, smaller values are selected. Adjust the individual search step size;
[0041] S320.3, Layered Constraint Judgment: The first layer constraint is the physical safety limit of the equipment. This includes the upper limit of transformer load rate. Rated current of the line The second layer of constraints consists of system stability indicators. Including voltage deviation range Frequency deviation range First verify during the solution process. If the condition is not met, discard the individual directly; if it is met, then re-verify. ;
[0042] S320.4 Optimal Scheduling Scheme Generation: Through multi-objective scheduling functions Calculate the fitness value of each individual in the population, iteratively update and retain the individual with the best fitness, and finally output the result that simultaneously satisfies... , The combination of constrained scheduling parameters forms a scheduling scheme.
[0043] As a further improvement to this technical solution, in S400, the real-time closed-loop adjustment process includes the following steps:
[0044] S410.1, high-frequency load monitoring and deviation rate calculation: according to the multi-dimensional sensing device acquisition dimension of S110.1, set the load monitoring frequency of not less than 12 times per hour, real-time acquisition of actual load data of power grid, combined with the corresponding period load prediction result output by S230.3, the relative deviation degree of actual load and predicted load is used to calculate the deviation rate;
[0045] S410.2, deviation range judgment: preset deviation reasonable range threshold, which is set according to the control logic of system operation stability index in S320.3, if the deviation rate is within the threshold range, maintain the current dispatching scheme generated by S300; if the deviation rate exceeds the threshold range, trigger the closed-loop adjustment mechanism;
[0046] S410.3, incremental data training: collect the environment side, equipment side and user side data collected by S110.1 to form an incremental data set, and use the incremental data set to supplement the training of the attention enhanced load prediction model in S200, update the model parameters to improve the short-term prediction accuracy;
[0047] S410.4, dispatching parameter fine tuning: based on the prediction result after incremental training, the power source node output distribution ratio involved in S310.1 is slightly corrected, the correction amplitude is increased accordingly with the increase of the deviation rate, and the corrected parameters need to meet the equipment physical safety limit requirements in S320.3;
[0048] S410.5, deviation backfall verification: according to the monitoring frequency set by S410.1, the actual load is reacquired and the deviation rate is calculated, the steps of S410.3 to S410.4 are repeated, until the deviation rate falls within the reasonable range threshold, and the final dispatching parameters and adjustment process data are synchronized to the time series database of S500.
[0049] As a further improvement of the technical solution, in the S500, the data archiving and tracing process includes the following steps:
[0050] S510.1, time series database configuration and archiving classification design: use a time series database that supports high-frequency data writing and time series indexing, and divide the archiving categories according to the stages of S100-S400, wherein S100 archives the standardized data set after preprocessing and acquisition of abnormal records, S200 archives the load prediction model parameters, prediction results and accuracy verification data, S300 archives the multi-objective dispatching function parameters, optimization algorithm iteration records and final dispatching scheme, and S400 archives the load monitoring data, deviation rate calculation results and closed-loop adjustment records;
[0051] S510.2, data retention period setting: determine the retention period according to the demand of power grid equipment operation and maintenance traceability, wherein the associated data retention period of the transformer, power transmission line and switching device is not shorter than the design life of the corresponding device, the retention period of temporary scheduling adjustment record is not shorter than 3 power grid operation and maintenance cycles, ensuring that the device full life cycle traceability and operation and maintenance review requirements are covered;
[0052] S510.3, multi-dimensional search function configuration: build multi-dimensional search index in time series database, time dimension supports accurate to minute level time period search, region dimension is associated with power grid partition code, device number dimension maps the unique identification number of sensor device, transformer and line in S110.1, and the search response time is not more than 10 seconds;
[0053] S510.4, archive data traceability guarantee: regularly check the integrity of the archived data, compare the consistency of the archived data with the original generated data in S100-S400 stage, generate a check log and archive it; if data is missing or inconsistent, trigger the automatic supplement mechanism to extract data from the temporary storage node of S100-S400 to rearchive, ensure the accuracy and integrity of the traceability data.
[0054] Compared with the prior art, the beneficial effects of the present application are:
[0055] 1. The present application can systematically solve the problem of uneven quality of multi-source data by carrying out collection integrity monitoring, abnormal emergency treatment and structured identification temporary storage on multi-source original data of user side, device side and environment side, and combining the combined preprocessing process of "missing value classification filling + 3σ criterion abnormal value elimination + Z-Score standardization", effectively guaranteeing the consistency and availability of data, and providing reliable data support for subsequent load prediction and dispatching optimization;
[0056] 2. The present application can more accurately capture the load change law in different periods and improve the adaptability of load prediction results to actual power grid operation by constructing an attention-enhanced load prediction model, using LSTM network to extract time series load characteristics, random forest to select key influence factor characteristics, and then introducing attention mechanism to highlight the feature weight of peak electricity consumption period, and combining Adam optimizer training and precision verification;
[0057] 3. The present application can meet the power supply and demand balance while taking into account energy consumption optimization by constructing a multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization auxiliary", using an improved whale optimization algorithm with adaptive weight factor to solve, and ensuring the physical safety limit of the device by hierarchical constraint priority, so as to ensure that the scheduling scheme meets the safety requirements of the device, and improves the rationality and safety of power grid scheduling;
[0058] 4. The application calculates the deviation rate through high-frequency load monitoring. When the deviation exceeds the reasonable range, the closed-loop adjustment is realized in the mode of "incremental data training and updating prediction model + fine-tuning scheduling parameters", which can timely correct the deviation between prediction and actual operation, reduce the influence of deviation on power grid operation, and maintain the stability of power grid operation;
[0059] 5. The application classifies and archives the key data of the whole process through the time series database, sets the data retention period in line with the operation and maintenance requirements, configures the multi-dimensional retrieval function and carries out the regular integrity check and automatic supplement, which can realize the traceability of the whole process data and provide data support for the operation and maintenance review and technical optimization of the power grid equipment. BRIEF DESCRIPTION OF DRAWINGS
[0060] Figure 1 The figure is a schematic diagram of the overall method steps of the application;
[0061] Figure 2 The figure is a schematic diagram of the step of collecting raw data in the application;
[0062] Figure 3 The figure is a schematic diagram of the step of data processing in the application;
[0063] Figure 4 The figure is a schematic diagram of the step of real-time closed-loop adjustment in the application;
[0064] Figure 5 The figure is a schematic diagram of the step of data archiving and tracing in the application. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0066] As shown in the figure, the embodiment provides a power grid load prediction and dynamic scheduling optimization method based on big data, which comprises: Figures 1-3
[0067] S100, power grid multi-source data preprocessing: multi-dimensional sensing equipment is used to collect user side, equipment side and environment side raw data, and a combined process of "missing value classification and filling + 3 sigma rule for removing outliers + Z-Score standardization" is used for processing;
[0068] In this step, in the S100, the process of collecting user side, equipment side and environment side raw data by multi-dimensional sensing equipment comprises the following steps:
[0069] S110.1, selection and matching of collection equipment: for user-side raw data collection, multi-dimensional sensing equipment suitable for different power consumption scenarios is selected to collect real-time active power, reactive power, and cumulative power consumption; for device-side raw data collection, sensing equipment matching the types of transformers, transmission lines, and switch devices is selected to collect transformer oil temperature, line current, line voltage, and switch opening and closing state; for environment-side raw data collection, sensing equipment suitable for outdoor power grid scenarios is selected to collect air temperature, relative humidity, and wind speed;
[0070] As a further illustration of this step, the selection and matching of collection equipment in this step specifically includes:
[0071] User-side sensing equipment: for residential scenarios, intelligent electric meters with RS485 interfaces are selected (supporting active power / reactive power collection, with a sampling frequency of 15 minutes / time); for industrial and commercial scenarios, industrial-grade three-phase electric meters are selected (with a sampling frequency of 5 minutes / time, supporting multi-loop simultaneous collection).
[0072] Device-side sensing equipment: platinum resistance temperature sensors are used for transformer oil temperature collection; electronic current transformers and voltage transformers are used for line current / voltage collection, with a sampling frequency of 1 minute / time; passive contact sensors are used for switch opening and closing state collection, with real-time uploading when the state changes.
[0073] Environment-side sensing equipment: temperature and humidity sensors with rain and snowproof housings and ultrasonic wind speed sensors are selected, installed in outdoor areas of substations and transmission line towers, with a sampling frequency of 10 minutes / time.
[0074] S110.2, monitoring of data integrity during collection process: during the collection process, the data stream transmission state of the sensing equipment is monitored in real time, and when it is detected that a single device has not uploaded data for M consecutive times or the missing proportion of uploaded data fields exceeds the preset threshold, it is determined that there is an abnormality in collection, triggering a device self-checking program.
[0075] As a further illustration of this step, M is preferably 3 times (i.e., no data uploaded for 3 consecutive sampling periods), which takes into account the communication stability and real-time requirements of power grid data collection. Most sensing equipment has a sampling period of 5-15 minutes, and not uploading data for 3 consecutive times provides both a fault tolerance space for temporary communication fluctuations (such as network delays) and timely alarms in the event of persistent device failures, ensuring the continuity and reliability of data collection.
[0076] Further, the data field missing ratio preset threshold is preferably 30% (for example, if a single piece of data contains 10 fields, and more than 3 fields are missing, it is determined to be abnormal), from the balance between the importance and integrity of the fields of the power grid multi-source data, if a single piece of data is missing more than 30%, the remaining effective fields are difficult to support subsequent multi-dimensional analysis (such as load prediction requiring the linkage of multiple types of fields such as users, devices, and environment), this threshold can avoid over-determining abnormality due to a small amount of missing, and can ensure that the data quality meets the basic requirements of model training and dispatching decision-making;
[0077] Further, the data stream transmission state is realized through a heartbeat detection mechanism of the edge gateway, a state packet is sent every 5 seconds, and if there is no response after timeout, an abnormality determination is triggered.
[0078] S110.3, collecting abnormal emergency processing: if the self-checking procedure still cannot restore normal collection, a backup sensing device pre-deployed in the same region is enabled to continue collecting in place of the abnormal device, and the number of the abnormal device, the abnormal occurrence time and the abnormal type are recorded;
[0079] As a further description of this step, the collecting abnormal emergency processing mechanism in this step specifically includes:
[0080] The backup sensing device is deployed in a “1 main 1 backup” or “3 main 1 backup” ratio (the former is used in important load areas), the backup device and the main device collect parameters in the same way, and the backup device is in a hot standby state;
[0081] The abnormal type includes “communication interruption”, “sensor failure” and “data out of range”, which correspond to different alarm levels (for example, communication interruption is a secondary alarm, and sensor failure is a primary alarm), and the alarm information is synchronously uploaded to the power grid monitoring platform.
[0082] S110.4, raw data identification and temporary storage: the normally collected raw data are identified in a structured manner according to “collection object type-device unique number-collection time”, and are temporarily stored in the nearest edge storage node, and the data integrity is verified by using a data check code during the temporary storage process.
[0083] As a further description of this step, the temporary storage specification is specifically as follows: the edge storage node uses an industrial-grade SD card or a local hard disk, the storage capacity of a single node is greater than or equal to 1 TB, the temporary storage period is 24 hours, and after the period expires, the data are automatically compressed and uploaded to the regional power grid data center, and a local backup is kept for 30 days.
[0084] Further, the embodiment also provides an example of structured identification, for example:
[0085] User side: user side-residential area A-meter ID001-202310010800
[0086] Device side: Device side - transformer T102 - temperature sensor S203 - 202310010800
[0087] Environment side: Environment side - transmission line L506 - temperature and humidity sensor H701 - 202310010800.
[0088] In this step, in the S100, the "missing value classification filling + 3σ criterion outlier removal + Z-Score standardization" combined process includes the following steps:
[0089] S120.1, missing value classification filling: classify the raw data stored in S110.4 by user side, device side, and environment side. For missing values in each type of data, short-term continuous missing is filled by linear interpolation of adjacent valid data before and after the data, and long-term continuous missing is filled by the mean value of the same type of data in the same period.
[0090] As a further description of this step, missing value filling specifically includes:
[0091] Short-term continuous missing: refers to missing duration ≤1 hour (e.g. 15-minute sampling interval, 4 consecutive data missing). When filling, take the nearest two valid data before and after the missing data, and calculate the missing value in the middle according to the time interval ratio (e.g. 8:00 and 8:30 data valid, 8:15 missing value takes the middle value of the two).
[0092] Long-term continuous missing: refers to missing duration >1 hour. When filling, select historical data of the same type of equipment in the same period on the same date in the same region (e.g. 8:00 missing on October 1, take 8:00 data on September 28, 29, and October 2, 3), calculate the average value as the filling value.
[0093] S120.2, 3σ criterion outlier removal: for the data filled in S120.1, calculate the mean and standard deviation of each type of data by user side, device side, and environment side respectively, and judge the values deviating from the mean more than 3 times the standard deviation as outliers and remove them.
[0094] As a further description of this step, 3σ criterion outlier removal specifically includes:
[0095] Calculate the average value and fluctuation range (standard deviation) of each type of data by "day" as the unit, for example, the daily average value of a certain transformer oil temperature is 45℃, and the fluctuation range is 5℃, then the values exceeding 45±15℃ are judged as outliers and removed.
[0096] For data with large fluctuations such as wind speed, use a 24-hour sliding window to calculate (update the average value and fluctuation range every hour), to ensure that the outlier judgment adapts to the data characteristics.
[0097] The outliers removed are marked with the location and reason (such as "out of 3 sigma range") in the original data set, and the record is kept for subsequent data quality analysis.
[0098] S120.3, Z-Score standardization: for the data after removing outliers in S120.2, Z-Score standardization is performed according to the data type to eliminate the magnitude difference of different dimensions of data, so that the processed data is dimensionless.
[0099] As a further description of this step, Z-Score standardization specifically includes:
[0100] The magnitude difference of different data is eliminated during processing (such as power unit kW and temperature unit ℃), and the data is converted into dimensionless relative values. The converted data is usually distributed between -3 and 3, and the values outside this range are uniformly processed as boundary values (i.e. less than -3 is processed as -3, and greater than 3 is processed as 3), ensuring stable data distribution and facilitating subsequent model training.
[0101] S200, attention-enhanced load forecasting model construction: based on the standardized data set of S100, time series load features are extracted through an LSTM network, and key influence factor features are selected through a random forest; an attention mechanism is introduced to weight and fuse the two types of features, highlighting the feature weight of the peak electricity consumption period; an Adam optimizer is used to train the model, and the model accuracy is verified through the mean absolute percentage error and the root mean square error, and the load prediction result is output;
[0102] In this step, in the S200, the process of extracting time series load features through an LSTM network and selecting key influence factor features through a random forest includes the following steps:
[0103] S210.1, LSTM network feature extraction: the standardized data set of S100 is divided into continuous input units according to the time sequence , for time index, input LSTM network; the network controls new information inflow through input gate , forget gate to clear redundant information, output gate to filter output information, and generate hidden layer state , finally output time series load feature vector ;
[0104] As a further description of this step, in the S210.1 LSTM network feature extraction of the present embodiment, the calculation of the key gate and the hidden layer state specifically includes:
[0105] Input gate calculation: the input gate is used to control the inflow of new information into the LSTM cell, and the formula is: ;in, for The input gate output value at any time (range 0-1, where 0 means completely blocking information and 1 means completely allowing information to flow in). It is the sigmoid activation function; For input layer (input unit) The weight matrix from the input gate; for Time-based input unit (including vectors of standardized data from the user side, device side, and environment side); The hidden layer of the previous time step ( The weight matrix from the input gate; for Hidden layer state vector at any given time; The input gate bias term (a constant vector used to adjust the activation function offset); passed through the input unit Compared to the previous hidden layer state A linear combination of these values is mapped to a value between 0 and 1 via the sigmoid function, thereby controlling the intensity of the inflow of new information.
[0106] Forgotten Gate Calculation: The forget gate is used to remove redundant historical information within cells. The formula is: ;in, for Output value of the time-forget gate (range 0-1, 0 means completely retaining historical information, 1 means completely forgetting historical information). This is the weight matrix from the input layer to the forget gate; This is the weight matrix from the previous hidden layer to the forget gate; This is the forgetting gate bias term. Its computational logic is consistent with the input gate structure. Through linear combination and sigmoid mapping, it determines the proportion of historical information retained within the cell. For example, in load forecasting, the "historical load patterns during off-peak periods" can be obtained through... Forget about things appropriately to avoid interfering with current predictions;
[0107] Cell state update Calculation: Cell state is the core memory unit of LSTM, and the update formula is as follows: ;in, for Time-based cell state update value (range -1 to 1, representing newly generated memory information); It is the hyperbolic tangent activation function (used to map numerical values to -1-1, enhancing gradient propagation stability). This is the weight matrix from the input layer to the cell state; This is the weight matrix from the previous hidden layer to the cell state; This refers to cell state bias terms; new memory information is generated through linear combination, after... Function normalization provides a basis for subsequent cell state updates.
[0108] Cell state With hidden layer state Calculation: The formula for the final cell state and the hidden layer state is: , ;in, for The final cell state at any given moment (preserving core memory information); for Cell state at any given moment; for The output value of the time-limited output gate is calculated in the same way. The formula is ,in This is the weight matrix from the input layer to the output gate. This is the weight matrix from the previous hidden layer to the output gate. (for output gate bias terms). for Hidden layer state at any given time (i.e., time-series load feature vector) The core source); its computational logic is: first through the forget gate reserve Valid historical information in the input gate New information after filtering ,get Then through the output gate filter Key information in the middle, After normalization, we get Ultimately Organized into a 32-dimensional time-series load feature vector .
[0109] S210.2, Random Forest Feature Screening: Combining environmental features and equipment operation features into a feature set of influencing factors. With load data Input a random forest model and calculate the node splitting gain for each feature. Determine the contribution level, and select features that meet preset conditions to form a feature set of key influencing factors. ;
[0110] As a further explanation of this step, in the random forest feature selection in S210.2 of this embodiment, the node splitting gain... The calculation of (information gain ratio) specifically includes:
[0111] Information gain ratio is used to measure the impact of features on load data. The formula for the classification contribution is: ;
[0112] In this formula, information gain ,in For the feature set of influencing factors For sample set Information gain (the larger the value, the greater the influence factor feature set) The higher the contribution to load forecasting, the better. The sample set for the random forest (including and The corresponding data, with a sample size of ); The information entropy of sample set D (reflecting) (uncertainty) Features under conditions Conditional entropy (reflecting known conditions) back (remaining uncertainty);
[0113] Split Information ;in, Features The number of value categories, For sample set Chinese characteristics Take the first The sample subset when the value is 1, for The number of samples, for The total number of samples; Used to measure characteristics The complexity of the subsets after splitting can avoid information gain bias towards features with too many values (such as "device number" which has many values but no actual predictive significance).
[0114] Conditional entropy ;in, For the feature set of influencing factors The number of possible values (e.g., "air temperature" is divided into 5 values based on intervals) =5); for Characteristic set of influencing factors Take the first A subset of samples with values; for Information entropy (calculated in the same way) );
[0115] In this formula, splitting information ; wherein, for measuring the dispersion degree of the value, avoiding the information gain from being biased to the features with too many values (such as "device number" which has too many values but no actual predictive significance);
[0116] First, the information entropy of the sample set is calculated , and then the sample subsets are divided according to the values of the influencing factor feature set , the conditional entropy and the split information are calculated, and finally the information gain ratio is used to filter the features - the of all features are sorted, and the top 60% of the features are selected to form the key influencing factor feature set (such as selecting 4 types of features from 7 types of features).
[0117] S210.3, feature dimension alignment: based on the timestamp, the time series load feature vector is unified with the time scale of the key influencing factor feature set , and the dimensions are expanded or compressed to match the dimensions of the two, forming a set of features to be fused .
[0118] In this step, in the S200, the process of introducing an attention mechanism for weighted fusion of the two types of features includes the following steps:
[0119] S220.1, feature correlation analysis: calculate the correlation coefficient of each feature in the set of features to be fused and the actual load change , forming a correlation matrix ;
[0120] As a further description of this step, in the S220.1 feature correlation analysis of the present embodiment, the calculation of the correlation coefficient (Pearson correlation coefficient) specifically includes:
[0121] The Pearson correlation coefficient is used to measure the linear correlation degree of a feature in the set of features to be fused and the actual load change , and the formula is as follows:
[0122] ;
[0123] is the correlation coefficient (range -1-1, 1 indicates complete positive correlation, -1 indicates complete negative correlation, and 0 indicates no linear correlation); For the number of samples (e.g., 30 days of 15-minute interval data, For the value of the i-th sample of a certain feature (e.g., the 15-minute average of "air temperature") in For the sample average of For the actual load change amount of the i-th sample, For the sample average of
[0124] The specific calculation logic is: through the numerator to calculate the covariance of the feature and ; arrange the values of all 32 features in in a "feature-feature" matrix to form a 32x32 correlation matrix (e.g., represents the value of the i-th feature and the j-th feature).
[0125] S220.2, attention weight allocation: based on the correlation matrix , the features in the power peak period are assigned weights , the features in the non-peak period are assigned weights , wherein , forming a dynamic weight matrix
[0126] As a further explanation of this step, in the attention weight allocation of S220.2 of the present embodiment, the calculation of the dynamic weights (high peak period) and (non-peak period) specifically includes:
[0127] Peak period weight calculation: ; wherein, is the power peak period (08:00-11:00, 18:00-22:00) The weights of each feature (ranging from 0 to 1, and the weights of all features) (sum of 1) The first during peak hours Features and The correlation coefficient (taken from) The corresponding time period in ) value); for Feature dimensions ( =32); For all 32 features during peak hours The sum of the absolute values;
[0128] Off-peak period weighting calculate: ;in, Off-peak hours (remove (outside of the time period) The weights of each feature (ranging from 0 to 1, and the weights of all features) (sum of 1) During off-peak hours Features and The correlation coefficient; For all 32 features during off-peak hours The sum of the absolute values;
[0129] Specifically, the calculation logic is as follows: by taking the absolute value of the correlation coefficient (ignoring the positive and negative correlation directions and only focusing on the correlation strength), and then performing normalization (dividing by all features). (The sum of the absolute values), ensuring that the sum of all feature weights is 1 within the same time period; due to peak periods Greater fluctuations The absolute value is usually higher than Therefore (such as a certain feature) =0.7, =0.2), ultimately will Arranged by time index, forming a 32×32 dynamic weight matrix. ( express Time of the first (Weights of each feature).
[0130] S220.3, Feature-weighted fusion: This involves using matrix operations to combine the dynamic weight matrix... With the feature set to be fused Fusion to generate a comprehensive feature matrix .
[0131] As a further explanation of this step, and as a further explanation of this embodiment, in the feature weighted fusion of S220.3 of this embodiment, the comprehensive feature matrix... The calculations specifically include:
[0132] Element-wise multiplication is used to fuse dynamic weights and features to be fused, as shown in the formula: ; for Time of the first The comprehensive eigenvalue of dimension (i.e.) (core elements) for Time of the first Dynamic weights of dimension (taken from) ); for Time of the first The dimensional features to be fused (taken from) );
[0133] The specific calculation logic is: for each time point Each dimension of features Using weights For original features Scaling is applied to features with high weight during peak hours (such as "air temperature" and "transformer oil temperature"). The value is amplified, and the characteristic of low weight during off-peak hours is that... The values are reduced to highlight the impact of key characteristics during peak periods on load forecasting, ultimately resulting in a 32-dimensional value. matrix.
[0134] It is understandable that for dynamic weight matrices... Its dimensions are The rationale for setting this dimension lies in: comprehensive features It is a 32-dimensional vector. This describes the strength of the attention association between each feature within these 32 dimensions and all other features—that is, the first feature in the matrix. Line 1 The elements of the column represent the first element. dimensional features for the 1st Attention weights for dimensional features, through The matrix structure can fully characterize the pairwise attention relationships between 32-dimensional features, enabling refined modeling of the internal correlations of features.
[0135] In this step, the process of training the model using the Adam optimizer and verifying the model accuracy through mean absolute percentage error and root mean square error in S200 includes the following steps:
[0136] S230.1, model training configuration: the comprehensive feature matrix is divided into training set and validation set , the first moment coefficient of Adam optimizer is set , the second moment coefficient is set , and the model parameters are initialized ;
[0137] As a further description of this step, the comprehensive feature matrix obtained after preprocessing is randomly divided according to the ratio of training set: validation set: test set = 7:2:1; if the data volume exceeds 100,000, the ratio of 8:1:1 is used to balance the accuracy of model training sufficiency and generalization ability evaluation. Among them, the training set is used for model parameter learning, the validation set is used for iterative process of hyperparameter (such as LSTM network layer number, random forest decision tree number) adjustment, and the test set is used for independent verification of the final model performance.
[0138] S230.2, iterative training: input the model with the training set , update the model parameters through the optimizer iteration , calculate the loss value with the validation set after each iteration , stop training when the loss value does not decrease for a continuous preset number of times, and save the optimal parameters ;
[0139] As a further description of this step, in S230.1-S230.2 of the embodiment, the parameter update of Adam optimizer and the loss function calculation specifically include:
[0140] On the one hand, it includes Adam optimizer parameter update: Adam optimizer updates model parameters through first moment (momentum) and second moment (adaptive learning rate) , the steps are as follows;
[0141] First moment estimation (momentum): ; wherein, is the first moment of the model parameters at the moment (reflecting the moving average of gradient); is the first moment decay coefficient (preset to 0.9, controlling the retention ratio of historical momentum); is the first moment at the moment ; is the gradient of the model parameters at the moment (derived by taking the derivative of the loss function with respect to , reflecting the influence degree of parameters on prediction deviation);
[0142] Second moment estimation (adaptive learning rate): ; where, is the second moment of model parameters at time t (reflecting moving average of squared gradient); is the second moment decay coefficient (preset as 0.999, controlling the retention ratio of historical squared gradient); is the second moment at time t; is the squared gradient at time t (element-wise square); is the squared gradient at time t (element-wise square);
[0143] First moment bias correction: ; is the corrected first moment (eliminating the bias of first moment towards 0 at the beginning of training); is the power of ; is the current iteration round;
[0144] Second moment bias correction: ; where, is the corrected second moment (eliminating the bias of second moment towards 0 at the beginning of training); is the power of .
[0145] Parameter update: ; where, is the updated model parameter at time t (including LSTM weights, random forest split thresholds, etc.); is the model parameter at time t is the learning rate (preset as 0.001, controlling the parameter update step size); is the square root of the corrected second moment (providing adaptive learning rate); is a small value to prevent the denominator from being 0 (preset as 10 -8 );
[0146] By preserving the continuity of gradient direction through the first moment (avoiding parameter oscillation), the second moment adjusts the learning rate according to the gradient size (small learning rate for large gradient, large learning rate for small gradient), and the correction term eliminates the bias at the beginning of training, achieving stable and efficient parameter update.
[0147] On the other hand, the loss function (mean square error) calculation:
[0148] The loss function is used to measure the deviation between the model prediction value and the actual value, and the formula is ; is The loss value of each iteration (the smaller the value, the higher the model prediction accuracy); For the validation set The number of samples (e.g.) =2000); For the model on the validation set The predicted load value for each sample; For the verification set The actual load value of each sample;
[0149] The specific calculation logic is as follows: Square the prediction bias for each sample (amplifying the impact of larger biases), then calculate the average to obtain the loss value for that iteration; when there are 5 consecutive iterations... No decrease (i.e.) When ), stop training and save the data. The parameters of the wheel are used as optimal parameters. .
[0150] S230.3, Accuracy Verification: The test set... Input the optimal model and calculate the predicted load. Compared with actual load The mean absolute percentage error and root mean square error are used to output load forecast results that meet the error requirements. .
[0151] As a further explanation of this step, in the accuracy verification of S230.3 in this embodiment, the calculation of the mean absolute percentage error and the root mean square error specifically includes:
[0152] On the one hand, this includes the calculation of the mean absolute percentage error (MAPE):
[0153] MAPE is used to measure the proportion of the predicted value to the actual value. The formula is: ;in, The mean absolute percentage error (in %, the smaller the value, the smaller the relative deviation). For the test set The number of samples (e.g.) =1000); For the model on the test set The predicted load value for each sample; For the test set The actual load value of each sample ( (to avoid the denominator being 0) It is the absolute value symbol;
[0154] On the other hand, this includes the calculation of the root mean square error (RMSE):
[0155] RMSE is used to measure the absolute deviation between a predicted value and the actual value. The formula is: ; wherein, is the root mean square error (unit consistent with load, such as kW, the smaller the value, the smaller the absolute deviation);
[0156] The specific calculation logic is: first, calculate the square sum and average of the prediction deviation of each sample (i.e. mean square error MSE), then take the square root to get the error value consistent with the load unit, which meets the demand of the power grid for absolute accuracy of load prediction.
[0157] S300, hierarchical constraint improved whale optimization scheduling scheme generation: based on the load prediction result of S200 and the real-time operation parameter of the power grid, a multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization as auxiliary" is constructed; and an improved whale optimization algorithm is used for solving, the algorithm introduces an adaptive weight factor to adjust the search strategy, and the priority of meeting the physical safety limit of the equipment is judged through hierarchical constraint to generate a scheduling scheme;
[0158] In this step, in the S300, the process of constructing the multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization as auxiliary" includes the following steps:
[0159] S310.1, scheduling parameter definition: define the load prediction result output by S200 as the baseline load demand , the real-time operation parameters of the power grid include the output of each power source node , the transmission line loss coefficient , and the device energy consumption characteristic parameter ;
[0160] S310.2, supply-demand balance target construction: taking the matching degree of regional total power supply and baseline load demand as the core target, the matching degree is represented by the supply-demand deviation coefficient , and based on the deviation degree calculation of and , set the deviation threshold , when , it is determined that the basic supply-demand balance is met;
[0161] S310.3, energy consumption optimization target construction: taking the minimum of the comprehensive energy consumption of the power grid as the auxiliary target, based on , , and the scheduling duration , a mapping relationship is established, and the correlation between output and energy consumption is quantified through the energy consumption characteristic parameter;
[0162] S310.4, multi-objective integration: introduce priority weight , and The supply-demand balance target and the energy consumption optimization target are weighted and fused to form a multi-target scheduling function .
[0163] As a further description of this step, the multi-target scheduling function of this embodiment specifically includes: combining the core requirement of "supply-demand balance priority + energy consumption optimization as auxiliary", taking the benchmark load demand , the output of each power supply node , the transmission line loss coefficient , and the equipment energy consumption characteristic parameters as core parameters, constructing a multi-target scheduling function , and the formula is:
[0164] ;
[0165] Among them, is the comprehensive value of the multi-target scheduling function (the smaller the value, the better the scheduling scheme); , is the priority weight defined in the file, and (like =0.7, =0.3), which embodies "supply-demand balance priority"); is the load prediction result output by S200 (the benchmark load demand defined in S310.1), is the output of each power supply node in S310.1;
[0166] In the above formula:
[0167] is the supply-demand balance target function, and the calculation formula is ; wherein, is the supply-demand deviation coefficient defined in S310.2, and the calculation formula is: ; wherein, is the scheduling duration in S310.3 (such as 1 day divided by 15-minute intervals, =96); is the scheduling time step ( ); is the actual output of the th power supply node at time (a subdivided parameter belonging to ); is the benchmark load demand at time (the time sequence value of ); reflects the deviation degree of and at t time, which needs to meet in S310.2 For example, the deviation threshold, =0.05, meaning a deviation within 5% is considered to meet the basic supply and demand balance);
[0168] Through time-series averaging Quantify the supply-demand matching degree throughout the entire scheduling cycle. The smaller the value, the better. and The smaller the overall deviation, the better it aligns with the document's requirement of "supply and demand balance as the core objective";
[0169] The objective function for energy consumption optimization is: ;in, For the first Energy consumption characteristic parameters of power supply equipment (belonging to) Detailed parameters, such as those for coal-fired power units (Value higher than that of photovoltaic units)
[0170] pass Quantify the energy consumption of the equipment itself. Quantify the transmission loss of the lines, and sum the two to obtain the comprehensive energy consumption of the power grid over the entire cycle. , The smaller the value, the lower the energy consumption, which aligns with the positioning of "energy consumption optimization as a secondary measure".
[0171] Furthermore, and pass Weighted fusion ,because In algorithm iterations, it will preferentially reduce (Ensuring supply and demand balance), further optimization (Reduce energy consumption) to ultimately form a scheduling function that takes into account both core and auxiliary objectives. .
[0172] In this step, the process of using the improved whale optimization algorithm to solve the problem and generating a scheduling scheme through hierarchical constraints in S300 includes the following steps:
[0173] S320.1 Algorithm Parameter Initialization: Set the population size for the improved whale optimization algorithm. Maximum number of iterations Initialize population individuals (Each individual corresponds to a set of scheduling parameter combinations);
[0174] As a further explanation of this step, For population size (e.g.) Each individual in the population A corresponding set of scheduling parameters, i.e. , ); is the maximum iteration number (such as =100, the control algorithm solving time length);
[0175] Further, according to the actual number of power sources of the power grid (such as =5, including 2 thermal power units, 2 wind power plants and 1 photovoltaic power station), a reasonable range is set for in each (such as the thermal power unit , is the rated power of the power source), to ensure that meets the actual operation capacity of the equipment.
[0176] It can be understood that when the improved whale optimization algorithm is used to optimize the multi-objective scheduling function, the population size (in the field of intelligent optimization algorithm, the population size is a typical balance interval of "ensuring population diversity to avoid local optimum" and "controlling calculation cost", and combined with the calculation complexity of the power grid scheduling problem, is a conventional choice in this interval considering efficiency and effect); and the maximum iteration number (according to the general experience of intelligent optimization algorithm, when the iteration number reaches the order of hundreds, if the fitness function value has no significant change, the benefit of continuing iteration is very low, therefore is a reasonable iteration upper limit for such multi-constrained optimization problems).
[0177] S320.2, adaptive weight factor adjustment: introducing a weight factor that changes dynamically with the iteration process , taking a large value in the early stage of iteration to strengthen the global search ability, and taking a small value in the later stage of iteration to strengthen the local optimization accuracy, and adjusting the individual search step through ; ;
[0178] As a further description of this step, the calculation formula of the weight factor is as follows:
[0179] ;
[0180] wherein is the early iteration weight (such as =1.5, to strengthen the global search); is the late iteration weight (such as =0.5, to strengthen the local optimization); The current iteration number ( );
[0181] when (Early Iteration) near The algorithm covers more candidate scheduling schemes by increasing the individual search step size;
[0182] when (Late iteration) near The algorithm reduces the search step size and refines the optimization around the optimal solution. Balancing the capabilities of "global exploration" and "local development".
[0183] Understandably, in the weighting factor In the calculation, combining the optimization characteristics of the power grid dispatching problem (multiple constraints, nonlinearity) and the general practices in the field of intelligent optimization algorithms, we take... This range allows the algorithm to maintain a large weight in the early stages (when the number of iterations t is small) to achieve global solution space exploration, and in the later stages ( Approaching the maximum number of iterations Gradually reducing the weight and focusing on local refinement for optimization is a commonly used value range in the industry that takes into account both "global exploration and local development".
[0184] S320.3, Layered Constraint Judgment: The first layer constraint is the physical safety limit of the equipment. This includes the upper limit of transformer load rate. Rated current of the line The second layer of constraints consists of system stability indicators. Including voltage deviation range Frequency deviation range First verify during the solution process. If the condition is not met, discard the individual directly; if it is met, then re-verify. ;
[0185] As a further explanation of this step, the hierarchical constraint judgment in this step specifically includes:
[0186] First level of constraints (Equipment physical safety limits):
[0187] Transformer load factor constraint: ,in for The actual load rate of the transformer at any time. , This is the transformer terminal voltage. This is the load current; Rated apparent power is the apparent power that electrical equipment (such as transformers, generators, etc.) can provide during long-term safe operation under rated voltage and rated current. The unit is usually kilovolt-amperes (kVA), and it is a core rated parameter that reflects the capacity of the equipment.
[0188] Line current constraints: ,in for The actual current of the transmission line at any given time.
[0189] Second-level constraint C2 (system stability index):
[0190] Voltage deviation constraint: ,in The rated voltage of the busbar (e.g., 10kV);
[0191] Frequency deviation constraint: ,in This is the system's rated frequency.
[0192] Verification logic: First verify the physical safety limits of the device. ,like or Directly discard individuals in the current population ; Verify after satisfaction If the voltage or frequency exceeds , The scope is also discarded. Only retain those conditions where both constraints are satisfied. Moving on to the next iteration.
[0193] S320.4 Optimal Scheduling Scheme Generation: Through multi-objective scheduling functions Calculate the fitness value of each individual in the population, iteratively update and retain the individual with the best fitness, and finally output the result that simultaneously satisfies... , The combination of constrained scheduling parameters forms a scheduling scheme.
[0194] Specifically, the generation of the optimal scheduling scheme includes the following steps:
[0195] Fitness calculation: using a multi-objective scheduling function As an individual in a population fitness value, The smaller, The better the corresponding scheduling scheme;
[0196] Iterative update: Each iteration retains the best-fitting element. and based on Adjust other The search step size is used to generate a new population;
[0197] Solution output: When the number of iterations... When the iteration stops, output the final optimal value that is retained. ,Should corresponding Combination is to satisfy , The constrained scheduling scheme forms a power output allocation plan that can be directly executed.
[0198] like Figure 4 As shown, S400 real-time closed-loop adjustment: high-frequency load monitoring is used to calculate the deviation rate; when the deviation exceeds the reasonable range, the deviation is brought back to the reasonable range through a combination of "incremental data training + scheduling parameter fine-tuning";
[0199] In this step, the real-time closed-loop adjustment process in S400 includes the following steps:
[0200] S410.1 High-frequency load monitoring and deviation rate calculation: According to the multi-dimensional sensor data acquisition dimensions in S110.1, set the load monitoring frequency to no less than 12 times per hour, collect the actual load data of the power grid in real time, and combine it with the load prediction results of the corresponding time period output in S230.3 to calculate the deviation rate by the relative deviation between the actual load and the predicted load.
[0201] As a further explanation of this step, S410.1 high-frequency load monitoring and deviation rate calculation in this embodiment specifically includes:
[0202] Based on the multi-dimensional sensing device data acquisition dimensions (user side, device side, and environment side) defined in S110.1, the load monitoring frequency is specified as once every 5 minutes (12 times per hour, meeting the requirement of "no less than 12 times per hour"). The monitoring equipment reuses the sensing devices deployed in S110.1. On the user side, the total active power and reactive power of the zone are collected through smart meters. On the device side, the actual load current of the transmission line is collected through line current / voltage sensors. On the environment side, environmental parameters at the monitoring time are recorded synchronously through temperature and humidity sensors (for subsequent incremental training and correlation analysis). All monitoring data are transmitted to the power grid dispatching platform in real time through the edge gateway.
[0203] The deviation rate is calculated based on the "actual load at the monitoring time and the predicted load for the corresponding time period", and the formula is as follows: ;in, for Load deviation rate at any given time (reflecting the relative deviation between actual and predicted loads). for The actual load value of the power grid at any given time (taken from the sum of the zone loads of the above monitoring data, such as the actual load of a certain area at 08:05 is 120MW). For Load prediction result at the moment (taken from the corresponding timestamp prediction value output by S230.3, for example, the load at 08:05 is predicted to be 115MW).
[0204] Quantify the difference between prediction and actual by relative deviation rate, avoid deviation judgment error caused by absolute value of load, for example, 10MW absolute deviation is 10% in 100MW load scenario, and 5% in 200MW load scenario, which can more accurately reflect the severity of deviation under different load levels.
[0205] S410.2, deviation range judgment: preset deviation reasonable range threshold, which is set according to the control logic of system operation stability index in S320.3, if the deviation rate is within the threshold range, maintain the current dispatching scheme generated by S300; if the deviation rate exceeds the threshold range, trigger the closed-loop adjustment mechanism;
[0206] As a further description of this embodiment, S410.2 deviation range judgment of this embodiment specifically includes:
[0207] When the preset deviation reasonable range threshold is set, strictly follow the control logic of "system operation stability index priority" in S320.3, and set differentiated threshold according to the load sensitivity of different operation periods of power grid:
[0208] Peak load period (consistent with S200, i.e. 08:00-11:00, 18:00-22:00): because the load fluctuation has a greater impact on system stability, the reasonable threshold of deviation rate is set to ≤5%;
[0209] Off-peak period (other periods): the impact of load fluctuation is relatively small, and the reasonable threshold of deviation rate is set to ≤8%;
[0210] The threshold is set according to the stability requirement in "GB / T 15945-2008 Power Quality Power System Frequency Deviation" that "system frequency deviation needs to be controlled within ±0.2Hz", to ensure that the deviation threshold matches the system operation safety boundary.
[0211] The judgment logic is: the dispatching platform compares with the threshold value of the corresponding period in real time, if it is within the threshold range (such as peak period =3%), the current dispatching scheme generated by S300 is maintained, and no adjustment is needed; if it exceeds the threshold (such as peak period =7%), the closed-loop adjustment mechanism is triggered-the dispatching platform automatically sends "adjustment start signal" to the incremental training module and the parameter fine-tuning module, and records the triggering time, current load data and deviation rate at the same time, which provides basis for subsequent tracing.
[0212] S410.3, incremental data training: collect the environment side, device side, and user side data collected in S110.1 to form an incremental data set, and use the incremental data set to supplement the training of the attention-enhanced load prediction model in S200, update the model parameters to improve the short-term prediction accuracy;
[0213] As a further description of this embodiment, the incremental data training S410.3 of this embodiment specifically includes:
[0214] The construction of the incremental data set takes the data within 1 hour after the deviation rate exceeds the threshold value as the time range, and the data sources are completely consistent with the collection dimensions in S110.1. The specific composition includes:
[0215] Environment side data: air temperature, relative humidity, and wind speed collected every 5 minutes within 1 hour (a total of 12 groups of data);
[0216] Device side data: transformer oil temperature, line current, and bus voltage collected every 5 minutes within 1 hour (a total of 12 groups of data);
[0217] User side data: partition active power, reactive power, and cumulative power consumption collected every 5 minutes within 1 hour (a total of 12 groups of data);
[0218] All data are structured and sorted according to the "timestamp-data type-value" structure to ensure that the input format matches the attention-enhanced load prediction model in S200.
[0219] Further, the incremental training process adopts a "small batch iteration update" mode, and the specific parameters are set as follows: the learning rate is set to 0.0005 (lower than 0.001 of the initial training in S230 to avoid excessive updating of the original model parameters), the training batch size is 32, and the training rounds are 5-10 rounds (because the incremental data volume is small, the rounds do not need to be too many to prevent model overfitting). The training target is to "minimize the 15-minute short-term prediction deviation", and the model can learn the correlation between temperature sudden drop and load growth through incremental data to update the contribution of the environment temperature feature in the attention weight, and thus improve the prediction accuracy of the next 15-30 minutes.
[0220] Understandably, to ensure the stability of incremental model training, and referencing the typical time resolution of power grid load data (e.g., 5 minutes / sample) and the minimum data volume requirement for incremental machine learning, the minimum size of the incremental dataset must contain at least 12 sets of continuous valid data (corresponding to approximately 1 hour of sampling time, which can cover basic load fluctuation characteristics). When data is abnormal (e.g., missing data due to sensor failure, values exceeding the equipment's rated range, etc.), a historical data compensation strategy based on the same operating conditions is adopted: based on the three-dimensional label of "time period type (weekday / holiday) - environmental parameters (temperature, humidity) - equipment operating status (load rate, oil temperature)," the historical database is matched with the time period data most similar to the current operating conditions for filling. This strategy is a mature method for handling short-term data anomalies in the power grid field and has been validated in engineering scenarios such as distribution network measurement data repair and load curve completion.
[0221] S410.4 Fine-tuning of scheduling parameters: Based on the prediction results after incremental training, the output allocation ratio of the power nodes involved in S310.1 is slightly modified. The modification range increases accordingly as the deviation rate increases, and the modified parameters must meet the physical safety limit requirements of the equipment in S320.3.
[0222] As a further explanation of this step, the fine-tuning of scheduling parameters in S410.4 of this embodiment specifically includes: fine-tuning the scheduling parameters as defined in S310.1, "output of each power node". "With " as the core adjustment target, the correction magnitude and deviation rate For hooks, the principle of "the greater the deviation, the more appropriate the correction range should be, but not exceeding the equipment safety boundary" should be followed. The specific correction rules are as follows:
[0223] when Time (such as peak hours) =6%), with the correction range set to ≤5%;
[0224] when When this occurs, the correction range is set to ≤10%;
[0225] when When this occurs, the correction range is set to ≤15%;
[0226] Furthermore, the correction process requires real-time verification of the equipment's physical safety limits in S320.3: for example, the original output of thermal power units in a certain area. Deviation rate =8% (peak hours, exceeding the 5% threshold), the adjustment range according to the rules is 4%, proposed to be adjusted to... At this time, it is necessary to simultaneously calculate the transformer load rate corresponding to the unit. The calculation formula is: , This represents the actual apparent power of the adjusted transformer. For rated apparent power, if (S320.3 set ), the line current (S320.3 set ), the correction is allowed; if the correction exceeds the safety limit (such as ), the correction amplitude is reduced to 3%, the calculation is re-calculated and checked again until all device safety constraints are met.
[0227] In addition, fine-tuning preferentially selects power supply nodes with high regulation flexibility (such as gas turbine units, energy storage power stations), and the correction amplitude of slow-regulating power sources (such as coal-fired units) is controlled at ≤3%, avoiding frequent adjustments that affect device life, consistent with the principle of "device physical safety priority" in S300.
[0228] S410.5, deviation back-off verification: re-acquire actual load and calculate deviation rate according to the monitoring frequency set in S410.1, repeat the steps of S410.3 to S410.4 until the deviation rate falls within the reasonable range threshold, and synchronize the final scheduling parameters and adjustment process data to the time series database of S500.
[0229] As a further description of this step, S410.5 deviation back-off verification of the present embodiment specifically includes:
[0230] According to the monitoring frequency of "once every 5 minutes" set in S410.1, re-acquire the adjusted actual load data, and calculate the new deviation rate , the verification logic is as follows:
[0231] If falls within the reasonable threshold for the corresponding period (such as =4.5% during peak hours), stop closed-loop adjustment, and record the final scheduling parameters (each power supply node's corrected output, adjustment duration, deviation rate change curve);
[0232] If still exceeds the threshold (such as =5.8% during peak hours), repeat the steps of S410.3-S410.4: update the incremental data set to "data within 1 hour after the last adjustment", re-train the incremental training (at this time the model can learn the load response law after the last adjustment), and fine-tune the scheduling parameters again based on the new prediction results until requirements are met;
[0233] Finally, the "adjustment trigger time-increment dataset range-model update parameter-output after each round of correction-deviation rate change" and other full-process data are synchronized to the time series database of S500 (consistent with the data archiving format of S100) in a structured format of "timestamp-data category-specific value". The scheduling parameter data need to be associated with the corresponding power supply node number and device safety check results to ensure the traceability of the adjustment rationality and safety during subsequent operation and maintenance review.
[0234] As shown in Figure 5 S500, data archiving and traceability: the key data of S100-S400 is classified and archived by using a time series database; the data retention period meets the needs of power grid equipment operation and maintenance traceability, supports retrieval by time, region, and device number, and ensures data traceability.
[0235] In this step, in the S500, the data archiving and traceability process includes the following steps:
[0236] S510.1, time series database configuration and archiving classification design: a time series database supporting high-frequency data writing and time series indexing is used to divide the archiving categories according to the stages of S100-S400. S100 archives the standardized dataset and collection exception records after preprocessing, S200 archives the load prediction model parameters, prediction results, and precision verification data, S300 archives the multi-objective scheduling function parameters, optimization algorithm iteration records, and final scheduling scheme, and S400 archives the load monitoring data, deviation rate calculation results, and closed-loop adjustment records.
[0237] As a further description of this step, the S510.1 time series database configuration and archiving classification design of the present embodiment specifically includes: selecting an open-source time series database (such as InfluxDB or TimescaleDB) that supports high-frequency data writing and time series indexing optimization, using a "regional data center + edge node backup" architecture, configuring 2 master-slave servers (single storage capacity ≥100TB, supporting RAID5 redundant backup) in the regional data center, and using the edge node to temporarily store real-time data for 30 days in conjunction with the S100 edge storage node; the database core parameters are set to 8 writing threads, 1 minute time series indexing granularity, and LZ4 data compression algorithm;
[0238] Further, the archiving categories are divided according to the stages of S100-S400, specifically including:
[0239] S100 archives the standardized dataset containing "timestamp, data source, feature name, and standardized value", and the collection exception records containing "abnormal device number, abnormal time, type, processing measures, and recovery time";
[0240] S200 archives model parameters containing "training batch, iteration round, parameter name, parameter value", prediction results containing "prediction timestamp, prediction value, actual value, MAPE, RMSE", and precision verification data containing data set division range, iteration loss value, and error determination result;
[0241] S300 archives scheduling function parameters containing "effective time, parameter name, value, and setting basis", algorithm iteration records containing "iteration round, optimal population individual, value, constraint result", and final scheduling scheme containing "time index, power supply number, planned output, and verification details";
[0242] S400 archives monitoring data containing "monitoring timestamp, partition load, and sensor data", deviation rate results containing "calculation timestamp, 、 , deviation rate, threshold value", and closed-loop adjustment records containing "trigger time, incremental data range, model update parameters, correction comparison, and rollback results".
[0243] S510.2, data retention period setting: determine the retention period according to the power grid equipment operation and maintenance traceability requirements, wherein the retention period of the associated data of the transformer, transmission line, and switch equipment is not shorter than the design life of the corresponding equipment, the retention period of the temporary scheduling adjustment record is not shorter than 3 power grid operation and maintenance review periods, ensuring that the equipment full life cycle traceability and operation and maintenance review requirements are covered;
[0244] As a further description of this step, the S510.2 data retention period setting of the present embodiment is as follows:
[0245] Based on the power grid equipment operation and maintenance traceability requirements, the retention period of the transformer associated data (oil temperature, load rate, output distribution) is set to 30 years (not shorter than the longest 20-30 year design life), the retention period of the transmission line associated data (current, voltage, loss) is set to 40 years (not shorter than the longest 30-40 year design life), and the retention period of the switch equipment associated data (split state, operation record, fault alarm) is set to 20 years (not shorter than the longest 15-20 year design life);
[0246] The retention period of the temporary scheduling adjustment record (single closed-loop adjustment, short-term deviation processing) is set to 3 months (covering 3 one-month operation and maintenance review periods), and it can be archived to offline storage medium after 3 months; the retention period of the S100 standardized basic data set and the S200 model training basic data is set to 5 years, meeting the medium and long-term load regularity analysis requirements.
[0247] S510.3 Multi-dimensional search function configuration: Build a multi-dimensional search index in the time series database. The time dimension supports time period search accurate to the minute level, the region dimension is associated with the power grid partition code, and the device number dimension is mapped to the unique identification number of the sensing device, transformer and line in S110.1. The search response time does not exceed 10 seconds.
[0248] As a further explanation of this step, the S510.3 multi-dimensional search function configuration of this embodiment specifically includes:
[0249] A joint index of time, region, and device number is constructed in the time-series database. The time dimension is indexed at the "year-month-day-hour-minute" level, supporting minute-level time period retrieval in the format "YYYY-MM-DDHH:MM:00 to YYYY-MM-DDHH:MM:00". The region dimension is associated with the power grid partition code (e.g., E01 for the East, W02 for the West), and each archived data is labeled with its region code. The device number dimension is mapped to the unique identifier of S110.1 device according to the "equipment type-voltage level / scenario-serial number" rule (e.g., S-user side-001, T-10kV-001). The retrieval function provides a visual interface, supports single-dimensional or multi-dimensional combined retrieval, and the results are displayed in tables or curves and can be exported to CSV / Excel. The speed is optimized by "quarterly data sharding storage + data in memory preloading within 1 year", ensuring that the retrieval response time does not exceed 10 seconds.
[0250] S510.4, Archived Data Traceability Guarantee: The archived data is periodically checked for integrity, and the consistency between the archived data and the original data generated in the S100-S400 stages is compared. A verification log is generated and archived. If data is found to be missing or inconsistent, an automatic re-archiving mechanism is triggered to retrieve the data from the temporary storage nodes of S100-S400 and re-archive it to ensure the accuracy and integrity of the traceable data.
[0251] As a further explanation of this step, the traceability guarantee of archived data in S510.4 of this embodiment specifically includes:
[0252] The full-volume data verification is performed every day at 2:00 (low load period) for the previous 24 hours, and the monthly full-volume verification is performed at the end of each month. The MD5 hash value is used to compare the original data of the S100-S400 temporary storage nodes with the archived data of the time series database. If the two are consistent, it is determined that the data is complete. If the two are inconsistent or missing, it is marked as abnormal. The automatic archiving mechanism takes the S100-S400 temporary storage nodes (retaining 30 days of data) or the regional backup center as the data source. After detecting the abnormality, the archiving task is generated and an alarm is sent. The archiving program extracts the original data and writes it into the database according to the archiving format. After archiving, the hash value is verified again. If the verification fails, manual intervention is triggered. Each verification generates a log containing "verification time, data range, result, data volume, processing measures, and processing result". The log is archived in the time series database special table according to "year-month", and the retention period is consistent with the corresponding archived data, forming a full-link traceable closed loop.
[0253] Those of ordinary skill in the art can understand that the processes for implementing all or part of the steps of the above embodiments can be completed by hardware, or by programs instructing relevant hardware.
[0254] The basic principles, main features and advantages of the present application are shown and described above. Those skilled in the art should understand that the present application is not limited by the above embodiments, and the above embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A big data based power grid load forecasting and dynamic dispatch optimization method, characterized in that, The method comprises the following steps: S100, power grid multi-source data preprocessing: collecting user side, equipment side and environment side original data by using multi-dimensional sensing equipment, and processing through a combined process of "missing value classification filling + 3 sigma criterion abnormal value elimination + Z-Score standardization"; S200, attention-enhanced load forecasting model construction: based on the standardized data set of S100, extracting time sequence load characteristics through an LSTM network, and screening key influence factor characteristics through a random forest; Introducing an attention mechanism to weight and fuse the two types of characteristics, highlighting the feature weight of the power consumption peak period; Using an Adam optimizer to train the model, verifying the model accuracy by mean absolute percentage error and root mean square error, and outputting the load prediction result; In the S200, the process of introducing an attention mechanism to weight and fuse the two types of characteristics comprises the following steps: S220.1, Feature correlation analysis: calculate the correlation coefficient of each feature in the feature set to be fused with the actual load change amount ; S220.2, attention weight allocation: based on correlation matrix , assigning weights to features for peak hours , assigning weights to features for off-peak hours , assigning weights to features for peak hours , assigning weights to features for off-peak hours , wherein , forming a dynamic weight matrix ; S220.3, Feature Weighted Fusion: Through matrix operation, the dynamic weight matrix with the feature set to be fused fusion, generating a comprehensive feature matrix ; S300, hierarchical constraint improved whale optimization scheduling scheme generation: based on the load prediction result of S200 and the real-time operation parameters of the power grid, a multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization auxiliary" is constructed; and an improved whale optimization algorithm is used for solving, the algorithm introduces an adaptive weight factor to adjust the search strategy, and the priority of meeting the physical safety limit of the equipment is judged through hierarchical constraints to generate a scheduling scheme; In the S300, the process of using an improved whale optimization algorithm to solve and generating a scheduling scheme through hierarchical constraints comprises the following steps: S320.1, algorithm parameter initialization: set the population size of the improved whale optimization algorithm , the maximum number of iterations , initialize the population individuals ; S320.2 Adaptive Weight Factor Adjustment: Introducing a weight factor that dynamically changes with the iteration process. Early iteration hour To enhance global search capabilities, larger values are selected in later iterations. hour To improve the accuracy of local optimization, smaller values are selected. Adjust the individual search step size; S320.3, hierarchical constraint judgment: the first layer constraint is the equipment physical safety limit , including transformer load rate upper limit , line rated current ; the second layer constraint is the system operation stability index , including voltage deviation range , frequency deviation range ; in the solving process, first check , if not satisfied, directly discard the individual, and then check ; S320.4, optimal scheduling scheme generation: through multi-objective scheduling function The fitness value of the population individual is calculated, the optimal individual is updated iteratively, and the final output is the scheduling parameter combination that meets the 、 constraints, forming a scheduling scheme; S400, real-time closed-loop adjustment: using high-frequency load monitoring to calculate the deviation rate; when the deviation exceeds a reasonable range, using a combined method of "incremental data training + scheduling parameter fine-tuning" to make the deviation fall within a reasonable range; S500, data archiving and tracing: using a time series database to classify and archive the key data of S100-S400; the data retention period meets the power grid equipment operation and maintenance tracing requirements, supports time, region and equipment number retrieval, and ensures data traceability.
2. The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 1, characterized in that, In the S100, the process of collecting user side, equipment side and environment side original data by using multi-dimensional sensing equipment comprises the following steps: S110.1, selection and matching of collection equipment: for user side original data collection, multi-dimensional sensing equipment suitable for different power consumption scenarios is selected to collect real-time active power, reactive power and cumulative power consumption; for equipment side original data collection, sensing equipment matching the types of transformers, transmission lines and switch devices is selected to collect transformer oil temperature, line current, line voltage and switch opening and closing state; for environment side original data collection, sensing equipment suitable for outdoor power grid scenarios is selected to collect air temperature, relative humidity and wind speed; S110.2, collection process data integrity monitoring: during the collection process, the data stream transmission state of the sensing equipment is monitored in real time; when it is detected that a single device has not uploaded data for M times continuously or the missing data field ratio exceeds a preset threshold, it is determined that there is a collection anomaly, and a device self-checking program is triggered; S110.3, emergency treatment of collection anomaly: if the self-checking program still cannot restore normal collection, a backup sensing device previously deployed in the same region is enabled to continue collection in place of the abnormal device, and the number, abnormal time and type of the abnormal device are recorded; S110.4, original data identification and temporary storage: structurally identify the normally collected original data according to "acquisition object type-equipment unique number-acquisition time", and temporarily store them in the nearest edge storage node. Data check code is used to verify data integrity during the temporary storage. 3.The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 2, characterized in that, In the S100, the combined process of "missing value classification filling + 3σ criterion abnormal value elimination + Z-Score standardization" includes the following steps: S120.1, missing value classification filling: classify the original data temporarily stored in S110.4 according to the user side, the equipment side and the environment side. For the missing values in each type of data, linear interpolation method is used to fill the short-term continuous missing values with the adjacent effective data before and after the data, and the mean value of the same type of data in the same region in the same period is used to fill the long-term continuous missing values; S120.2, 3σ criterion abnormal value elimination: calculate the mean value and standard deviation of each type of data according to the user side, the equipment side and the environment side respectively for the data filled in S120.
1. The values deviating from the mean value by more than 3 times the standard deviation are determined as abnormal values and eliminated; S120.3, Z-Score standardization: perform Z-Score standardization processing on the data after the elimination of abnormal values in S120.2 according to the data type, eliminate the magnitude difference of different dimension data, and make the processed data dimensionless.
4. The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 3, characterized in that, In the S200, the process of extracting time series load features through the LSTM network and screening key influence factor features through the random forest includes the following steps: S210.1, LSTM network feature extraction: the standardized data set of S100 is divided into continuous input units according to time sequence , is a time index, input LSTM network; the network controls the inflow of new information through the input gate forget gate clear redundant information, output gate screen output information, generate hidden layer state, and finally output time sequence load feature vector ; S210.2, random forest feature screening: the influence factor feature set composed of environmental features and equipment operation features With load data Input the random forest model, calculate the node split gain of each feature Determine the contribution degree, screen the features with contribution degree meeting the preset condition to form the key influence factor feature set ; S210.3, Feature dimension alignment: align the time series load feature vector with the time scale of the key impact factor feature set based on the timestamp with the time scale of the key impact factor feature set , match the dimensions of both by dimension expansion or compression to form the feature set to be fused .
5. The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 4, characterized in that, In the S200, the process of training the model by using the Adam optimizer and verifying the model accuracy by the mean absolute percentage error and the root mean square error includes the following steps: S230.1, model training configuration: the comprehensive feature matrix is divided into a training set and a validation set , the first moment coefficient of the Adam optimizer is set , the second moment coefficient , the model parameters are initialized ; S230.2, iteration training: with training set input model, update model parameters by optimizer iteration , after each round of iteration, use validation set calculate loss value stop training when there is no loss value for a preset number of consecutive rounds, save the optimal parameters ; S230.3, precision verification: calculate the mean absolute percentage error and root mean square error of the predicted load and the actual load, output the load prediction result that meets the error requirement Input the optimal model, calculate the predicted load With the actual load The average absolute percentage error and root mean square error, output the load prediction result that meets the error requirement .
6. The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 5, characterized in that, In the S300, the process of constructing the multi-objective scheduling function of "supply-demand balance priority + energy consumption optimization as auxiliary" includes the following steps: S310.1, scheduling parameter definition: define the load prediction result outputted by S200 as the reference load demand The real-time operation parameters of the power grid include the output of each power supply node The transmission line loss coefficient The equipment energy consumption characteristic parameters ; S310.2, supply and demand balance target construction: taking the matching degree of the total power supply and the benchmark load demand in the region as the core target, the matching degree is calculated through the supply and demand deviation coefficient characterized, and based on deviation from the degree of deviation, set the deviation threshold , when the basic supply and demand balance is satisfied; S310.3, energy consumption optimization target construction: based on the comprehensive energy consumption of the power grid The minimum is an auxiliary target, Based on , , And the scheduling duration Establish a mapping relationship to quantify the correlation between output and energy consumption through energy consumption characteristics parameters; S310.4, Multi-objective Integration: Introducing Priority Weight , and , the supply-demand balance objective and the energy consumption optimization objective are weighted and integrated to form a multi-objective scheduling function .
7. The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 6, characterized in that, In the S400, the process of real-time closed-loop adjustment includes the following steps: S410.1, high-frequency load monitoring and deviation rate calculation: set the load monitoring frequency to not less than 12 times per hour according to the multi-dimensional sensing equipment acquisition dimension of S110.1, real-time collect the actual load data of the power grid, combine the corresponding period load prediction results output by S230.3, and calculate the deviation rate through the relative deviation degree of the actual load and the predicted load; S410.2, deviation range judgment: preset the deviation reasonable range threshold, which is set according to the control logic of the system operation stability index in S320.
3. If the deviation rate is within the threshold range, maintain the current scheduling scheme generated by S300. If the deviation rate exceeds the threshold range, trigger the closed-loop adjustment mechanism; S410.3, incremental data training: collect the environment side, equipment side and user side data collected by S110.1 to form an incremental data set, use the incremental data set to supplement the training of the attention-enhanced load prediction model in S200, update the model parameters to improve the short-term prediction accuracy; S410.4, Fine-tuning of scheduling parameters: Based on the prediction results after incremental training, the output distribution ratio of the power supply nodes involved in S310.1 is slightly corrected. The correction amplitude increases with the deviation rate, and the corrected parameters need to meet the physical safety limit requirements of the equipment in S320.3; S410.5, Deviation backfall verification: According to the monitoring frequency set in S410.1, the actual load is re-collected and the deviation rate is calculated. Repeat the steps of S410.3 to S410.4 until the deviation rate falls within the reasonable range threshold. The final scheduling parameters and adjustment process data are synchronized to the time series database of S500. 8.The big data based power grid load forecasting and dynamic dispatch optimization method according to claim 7, characterized in that, In the S500, the process of data archiving and tracing includes the following steps: S510.1, Time series database configuration and archiving classification design: Use a time series database that supports high-frequency data writing and time series indexing. According to the stage division of S100-S400, the archiving categories are divided into S100, which archives the standardized data set after preprocessing and the collection of abnormal records, S200, which archives the load prediction model parameters, prediction results and precision verification data, S300, which archives the multi-objective scheduling function parameters, optimization algorithm iteration records and final scheduling scheme, and S400, which archives the load monitoring data, deviation rate calculation results and closed-loop adjustment records; S510.2, Data retention period setting: Determine the retention period according to the grid equipment operation and maintenance traceability requirements. The associated data retention period of transformers, transmission lines and switch devices should not be shorter than the design life of the corresponding equipment. The temporary scheduling adjustment record retention period should not be shorter than 3 grid operation and maintenance review periods to ensure that the equipment life cycle traceability and operation review requirements are met; S510.3, Multi-dimensional search function configuration: Build multi-dimensional search indexes in the time series database. The time dimension supports accurate minute-level period search. The regional dimension is associated with the grid partition code. The equipment number dimension maps the unique identification number of the sensor equipment, transformer and line in S110.1, and the search response time is not more than 10 seconds; S510.4, Archiving data traceability guarantee: Regularly check the integrity of the archived data, compare the consistency of the archived data with the original generated data in S100-S400, generate a verification log and archive it. If data is missing or inconsistent, trigger the automatic supplement mechanism to extract data from the temporary storage node of S100-S400 for re-archiving, ensuring the accuracy and integrity of the traceability data.
Citation Information
Patent Citations
A power load forecasting method and system based on big data drive
CN118889402B
Power grid load prediction improving method and device based on AI intelligent algorithm
CN119921324A
Electric vehicle charging load prediction method based on user behavior analysis
CN120280901A
Intelligent risk early warning method, device and equipment for power distribution network and medium
CN120430612A