Power grid load prediction and dynamic scheduling optimization method based on big data
By employing multi-dimensional data acquisition and processing, attention-enhanced prediction models, and multi-objective scheduling optimization, the problems of multi-source data processing and insufficient prediction models in power grid load forecasting have been solved. This has enabled the accuracy of power grid load forecasting and the rationality of scheduling schemes, ensuring the stability of power grid operation and the traceability of data.
Patent Information
- Application Number
- CN202511357593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-23
AI Technical Summary
Existing technologies for power grid load forecasting suffer from several problems, including a lack of systematic and comprehensive multi-source data processing capabilities, insufficient feature fusion depth in forecasting models and a lack of time-specific targeting, as well as the absence of a complete optimization system encompassing forecasting, scheduling, adjustment, and tracing.
Raw data from the user side, equipment side, and environment side are collected using multi-dimensional sensing devices. Through processing of missing value classification and imputation, outlier removal by the 3σ criterion, and Z-Score standardization, an attention-enhanced load forecasting model is constructed. The key influencing factors are screened by combining LSTM network and random forest, and an attention mechanism is introduced for weighted fusion to construct a multi-objective scheduling function of "supply and demand balance first, energy consumption optimization second". An improved whale optimization algorithm is used to solve the function, and real-time closed-loop adjustment is achieved through high-frequency load monitoring. The data is archived to a time-series database to support traceability.
It achieves consistency and availability of multi-source data quality, improves the accuracy of load forecasting and the rationality of scheduling schemes, ensures the stability and security of power grid operation, supports full-process data traceability, and meets the needs of power grid operation.
Smart Images

Figure CN120855327A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid load forecasting technology, and more specifically, to a power grid load forecasting and dynamic scheduling optimization method based on big data. Background Technology
[0002] With the increasing proportion of renewable energy grid connection and the diversification of user-side electricity load (such as electric vehicle charging and flexible industrial and commercial loads), the volatility and uncertainty of power grid load have increased significantly. Accurate load forecasting and dynamic scheduling optimization have become core requirements for ensuring the safe and stable operation of the power grid and improving energy efficiency. In current power grid operation, the accuracy of load forecasting results directly affects power output allocation and equipment operation and maintenance planning, while the rationality of scheduling schemes is related to the physical safety of equipment (such as transformer load rate and line current control) and the economic efficiency of system operation (such as energy consumption optimization). Therefore, it is urgent to build a complete technical solution covering the entire process of "data processing - load forecasting - scheduling optimization - closed-loop adjustment - data traceability".
[0003] For example, Chinese patent CN202411002209.1 discloses a power load forecasting method and system based on big data. This method acquires regional power station data, classifies regional electricity consumption types, and assesses regional power complexity. Then, based on a preset complexity threshold, it performs time-series load analysis to generate high-complexity regional power peak load data and low-complexity regional power base load data. The core is to improve forecast reliability by combining data classification and complexity assessment with historical real-time data. Another example is Chinese patent CN202510413310.4, which discloses a power grid load forecasting improvement method and device based on AI intelligent algorithms. This method constructs an AI intelligent algorithm library, clusters power grid loads by industry, selects algorithms for load forecasting by industry, and then corrects the industry-specific load forecasting data in real time by constructing a temperature correction model and combining it with forecast meteorological data. The key point is to use industry classification and environmental factors to correct and optimize the AI forecasting effect.
[0004] While the aforementioned technical solutions possess design advantages in data classification and algorithm adaptation for power load forecasting, they also suffer from the following technical shortcomings: First, multi-source data processing lacks systematic and comprehensive coverage: Chinese patent CN202411002209.1 only classifies and assesses the complexity of regional power station data, failing to incorporate key data such as user-side electricity consumption parameters and equipment-side operating status, and lacks a classification and filling strategy for missing data and a standardized mechanism for removing outliers; although Chinese patent CN202510413310.4 introduces environmental factors such as temperature for correction, it still fails to form a comprehensive data processing flow covering the user side, equipment side, and environment side, making it difficult to guarantee the quality consistency and usability of multi-source data; Second, the feature fusion depth of the prediction model is insufficient and lacks time-specificity: Chinese patent CN202411002209.1 does not... The first two patents, CN202510413310.4 and CN202510413310.4, only optimize predictions through industry clustering and temperature correction, failing to effectively integrate time-series characteristics with influencing factors such as equipment operation and user behavior, thus failing to accurately capture the strong fluctuations in load during peak hours. Secondly, they lack a complete "prediction-scheduling-adjustment-tracing" optimization system: both patents focus solely on load prediction, neglecting to extend to the design of dynamic grid dispatching schemes and lacking multi-objective dispatching optimization logic prioritizing equipment physical safety. Furthermore, they lack a real-time closed-loop adjustment mechanism after prediction deviations occur and fail to consider archiving and tracing key data throughout the entire process, thus failing to meet the full-scenario operational needs of the power grid from load prediction to dispatch execution, deviation correction, and data tracing. Therefore, we propose a big data-based power grid load prediction and dynamic dispatching optimization method. Summary of the Invention
[0005] The purpose of this invention is to provide a big data-based method for power grid load forecasting and dynamic scheduling optimization, in order to solve the problems mentioned in the background art, such as the lack of systematic and full-dimensional coverage of multi-source data processing, insufficient feature fusion depth of the prediction model and lack of time-specificity, and the failure to construct a full-chain optimization system of "prediction-scheduling-adjustment-tracing".
[0006] To address the aforementioned technical problems, this invention provides a method for power grid load forecasting and dynamic scheduling optimization based on big data, comprising the following steps: S100, Power Grid Multi-Source Data Preprocessing: Use multi-dimensional sensing devices to collect raw data from the user side, equipment side, and environment side, and process it through a combination of "missing value classification and imputation + 3σ criterion outlier removal + Z-Score standardization"; S200 Attention-Enhanced Load Forecasting Model Construction: Based on the standardized S100 dataset, time-series load features are extracted using an LSTM network, and key influencing factor features are selected using a random forest. An attention mechanism is introduced to weight and fuse the two types of features, highlighting the feature weights during peak electricity consumption periods. The model is trained using the Adam optimizer, and the model accuracy is verified using mean absolute percentage error and root mean square error. Load forecasting results are then output. S300, generation of improved whale optimization scheduling scheme with hierarchical constraints: Based on the load forecast results of S200 and the real-time operating parameters of the power grid, a multi-objective scheduling function with "supply and demand balance as the priority and energy consumption optimization as the auxiliary" is constructed; and an improved whale optimization algorithm is used to solve it. The algorithm introduces an adaptive weight factor to adjust the search strategy, and determines the priority to meet the physical safety limits of the equipment through hierarchical constraints to generate a scheduling scheme. S400, Real-time Closed-Loop Adjustment: High-frequency load monitoring is used to calculate the deviation rate; when the deviation exceeds the reasonable range, the deviation is brought back to the reasonable range through a combination of "incremental data training + scheduling parameter fine-tuning"; S500 Data Archiving and Traceability: Key data from S100 to S400 are archived using a time-series database; the data retention period meets the traceability requirements of power grid equipment operation and maintenance, and supports retrieval by time, region, and equipment number to ensure data traceability.
[0007] As a further improvement to this technical solution, in step S100, the process of collecting raw data from the user side, device side, and environment side using multi-dimensional sensing devices includes the following steps: S110.1 Selection and Matching of Data Acquisition Equipment: For raw data acquisition on the user side, select multi-dimensional sensing equipment adapted to different power consumption scenarios to collect real-time active power, reactive power, and cumulative power consumption; for raw data acquisition on the equipment side, select sensing equipment matched to the type of transformer, transmission line, and switching equipment to collect transformer oil temperature, line current, line voltage, and switch open / closed status; for raw data acquisition on the environmental side, select sensing equipment adapted to outdoor power grid scenarios to collect air temperature, relative humidity, and wind speed. S110.2 Data integrity monitoring during the acquisition process: During the acquisition process, the data stream transmission status of the sensing device is monitored in real time. When it is detected that a single device has failed to upload data for M consecutive times or the proportion of missing fields in the uploaded data exceeds the preset threshold, it is determined to be an acquisition anomaly and the device self-test program is triggered. S110.3 Emergency handling of abnormal data acquisition: If the self-test procedure still cannot restore normal data acquisition, activate the backup sensor device pre-deployed in the same area to take over the abnormal device and continue data acquisition. At the same time, record the abnormal device number, the time of the abnormality and the type of abnormality. S110.4 Raw Data Identification and Temporary Storage: The raw data collected normally is structurally identified according to "collection object type - device unique number - collection time" and temporarily stored in the nearest edge storage node. During the temporary storage process, data verification codes are used to verify data integrity.
[0008] As a further improvement to this technical solution, in S100, the combined process of "missing value classification and imputation + 3σ criterion outlier removal + Z-Score normalization" includes the following steps: S120.1 Missing Value Classification and Imputation: The original data temporarily stored in S110.4 is classified into user side, device side and environment side. For missing values in each type of data, if there are consecutive missing values in a short period of time, the linear interpolation method of the adjacent valid data before and after the missing value is used to imput them. If there are consecutive missing values in a long period of time, the mean of the same type of data in the same area in the same period of time is used to imput them. S120.2, 3σ Criterion Outlier Removal: For the data imputed by S120.1, calculate the mean and standard deviation of each data type according to the user side, equipment side and environment side respectively. Values that deviate from the mean by more than 3 times the standard deviation are identified as outliers and removed. S120.3 Z-Score Standardization: After outliers are removed by S120.2, Z-Score standardization is performed on the data according to the data type to eliminate the difference in magnitude between different dimensions of data and make the processed data dimensionless.
[0009] As a further improvement to this technical solution, in step S200, the process of extracting time-series load features through an LSTM network and screening key influencing factor features through a random forest includes the following steps: S210.1, LSTM network feature extraction: Divide the standardized dataset of S100 into continuous input units according to the time series. , For time indexing, input to the LSTM network; the network passes through the input gate. Controlling the inflow of new information, forget gate Clear redundant information, output gate Filter the output information and generate the hidden layer state. The final output is a time-series load feature vector. ; S210.2, Random Forest Feature Screening: Combining environmental features and equipment operation features into a feature set of influencing factors. With load data Input a random forest model and calculate the node splitting gain for each feature. Determine the contribution level, and select features that meet preset conditions to form a feature set of key influencing factors. ; S210.3 Feature Dimension Alignment: Aligning Time-Series Load Feature Vectors Based on Timestamps Key Influencing Factors Feature Set By unifying the time scales and matching the dimensions of the two sets through dimensional expansion or compression, a set of features to be fused is formed. .
[0010] As a further improvement to this technical solution, in step S200, the process of introducing an attention mechanism to weightedly fuse the two types of features includes the following steps: S220.1 Feature Correlation Analysis: Calculate the set of features to be fused Various characteristics and actual load changes correlation coefficient To form a correlation matrix ; S220.2 Attention Weight Allocation: Based on the Relevance Matrix During peak electricity consumption periods Feature assignment weights For off-peak hours Feature assignment weights ,in To form a dynamic weight matrix ; S220.3, Feature-weighted fusion: This involves using matrix operations to combine the dynamic weight matrix... With the feature set to be fused Fusion to generate a comprehensive feature matrix .
[0011] As a further improvement to this technical solution, in step S200, the process of training the model using the Adam optimizer and verifying the model accuracy through mean absolute percentage error and root mean square error includes the following steps: S230.1, Model Training Configuration: Integrating the Feature Matrix Divided into training set and verification set Set the first moment coefficients of the Adam optimizer Second-order moment coefficients Initialize model parameters ; S230.2, Iterative Training: Using the training set Input the model, and iteratively update the model parameters through the optimizer. After each iteration, the validation set is used. Calculate the loss value Training stops when the loss value does not decrease after a preset number of rounds, and the optimal parameters are saved. ; S230.3, Accuracy Verification: The test set... Input the optimal model and calculate the predicted load. Compared with actual load The mean absolute percentage error and root mean square error are used to output load forecast results that meet the error requirements. .
[0012] As a further improvement to this technical solution, in S300, the process of constructing a multi-objective scheduling function that prioritizes supply and demand balance and is supplemented by energy consumption optimization includes the following steps: S310.1, Scheduling Parameter Definition: Define the load forecast result output by S200 as the baseline load demand. Real-time power grid operating parameters include the output of each power source node. Transmission line loss coefficient Equipment energy consumption characteristics parameters ; S310.2, Construction of Supply and Demand Balance Target: The core target is the matching degree between the total regional power supply and the baseline load demand. The matching degree is measured by the supply and demand deviation coefficient. Characterization, and based on and Calculate the degree of deviation and set a deviation threshold. ,when The time frame is determined to be in a state of basic supply and demand balance; S310.3, Energy Consumption Optimization Target Construction: Based on the comprehensive energy consumption of the power grid The smallest is the auxiliary target. based on , , and scheduling duration Establish a mapping relationship and quantify the correlation between output power and energy consumption through energy consumption characteristic parameters; S310.4 Multi-objective integration: Introducing priority weights , and The supply and demand balance objective and the energy consumption optimization objective are weighted and integrated to form a multi-objective scheduling function. .
[0013] As a further improvement to this technical solution, in step S300, the process of using the improved whale optimization algorithm to solve the problem and generating a scheduling scheme through hierarchical constraints includes the following steps: S320.1 Algorithm Parameter Initialization: Set the population size for the improved whale optimization algorithm. Maximum number of iterations Initialize population individuals (Each individual corresponds to a set of scheduling parameter combinations); S320.2 Adaptive Weight Factor Adjustment: Introducing a weight factor that dynamically changes with the iteration process. Early iteration hour To enhance global search capabilities, larger values are selected in later iterations. hour To improve the accuracy of local optimization, smaller values are selected. Adjust the individual search step size; S320.3, Layered Constraint Judgment: The first layer constraint is the physical safety limit of the equipment. This includes the upper limit of transformer load rate. Rated current of the line The second layer of constraints consists of system stability indicators. Including voltage deviation range Frequency deviation range First verify during the solution process. If the condition is not met, discard the individual directly; if it is met, then re-verify. ; S320.4 Optimal Scheduling Scheme Generation: Through multi-objective scheduling functions Calculate the fitness value of each individual in the population, iteratively update and retain the individual with the best fitness, and finally output the result that simultaneously satisfies... , The combination of constrained scheduling parameters forms a scheduling scheme.
[0014] As a further improvement to this technical solution, in S400, the real-time closed-loop adjustment process includes the following steps: S410.1 High-frequency load monitoring and deviation rate calculation: According to the multi-dimensional sensor data acquisition dimensions in S110.1, set the load monitoring frequency to no less than 12 times per hour, collect the actual load data of the power grid in real time, and combine it with the load prediction results of the corresponding time period output in S230.3 to calculate the deviation rate by the relative deviation between the actual load and the predicted load. S410.2 Deviation Range Judgment: A reasonable deviation range threshold is preset. This threshold is set according to the control logic of the system operation stability index in S320.3. If the deviation rate is within the threshold range, the current scheduling scheme generated by S300 is maintained; if the deviation rate exceeds the threshold range, the closed-loop adjustment mechanism is triggered. S410.3 Incremental Data Training: Collect environmental, equipment, and user data collected in S110.1 to form an incremental dataset. Use the incremental dataset to supplement the training of the attention-enhanced load prediction model in S200, update the model parameters to improve short-term prediction accuracy. S410.4 Fine-tuning of scheduling parameters: Based on the prediction results after incremental training, the output allocation ratio of the power nodes involved in S310.1 is slightly modified. The modification range increases accordingly as the deviation rate increases, and the modified parameters must meet the physical safety limit requirements of the equipment in S320.3. S410.5 Deviation Reduction Verification: Re-collect the actual load and calculate the deviation rate according to the monitoring frequency set in S410.1. Repeat steps S410.3 to S410.4 until the deviation rate falls back to the reasonable threshold range. Then synchronize the final scheduling parameters and adjustment process data to the time series database of S500.
[0015] As a further improvement to this technical solution, in S500, the data archiving and traceability process includes the following steps: S510.1 Time Series Database Configuration and Archive Classification Design: A time series database supporting high-frequency data writing and time series indexing is adopted. The archive categories are divided into stages S100-S400. S100 archives the preprocessed standardized dataset and collection anomaly records; S200 archives the load prediction model parameters, prediction results and accuracy verification data; S300 archives the multi-objective scheduling function parameters, optimization algorithm iteration records and final scheduling scheme; S400 archives the load monitoring data, deviation rate calculation results and closed-loop adjustment records. S510.2 Data Retention Period Setting: The retention period shall be determined according to the power grid equipment operation and maintenance traceability requirements. The retention period for associated data of transformers, transmission lines and switching equipment shall not be less than the design life of the corresponding equipment, and the retention period for temporary dispatch adjustment records shall not be less than 3 power grid operation and maintenance assessment cycles, so as to ensure coverage of the equipment's full life cycle traceability and operation and maintenance review requirements. S510.3 Multi-dimensional search function configuration: Build a multi-dimensional search index in the time series database. The time dimension supports time period search accurate to the minute level, the region dimension is associated with the power grid partition code, and the device number dimension is mapped to the unique identification number of the sensing device, transformer and line in S110.1. The search response time does not exceed 10 seconds. S510.4, Archived Data Traceability Guarantee: The archived data is periodically checked for integrity, and the consistency between the archived data and the original data generated in the S100-S400 stages is compared. A verification log is generated and archived. If data is found to be missing or inconsistent, an automatic re-archiving mechanism is triggered to retrieve the data from the temporary storage nodes of S100-S400 and re-archive it to ensure the accuracy and integrity of the traceable data.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention systematically solves the problem of uneven quality of multi-source data by carrying out collection integrity monitoring, anomaly emergency handling and structured identification and temporary storage of raw data from multiple sources such as user side, equipment side and environment side, and combining it with a combined preprocessing process of "missing value classification and filling + 3σ criterion outlier removal + Z-Score standardization", which can effectively ensure the consistency and availability of data and provide reliable data support for subsequent load forecasting and scheduling optimization. 2. This invention constructs an attention-enhanced load forecasting model, uses an LSTM network to extract time-series load features and a random forest to screen key influencing factors, and then introduces an attention mechanism to highlight the feature weights of peak electricity consumption periods. Combined with Adam optimizer training and accuracy verification, it can more accurately capture the load change patterns in different time periods and improve the adaptability of load forecasting results to the actual power grid operation. 3. This invention constructs a multi-objective scheduling function that prioritizes supply and demand balance and is supplemented by energy consumption optimization. It solves the function using an improved whale optimization algorithm with adaptive weighting factors and prioritizes the physical safety limits of equipment through hierarchical constraints. This allows the invention to meet the power grid supply and demand balance while also taking into account energy consumption optimization, ensuring that the scheduling scheme meets equipment safety requirements and improving the rationality and safety of power grid scheduling. 4. This invention calculates the deviation rate through high-frequency load monitoring. When the deviation exceeds a reasonable range, it achieves closed-loop adjustment by "incremental data training and updating the prediction model + fine-tuning the scheduling parameters". This can promptly correct the deviation between the prediction and the actual operation, reduce the impact of the deviation on the power grid operation, and maintain the stability of the power grid operation. 5. This invention classifies and archives key data throughout the entire process using a time-series database, sets a data retention period that meets operation and maintenance requirements, configures multi-dimensional search functions, and performs timed integrity checks and automatic data replenishment. This enables full-process data traceability and provides data support for power grid equipment operation and maintenance review and technical optimization. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall method steps of the present invention; Figure 2 This is a schematic diagram illustrating the steps of collecting raw data in this invention; Figure 3 This is a schematic diagram illustrating the data processing steps in this invention; Figure 4 This is a schematic diagram of the real-time closed-loop adjustment steps in this invention; Figure 5 This is a schematic diagram illustrating the data archiving and traceability steps in this invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] like Figures 1-3 As shown, this embodiment provides a method for power grid load forecasting and dynamic scheduling optimization based on big data, including: S100, Power Grid Multi-Source Data Preprocessing: Use multi-dimensional sensing devices to collect raw data from the user side, equipment side, and environment side, and process it through a combination of "missing value classification and imputation + 3σ criterion outlier removal + Z-Score standardization"; In this step, the process of collecting raw data from the user side, device side, and environment side using multi-dimensional sensing devices in S100 includes the following steps: S110.1 Selection and Matching of Data Acquisition Equipment: For raw data acquisition on the user side, select multi-dimensional sensing equipment adapted to different power consumption scenarios to collect real-time active power, reactive power, and cumulative power consumption; for raw data acquisition on the equipment side, select sensing equipment matched to the type of transformer, transmission line, and switching equipment to collect transformer oil temperature, line current, line voltage, and switch open / closed status; for raw data acquisition on the environmental side, select sensing equipment adapted to outdoor power grid scenarios to collect air temperature, relative humidity, and wind speed. As a further explanation of this step, the selection and matching of data acquisition equipment in this step specifically includes: User-side sensing equipment: For residential scenarios, smart meters with RS485 interfaces are selected (supporting active / reactive power acquisition, with a sampling frequency that can be designed to be 15 minutes / time); for industrial and commercial scenarios, industrial-grade three-phase energy meters are selected (with a sampling frequency that can be designed to be 5 minutes / time, supporting simultaneous acquisition of multiple circuits).
[0020] Equipment-side sensing devices: Transformer oil temperature acquisition uses a platinum resistance temperature sensor; line current / voltage acquisition uses electronic current transformers and voltage transformers, with a sampling frequency that can be designed to be 1 minute / time; switch opening / closing status acquisition uses a passive contact sensor, which uploads data in real time when the status changes.
[0021] Environmental sensing equipment: Temperature and humidity sensors and ultrasonic wind speed sensors with rain and snow resistant housings are selected and installed in the outdoor area of the substation and transmission line towers. The sampling frequency can be designed to be 10 minutes / time.
[0022] S110.2 Data integrity monitoring during the acquisition process: During the acquisition process, the data stream transmission status of the sensing device is monitored in real time. When it is detected that a single device has failed to upload data for M consecutive times or the proportion of missing fields in the uploaded data exceeds the preset threshold, it is determined to be an acquisition anomaly and the device self-test program is triggered. As a further explanation of this step, the preferred value of M is 3 times (i.e. no data is uploaded for 3 consecutive sampling cycles). This value is combined with the communication stability and real-time requirements of power grid data acquisition. Most sensing devices have a sampling cycle of 5-15 minutes. The failure to upload for 3 consecutive times not only reserves fault tolerance space for short-term communication fluctuations (such as network delays) but also can promptly alarm when the device continues to fail, ensuring the continuity and reliability of data acquisition. Furthermore, the preset threshold for the proportion of missing data fields is preferably 30% (if a single data entry contains 10 fields, missing ≥3 fields is considered an anomaly). From the perspective of balancing the importance and completeness of fields in multi-source power grid data, if a single data entry is missing more than 30%, the remaining valid fields will be insufficient to support subsequent multi-dimensional analysis (such as load forecasting which requires the linkage of multiple fields such as users, equipment, and environment). This threshold avoids excessive judgment of anomalies due to a small number of missing fields, while ensuring that the data quality meets the basic requirements for model training and scheduling decisions. Furthermore, the data stream transmission status is monitored through the heartbeat detection mechanism of the edge gateway, which sends a status message every 5 seconds. If no response is received within the timeout period, an anomaly is detected.
[0023] S110.3 Emergency handling of abnormal data acquisition: If the self-test procedure still cannot restore normal data acquisition, activate the backup sensor device pre-deployed in the same area to take over the abnormal device and continue data acquisition. At the same time, record the abnormal device number, the time of the abnormality and the type of abnormality. As a further explanation of this step, the emergency handling mechanism for data acquisition anomalies in this step specifically includes: Backup sensing devices are deployed in a ratio of "1 main and 1 backup" or "3 main and 1 backup" (the former is used in important load areas). The backup devices collect the same parameters as the main devices and are in a hot standby state. The anomaly types include "communication interruption", "sensor failure" and "data over-range", each corresponding to a different alarm level (e.g., communication interruption is a level 2 alarm and sensor failure is a level 1 alarm). The alarm information is synchronously uploaded to the power grid monitoring platform.
[0024] S110.4 Raw Data Identification and Temporary Storage: The raw data collected normally is structurally identified according to "collection object type - device unique number - collection time" and temporarily stored in the nearest edge storage node. During the temporary storage process, data verification codes are used to verify data integrity.
[0025] As a further explanation of this step, the temporary storage specifications are as follows: edge storage nodes use industrial-grade SD cards or local hard drives, with a single node storage capacity of ≥1TB, a temporary storage period of 24 hours, and automatic compression and uploading to the regional power grid data center after expiration, with a 30-day backup retained locally.
[0026] Furthermore, this embodiment also provides examples of structured identifiers, such as: User side: User side - Residential area A - Meter ID 001 - 202310010800 Equipment side: Equipment side - Transformer T102 - Temperature sensor S203-202310010800 Environmental side: Environmental side - power transmission line L506 - temperature and humidity sensor H701-202310010800.
[0027] In this step, the combined process of "missing value classification and imputation + 3σ criterion outlier removal + Z-Score standardization" in S100 includes the following steps: S120.1 Missing Value Classification and Imputation: The original data temporarily stored in S110.4 is classified into user side, device side and environment side. For missing values in each type of data, if there are consecutive missing values in a short period of time, the linear interpolation method of the adjacent valid data before and after the missing value is used to imput them. If there are consecutive missing values in a long period of time, the mean of the same type of data in the same area in the same period of time is used to imput them. As a further explanation of this step, missing value imputation specifically includes: Short-term consecutive missing data: refers to missing data with a duration of ≤1 hour (e.g., 4 consecutive missing data points at a 15-minute sampling interval). When filling in missing data, take the two most recent valid data points before and after the missing data, and calculate the intermediate missing value according to the time interval ratio (e.g., if the data at 8:00 and 8:30 are valid, the missing value at 8:15 is the midpoint between the two).
[0028] Long-term continuous missing data: refers to missing data for a duration greater than 1 hour. When filling missing data, select historical data from the same region and the same type of equipment at the same time on a similar date (e.g., if missing data is at 8:00 on October 1, take the data at 8:00 on September 28, 29 and October 2 and 3), and calculate the average value as the filling value.
[0029] S120.2, 3σ Criterion Outlier Removal: For the data imputed by S120.1, calculate the mean and standard deviation of each data type according to the user side, equipment side and environment side respectively. Values that deviate from the mean by more than 3 times the standard deviation are identified as outliers and removed. As a further explanation of this step, the 3σ criterion for outlier removal specifically includes: The average value and fluctuation range (standard deviation) of each type of data are calculated on a daily basis. For example, if the daily average value of a transformer oil temperature is 45℃ and the fluctuation range is 5℃, then values exceeding 45±15℃ are judged as abnormal and removed.
[0030] For highly volatile data such as wind speed, a 24-hour sliding window is used for calculation (the average value and fluctuation range are updated every hour) to ensure that anomaly detection is adapted to the data characteristics.
[0031] Outliers that are removed are marked with their location and reason in the original dataset (e.g., "out of 3σ range"), and the record is retained for subsequent data quality analysis.
[0032] S120.3 Z-Score Standardization: After outliers are removed by S120.2, Z-Score standardization is performed on the data according to the data type to eliminate the difference in magnitude between different dimensions of data and make the processed data dimensionless.
[0033] As a further explanation of this step, Z-Score standardization specifically includes: During processing, the magnitude differences between different data (e.g., power is in kW, temperature is in °C) are eliminated, and the data is converted into dimensionless relative values. The converted data are usually distributed between -3 and 3. Values outside this range are uniformly treated as boundary values (i.e., less than -3 is treated as -3, and greater than 3 is treated as 3) to ensure stable data distribution and facilitate subsequent model training.
[0034] S200 Attention-Enhanced Load Forecasting Model Construction: Based on the standardized S100 dataset, time-series load features are extracted using an LSTM network, and key influencing factor features are selected using a random forest. An attention mechanism is introduced to weight and fuse the two types of features, highlighting the feature weights during peak electricity consumption periods. The model is trained using the Adam optimizer, and the model accuracy is verified using mean absolute percentage error and root mean square error. Load forecasting results are then output. In this step, the process of extracting time-series load features through an LSTM network and screening key influencing factor features through a random forest in S200 includes the following steps: S210.1, LSTM network feature extraction: Divide the standardized dataset of S100 into continuous input units according to the time series. , For time indexing, input to the LSTM network; the network passes through the input gate. Controlling the inflow of new information, forget gate Clear redundant information, output gate Filter the output information and generate the hidden layer state. The final output is a time-series load feature vector. ; As a further explanation of this step, in the feature extraction of the S210.1LSTM network in this embodiment, the calculation of key gating and hidden layer states specifically includes: Input gate Calculation: The input gate is used to control the inflow of new information into the LSTM cell. The formula is: ;in, for The input gate output value at any time (range 0-1, where 0 means completely blocking information and 1 means completely allowing information to flow in). It is the sigmoid activation function; For input layer (input unit) The weight matrix from the input gate; for Time-based input unit (including vectors of standardized data from the user side, device side, and environment side); The hidden layer of the previous time step ( The weight matrix from the input gate; for Hidden layer state vector at any given time; The input gate bias term (a constant vector used to adjust the activation function offset); passed through the input unit Compared to the previous hidden layer state A linear combination of these values is mapped to a value between 0 and 1 via the sigmoid function, thereby controlling the intensity of the inflow of new information. Forgotten Gate Calculation: The forget gate is used to remove redundant historical information within cells. The formula is: ;in, for Output value of the time-forget gate (range 0-1, 0 means completely retaining historical information, 1 means completely forgetting historical information). This is the weight matrix from the input layer to the forget gate; This is the weight matrix from the previous hidden layer to the forget gate; This is the forgetting gate bias term. Its computational logic is consistent with the input gate structure. Through linear combination and sigmoid mapping, it determines the proportion of historical information retained within the cell. For example, in load forecasting, the "historical load patterns during off-peak periods" can be obtained through... Forget about things appropriately to avoid interfering with current predictions; Cell state update Calculation: Cell state is the core memory unit of LSTM, and the update formula is as follows: ;in, for Time-based cell state update value (range -1 to 1, representing newly generated memory information); It is the hyperbolic tangent activation function (used to map numerical values to -1-1, enhancing gradient propagation stability). This is the weight matrix from the input layer to the cell state; This is the weight matrix from the previous hidden layer to the cell state; This refers to cell state bias terms; new memory information is generated through linear combination, after... Function normalization provides a basis for subsequent cell state updates.
[0035] Cell state With hidden layer state Calculation: The formula for the final cell state and the hidden layer state is: , ;in, for The final cell state at any given moment (preserving core memory information); for Cell state at any given moment; for The output value of the time-limited output gate is calculated in the same way. The formula is ,in This is the weight matrix from the input layer to the output gate. This is the weight matrix from the previous hidden layer to the output gate. (for output gate bias terms). for Hidden layer state at any given time (i.e., time-series load feature vector) The core source); its computational logic is: first through the forget gate reserve Valid historical information in the input gate New information after filtering ,get Then through the output gate filter Key information in the middle, After normalization, we get Ultimately Organized into a 32-dimensional time-series load feature vector .
[0036] S210.2, Random Forest Feature Screening: Combining environmental features and equipment operation features into a feature set of influencing factors. With load data Input a random forest model and calculate the node splitting gain for each feature. Determine the contribution level, and select features that meet preset conditions to form a feature set of key influencing factors. ; As a further explanation of this step, in the random forest feature selection in S210.2 of this embodiment, the node splitting gain... The calculation of (information gain ratio) specifically includes: Information gain ratio is used to measure the impact of features on load data. The formula for the classification contribution is: ; In this formula, information gain ,in For the feature set of influencing factors For sample set Information gain (the larger the value, the greater the influence factor feature set) The higher the contribution to load forecasting, the better. The sample set for the random forest (including and The corresponding data, with a sample size of ); The information entropy of sample set D (reflecting) (uncertainty) Features under conditions Conditional entropy (reflecting known conditions) back (remaining uncertainty); Split Information ;in, Features The number of value categories, For sample set Chinese characteristics Take the first The sample subset when the value is 1, for The number of samples, for The total number of samples; Used to measure characteristics The complexity of the subsets after splitting can avoid information gain bias towards features with too many values (such as "device number" which has many values but no actual predictive significance).
[0037] Conditional entropy ;in, For the feature set of influencing factors The number of possible values (e.g., "air temperature" is divided into 5 values based on intervals) =5); for Characteristic set of influencing factors Take the first A subset of samples with values; for Information entropy (calculated in the same way) ); In this formula, splitting information ;in, Used for measurement The degree of dispersion of values should be considered to avoid information gain bias towards features with too many values (such as "device number", which has many values but no actual predictive significance). First calculate the sample set Information entropy Then, based on the characteristic set of influencing factors Divide the sample into subsets based on the value of and calculate the conditional entropy. With split information Ultimately, through information gain ratio Filter features — select all features Sort the features and select the top 60% to form the key influencing factor feature set. (For example, selecting 4 categories from 7 categories of features).
[0038] S210.3 Feature Dimension Alignment: Aligning Time-Series Load Feature Vectors Based on Timestamps Key Influencing Factors Feature Set By unifying the time scales and matching the dimensions of the two sets through dimensional expansion or compression, a set of features to be fused is formed. .
[0039] In this step, the process of introducing an attention mechanism to weightedly fuse the two types of features in S200 includes the following steps: S220.1 Feature Correlation Analysis: Calculate the set of features to be fused Various characteristics and actual load changes correlation coefficient To form a correlation matrix ; As a further explanation of this step, in the feature correlation analysis of S220.1 in this embodiment, the correlation coefficient... The calculation of the Pearson correlation coefficient specifically includes: Pearson correlation coefficient is used to measure the set of features to be fused. A certain characteristic and the actual load change The degree of linear correlation is given by the following formula: ; The correlation coefficient (range -1 to 1, where 1 indicates perfect positive correlation, -1 indicates perfect negative correlation, and 0 indicates no correlation). For the sample size (e.g., selecting 15-minute interval data over 30 days), ); for The first of a certain feature The first sample value (e.g., the first sample value of "air temperature") (15-minute average) This is the sample mean of this feature. ; For the first The actual load change of each sample ,in For the first The actual load value of each sample For the first The actual load value of each sample; for The sample mean, ; The specific calculation logic is as follows: Through molecular calculation characteristics and... The covariance (reflecting the common trend of both) is calculated, and the product of their standard deviations (reflecting their respective degrees of dispersion) is calculated in the denominator to obtain the standardized correlation coefficient. ;Will All 32 features The values are arranged in a feature-feature matrix to form a 32×32 correlation matrix. (like Indicates the first The first feature and the second Features value).
[0040] S220.2 Attention Weight Allocation: Based on the Relevance Matrix During peak electricity consumption periods Feature assignment weights For off-peak hours Feature assignment weights ,in To form a dynamic weight matrix ; As a further explanation of this step, in the attention weight allocation of S220.2 in this embodiment, the dynamic weight... (Peak hours) and The calculation for (off-peak hours) specifically includes: Peak period weighting calculate: ;in, Peak electricity consumption period (08:00-11:00, 18:00-22:00) The weights of each feature (ranging from 0 to 1, and the weights of all features) (sum of 1) The first during peak hours Features and The correlation coefficient (taken from) The corresponding time period in ) value); for Feature dimensions ( =32); For all 32 features during peak hours The sum of the absolute values; Off-peak period weighting calculate: ;in, Off-peak hours (remove (outside of the time period) The weights of each feature (ranging from 0 to 1, and the weights of all features) (sum of 1) During off-peak hours Features and The correlation coefficient; For all 32 features during off-peak hours The sum of the absolute values; Specifically, the calculation logic is as follows: by taking the absolute value of the correlation coefficient (ignoring the positive and negative correlation directions and only focusing on the correlation strength), and then performing normalization (dividing by all features). (The sum of the absolute values), ensuring that the sum of all feature weights is 1 within the same time period; due to peak periods Greater fluctuations The absolute value is usually higher than Therefore (such as a certain feature) =0.7, =0.2), ultimately will Arranged by time index, forming a 32×32 dynamic weight matrix. ( express Time of the first (Weights of each feature).
[0041] S220.3, Feature-weighted fusion: This involves using matrix operations to combine the dynamic weight matrix... With the feature set to be fused Fusion to generate a comprehensive feature matrix .
[0042] As a further explanation of this step, and as a further explanation of this embodiment, in the feature weighted fusion of S220.3 of this embodiment, the comprehensive feature matrix... The calculations specifically include: Element-wise multiplication is used to fuse dynamic weights and features to be fused, as shown in the formula: ; for Time of the first The comprehensive eigenvalue of dimension (i.e.) (core elements) for Time of the first Dynamic weights of dimension (taken from) ); for Time of the first The dimensional features to be fused (taken from) ); The specific calculation logic is: for each time point Each dimension of features Using weights For original features Scaling is applied to features with high weight during peak hours (such as "air temperature" and "transformer oil temperature"). The value is amplified, and the characteristic of low weight during off-peak hours is that... The values are reduced to highlight the impact of key characteristics during peak periods on load forecasting, ultimately resulting in a 32-dimensional value. matrix.
[0043] It is understandable that for dynamic weight matrices... Its dimensions are The rationale for setting this dimension lies in: comprehensive features It is a 32-dimensional vector. This describes the strength of the attention association between each feature within these 32 dimensions and all other features—that is, the first feature in the matrix. Line number The elements of the column represent the first element. dimensional features for the 1st Attention weights for dimensional features, through The matrix structure can fully characterize the pairwise attention relationships between 32-dimensional features, enabling refined modeling of the internal correlations of features.
[0044] In this step, the process of training the model using the Adam optimizer and verifying the model accuracy through mean absolute percentage error and root mean square error in S200 includes the following steps: S230.1, Model Training Configuration: Integrating the Feature Matrix Divided into training set and verification set Set the first moment coefficients of the Adam optimizer Second-order moment coefficients Initialize model parameters ; As a further explanation of this step, the preprocessed comprehensive feature matrix will be... The dataset is randomly partitioned according to a 7:2:1 ratio of training set:validation set:test set. If the dataset exceeds 100,000 records, an 8:1:1 ratio is used to balance the sufficiency of model training with the accuracy of generalization assessment. The training set is used for learning model parameters, the validation set is used for adjusting hyperparameters (such as the number of LSTM network layers and the number of random forest decision trees) during iterations, and the test set is used for independent validation of the final model performance.
[0045] S230.2, Iterative Training: Using the training set Input the model, and iteratively update the model parameters through the optimizer. After each iteration, the validation set is used. Calculate the loss value Training stops when the loss value does not decrease after a preset number of rounds, and the optimal parameters are saved. ; As a further explanation of this step, in S230.1-S230.2 of this embodiment, the parameter update and loss function calculation of the Adam optimizer specifically include: On one hand, this includes Adam optimizer parameter updates: the Adam optimizer updates model parameters through first moment (momentum) and second moment (adaptive learning rate). The steps are as follows; First moment estimate (momentum): ;in, for The first moment of the model parameters at each time step (a moving average reflecting the gradient); The first-order moment decay coefficient (preset to 0.9, controlling the retention ratio of historical momentum); for The first moment at time; for The gradient of the model parameters at time step (via the loss function) The derivative is obtained, reflecting the degree of influence of the parameters on the prediction bias. Second-order moment estimation (adaptive learning rate): ;in, for The second moment of the model parameters at each time step (a moving average reflecting the squared gradient). This is the second-order moment decay coefficient (preset to 0.999, controlling the retention ratio of the squared historical gradient). for The second moment at time t; for The square of the gradient at time step (element-wise square); First-order moment deviation correction: ; The corrected first moment (eliminating the bias of the first moment towards 0 in the early stage of training); for of power ( (current iteration round); Second-order moment deviation correction: ;in, The corrected second moment (eliminating the bias of the second moment towards 0 in the early stage of training); for of Power of 1.
[0046] Parameter update: ;in, for The updated model parameters (including LSTM weights, random forest splitting threshold, etc.) at each time step. for Model parameters at time The learning rate (default is 0.001, controlling the step size of parameter updates); The square root of the corrected second moment (providing an adaptive learning rate); To prevent small values with a denominator of 0 (default is 10) -8 ); The first moment preserves the continuity of the gradient direction (avoiding parameter oscillations), the second moment adjusts the learning rate according to the gradient magnitude (smaller learning rate for parameters with larger gradients, and larger learning rate for parameters with smaller gradients), and the correction term eliminates the bias in the early stage of training, achieving stable and efficient parameter updates.
[0047] On the other hand, including the loss function (Mean Squared Error) Calculation: The loss function measures the deviation between the model's predicted values and the actual values; the formula is as follows: ; for The loss value of each iteration (the smaller the value, the higher the model prediction accuracy); For the validation set The number of samples (e.g.) =2000); For the model on the validation set The predicted load value for each sample; For the verification set The actual load value of each sample; The specific calculation logic is as follows: Square the prediction bias for each sample (amplifying the impact of larger biases), then calculate the average to obtain the loss value for that iteration; when there are 5 consecutive iterations... No decrease (i.e.) When ), stop training and save the first (number of) data. The parameters of the wheel are used as optimal parameters. .
[0048] S230.3, Accuracy Verification: The test set... Input the optimal model and calculate the predicted load. Compared with actual load The mean absolute percentage error and root mean square error are used to output load forecast results that meet the error requirements. .
[0049] As a further explanation of this step, in the accuracy verification of S230.3 in this embodiment, the calculation of the mean absolute percentage error and the root mean square error specifically includes: On the one hand, this includes the calculation of the mean absolute percentage error (MAPE): MAPE is used to measure the proportion of the predicted value to the actual value. The formula is: ;in, The mean absolute percentage error (in %, the smaller the value, the smaller the relative deviation). For the test set The number of samples (e.g.) =1000); For the model on the test set The predicted load value for each sample; For the test set The actual load value of each sample ( (to avoid a denominator of 0) It is the absolute value symbol; On the other hand, this includes the calculation of the root mean square error (RMSE): RMSE is used to measure the absolute deviation between a predicted value and the actual value. The formula is: ;in, This is the root mean square error (units consistent with the load, such as kW; the smaller the value, the smaller the absolute deviation). The specific calculation logic is as follows: First, calculate the sum of squares and average values of the prediction deviations for each sample (i.e., mean square error MSE), then take the square root of the mean square error to obtain an error value consistent with the load unit, thus adapting to the power grid's requirement for absolute accuracy in load prediction.
[0050] S300, generation of improved whale optimization scheduling scheme with hierarchical constraints: Based on the load forecast results of S200 and the real-time operating parameters of the power grid, a multi-objective scheduling function with "supply and demand balance as the priority and energy consumption optimization as the auxiliary" is constructed; and an improved whale optimization algorithm is used to solve it. The algorithm introduces an adaptive weight factor to adjust the search strategy, and determines the priority to meet the physical safety limits of the equipment through hierarchical constraints to generate a scheduling scheme. In this step, the process of constructing a multi-objective scheduling function with "supply and demand balance as the priority and energy consumption optimization as the secondary objective" in S300 includes the following steps: S310.1, Scheduling Parameter Definition: Define the load forecast result output by S200 as the baseline load demand. Real-time power grid operating parameters include the output of each power source node. Transmission line loss coefficient Equipment energy consumption characteristics parameters ; S310.2, Construction of Supply and Demand Balance Target: The core target is the matching degree between the total regional power supply and the baseline load demand. The matching degree is measured by the supply and demand deviation coefficient. Characterization, and based on and Calculate the degree of deviation and set a deviation threshold. ,when The time frame is determined to meet the basic supply and demand balance; S310.3, Energy Consumption Optimization Target Construction: Based on the comprehensive energy consumption of the power grid The smallest is the auxiliary target. based on , , and scheduling duration Establish a mapping relationship and quantify the correlation between output power and energy consumption through energy consumption characteristic parameters; S310.4 Multi-objective integration: Introducing priority weights , and The supply and demand balance objective and the energy consumption optimization objective are weighted and integrated to form a multi-objective scheduling function. .
[0051] As a further explanation of this step, the construction of the multi-objective scheduling function in this embodiment specifically includes: combining the core requirements of "prioritizing supply and demand balance + supplementing with energy consumption optimization", based on the baseline load demand. Output of each power node Transmission line loss coefficient Equipment energy consumption characteristics parameters Using these as core parameters, construct a multi-objective scheduling function. The formula is: ; in, This is the combined value of the multi-objective scheduling function (the smaller the value, the better the scheduling scheme). , The priority weights defined in the file, and (like =0.7、 =0.3), reflecting the principle of "prioritizing supply and demand balance"); The load forecast result output by S200 (baseline load demand defined by S310.1). It provides power to each power node in S310.1; In the above formula: The objective function for supply and demand balance is calculated using the following formula: ;in, The supply-demand deviation coefficient defined in S310.2 is calculated using the following formula: In the formula, The scheduling duration in S310.3 (e.g., a day divided into 15-minute intervals), =96); For scheduling time step ( ); for Time of the first The actual output of each power node (belonging to) (detailed parameters); for Baseline load demand at any given time ( (time series values) Reflecting time t and The degree of deviation must meet the requirements in S310.2. ( For example, the deviation threshold, =0.05, meaning a deviation within 5% is considered to meet the basic supply and demand balance); Through time-series averaging Quantify the supply-demand matching degree throughout the entire scheduling cycle. The smaller the value, the better. and The smaller the overall deviation, the better it aligns with the document's requirement of "supply and demand balance as the core objective"; The objective function for energy consumption optimization is: ;in, For the first Energy consumption characteristic parameters of power supply equipment (belonging to) Detailed parameters, such as those for coal-fired power units (Value higher than that of photovoltaic units) pass Quantify the energy consumption of the equipment itself. Quantify the transmission loss of the lines, and sum the two to obtain the comprehensive energy consumption of the power grid over the entire cycle. , The smaller the value, the lower the energy consumption, which aligns with the positioning of "energy consumption optimization as a secondary measure".
[0052] Furthermore, and pass Weighted fusion ,because In algorithm iterations, it will preferentially reduce (Ensuring supply and demand balance), further optimization (Reduce energy consumption) to ultimately form a scheduling function that takes into account both core and auxiliary objectives. .
[0053] In this step, the process of using the improved whale optimization algorithm to solve the problem and generating a scheduling scheme through hierarchical constraints in S300 includes the following steps: S320.1 Algorithm Parameter Initialization: Set the population size for the improved whale optimization algorithm. Maximum number of iterations Initialize population individuals (Each individual corresponds to a set of scheduling parameter combinations); As a further explanation of this step, For population size (e.g.) Each individual in the population A corresponding set of scheduling parameters, i.e. , ); Maximum number of iterations (e.g.) =100), controlling the algorithm's solution time); Furthermore, based on the actual number of power sources in the power grid (like =5, including 2 thermal power units, 2 wind farms, and 1 photovoltaic power station), for each in Set a reasonable range (such as thermal power units) , (Rated output of the power supply) to ensure It matches the actual operating capacity of the equipment.
[0054] It is understandable that when using the improved whale optimization algorithm to optimize multi-objective scheduling functions, the population size is set. (In the field of intelligent optimization algorithms, population size) This represents a typical balance between "ensuring population diversity to avoid local optima" and "controlling computational costs," considering the computational complexity of power grid dispatching problems. (This is a common choice that balances efficiency and effectiveness within this range); set the maximum number of iterations. (Based on general experience with intelligent optimization algorithms, when the number of iterations reaches the hundreds, if the fitness function value does not change significantly, the benefit of continuing iterations is extremely low. Therefore...) This is a reasonable upper limit for iterations in this type of multi-constraint optimization problem.
[0055] S320.2 Adaptive Weight Factor Adjustment: Introducing a weight factor that dynamically changes with the iteration process. Early iteration hour To enhance global search capabilities, larger values are selected in later iterations. hour To improve the accuracy of local optimization, smaller values are selected. Adjust the individual search step size; As a further explanation of this step, weighting factors The calculation formula is: ; in, For the weights in the early stages of iteration (e.g.) =1.5, enhancing global search); For later iteration weights (e.g.) =0.5, strengthening local optimization); The current iteration number ( ); when (Early Iteration) near The algorithm covers more candidate scheduling schemes by increasing the individual search step size; when (Late iteration) near The algorithm reduces the search step size and refines the optimization around the optimal solution. Balancing the capabilities of "global exploration" and "local development".
[0056] Understandably, in the weighting factor In the calculation, combining the optimization characteristics of the power grid dispatching problem (multiple constraints, nonlinearity) and the general practices in the field of intelligent optimization algorithms, we take... This range allows the algorithm to maintain a large weight in the early stages (when the number of iterations t is small) to achieve global solution space exploration, and in the later stages ( Approaching the maximum number of iterations Gradually reducing the weight and focusing on local refinement for optimization is a commonly used value range in the industry that takes into account both "global exploration and local development".
[0057] S320.3, Layered Constraint Judgment: The first layer constraint is the physical safety limit of the equipment. This includes the upper limit of transformer load rate. Rated current of the line The second layer of constraints consists of system stability indicators. Including voltage deviation range Frequency deviation range First verify during the solution process. If the condition is not met, discard the individual directly; if it is met, then re-verify. ; As a further explanation of this step, the hierarchical constraint judgment in this step specifically includes: First level of constraints (Equipment physical safety limits): Transformer load factor constraint: ,in for Real-time transformer load rate , This is the transformer terminal voltage. This is the load current; Rated apparent power is the apparent power that electrical equipment (such as transformers, generators, etc.) can provide during long-term safe operation under rated voltage and rated current. The unit is usually kilovolt-amperes (kVA), and it is a core rated parameter that reflects the capacity of the equipment.
[0058] Line current constraints: ,in for The actual current of the transmission line at any given time.
[0059] Second-level constraint C2 (system stability index): Voltage deviation constraint: ,in The rated voltage of the busbar (e.g., 10kV); Frequency deviation constraint: ,in This is the system's rated frequency.
[0060] Verification logic: First verify the physical safety limits of the device. ,like or Directly discard individuals in the current population ; Verify after satisfaction If the voltage or frequency exceeds , The scope is also discarded. Only retain those conditions where both constraints are satisfied. Moving on to the next iteration.
[0061] S320.4 Optimal Scheduling Scheme Generation: Through multi-objective scheduling functions Calculate the fitness value of each individual in the population, iteratively update and retain the individual with the best fitness, and finally output the result that simultaneously satisfies... , The combination of constrained scheduling parameters forms a scheduling scheme.
[0062] Specifically, the generation of the optimal scheduling scheme includes the following steps: Fitness calculation: using a multi-objective scheduling function As an individual in a population fitness value, The smaller, The better the corresponding scheduling scheme; Iterative update: Each iteration retains the best-fitting element. and based on Adjust other The search step size is used to generate a new population; Solution output: When the number of iterations... When the iteration stops, output the final optimal value that is retained. ,Should corresponding Combination is to satisfy , The constrained scheduling scheme forms a power output allocation plan that can be directly executed.
[0063] like Figure 4 As shown, S400 real-time closed-loop adjustment: high-frequency load monitoring is used to calculate the deviation rate; when the deviation exceeds the reasonable range, the deviation is brought back to the reasonable range through a combination of "incremental data training + scheduling parameter fine-tuning"; In this step, the real-time closed-loop adjustment process in S400 includes the following steps: S410.1 High-frequency load monitoring and deviation rate calculation: According to the multi-dimensional sensor data acquisition dimensions in S110.1, set the load monitoring frequency to no less than 12 times per hour, collect the actual load data of the power grid in real time, and combine it with the load prediction results of the corresponding time period output in S230.3 to calculate the deviation rate by the relative deviation between the actual load and the predicted load. As a further explanation of this step, S410.1 high-frequency load monitoring and deviation rate calculation in this embodiment specifically includes: Based on the multi-dimensional sensing device data acquisition dimensions (user side, device side, and environment side) defined in S110.1, the load monitoring frequency is specified as once every 5 minutes (12 times per hour, meeting the requirement of "no less than 12 times per hour"). The monitoring equipment reuses the sensing devices deployed in S110.1. On the user side, the total active power and reactive power of the zone are collected through smart meters. On the device side, the actual load current of the transmission line is collected through line current / voltage sensors. On the environment side, environmental parameters at the monitoring time are recorded synchronously through temperature and humidity sensors (for subsequent incremental training and correlation analysis). All monitoring data are transmitted to the power grid dispatching platform in real time through the edge gateway.
[0064] The deviation rate is calculated based on the "actual load at the monitoring time and the predicted load for the corresponding time period", and the formula is as follows: ;in, for Load deviation rate at any given time (reflecting the relative deviation between actual and predicted loads). for The actual load value of the power grid at any given time (taken from the sum of the zone loads of the above monitoring data, such as the actual load of a certain area at 08:05 is 120MW). for The load forecast result at the specified time (taken from the corresponding timestamp forecast value output by S230.3, such as the predicted load at 08:05 is 115MW).
[0065] By quantifying the difference between the prediction and the actual load using the relative deviation rate, errors in deviation judgment caused by the absolute value of the load can be avoided. For example, the absolute deviation of 10MW has a deviation rate of 10% in a 100MW load scenario and a deviation rate of 5% in a 200MW load scenario, which can more accurately reflect the severity of deviation under different load levels.
[0066] S410.2 Deviation Range Judgment: A reasonable deviation range threshold is preset. This threshold is set according to the control logic of the system operation stability index in S320.3. If the deviation rate is within the threshold range, the current scheduling scheme generated by S300 is maintained; if the deviation rate exceeds the threshold range, the closed-loop adjustment mechanism is triggered. As a further explanation of this embodiment, the deviation range determination in S410.2 of this embodiment specifically includes: When setting a reasonable range threshold for the preset deviation, strictly follow the control logic of "prioritizing system operation stability indicators" in S320.3, and set differentiated thresholds based on the load sensitivity of the power grid during different operating periods: Peak electricity consumption periods (consistent with the S200 definition, i.e., 08:00-11:00 and 18:00-22:00): Due to the significant impact of load fluctuations on system stability, the reasonable threshold for deviation rate is set at ≤5%; Off-peak hours (other times): The impact of load fluctuations is relatively small, and the reasonable threshold for deviation rate is set at ≤8%; The threshold setting refers to the stability requirement in "GB / T 15945-2008 Power Quality Power System Frequency Deviation" that "the system frequency deviation must be controlled within ±0.2Hz" to ensure that the deviation threshold matches the system's operational safety boundary.
[0067] The judgment logic is: real-time comparison by the scheduling platform. Compared with the corresponding time period threshold, if Within the threshold range (such as during peak hours) If the percentage is 3%, then the current scheduling scheme generated by S300 will remain unchanged; if... Exceeding the threshold (e.g., during peak hours) If the deviation rate is 7%, the closed-loop adjustment mechanism is triggered—the scheduling platform automatically sends an "adjustment start signal" to the incremental training module and the parameter fine-tuning module, and simultaneously records the trigger time, current load data and deviation rate, which serves as a basis for subsequent traceability.
[0068] S410.3 Incremental Data Training: Collect environmental, equipment, and user data collected in S110.1 to form an incremental dataset. Use the incremental dataset to supplement the training of the attention-enhanced load prediction model in S200, update the model parameters to improve short-term prediction accuracy. As a further explanation of this embodiment, the incremental data training in S410.3 of this embodiment specifically includes: The incremental dataset is constructed with a time frame of "data within one hour after the deviation rate exceeds the threshold". The data source is completely consistent with the data collection dimensions of S110.1, and its specific components include: Environmental data: Air temperature, relative humidity, and wind speed collected every 5 minutes within 1 hour (a total of 12 sets of data); Equipment-side data: Transformer oil temperature, line current, and bus voltage collected every 5 minutes within 1 hour (a total of 12 sets of data); User-side data: Active power, reactive power, and cumulative electricity consumption of each zone are collected every 5 minutes within 1 hour (12 sets of data in total). All data is structured in the format of "timestamp-data type-value" to ensure that it matches the input format of the attention-enhanced load forecasting model in S200.
[0069] Furthermore, the incremental training process adopts a "small batch iterative update" mode, with specific parameters set as follows: the learning rate is set to 0.0005 (lower than the initial training of S230 at 0.001, to avoid over-updating the original model parameters), the training batch size is 32, and the training epochs are 5-10 (due to the small amount of incremental data, the number of epochs does not need to be too many to prevent model overfitting). The training objective is to "minimize the 15-minute short-term prediction bias," correcting the model's understanding of the current load fluctuation pattern through incremental data—for example, if the bias is caused by a sudden drop in temperature (a sudden drop of 5°C in ambient temperature), the model can learn the "correlation between sudden temperature drop and load growth" through incremental data, update the contribution of ambient temperature features in the attention weights, and thus improve the prediction accuracy for the following 15-30 minutes.
[0070] Understandably, to ensure the stability of incremental model training, and referencing the typical time resolution of power grid load data (e.g., 5 minutes / sample) and the minimum data volume requirement for incremental machine learning, the minimum size of the incremental dataset must contain at least 12 sets of continuous valid data (corresponding to approximately 1 hour of sampling time, which can cover basic load fluctuation characteristics). When data is abnormal (e.g., missing data due to sensor failure, values exceeding the equipment's rated range, etc.), a historical data compensation strategy based on the same operating conditions is adopted: based on the three-dimensional label of "time period type (weekday / holiday) - environmental parameters (temperature, humidity) - equipment operating status (load rate, oil temperature)," the historical database is matched with the time period data most similar to the current operating conditions for filling. This strategy is a mature method for handling short-term data anomalies in the power grid field and has been validated in engineering scenarios such as distribution network measurement data repair and load curve completion.
[0071] S410.4 Fine-tuning of scheduling parameters: Based on the prediction results after incremental training, the output allocation ratio of the power nodes involved in S310.1 is slightly modified. The modification range increases accordingly as the deviation rate increases, and the modified parameters must meet the physical safety limit requirements of the equipment in S320.3. As a further explanation of this step, the fine-tuning of scheduling parameters in S410.4 of this embodiment specifically includes: fine-tuning the scheduling parameters as defined in S310.1, "output of each power node". "With " as the core adjustment target, the correction magnitude and deviation rate For hooks, the principle of "the greater the deviation, the more appropriate the correction range should be, but not exceeding the equipment safety boundary" should be followed. The specific correction rules are as follows: when Time (such as peak hours) =6%), with the correction range set to ≤5%; when When this occurs, the correction range is set to ≤10%; when When this occurs, the correction range is set to ≤15%; Furthermore, the correction process requires real-time verification of the equipment's physical safety limits in S320.3: for example, the original output of thermal power units in a certain area. Deviation rate =8% (peak hours, exceeding the 5% threshold), the adjustment range according to the rules is 4%, proposed to be adjusted to... At this time, it is necessary to simultaneously calculate the transformer load rate corresponding to the unit. The calculation formula is: , This represents the actual apparent power of the adjusted transformer. For the rated apparent power, if (Set by S320.3) ), line current (Set by S320.3) If the modification exceeds the safety limits (e.g.), then the modification is allowed; if the modification results in the modification exceeding the safety limits (e.g.), then the modification is allowed. If so, the correction range will be reduced to 3%, and the calculation will be repeated. And verify again until all device safety constraints are met.
[0072] In addition, fine-tuning prioritizes power nodes with high adjustment flexibility (such as gas turbine units and energy storage power stations), while the correction range for power sources with slow adjustment speed (such as coal-fired units) is controlled within ≤3% to avoid frequent adjustments affecting equipment lifespan, which is consistent with the principle of "prioritizing physical safety of equipment" in S300.
[0073] S410.5 Deviation Reduction Verification: Re-collect the actual load and calculate the deviation rate according to the monitoring frequency set in S410.1. Repeat steps S410.3 to S410.4 until the deviation rate falls back to the reasonable threshold range. Then synchronize the final scheduling parameters and adjustment process data to the time series database of S500.
[0074] As a further explanation of this step, the deviation fallback verification in S410.5 of this embodiment specifically includes: According to the monitoring frequency of "once every 5 minutes" set in S410.1, the adjusted actual load data is re-collected, and the new deviation rate is calculated. The verification logic is as follows: like The price has fallen back to within a reasonable threshold for the corresponding time period (such as peak hours). If the deviation rate is 4.5%, then stop the closed-loop adjustment and record the final scheduling parameters (output, adjustment time, and deviation rate change curves of each power node after correction). like Still exceeds the threshold (e.g., during peak hours) =5.8%), then repeat steps S410.3-S410.4: update the incremental dataset to "data within 1 hour after the last adjustment", retrain the incremental dataset (at this time the model can learn the load response pattern after the previous adjustment), and fine-tune the scheduling parameters again based on the new prediction results until... The requirements are met; Finally, the entire process data, including "adjustment trigger time - incremental dataset range - model update parameters - output after each round of correction - deviation rate change", will be synchronized to the S500 time series database (consistent with the S100 data archiving format) in a structured format of "timestamp - data category - specific value". The scheduling parameter data must be associated with the corresponding power node number and equipment safety verification results to ensure the rationality and safety of adjustments that can be traced during subsequent operation and maintenance review.
[0075] like Figure 5 As shown, S500 data archiving and traceability: key data from S100 to S400 are archived using a time-series database; the data retention period meets the traceability requirements of power grid equipment operation and maintenance, and supports retrieval by time, region, and equipment number to ensure data traceability.
[0076] In this step, the data archiving and traceability process in S500 includes the following steps: S510.1 Time Series Database Configuration and Archive Classification Design: A time series database supporting high-frequency data writing and time series indexing is adopted. The archive categories are divided into stages S100-S400. S100 archives the preprocessed standardized dataset and collection anomaly records; S200 archives the load prediction model parameters, prediction results and accuracy verification data; S300 archives the multi-objective scheduling function parameters, optimization algorithm iteration records and final scheduling scheme; S400 archives the load monitoring data, deviation rate calculation results and closed-loop adjustment records. As a further explanation of this step, the S510.1 time-series database configuration and archiving classification design in this embodiment specifically includes: selecting an open-source time-series database (such as InfluxDB or TimescaleDB) that supports high-frequency data writing and time-series index optimization; adopting a "regional data center + edge node backup" architecture; configuring two master-slave servers in the regional data center (single storage capacity ≥100TB, supporting RAID5 redundancy backup); and linking the edge nodes with the S100 edge storage nodes to temporarily store real-time data within 30 days; setting the core database parameters to 8 write threads, 1 minute time-series index granularity, and LZ4 data compression algorithm. Furthermore, the archiving categories are divided according to the S100-S400 stages, specifically including: The S100 archive contains a standardized dataset with "timestamp, data source, feature name, and standardized value" and a collection anomaly record with "abnormal device number, abnormal time, type, handling measures, and recovery time". The S200 archive contains model parameters including "training batch, iteration round, parameter name, and parameter value", prediction results including "prediction timestamp, predicted value, actual value, MAPE, and RMSE", and accuracy verification data including dataset partitioning range, iteration loss value, and error judgment results. The S300 archive contains scheduling function parameters including "effective time, parameter name, value, and setting basis," and includes "iteration round, optimal population individual, ... The algorithm iteration record of "value and constraint result" and the final scheduling scheme containing "time index, power supply number, planned output and verification details"; The S400 archive contains monitoring data including "monitoring timestamp, zone load, and sensor data", and includes "calculation timestamp, , The deviation rate results of "deviation rate, threshold" and the closed-loop adjustment record containing "trigger time, incremental data range, model update parameters, correction comparison, and fallback results".
[0077] S510.2 Data Retention Period Setting: The retention period shall be determined according to the power grid equipment operation and maintenance traceability requirements. The retention period for associated data of transformers, transmission lines and switching equipment shall not be less than the design life of the corresponding equipment, and the retention period for temporary dispatch adjustment records shall not be less than 3 power grid operation and maintenance assessment cycles, so as to ensure coverage of the equipment's full life cycle traceability and operation and maintenance review requirements. As a further explanation of this step, the specific data retention period setting in S510.2 of this embodiment is as follows: Based on the need for power grid equipment operation and maintenance traceability, the retention period for transformer-related data (oil temperature, load rate, output distribution) is set at 30 years (not less than the maximum design life of 20-30 years), the retention period for transmission line-related data (current, voltage, loss) is set at 40 years (not less than the maximum design life of 30-40 years), and the retention period for switchgear-related data (opening and closing status, operation records, fault alarms) is set at 20 years (not less than the maximum design life of 15-20 years). Temporary scheduling adjustment records (single closed-loop adjustment, short-term deviation handling) are retained for 3 months (covering 3 one-month operation and maintenance assessment cycles), and can be archived to offline storage media after 3 months; the S100 standardized basic dataset and S200 model training basic data are retained for 5 years to meet the needs of medium and long-term load pattern analysis.
[0078] S510.3 Multi-dimensional search function configuration: Build a multi-dimensional search index in the time series database. The time dimension supports time period search accurate to the minute level, the region dimension is associated with the power grid partition code, and the device number dimension is mapped to the unique identification number of the sensing device, transformer and line in S110.1. The search response time does not exceed 10 seconds. As a further explanation of this step, the S510.3 multi-dimensional search function configuration of this embodiment specifically includes: A joint index of time, region, and device number is constructed in the time-series database. The time dimension is indexed at the "year-month-day-hour-minute" level, supporting minute-level time period retrieval in the format "YYYY-MM-DDHH:MM:00 to YYYY-MM-DDHH:MM:00". The region dimension is associated with the power grid partition code (e.g., E01 for the East, W02 for the West), and each archived data is labeled with its region code. The device number dimension is mapped to the unique identifier of S110.1 device according to the "equipment type-voltage level / scenario-serial number" rule (e.g., S-user side-001, T-10kV-001). The retrieval function provides a visual interface, supports single-dimensional or multi-dimensional combined retrieval, and the results are displayed in tables or curves and can be exported to CSV / Excel. The speed is optimized by "quarterly data sharding storage + data in memory preloading within 1 year", ensuring that the retrieval response time does not exceed 10 seconds.
[0079] S510.4, Archived Data Traceability Guarantee: The archived data is periodically checked for integrity, and the consistency between the archived data and the original data generated in the S100-S400 stages is compared. A verification log is generated and archived. If data is found to be missing or inconsistent, an automatic re-archiving mechanism is triggered to retrieve the data from the temporary storage nodes of S100-S400 and re-archive it to ensure the accuracy and integrity of the traceable data.
[0080] As a further explanation of this step, the traceability guarantee of archived data in S510.4 of this embodiment specifically includes: Every day at 2:00 AM (off-peak hours), a full verification of the previous 24 hours is performed, and a monthly full verification is performed at the end of each month. The MD5 hash value is used to compare the original data of the S100-S400 temporary storage nodes with the archived data of the time series database. If they match, the data is considered complete; otherwise, they are marked as abnormal. The automatic archiving mechanism uses the S100-S400 temporary storage nodes (retaining 30 days of data) or the regional backup center as the data source. After detecting anomalies, an archiving task is generated and an alarm is triggered. The archiving program extracts the original data and writes it into the database according to the archiving format. After archiving, the hash value is verified again. If it fails, manual intervention is triggered. Each verification generates a log containing "verification time, data range, result, data volume, handling measures, and handling result", which is archived to a dedicated table in the time series database by "year-month". The retention period is consistent with the corresponding archived data, forming a closed loop of full-link traceability.
[0081] Those skilled in the art will understand that the process of implementing all or part of the steps of the above embodiments can be carried out by hardware or by a program instructing the relevant hardware.
[0082] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for power grid load forecasting and dynamic scheduling optimization based on big data, characterized in that, Includes the following steps: S100, Power Grid Multi-Source Data Preprocessing: Use multi-dimensional sensing devices to collect raw data from the user side, equipment side, and environment side, and process it through a combined process of "missing value classification and imputation + 3σ criterion outlier removal + Z-Score standardization"; S200 and Attention-Enhanced Load Prediction Model Construction: Based on the standardized dataset of S100, time-series load features are extracted through LSTM network, and key influencing factor features are selected through random forest; An attention mechanism is introduced to weight and fuse the two types of features, highlighting the feature weights during peak electricity consumption periods; The model is trained using the Adam optimizer, and the model accuracy is verified by mean absolute percentage error and root mean square error. The load prediction results are then output. S300, generation of improved whale optimization scheduling scheme with hierarchical constraints: Based on the load forecast results of S200 and the real-time operating parameters of the power grid, a multi-objective scheduling function with "supply and demand balance priority + energy consumption optimization as a supplement" is constructed; and an improved whale optimization algorithm is used to solve it. The algorithm introduces an adaptive weight factor to adjust the search strategy, and determines the priority to meet the physical safety limits of the equipment through hierarchical constraints to generate a scheduling scheme. S400, Real-time Closed-Loop Adjustment: High-frequency load monitoring is used to calculate the deviation rate; when the deviation exceeds the reasonable range, the deviation is brought back to the reasonable range through a combination of "incremental data training + fine-tuning of scheduling parameters"; S500 Data Archiving and Traceability: Key data from S100 to S400 are archived using a time-series database; the data retention period meets the traceability requirements of power grid equipment operation and maintenance, and supports retrieval by time, region, and equipment number to ensure data traceability.
2. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 1, characterized in that, In step S100, the process of collecting raw data from the user side, device side, and environment side using multi-dimensional sensing devices includes the following steps: S110.1 Selection and Matching of Data Acquisition Equipment: For raw data acquisition on the user side, select multi-dimensional sensing equipment adapted to different power consumption scenarios to collect real-time active power, reactive power, and cumulative power consumption; for raw data acquisition on the equipment side, select sensing equipment matched to the type of transformer, transmission line, and switching equipment to collect transformer oil temperature, line current, line voltage, and switch open / closed status; for raw data acquisition on the environmental side, select sensing equipment adapted to outdoor power grid scenarios to collect air temperature, relative humidity, and wind speed. S110.2 Data integrity monitoring during the acquisition process: During the acquisition process, the data stream transmission status of the sensing device is monitored in real time. When it is detected that a single device has failed to upload data for M consecutive times or the proportion of missing fields in the uploaded data exceeds the preset threshold, it is determined to be an acquisition anomaly and the device self-test program is triggered. S110.3 Emergency handling of abnormal data acquisition: If the self-test procedure still cannot restore normal data acquisition, activate the backup sensor device pre-deployed in the same area to take over the abnormal device and continue data acquisition. At the same time, record the abnormal device number, the time of the abnormality and the type of abnormality. S110.4 Raw Data Identification and Temporary Storage: The raw data collected normally is structurally identified according to "collection object type - device unique number - collection time" and temporarily stored in the nearest edge storage node. During the temporary storage process, data verification codes are used to verify data integrity.
3. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 2, characterized in that, In S100, the combined process of "missing value classification and imputation + 3σ criterion outlier removal + Z-Score standardization" includes the following steps: S120.1 Missing Value Classification and Imputation: The original data temporarily stored in S110.4 is classified into user side, device side and environment side. For missing values in each type of data, if there are consecutive missing values in a short period of time, the linear interpolation method of the adjacent valid data before and after the missing value is used to imput them. If there are consecutive missing values in a long period of time, the mean of the same type of data in the same area in the same period of time is used to imput them. S120.2, 3σ Criterion Outlier Removal: For the data imputed by S120.1, calculate the mean and standard deviation of each data type according to the user side, equipment side and environment side respectively. Values that deviate from the mean by more than 3 times the standard deviation are identified as outliers and removed. S120.3 Z-Score Standardization: After outliers are removed by S120.2, Z-Score standardization is performed on the data according to the data type to eliminate the difference in magnitude between different dimensions of data and make the processed data dimensionless.
4. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 3, characterized in that, In S200, the process of extracting time-series load features through an LSTM network and screening key influencing factor features through a random forest includes the following steps: S210.1, LSTM network feature extraction: Divide the standardized dataset of S100 into continuous input units according to the time series. , For time indexing, input to the LSTM network; the network passes through the input gate. Controlling the inflow of new information, forget gate Clear redundant information, output gate Filter the output information and generate the hidden layer state. The final output is a time-series load feature vector. ; S210.2, Random Forest Feature Screening: Combining environmental features and equipment operation features into a feature set of influencing factors. With load data Input a random forest model and calculate the node splitting gain for each feature. Determine the contribution level, and select features that meet preset conditions to form a feature set of key influencing factors. ; S210.3 Feature Dimension Alignment: Aligning Time-Series Load Feature Vectors Based on Timestamps Key Influencing Factors Feature Set By unifying the time scales and matching the dimensions of the two sets through dimensional expansion or compression, a set of features to be fused is formed. .
5. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 4, characterized in that, In S200, the process of introducing an attention mechanism to weightedly fuse the two types of features includes the following steps: S220.1 Feature Correlation Analysis: Calculate the set of features to be fused Various characteristics and actual load changes correlation coefficient To form a correlation matrix ; S220.2 Attention Weight Allocation: Based on the Relevance Matrix During peak electricity consumption periods Feature assignment weights For off-peak hours Feature assignment weights ,in To form a dynamic weight matrix ; S220.3, Feature-weighted fusion: This involves using matrix operations to combine the dynamic weight matrix... With the feature set to be fused Fusion to generate a comprehensive feature matrix .
6. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 5, characterized in that, In step S200, the process of training the model using the Adam optimizer and verifying the model accuracy through mean absolute percentage error and root mean square error includes the following steps: S230.1, Model Training Configuration: Integrating the Feature Matrix Divided into training set and verification set Set the first moment coefficients of the Adam optimizer Second-order moment coefficients Initialize model parameters ; S230.2, Iterative Training: Using the training set Input the model, and iteratively update the model parameters through the optimizer. After each iteration, the validation set is used. Calculate the loss value Training stops when the loss value does not decrease after a preset number of rounds, and the optimal parameters are saved. ; S230.3, Accuracy Verification: The test set... Input the optimal model and calculate the predicted load. Compared with actual load The mean absolute percentage error and root mean square error are used to output load forecast results that meet the error requirements. .
7. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 6, characterized in that, In S300, the process of constructing a multi-objective scheduling function that prioritizes supply and demand balance and is supplemented by energy consumption optimization includes the following steps: S310.1, Scheduling Parameter Definition: Define the load forecast result output by S200 as the baseline load demand. Real-time power grid operating parameters include the output of each power source node. Transmission line loss coefficient Equipment energy consumption characteristics parameters ; S310.2, Construction of Supply and Demand Balance Target: The core target is the matching degree between the total regional power supply and the baseline load demand. The matching degree is measured by the supply and demand deviation coefficient. Characterization, and based on and Calculate the degree of deviation and set a deviation threshold. ,when The time frame is determined to be in a state of basic supply and demand balance; S310.3, Energy Consumption Optimization Target Construction: Based on the comprehensive energy consumption of the power grid The smallest is the auxiliary target. based on , , and scheduling duration Establish a mapping relationship and quantify the correlation between output power and energy consumption through energy consumption characteristic parameters; S310.4 Multi-objective integration: Introducing priority weights , and The supply and demand balance objective and the energy consumption optimization objective are weighted and integrated to form a multi-objective scheduling function. .
8. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 7, characterized in that, In step S300, the process of solving the problem using the improved whale optimization algorithm and generating a scheduling scheme by determining hierarchical constraints includes the following steps: S320.1 Algorithm Parameter Initialization: Set the population size for the improved whale optimization algorithm. Maximum number of iterations Initialize population individuals ; S320.2 Adaptive Weight Factor Adjustment: Introducing a weight factor that dynamically changes with the iteration process. Early iteration hour To enhance global search capabilities, larger values are selected in later iterations. hour To improve the accuracy of local optimization, smaller values are selected. Adjust the individual search step size; S320.3, Layered Constraint Judgment: The first layer constraint is the physical safety limit of the equipment. This includes the upper limit of transformer load rate. Rated current of the line The second layer of constraints consists of system stability indicators. Including voltage deviation range Frequency deviation range First verify during the solution process. If the condition is not met, discard the individual directly; if it is met, then re-verify. ; S320.4 Optimal Scheduling Scheme Generation: Through multi-objective scheduling functions Calculate the fitness value of each individual in the population, iteratively update and retain the individual with the best fitness, and finally output the result that simultaneously satisfies... , The combination of constrained scheduling parameters forms a scheduling scheme.
9. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 8, characterized in that, In the S400, the real-time closed-loop adjustment process includes the following steps: S410.1 High-frequency load monitoring and deviation rate calculation: According to the multi-dimensional sensor data acquisition dimensions in S110.1, set the load monitoring frequency to no less than 12 times per hour, collect the actual load data of the power grid in real time, and combine it with the load prediction results of the corresponding time period output in S230.3 to calculate the deviation rate by the relative deviation between the actual load and the predicted load. S410.2 Deviation Range Judgment: A reasonable deviation range threshold is preset. This threshold is set according to the control logic of the system operation stability index in S320.
3. If the deviation rate is within the threshold range, the current scheduling scheme generated by S300 is maintained; if the deviation rate exceeds the threshold range, the closed-loop adjustment mechanism is triggered. S410.3 Incremental Data Training: Collect environmental, equipment, and user data collected in S110.1 to form an incremental dataset. Use the incremental dataset to supplement the training of the attention-enhanced load prediction model in S200, update the model parameters to improve short-term prediction accuracy. S410.4 Fine-tuning of scheduling parameters: Based on the prediction results after incremental training, the output allocation ratio of the power nodes involved in S310.1 is slightly modified. The modification range increases accordingly as the deviation rate increases, and the modified parameters must meet the physical safety limit requirements of the equipment in S320.
3. S410.5 Deviation Reduction Verification: Re-collect the actual load and calculate the deviation rate according to the monitoring frequency set in S410.
1. Repeat steps S410.3 to S410.4 until the deviation rate falls back to the reasonable threshold range. Then synchronize the final scheduling parameters and adjustment process data to the time series database of S500.
10. The power grid load forecasting and dynamic scheduling optimization method based on big data according to claim 9, characterized in that, In the S500, the data archiving and traceability process includes the following steps: S510.1 Time Series Database Configuration and Archive Classification Design: A time series database supporting high-frequency data writing and time series indexing is adopted. The archive categories are divided into stages S100-S400. S100 archives the preprocessed standardized dataset and collection anomaly records; S200 archives the load prediction model parameters, prediction results and accuracy verification data; S300 archives the multi-objective scheduling function parameters, optimization algorithm iteration records and final scheduling scheme; S400 archives the load monitoring data, deviation rate calculation results and closed-loop adjustment records. S510.2 Data Retention Period Setting: The retention period shall be determined according to the power grid equipment operation and maintenance traceability requirements. The retention period for associated data of transformers, transmission lines and switching equipment shall not be less than the design life of the corresponding equipment, and the retention period for temporary dispatch adjustment records shall not be less than 3 power grid operation and maintenance assessment cycles, so as to ensure coverage of the equipment's full life cycle traceability and operation and maintenance review requirements. S510.3 Multi-dimensional search function configuration: Build a multi-dimensional search index in the time series database. The time dimension supports time period search accurate to the minute level, the region dimension is associated with the power grid partition code, and the device number dimension is mapped to the unique identification number of the sensing device, transformer and line in S110.
1. The search response time does not exceed 10 seconds. S510.4, Archived Data Traceability Guarantee: The archived data is periodically checked for integrity, and the consistency between the archived data and the original data generated in the S100-S400 stages is compared. A verification log is generated and archived. If data is found to be missing or inconsistent, an automatic re-archiving mechanism is triggered to retrieve the data from the temporary storage nodes of S100-S400 and re-archive it to ensure the accuracy and integrity of the traceable data.
Citation Information
Patent Citations
A power load forecasting method and system based on big data drive
CN118889402B
Power grid load prediction improving method and device based on AI intelligent algorithm
CN119921324A
Power distribution network load prediction and electric quantity balance optimization method and system based on big data
CN118889419A
Intelligent power distribution load prediction and adaptive scheduling method
CN119994909A
Electric vehicle charging load prediction method based on user behavior analysis
CN120280901A
Cited By
Distributed power distribution cabinet group load coordinated regulation and control method and system based on Internet of Things
CN121036058A
Distribution network user side resource optimization scheduling method for multi-source data fusion
CN121436590A
A method for optimizing and scheduling user-side resources in distribution networks based on multi-source data fusion
CN121436590B
Electric vehicle charging load prediction method based on pre-training language model
CN121834756A
Thermal power plant energy consumption data analysis method
CN122292508A