Multi-source data driven source-load prediction optimization scheduling method and device
By constructing a multi-source data-driven source-load prediction and optimization scheduling method, and using a prediction model to generate confidence intervals and combine them with a deep reinforcement learning decision model, the real-time performance and robustness of the power system under multi-regional source-load uncertainty are solved, thereby improving the economic operation and safety stability of the system.
Patent Information
- Application Number
- CN202610360241.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-23
Smart Images

Figure CN122267791A_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of power system operation control and artificial intelligence application technology, and relates to a multi-source data-driven source-load prediction optimization scheduling method, device, and electronic equipment. Specifically, it involves technical means such as constructing a Transformer prediction model based on decision perception loss and constructing a deep reinforcement learning decision model based on near-end policy optimization, to achieve high-precision perception and collaborative control of source-load uncertainties in multiple regions. Background Technology
[0002] With the escalating global energy crisis and the advancement of the "dual carbon" target, building a new power system dominated by new energy sources has become an inevitable trend in energy transformation. The large-scale integration of renewable energy sources such as wind and solar power has resulted in highly random and volatile power supply-side characteristics. Simultaneously, the widespread adoption of flexible loads such as electric vehicles and distributed energy storage has made demand-side load characteristics increasingly complex and volatile. This dual uncertainty on both the power source and load sides poses a severe challenge to the real-time power balance and safe, stable operation of the power system.
[0003] Source-load coordinated dispatch refers to the process of minimizing system operating costs or maximizing renewable energy consumption by coordinating various resources such as power sources, grids, loads, and storage, while meeting the physical constraints of the power system. However, existing dispatch methods face significant bottlenecks when dealing with massive amounts of multi-source heterogeneous data and complex interconnected systems.
[0004] First, at the forecasting level, traditional time series forecasting methods typically aim to minimize forecast error, neglecting the impact of forecast error on the cost of subsequent scheduling decisions. For example, underestimating loads during peak periods can lead to severe load shedding incidents, the cost of which is far greater than forecasting errors during off-peak periods, but traditional loss functions cannot distinguish this difference. Furthermore, multi-regional source-load data exhibits long-range dependencies and strong dynamic fluctuations, making it difficult for existing models to simultaneously achieve deterministic forecasting and robust interval quantization.
[0005] Secondly, at the scheduling level, multi-region interconnected systems face high model complexity and strict physical security constraints. Traditional mathematical programming methods are computationally time-consuming when dealing with nonlinear constraints, making it difficult to meet the timeliness requirements of real-time scheduling; heuristic algorithms are prone to getting trapped in local optima, making it difficult to ensure the robustness of decisions in dynamically changing environments.
[0006] Therefore, there is an urgent need to construct an optimization scheduling framework with time-series awareness capabilities and the ability to effectively cope with strong uncertainties on both the source and load sides. This framework aims to overcome the limitations of traditional decision-making models in long-range time-series feature mining. By using interval prediction results containing time-series information as key state inputs, the model can predict future dynamic evolution trends and risk boundaries, thereby achieving optimized scheduling that balances real-time performance and robustness. Summary of the Invention
[0007] This disclosure provides a multi-source data-driven source-load prediction and optimization scheduling method. The method is based on an optimization scheduling model, which includes a prediction model and a decision model. The method includes: acquiring real-time multi-source data of the power system to be scheduled, and determining the current operating state of the power system based on the real-time multi-source data. The multi-source data consists of multi-dimensional heterogeneous operating data characterizing the source-side, load-side, and system-side features of the power system to be scheduled; performing source-load prediction based on the current operating state using the prediction model, whereby the source-load prediction represents the predicted operating state of the power system to be scheduled at the next time step and the corresponding confidence interval; and generating scheduling actions based on the current operating state, the predicted operating state, and the confidence interval using the decision model. The prediction model and decision model are trained through the following steps: acquiring historical multi-source data of the power system to be dispatched, and constructing multi-dimensional time-series feature samples based on the historical multi-source data; using the prediction model, generating the predicted operating state and corresponding confidence interval at time t+1 based on the features at time t in the multi-dimensional time-series feature samples, constructing a composite loss function composed of a prediction error term and a dispatch risk cost term, and training the prediction model based on the composite loss function, wherein the prediction error term is used to measure the deviation between the predicted operating state at time t+1 and the actual operating state at time t+1 in the multi-dimensional time-series feature samples; the dispatch risk cost term is used to measure the deviation of system reserve caused by the deviation when dispatching based on the predicted operating state, and assigning a higher penalty weight to the deviation that leads to insufficient system reserve than to the deviation that leads to excess reserve; using the operating state represented by the features at time t, the predicted operating state at time t+1, and the corresponding confidence interval as the state input of the decision model, generating dispatch actions through the decision model; applying the dispatch actions to the power system environment constructed based on the operating constraints of the power system to be dispatched, obtaining a reward signal, and updating the policy parameters of the decision model based on the reward signal.
[0008] According to one or more embodiments of this disclosure, the formula for calculating the composite loss function is as follows: ,in, The prediction error term is calculated using the mean square error. This is a scheduling risk cost item; This is a balancing coefficient used to adjust the weights of the two loss terms. In some embodiments, It can be a predetermined value, for example , , or .
[0009] According to one or more embodiments of this disclosure, confidence intervals are generated by a prediction model using quantile regression.
[0010] According to one or more embodiments of this disclosure, the decision model is a reinforcement learning model configured to employ an enforcer-commentator architecture.
[0011] According to one or more embodiments of this disclosure, when updating the policy parameters of a decision model, a proximal policy optimization algorithm is used to update the parameters based on a pruning objective function that limits the update step size.
[0012] According to one or more embodiments of this disclosure, the reward signal includes at least one of the following: clean energy utilization rate reward, system operating cost penalty, energy balance penalty, ramp-up penalty, carbon emission cost penalty, and interconnection aggregation penalty.
[0013] According to one or more embodiments of this disclosure, the dispatching action includes output allocation instructions for each unit in the power system to be dispatched and charging / discharging power instructions for each energy storage device.
[0014] According to one or more embodiments of this disclosure, the operating constraints of the power system environment include: node power flow constraints, unit ramp rate constraints, and energy storage state of charge constraints.
[0015] This disclosure also provides a multi-source data-driven source-load prediction optimization scheduling device, comprising: an input module configured to acquire real-time multi-source data of the power system to be scheduled and determine the current operating state based on the real-time multi-source data, wherein the multi-source data is multi-dimensional heterogeneous operating data used to characterize the source-side, load-side, and system-side features of the power system to be scheduled; a prediction module including a prediction model configured to perform source-load prediction based on the current operating state, wherein the source-load prediction represents the predicted operating state of the power system to be scheduled at the next time step and the corresponding confidence interval; a decision module including a decision model configured to generate a scheduling action based on the current operating state, the predicted operating state, and the confidence interval; an output module configured to output the scheduling action; and a training module configured to acquire historical multi-source data of the power system to be scheduled and construct multi-dimensional time-series feature samples based on the historical multi-source data; and, through the prediction model, generate a scheduling action based on the t-th time-series feature samples. The predicted operating state and corresponding confidence interval at time t+1 are generated using the features of time t. A composite loss function consisting of a prediction error term and a scheduling risk cost term is constructed, and a prediction model is trained based on the composite loss function. The prediction error term measures the deviation between the predicted operating state at time t+1 and the actual operating state at time t+1 in the multi-dimensional time-series feature samples. The scheduling risk cost term measures the deviation of system reserve caused by the deviation when scheduling is based on the predicted operating state, and assigns a higher penalty weight to deviations that lead to insufficient system reserve than deviations that lead to excess reserve. The operating state represented by the features at time t, the predicted operating state at time t+1, and the corresponding confidence interval are used as the state input of the decision model, and the decision model generates scheduling actions. The scheduling actions are applied to the power system environment constructed based on the operating constraints of the power system to be scheduled, and a reward signal is obtained. The policy parameters of the decision model are updated based on the reward signal.
[0016] This disclosure also provides an electronic device, which includes: a memory for storing a computer program; and a processor for executing the computer program; wherein, when the processor executes the computer program, it implements the source-load prediction optimization scheduling method as described above.
[0017] This disclosure presents a prediction-scheduling closed-loop feedback mechanism formed by constructing a decision-aware prediction model (i.e., the prediction model) and a deep reinforcement learning scheduling model (i.e., the decision model). This mechanism backpropagates the cost gradient of scheduling decisions to the prediction layer and utilizes reinforcement learning to proactively adapt to the uncertainty of predictions during interaction with the environment. This allows the scheduling system to fully utilize prior information from the prediction data while flexibly responding to random disturbances in actual operation. This significantly improves the system's economic operation level and safety margin in environments with high penetration rates of new energy sources and strong load fluctuations. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0019] Figure 1 A flowchart of a multi-source data-driven source load prediction optimization scheduling method according to this disclosure is shown.
[0020] Figure 2 A diagram illustrating the overall algorithm architecture according to this disclosure is shown.
[0021] Figure 3 A block diagram of a multi-source data-driven source load prediction optimization scheduling apparatus according to the present disclosure is shown.
[0022] Figure 4 This is a block diagram of the electronic device according to the present disclosure. Detailed Implementation
[0023] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0024] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such order can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0025] It should be noted that the phrase "at least one of several items" in this disclosure refers to three parallel cases: "any one of the several items", "a combination of any number of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.
[0026] Various embodiments of this disclosure will now be described in detail. Numerous specific details are set forth in the following description in order to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure can be practiced without some of these specific details.
[0027] According to one or more embodiments of this disclosure, a multi-source data-driven source-load prediction and optimization scheduling method is implemented based on an optimization scheduling model, which includes a prediction model and a decision model. The method includes an online scheduling phase and a model training phase.
[0028] Figure 1 A flowchart of a multi-source data-driven source load prediction optimization scheduling method according to this disclosure is shown. Figure 2 A diagram illustrating the overall algorithm architecture according to this disclosure is shown.
[0029] Reference Figure 1 and Figure 2 In step S100, real-time multi-source data of the power system to be dispatched is acquired, and the current operating status of the power system to be dispatched is determined based on the real-time multi-source data. The multi-source data is multi-dimensional heterogeneous operating data used to characterize the characteristics of the source side, load side, and system side of the power system to be dispatched. As an example, the multi-source data includes at least one of the following: new energy unit output data and meteorological environment data characterizing the characteristics of the source side; user-side load data characterizing the characteristics of the load side; and unit operating status data, power flow data, energy storage state of charge data, and load response information characterizing the characteristics of the system side.
[0030] In the embodiments of this disclosure, the construction of multi-source data (multi-dimensional time-series features) is implemented in a standardized data preprocessing environment. A data cleaning script converts collected meteorological data (e.g., meteorological environmental data (irradiance, wind speed, temperature)), load data (output of new energy units (wind power, photovoltaic)), and unit operating status data into tensor formats required by the model. Subsequently, a pre-defined feature engineering module automatically performs missing value imputation and normalization, constructing a multi-dimensional feature matrix containing time lag features and periodic factors, which serves as the standard input for subsequent models. For example, the power system to be dispatched includes source, grid, load, and storage resources in multiple regions. Sources include thermal power units, wind power units, and photovoltaic units; the grid includes tie lines and bus nodes; the load includes traditional loads and flexible loads; and the storage includes distributed energy storage and centralized energy storage. The system can collect power, power flow, unit status, energy storage state of charge, and load response information in real time through sensors, metering devices, and regional controllers to construct the current operating status.
[0031] In step S200, source load prediction is performed based on the current operating state using a prediction model. The source load prediction represents the predicted operating state of the power system to be dispatched at the next moment and the corresponding confidence interval.
[0032] According to one or more embodiments of this disclosure, confidence intervals are generated by a prediction model using quantile regression.
[0033] In the embodiments of this disclosure, the probability prediction function is implemented by a prediction module built on a deep learning framework. This module receives the feature matrix output from step S100 and internally runs a Transformer network that incorporates a decision-aware loss function. Utilizing a multi-head attention mechanism, this module can compute the spatiotemporal correlation weights between different variables in parallel. The output not only includes the point prediction value of the power but also the upper and lower bounds of the confidence interval calculated based on quantile regression. These two parts of data are encapsulated into a standard state message and passed to the subsequent scheduling module.
[0034] The prediction model receives the current operating status and time-series features extracted from historical observation windows, and outputs dual results that take into account both deterministic point prediction and robust interval prediction. By extracting the spatiotemporal correlation features on both the source and load sides, the prediction model performs point prediction and uncertainty range quantification of power fluctuations, and encapsulates the point prediction values and upper and lower bounds of the confidence interval into a status message for subsequent scheduling.
[0035] In step S300, a scheduling action is generated based on the current operating state, the predicted operating state, and the confidence interval by a decision model.
[0036] According to one or more embodiments of this disclosure, the dispatching action includes output allocation instructions for each unit in the power system to be dispatched and charging / discharging power instructions for each energy storage device.
[0037] According to one or more embodiments of this disclosure, the operating constraints of the power system environment include node power flow constraints, unit ramp rate constraints, and energy storage state of charge constraints.
[0038] In the embodiments of this disclosure, the decision model perceives future risk boundaries based on the current operating status, point prediction values, and interval prediction results. When the uncertainty in the prediction interval increases, the decision model can automatically adjust its risk preference, such as increasing reserve capacity or enhancing the energy storage regulation weight. The generated scheduling actions may also include flexible load response quantities, thereby forming joint control over the output of distributed power sources in multiple regions, the charging and discharging power of energy storage, and adjustable resources on the load side.
[0039] In some embodiments, the collaborative control strategy can be transformed into physical execution instructions through the interaction and control module and sent to each area controller. At the same time, economic indicators, safety margins and clean disposal status can be displayed through a visual interface to support online operation monitoring and manual review.
[0040] Reference Figure 1The multi-source data-driven source load prediction optimization scheduling method according to embodiments of this disclosure further includes a step S400 of training a prediction model and a decision model. In such an embodiment, step S400 can be performed first to train the prediction model and the decision model, and then steps S100-S300 described above can be performed.
[0041] In step S400, historical multi-source data of the power system to be dispatched can be obtained, and multi-dimensional time-series feature samples can be constructed based on the historical multi-source data.
[0042] In the embodiments of this disclosure, historical multi-source data includes historical power output of new energy units (wind power, photovoltaic), meteorological environment data (irradiance, wind speed, temperature), historical load data on the user side, and may further include electricity price data (also known as price data) and unit operating status data. To eliminate differences between data of different dimensions (such as power in MW and temperature in °C) and accelerate neural network convergence, the original data can be cleaned, missing value imputation, and max-min normalization processing can be performed. The normalization formula can be expressed as:
[0043] in Indicates the first i Class features in t The original observations at time [time]. and These represent the maximum and minimum values of the feature within the historical time window, respectively. These are the normalized eigenvalues. Subsequently, based on the normalized data, a multidimensional time-series feature matrix is constructed.
[0044] In the embodiments of this disclosure, a sliding window technique is used to divide the continuous time series into fixed-length sample pairs and construct an input feature matrix that includes time lag features and periodic factors:
[0045] in, The length of the historical observation window. This represents the total number of feature dimensions for wind, solar, load, and meteorological factors. This step outputs a series of standardized tensors, which serve as the standard input for subsequent training of the prediction model.
[0046] In embodiments of this disclosure, reference is made to Figure 2 This shows the power system environment, where historical source-load data, wind speed, solar radiation, and temperature information can be used together as historical operating data sources and further constitute multi-dimensional time-series feature samples. Figure 2 The energy supply, power generation and storage equipment, and load-side objects shown can be used together as a concrete example of a power system environment to be dispatched.
[0047] Return to reference Figure 1 In step S400, the predicted operating state and corresponding confidence interval at time t+1 can be generated by the prediction model based on the features at time t in the multi-dimensional time series feature samples. A composite loss function composed of a prediction error term and a scheduling risk cost term is constructed, and the prediction model is trained based on the composite loss function. The prediction error term is used to measure the deviation between the predicted operating state at time t+1 and the actual operating state at time t+1 in the multi-dimensional time series feature samples. The scheduling risk cost term is used to measure the deviation of the system reserve caused by the deviation when scheduling is based on the predicted operating state, and the deviation that leads to insufficient system reserve is given a higher penalty weight than the deviation that leads to excess reserve.
[0048] According to one or more embodiments of this disclosure, the formula for calculating the composite loss function is as follows: ,in, The prediction error term is calculated using the mean square error. This is a scheduling risk cost item; This is a balancing coefficient used to adjust the weights of the two loss terms. In some embodiments, It can be a predetermined value, for example , , or This is to balance the contribution of prediction accuracy with the cost of scheduling risks.
[0049] In embodiments of this disclosure, reference is made to Figure 2 The diagram illustrates the training method, and the prediction model can also be called the source-load prediction module. Multidimensional temporal feature samples can include historical sequences and prior sequences. Input embedding, position encoding, Transformer encoder, and Transformer decoder together constitute the network structure of the prediction model. The prediction model outputs the predicted running state and the corresponding confidence interval (also known as uncertainty estimation). The composite loss function is also called decision-aware loss (including prediction error term and scheduling risk cost term).
[0050] In the embodiments of this disclosure, the prediction model may employ a Transformer architecture, with the network primarily consisting of a positional encoding layer, a multi-head self-attention layer, and a feedforward neural network layer. For the input multi-dimensional matrix feature samples, positional encoding vectors can be superimposed to preserve the temporal features of the time series; the multi-head self-attention layer maps the input to a query matrix. Key matrix Sum matrix It calculates the interdependencies between variables and uses a multi-head mechanism to capture the spatiotemporal correlation weights of multiple variables in parallel.
[0051]
[0052] in, Let be the dimension of the key vector. is a scaling factor used to prevent the gradient from vanishing due to excessively large dot product results. To capture multivariate coupling features in different subspaces (such as the strong correlation between photovoltaic output and irradiance, or the nonlinear relationship between load and temperature), the model employs a multi-head mechanism for parallel computation. The outputs of h heads are concatenated and then linearly transformed to obtain the fused feature representation:
[0053] in, in, Let be the projection weight matrix of the i-th head. This is the output transformation matrix.
[0054] In the embodiments of this disclosure, to achieve decision-aware training, the decision costs of subsequent scheduling tasks can be fed back into the prediction training process, enabling the prediction model to not only pursue numerical fitting accuracy but also optimize subsequent scheduling performance. The total loss function can be defined as:
[0055]
[0056] in, For the true value, This is a predicted value. Scheduling risk cost item. An asymmetric penalty mechanism is employed, imposing a higher penalty on prediction errors leading to insufficient system reserves (high risk) than on those leading to excess reserves (low risk), thereby guiding the model to output predictions that are more conducive to system safety. The trained prediction model can also incorporate quantile regression to output confidence intervals, providing robust boundary inputs for subsequent scheduling.
[0057] in, At the significance level, the coverage rate is... The confidence interval represents the future actual source load output observation. The probability falls within the upper and lower bounds of the interval, so as to reserve a probabilistically guaranteed adjustment reserve for subsequent training and scheduling.
[0058] in, It can be represented as:
[0059] in, To meet the actual net load requirements, Forecast net load requirements. When When the predicted value is too low, resulting in insufficient system reserves, a penalty weight is applied. ;when When the predicted value is too high, it indicates that the system has excess reserves, and a penalty weight is applied. .satisfy or at least .
[0060] In step S400, the operating state characterized by features at time t, the predicted operating state at time t+1, and the corresponding confidence interval can also be used as the state input of the decision model to generate scheduling actions. The scheduling actions are then applied to the power system environment constructed based on the operating constraints of the power system to be scheduled to obtain reward signals, and the strategy parameters of the decision model are updated based on the reward signals.
[0061] According to one or more embodiments of this disclosure, the decision model is a reinforcement learning model configured to employ an enforcer-commentator architecture.
[0062] According to one or more embodiments of this disclosure, when updating the policy parameters of a decision model, a proximal policy optimization algorithm is used to update the parameters based on a pruning objective function that limits the update step size.
[0063] According to one or more embodiments of this disclosure, the reward signal includes at least one of the following: clean energy utilization rate reward, system operating cost penalty, energy balance penalty, ramp-up penalty, carbon emission cost penalty, and interconnection aggregation penalty.
[0064] In embodiments of this disclosure, reference is made to Figure 2 The decision model, also known as the scheduling decision module, has a policy update structure that may include a temporary replay buffer, a PPO truncation loss, and agents (including actors (also known as the actor network) and critics (also known as the value network)). The decision model employs a proximal policy optimization (PPO) model with an actor-critic architecture. (State vector) It includes not only real-time power flow data of the power grid, but also incorporates predicted operating status (also known as point prediction values). confidence interval (also known as uncertainty interval) This enables intelligent agents to explicitly perceive potential risks such as future power exceeding limits and insufficient reserves, and to dynamically adjust policy preferences as the prediction interval increases, thereby providing an information basis for generating robust scheduling policies.
[0065] In embodiments of this disclosure, the state vector It not only includes real-time power flow data of the power grid (node voltage, line load), but more importantly, it also integrates the dual prediction results output by the prediction model (i.e., point prediction values). With uncertainty range By explicitly incorporating the prediction interval into the state space, the agent can perceive potential risks such as power overruns or insufficient reserves in the system at future moments, thus providing an information basis for generating robust policies.
[0066] Actor Network (Policy Network) According to the status Output action distribution and sample to generate scheduling instructions. (Including the output of each unit and the charging and discharging power of energy storage). This disclosure uses the core of the PPO algorithm—the ClippedSurrogate Objective—to update the network parameters. This ensures the monotonicity and stability of policy updates. Policy updates can utilize the pruning objective function of the PPO:
[0067] in, The advantage function is used to evaluate the advantage of the current action relative to the average policy. The probability ratio between the old and new strategies. The hyperparameters are used to limit the policy update step size and suppress training oscillations that could cause the scheduling policy to fail.
[0068] In embodiments of this disclosure, the Critic network is responsible for calculating the state value and processing the reward signal from the environment. The reward signal is passed to the Actor network to guide policy optimization. The reward signal may consist of at least one of the foregoing terms; in one specific embodiment, the reward function may be expressed as a weighted summation as follows:
[0069] in, These are the weighting coefficients for each item. The specific definitions of each item are as follows: (Clean Energy Utilization Rate): As a positive reward, it represents the proportion of wind and solar power absorbed by the system in the current period to the total load, aiming to guide the model to maximize the absorption of renewable energy. (Operating Cost): The total operating cost of the system is calculated based on the real-time output of each unit and the corresponding electricity price (such as the fuel cost of thermal power units or the real-time clearing price of the electricity market). This cost serves as a negative reward to guide the model in finding the most economically optimal scheduling strategy. (Energy Balance Penalty): Represents the real-time power imbalance (EBS) between the source and load sides. When the total power generation cannot meet the load demand or exceeds the demand, a high penalty is imposed to ensure the physical supply and demand balance of the system through strong constraints. (Ramp Constraint Penalty): A penalty that is triggered when the unit's output adjustment exceeds its physical ramp rate limit. (Carbon emissions): The total carbon emissions calculated based on the real-time output of various units and the corresponding carbon emission factors are used as a negative reward to achieve low-carbon dispatch. (Tether Aggregation Penalty): For multi-region collaborative scenarios, calculate the deviation between the sum of the power exchanged by the tie lines between each region and the planned value. This item is used to constrain the power interaction behavior between multiple regions and ensure the safe and stable operation of the tie lines between regions.
[0070] In some embodiments, such as Figure 2 As shown, historical source-load data, wind speed, solar radiation, temperature, as well as energy supply, power generation and storage equipment, and load-side objects can all participate in the construction of the reinforcement learning environment. To address the model complexity of multi-region interconnected systems, distributed cooperative scheduling algorithms can solve the global optimization problem through decentralized computation, and the resulting strategies can be evaluated based on theoretical proofs or verification results of system safety operation.
[0071] Figure 3 A block diagram of a multi-source data-driven source load prediction optimization scheduling apparatus according to the present disclosure is shown.
[0072] Reference Figure 3A block diagram of a multi-source data-driven source-load prediction optimization scheduling device 300 is shown. The optimization scheduling device 300 includes an input module 310, a prediction module 320, a decision module 330, an output module 340, and a training module 350. The input module 310 is used to acquire real-time multi-source data of the power system to be scheduled and determine the current operating state of the power system to be scheduled. The multi-source data is multi-dimensional heterogeneous operating data used to characterize the source-side, load-side, and system-side features of the power system to be scheduled. The prediction module 320 includes a prediction model and is configured to perform source-load prediction based on the current operating state. The source-load prediction represents the predicted operating state and the corresponding confidence interval at the next moment. The decision module 330 includes a decision model and is configured to generate scheduling actions based on the current operating state, the predicted operating state, and the confidence interval. The output module 340 is used to output the scheduling actions. The training module 350 is configured to: acquire historical multi-source data of the power system to be dispatched and construct multi-dimensional time-series feature samples; generate the predicted operating state and corresponding confidence interval at time t+1 based on the features at time t in the multi-dimensional time-series feature samples through the prediction model, and complete the training using a composite loss function composed of prediction error term and dispatch risk cost term; the decision model can use the operating state at time t, the predicted operating state at time t+1 and the corresponding confidence interval as state input to generate dispatch actions, apply the dispatch actions to the power system environment to obtain reward signals, and update the policy parameters of the decision model based on the reward signals.
[0073] In embodiments of this disclosure, the input module 310 can be used to perform a reference. Figure 1 The step S100 shown; the prediction module 320 can be used to perform the reference Figure 1 The step S200 shown; the decision module 330 can be used to execute the reference. Figure 1 The step S300 is shown; the output module 340 can be used to output scheduling actions, and the training module 350 can be used to execute the reference. Figure 1 Step S400 is shown. Redundant descriptions are omitted here.
[0074] Figure 4 This is a block diagram of the electronic device according to the present disclosure.
[0075] Reference Figure 4An electronic device 400 according to embodiments of the present disclosure may include a processor 410 and a memory 420. The processor 410 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field-programmable gate array (FPGA), a system-on-a-chip (SoC), a microprocessor, an application-specific integrated circuit (ASIC), etc. The memory 420 may store computer programs to be executed by the processor 410. The memory 420 includes high-speed random access memory and / or non-volatile computer-readable storage media. When the processor 410 executes the computer program stored in the memory 420, the multi-source data-driven source-load prediction optimization scheduling method described above can be implemented.
[0076] Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store computer programs and any associated data, data files, and data structures in a non-transitory manner and to provide the computer programs and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer programs. In one example, the computer programs and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer programs and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0077] While some embodiments of this disclosure have been shown and described, those skilled in the art will understand that modifications may be made to these embodiments without departing from the principles and spirit of this disclosure, which are defined by the claims and their equivalents.
Claims
1. A multi-source data-driven source-load prediction and optimization scheduling method, characterized in that, The source-load prediction and optimization scheduling method is based on an optimization scheduling model, which includes a prediction model and a decision model. The source load prediction and optimization scheduling method includes: The system acquires real-time multi-source data of the power system to be dispatched and determines the current operating status of the power system to be dispatched based on the real-time multi-source data. The multi-source data is multi-dimensional heterogeneous operating data used to characterize the source side, load side and system side of the power system to be dispatched. The source-load prediction is performed by the prediction model based on the current operating state. The source-load prediction represents the predicted operating state of the power system to be dispatched at the next moment and the corresponding confidence interval. The decision model generates scheduling actions based on the current operating state, the predicted operating state, and the confidence interval. The prediction model and the decision model are trained through the following steps: Historical multi-source data of the power system to be dispatched is obtained, and multi-dimensional time-series feature samples are constructed based on the historical multi-source data; The prediction model generates the predicted operating state and corresponding confidence interval at time t+1 based on the features at time t in the multi-dimensional time-series feature samples. A composite loss function composed of a prediction error term and a scheduling risk cost term is constructed, and the prediction model is trained based on the composite loss function. The prediction error term measures the deviation between the predicted operating state at time t+1 and the actual operating state at time t+1 in the multi-dimensional time-series feature samples. The scheduling risk cost term measures the deviation in system reserve caused by the deviation when scheduling is performed based on the predicted operating state, and assigns a higher penalty weight to deviations that lead to insufficient system reserve than to deviations that lead to excess reserve. The operating state characterized by the features at time t, the predicted operating state at time t+1, and the corresponding confidence interval are used as the state input of the decision model. The decision model generates scheduling actions. The scheduling actions are applied to the power system environment constructed based on the operating constraints of the power system to be scheduled to obtain a reward signal. The strategy parameters of the decision model are updated based on the reward signal.
2. The source-load prediction and optimization scheduling method according to claim 1, characterized in that, The formula for calculating the composite loss function is as follows: in, The prediction error term is calculated using the mean square error. This refers to the scheduling risk cost item; This is the balance coefficient.
3. The source-load prediction and optimization scheduling method according to claim 1, characterized in that, The confidence interval is generated by the prediction model using quantile regression.
4. The source-load prediction and optimization scheduling method according to claim 1, characterized in that, The decision-making model is a reinforcement learning model, configured to employ an executor-critic architecture.
5. The source-load prediction and optimization scheduling method according to claim 4, characterized in that, When updating the policy parameters of the decision model, a near-end policy optimization algorithm is used to update the parameters based on a pruning objective function that limits the update step size.
6. The source-load prediction and optimization scheduling method according to claim 1, characterized in that, The reward signals include at least one of the following: clean energy utilization rate reward, system operating cost penalty, energy balance penalty, ramping and exceeding limit penalty, carbon emission cost penalty, and tie-line aggregation penalty.
7. The source-load prediction and optimization scheduling method according to claim 1, characterized in that, The dispatching actions include output allocation instructions for each generating unit within the power system to be dispatched, as well as charging and discharging power instructions for each energy storage device.
8. The source-load prediction and optimization scheduling method according to claim 1, characterized in that, The operational constraints of the power system environment include: node power flow constraints, unit ramp rate constraints, and energy storage state of charge constraints.
9. A multi-source data-driven source load prediction and optimization scheduling device, characterized in that, The source load prediction and optimization scheduling device includes: The input module is configured to acquire real-time multi-source data of the power system to be dispatched, and determine the current operating status based on the real-time multi-source data. The multi-source data is multi-dimensional heterogeneous operating data used to characterize the source side, load side and system side of the power system to be dispatched. The prediction module, including a prediction model, is configured to perform source load prediction based on the current operating state, wherein the source load prediction represents the predicted operating state of the power system to be dispatched at the next moment and the corresponding confidence interval. The decision module, including a decision model, is configured to generate scheduling actions based on the current operating state, the predicted operating state, and the confidence interval. The output module is configured to output the scheduling action; and The training module is configured to acquire historical multi-source data of the power system to be dispatched, and construct multi-dimensional time-series feature samples based on the historical multi-source data; The prediction model generates the predicted operating state and corresponding confidence interval at time t+1 based on the features at time t in the multi-dimensional time-series feature samples. A composite loss function composed of a prediction error term and a scheduling risk cost term is constructed, and the prediction model is trained based on the composite loss function. The prediction error term measures the deviation between the predicted operating state at time t+1 and the actual operating state at time t+1 in the multi-dimensional time-series feature samples. The scheduling risk cost term measures the deviation in system reserve caused by the deviation when scheduling is performed based on the predicted operating state, and assigns a higher penalty weight to deviations that lead to insufficient system reserve than to deviations that lead to excess reserve. The operating state characterized by the features at time t, the predicted operating state at time t+1, and the corresponding confidence interval are used as the state input of the decision model. The decision model generates scheduling actions. The scheduling actions are applied to the power system environment constructed based on the operating constraints of the power system to be scheduled to obtain a reward signal. The strategy parameters of the decision model are updated based on the reward signal.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program; wherein, when the processor executes the computer program, it implements the source-load prediction optimization scheduling method as described in any one of claims 1-8.