Energy Storage Scheduling Model Training for Precise PPO Convergence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional scheduling models for energy storage systems face challenges in converging with high precision due to complex operations, including interference from charging-discharging actions, electricity prices, and state of charge variations, leading to low precision and speed in convergence.
Innovation Solution
A method and apparatus for training a scheduling model that determines main reward data based on payment data and branch penalty data from deviation data of charging-discharging actions, using proximal policy optimization and policy gradient and temporal difference learning to update parameters, ensuring the model converges with high precision quickly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If proximal policy optimization is used to construct the scheduling model, then global search capability is achieved, but convergence precision is insufficient due to complex operations
Solution Approach 1:
The reward function is segmented into multiple independent components: main reward data based on payment data, and branch penalty data based on deviation data of charging-discharging actions. This segmentation allows each component to be optimized independently while maintaining global search capability, thereby improving convergence precision without sacrificing adaptability.
Solution Approach 2:
The patent introduces multiple parameter dimensions including payment data parameters, deviation data parameters, and state variable parameters. By changing and optimizing these parameters separately through multi-objective optimization, the model achieves both global search capability and high convergence precision in complex operating conditions.
2Adaptability or versatility
If the scheduling model considers multiple factors including charging-discharging actions, electricity prices, and state of charge variations, then adaptability to complex operations is improved, but convergence speed and precision deteriorate
Solution Approach 1:
The patent segments the complex operation factors into distinct components: charging-discharging actions are evaluated through branch penalty data, electricity prices through main reward data, and state of charge variations through state variable data. This segmentation enables the model to maintain high adaptability to complex operations while achieving precise convergence by optimizing each segment independently.
Solution Approach 2:
Different evaluation criteria are applied to different operational aspects: payment data determines main reward quality, deviation data determines branch penalty quality, and state variable changes determine action validity. This local quality approach allows the model to adapt to various complex operations with appropriate precision for each specific operation type.
3Ease of manufacture
If conventional reward functions are used, then implementation simplicity is maintained, but convergence precision is insufficient for complex energy storage operations
Solution Approach 1:
The reward function is segmented into computationally independent main reward and branch penalty components. Each segment processes specific data types (payment data for main reward, deviation data for branch penalty), maintaining implementation simplicity through modular design while achieving high convergence precision through comprehensive multi-dimensional optimization.
Data Source
Figure 1a~1b
Figure 2~3
Figure 4
AI summary
A method and an apparatus for training a scheduling model of an energy storage system, an electronic device and a storage medium are provided. The method includes: determining main reward data of the scheduling model based on payment data of the energy storage system for electricity in a period from a first moment to a second moment, where the first moment is earlier than the second moment; determining branch penalty data of the scheduling model based on deviation data of a charging-discharging action of the energy storage system in the period, where the deviation data of the charging-discharging action indicates to what extent the charging-discharging action matches a change in a state variable of the energy storage system; and updating parameters of the scheduling model based on the main reward data and the branch penalty data, to obtain a target scheduling model. The scheduling model trained in this way can converge with high precision.