Energy Storage Scheduling Model Training for Precise PPO Convergence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional scheduling models for energy storage systems face challenges in converging with high precision due to complex operations, including interference from charging-discharging actions, electricity prices, and state of charge variations, leading to low precision and speed in convergence.

Innovation Solution

A method and apparatus for training a scheduling model that determines main reward data based on payment data and branch penalty data from deviation data of charging-discharging actions, using proximal policy optimization and policy gradient and temporal difference learning to update parameters, ensuring the model converges with high precision quickly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If proximal policy optimization is used to construct the scheduling model, then global search capability is achieved, but convergence precision is insufficient due to complex operations

Engineering Contradiction:
Improveglobal search capabilityVSAvoidconvergence precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The reward function is segmented into multiple independent components: main reward data based on payment data, and branch penalty data based on deviation data of charging-discharging actions. This segmentation allows each component to be optimized independently while maintaining global search capability, thereby improving convergence precision without sacrificing adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple parameter dimensions including payment data parameters, deviation data parameters, and state variable parameters. By changing and optimizing these parameters separately through multi-objective optimization, the model achieves both global search capability and high convergence precision in complex operating conditions.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the scheduling model considers multiple factors including charging-discharging actions, electricity prices, and state of charge variations, then adaptability to complex operations is improved, but convergence speed and precision deteriorate

Engineering Contradiction:
Improveadaptability to complex operationsVSAvoidconvergence precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the complex operation factors into distinct components: charging-discharging actions are evaluated through branch penalty data, electricity prices through main reward data, and state of charge variations through state variable data. This segmentation enables the model to maintain high adaptability to complex operations while achieving precise convergence by optimizing each segment independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different evaluation criteria are applied to different operational aspects: payment data determines main reward quality, deviation data determines branch penalty quality, and state variable changes determine action validity. This local quality approach allows the model to adapt to various complex operations with appropriate precision for each specific operation type.

Inventive Principle:
Principle #3Local quality

3Ease of manufacture

If conventional reward functions are used, then implementation simplicity is maintained, but convergence precision is insufficient for complex energy storage operations

Engineering Contradiction:
Improveimplementation simplicityVSAvoidconvergence precision
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The reward function is segmented into computationally independent main reward and branch penalty components. Each segment processes specific data types (payment data for main reward, deviation data for branch penalty), maintaining implementation simplicity through modular design while achieving high convergence precision through comprehensive multi-dimensional optimization.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4439411A1Method and apparatus for training scheduling model, electric device and storage medium
Publication Date: 2024.10.02 SUNGROW (SHANGHAI) CO LTD
  • EP4439411A1 patent drawingFigure 1a~1b
  • EP4439411A1 patent drawingFigure 2~3
  • EP4439411A1 patent drawingFigure 4

AI summary

A method and an apparatus for training a scheduling model of an energy storage system, an electronic device and a storage medium are provided. The method includes: determining main reward data of the scheduling model based on payment data of the energy storage system for electricity in a period from a first moment to a second moment, where the first moment is earlier than the second moment; determining branch penalty data of the scheduling model based on deviation data of a charging-discharging action of the energy storage system in the period, where the deviation data of the charging-discharging action indicates to what extent the charging-discharging action matches a change in a state variable of the energy storage system; and updating parameters of the scheduling model based on the main reward data and the branch penalty data, to obtain a target scheduling model. The scheduling model trained in this way can converge with high precision.