Artificial intelligence-based energy storage system load prediction and optimization management system
By constructing a multi-timescale prediction engine and a hierarchical reinforcement learning decision network, the problem of lacking long-term policy and market trend perception in the optimization management of energy storage systems is solved. This enables forward-looking insights and dynamic adaptive optimization of energy storage systems, improving the return on investment and the continuous optimality of strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN FLUORIDE NEW ENERGY TECH CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-19
AI Technical Summary
The existing energy storage system optimization management lacks the ability to dynamically perceive and integrate long-term policies and market trends, resulting in the failure of medium- and long-term optimization strategies and low return on investment.
We will construct an AI-based energy storage system load forecasting and optimization management system, including a data fusion perception layer, a multi-timescale forecasting engine, a dynamic strategy generator, and an adaptive execution and feedback loop. We will use deep learning and reinforcement learning technologies to achieve prediction and dynamic adaptive optimization of long-term policies and market trends.
It improves the scientific nature of investment decisions and return on assets for energy storage systems throughout their entire life cycle, enhances the system's robustness and adaptability in the face of market fluctuations and changes in internal conditions, and ensures the continued optimality and vitality of the strategy.
Smart Images

Figure CN122246688A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine learning and energy storage system control technology, specifically relating to an artificial intelligence-based energy storage system load prediction and optimization management system. Background Technology
[0002] Energy storage systems, as a key technology for balancing electricity supply and demand, improving grid stability, and promoting the consumption of renewable energy, are playing an increasingly important role in the fields of smart grids and the energy internet. The economics and operational efficiency of energy storage systems are highly dependent on the optimization of their charging and discharging strategies, which involves accurate prediction and coordinated decision-making on multiple factors such as load, electricity prices, and renewable energy output.
[0003] Load forecasting and optimization management of energy storage systems is a core technological direction for improving overall efficiency. This technology aims to predict future power load changes by analyzing historical and real-time data, and formulate optimal charging and discharging plans accordingly to achieve goals such as peak shaving and valley filling, reducing electricity costs, or participating in the ancillary services market.
[0004] Existing technologies typically rely on historical load data and current market electricity price information to construct static or short-term rule-based optimization models. However, these methods suffer from several problems: their optimization decisions are highly dependent on pre-set and relatively fixed electricity price curves and market rules, lacking the ability to perceive and integrate external dynamic factors such as the macroeconomic policy environment and long-term market trends. Because they cannot proactively anticipate and adapt to the long-term risks and opportunities brought about by future energy policy adjustments and market mechanism reforms, the optimization strategies formulated by the system may fail in the medium to long term, leading to lower-than-expected returns on investment or missed market opportunities. Therefore, how to construct an energy storage system management solution that can integrate long-term policy and market trend predictions and achieve dynamic adaptive optimization has become an urgent technical challenge. Summary of the Invention
[0005] The purpose of this invention is to provide an artificial intelligence-based load forecasting and optimization management system for energy storage systems, in order to solve the problem that the optimization management of energy storage systems in the prior art lacks the ability to dynamically perceive and integrate long-term policies and market trends, resulting in the failure of medium- and long-term optimization strategies and a decrease in the rate of return on investment.
[0006] This invention provides an artificial intelligence-based load forecasting and optimization management system for energy storage systems, comprising: The data fusion perception layer is used to collect and structure multi-source heterogeneous data in real time. The data fusion perception layer includes a load and operation data collection module, a market and policy data collection module, and an external environment data collection module. A multi-timescale prediction engine is used to run a long-term trend prediction model, a medium-term situation prediction model, and a short-term accurate prediction model in parallel based on the structured data provided by the data fusion perception layer. A dynamic policy generator is used to receive all prediction results output by the multi-timescale prediction engine and generate collaborative optimization policies covering different timescales based on a hierarchical reinforcement learning decision network. An adaptive execution and feedback loop is used to send the final control command output by the dynamic strategy generator to the power conversion system of the energy storage system for execution, and to build a closed-loop learning mechanism.
[0007] Preferably, the load and operation data acquisition module is used to acquire historical load curves, real-time load data, state of charge, charging and discharging power, efficiency curves, and equipment health status parameters of the energy storage system; The market and policy data acquisition module is used to collect real-time electricity price data, day-ahead market and real-time market clearing prices, ancillary services market rule texts, energy policy documents, carbon emission quota prices, and renewable energy quota system requirements. The external environment data acquisition module is used to collect meteorological forecast data and macroeconomic prosperity index.
[0008] Preferably, the long-term trend prediction model uses an annual forecast period, with inputs including a compilation of historical policy documents, carbon emission price series, macroeconomic data, and technology cost decline curves. It employs a bidirectional long short-term memory network architecture based on an attention mechanism to output future base load growth trend predictions, probability distributions of time-of-use electricity price structure evolution, and potential policy incentive windows. The medium-term trend prediction model uses a monthly forecast period. Its inputs are historical load data, market electricity price data, meteorological data, and trend guidance signals output by the long-term trend prediction model. It adopts a hybrid architecture of temporal convolutional network and gated recurrent unit to output the forecast of future monthly typical daily load curve, monthly average electricity price forecast, and major ancillary service demand forecast. The short-term accurate forecasting model uses hourly forecasting periods. Its inputs include refined load data, real-time electricity prices, ultra-short-term weather forecasts, and the baseline curve output by the medium-term situation forecasting model. It employs a gradient boosting decision tree algorithm to continuously forecast future load values and electricity price fluctuation ranges.
[0009] Preferably, the hierarchical reinforcement learning decision network includes a meta-policy layer and a sub-policy layer; The meta-strategy layer aims to maximize long-term trend prediction results and long-term system returns. Its state space is defined as the intensity of long-term policy guidance, the stage of market mechanism evolution, and the depreciation status of system assets. Its action space consists of weight allocation instructions for medium- and short-term optimization objectives. It is trained through a deep deterministic strategy gradient algorithm and outputs strategic-level decision parameters. The sub-strategy layer includes a mid-term scheduling sub-strategy and a real-time control sub-strategy; the mid-term scheduling sub-strategy operates under the target weights set by the meta-strategy layer, with the mid-term situation prediction results and the expected returns for the next week as its objectives. The state space consists of the predicted load curve, predicted electricity price curve, and energy storage system status for the next 7 days, while the action space consists of the day-ahead market bidding plan and energy storage system charge and discharge plan framework for the next 24 hours to 7 days. The real-time control sub-strategy operates within the planning framework provided by the medium-term scheduling sub-strategy, with the goal of maximizing short-term accurate prediction results and real-time revenue. Its state space is the load and electricity price prediction for the next 15 minutes to 2 hours and the real-time charge status of the energy storage system. Its action space is the specific charging and discharging power command for the next time interval. The dynamic strategy generator also includes a strategy coordinator, which is used to monitor the deviation between short-term execution results and medium-term plans in real time, and trigger the online replanning of the medium-term scheduling sub-strategy when the deviation exceeds a preset tolerance threshold, and feed back the replanning information to the meta-strategy layer.
[0010] Preferably, the adaptive execution and feedback loop includes an instruction compilation and distribution module, an execution status monitoring module, and a multi-dimensional feedback learning module; The instruction compilation and distribution module is used to compile the charging and discharging power instructions output by the real-time control sub-strategy into specific control commands that conform to the energy storage system communication protocol and then distribute them. The execution status monitoring module is used to collect the actual execution results of the instructions in real time and compare them with the expected values of the instructions to generate an execution deviation report; The multi-dimensional feedback learning module receives the execution deviation report, actual market settlement data, and policy change information. It uses actual revenue, policy fit, and equipment loss as multi-dimensional reward signals, and execution deviation and strategy failure events as penalty signals to form a composite reward function. The composite reward function is then synchronously fed back to the experience replay buffer of each strategy layer in the hierarchical reinforcement learning decision network to drive the network to perform periodic offline retraining and online parameter fine-tuning.
[0011] Preferably, the policy text analysis process in the long-term trend prediction model is as follows: The policy documents are segmented and vectorized using a pre-trained language model, transforming each policy sentence into a high-dimensional semantic vector; the vectors are then categorized into several predefined policy dimensions using a policy topic classifier. For each policy dimension, a time series regression model is used to analyze the trend of policy intensity over time, and the slope and curvature of the trend are extracted as features. The policy characteristics and the economic time-series data of the same period are input into a bidirectional long short-term memory network for joint prediction. The predefined policy dimensions include subsidy intensity, carbon emission constraints, market access, and technical standards.
[0012] Preferably, the strategic decision-making parameters output by the meta-strategy layer specifically include the weight of medium-term economic benefit targets, the weight of long-term policy risk aversion, and the weight of equipment life maintenance. The reward function of the intermediate scheduling sub-strategy is the weighted sum of each weight and the corresponding sub-reward set by the meta-strategy layer; In addition to immediate economic gains, the reward function of the real-time control sub-strategy also includes a penalty term for deviation from the medium-term plan. The coefficient of this penalty term is dynamically adjusted by the strategy coordinator based on the current deviation.
[0013] Preferably, the working logic of the policy coordinator is as follows: The cumulative deviation between the short-term actual execution trajectory and the medium-term scheduling plan trajectory at key time points is continuously calculated. These key time points include peak electricity price periods, peak load periods, and planned charging / discharging state switching points. When the cumulative deviation at any critical time point exceeds the preset dynamic tolerance threshold for that critical time point, a replanning trigger signal is sent to the mid-term scheduling sub-strategy, along with the latest ultra-short-term forecast data and actual state data. After receiving the signal, the intermediate scheduling sub-strategy quickly re-optimizes the remaining planning period starting from the current time, generates a revised scheduling plan, and updates and sends it to the real-time control sub-strategy. The replanning event is recorded and stored as a special experience in the experience replay buffer of the meta-policy layer; The cumulative deviation is calculated using the Euclidean distance formula.
[0014] Preferably, the method for constructing the composite reward function in the multi-dimensional feedback learning module is as follows: The economic benefit reward is defined as the normalized value of the sum of actual electricity cost savings and market ancillary service revenue; The policy alignment reward is defined as a function that is directly proportional to the degree to which the system's behavior conforms to the latest policy direction; The device loss penalty is defined as the negative value of the battery life degradation cost estimated based on the actual charge-discharge cycle depth and number of cycles. The composite reward function is a linear combination of economic benefit reward, policy fit reward, and equipment loss penalty. The combination coefficients of the composite reward function are set by the operation and maintenance personnel through the management interface or output by the meta-strategy layer, depending on the different stages of system operation.
[0015] Preferably, the system runs on an integrated hardware and software platform; The hardware platform includes an industrial server, a real-time data acquisition device, a security isolation device, and a protocol conversion gateway. The software platform adopts a microservice architecture, encapsulating the data fusion perception layer, the multi-timescale prediction engine, the dynamic strategy generator, and the adaptive execution and feedback loop as independent microservices. The microservices exchange data and are event-driven through message middleware, and are deployed and managed through containerization technology.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs a multi-timescale prediction engine encompassing long-term, medium-term, and short-term perspectives, and in particular, introduces a deep learning-based long-term policy and market trend prediction model, enabling the system to possess forward-looking insight capabilities. The system can proactively predict the medium- and long-term impacts of energy policy adjustments and market mechanism evolution, and integrate this trend judgment as a fundamental constraint into optimization decisions. This overcomes the shortcomings of existing technologies that rely on static rules and short-term data, leading to medium- and long-term strategy failures, and improves the scientific nature of investment decisions and asset returns for energy storage systems throughout their entire lifecycle.
[0017] 2. This invention employs a dynamic policy generator based on hierarchical reinforcement learning, innovatively designing a collaborative architecture between the meta-policy layer and sub-policy layers. The meta-policy layer focuses on long-term value and strategic objectives, while the sub-policy layers handle specific short- to medium-term scheduling and control. This hierarchical decision-making mechanism enables the system to grasp both macro-level direction and perform refined operations. Simultaneously, the introduction of a policy coordinator achieves dynamic calibration and closed-loop management of policies across time scales, ensuring that short-term execution remains consistent with medium- to long-term goals, thus enhancing the system's overall robustness and adaptability to market fluctuations and internal state changes.
[0018] 3. This invention establishes a complete adaptive execution and feedback loop, integrating strategy execution, status monitoring, and multi-dimensional feedback learning. The system not only executes instructions but also continuously learns from actual operational results through a composite reward function, including multiple objectives such as economic benefits, policy alignment, and equipment wear and tear. This data-driven closed-loop learning mechanism enables the system to continuously optimize and evolve, automatically adapting to the continuous changes in the external policy and market environment as well as the aging and degradation of internal equipment. It achieves a leap from static optimization to dynamic adaptation, ensuring the continuous optimality and vitality of the energy storage system management strategy. Attached Figure Description
[0019] Figure 1 This is a block diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the multi-timescale prediction engine in this invention; Figure 3 This is a detailed block diagram of the multi-timescale prediction engine of the present invention; Figure 4 This is a schematic diagram of the closed-loop interaction and data flow between the dynamic strategy generator and the adaptive execution feedback loop in this invention; Figure 5 This is a flowchart illustrating the logical flow of the policy coordinator triggering cross-timescale policy replanning in this invention. Detailed Implementation
[0020] Example 1: The overall technical architecture of the artificial intelligence-based energy storage system load forecasting and optimization management system proposed in this invention is shown in the attached figure. Figures 1 to 5 As shown, it comprises four main functional modules: a data fusion perception layer, a multi-timescale prediction engine, a dynamic strategy generator, and an adaptive execution and feedback loop. These modules collaborate with each other through standardized data interfaces and event-driven mechanisms, achieving high cohesion and low coupling, forming an intelligent optimization management system with forward-looking prediction, hierarchical decision-making, closed-loop execution, and continuous learning capabilities. The following will be combined with the appendix... Figure 1 To be continued Figure 5 The specific implementation methods of each component of the system are described in detail.
[0021] The data fusion perception layer, serving as the system's sensing front end, performs real-time acquisition, cleaning, structuring, and feature extraction of multi-source heterogeneous data from the physical world, market environment, and policy system. This data fusion perception layer comprises three parallel sub-modules: a load and operation data acquisition module, a market and policy data acquisition module, and an external environment data acquisition module.
[0022] The load and operation data acquisition module is directly connected to the battery management system, power conversion system, and smart meters deployed on the user side of the energy storage system via industrial Ethernet or fieldbus. It continuously collects key operating parameters such as historical load curves, real-time load power, energy storage system state of charge, charge / discharge power, round-trip efficiency curves, battery internal resistance change rate, temperature distribution matrix, and equipment health status index at a time granularity of 1 second to 15 minutes. All raw data undergoes data verification, outlier removal, and timestamp alignment before entering the system to ensure data integrity and temporal consistency.
[0023] The market and policy data acquisition module uses a web crawler deployed behind a secure isolation device and application programming interfaces (APIs) provided by authoritative institutions such as the power trading center, the National Energy Administration, and the carbon emission exchange to periodically capture structured and unstructured data. Structured data includes clearing price sequences for day-ahead and real-time markets, ancillary service bid prices, nodal tariffs, and time-of-use tariff tables. Unstructured data covers the full text of the latest energy policy documents, carbon emission quota allocation schemes, renewable energy power consumption responsibility weight notices, and revised drafts of grid dispatch rules. All text data is uniformly encoded in UTF-8 format and annotated with metadata according to publication date, issuing institution, and policy type.
[0024] The external environment data acquisition module obtains hourly weather forecast data for the next 72 hours through the application programming interface of the meteorological service platform. This data includes elements such as temperature, relative humidity, solar irradiance, wind speed, and precipitation probability. It also simultaneously accesses macroeconomic indicators such as the monthly macroeconomic prosperity index and the manufacturing purchasing managers' index released by the National Bureau of Statistics. All collected raw data is uniformly converted into JSON format via a protocol conversion gateway and pushed to subsequent processing units through a message middleware, completing the transformation from physical signals to structured digital assets.
[0025] The multi-timescale prediction engine receives structured data streams from the data fusion perception layer and runs long-term trend prediction models, medium-term situation prediction models, and short-term accurate prediction models in parallel according to different time dimensions and prediction objectives.
[0026] Before inputting policy documents into the model, this application employs the following quantitative techniques: Text structuring: Natural language processing technology is used to extract keywords from "historical policy documents of the past 5 years" (such as "subsidy standards", "peak-valley price difference", "carbon emission quota"), and the TF-IDF algorithm is used to calculate word frequency weights; Vectorized representation: The extracted text features are mapped into dense vectors of 512 dimensions using a pre-trained BERT model; Polarity score: Based on the degree of benefit / negative impact of policies on energy storage revenue, a pre-set sentiment analysis model provides a policy preference value between [-1, 1], which is then concatenated into the feature vector.
[0027] Dataset partitioning: The total number of samples collected in this application is divided into training set, validation set, and test set in a ratio of 7:2:1. The training set is used to fit the Bi-LSTM parameters, the validation set is used to adjust the number of hidden units and the learning rate, and the test set is used to evaluate the model's generalization ability in predicting medium- and long-term trends.
[0028] The overall logical framework of the multi-timescale prediction engine is attached. Figure 2 As shown, the long-term trend forecasting model uses an annual forecast period. Its input dataset includes a compilation of historical policy documents from the past five years, a monthly series of carbon emission allowance prices, a quarterly series of macroeconomic prosperity indices, and a curve showing the decline in the unit energy cost of lithium-ion batteries.
[0029] This long-term trend prediction model first performs deep semantic analysis on policy texts: Each policy document is segmented using a pre-trained large-scale language model, mapping each sentence to a 768-dimensional semantic vector; a policy topic classifier based on a convolutional neural network categorizes the vectors into four predefined policy dimensions—subsidy intensity, carbon emission constraints, market access thresholds, and technical standards requirements; for each dimension, a first-order difference time series regression model is used to fit the trajectory of policy intensity over time, extracting slope (reflecting the speed of policy tightening or loosening) and curvature (reflecting the acceleration of policy changes) as key features. These policy features, together with contemporaneous economic time-series data, constitute the input sequence of a bidirectional long short-term memory network.
[0030] This bidirectional long short-term memory network employs a two-layer stacked structure, with each layer containing 256 hidden units, and introduces a multi-head attention mechanism to dynamically weight the policy impact at different historical moments.
[0031] Specifically, in this embodiment, the combination of the bidirectional long short-term memory network (Bi-LSTM) and the multi-head attention mechanism is as follows: First, the quantified policy sequence vector Input the first layer of Bi-LSTM, They are the first The policy quantification feature vector at a monthly time step, with forward hidden state and backward hidden state Calculated using the following formula: ; ; This represents the Sigmoid activation function; The hidden weight matrix is used as input to the forward LSTM; The forward LSTM recurrent weight matrix; Forward LSTM bias term; The hidden weight matrix is used as the input for the backward LSTM; The weight matrix for the backward LSTM loop; Backward LSTM bias term; The hidden state at time t-1 of the forward LSTM; This represents the hidden state at time t+1 of the backward LSTM. By concatenating the states in both directions, we obtain a bidirectional hidden state concatenation vector. Multi-head attention mechanisms will Mapped to multiple subspaces: ; in, For the first Each attention head output, This represents the scaled dot product attention function. The first The query weight matrix, key weight matrix, and value weight matrix for each attention level; , The first Header query vector, key vector For transpose, The dimension of the key vector. Let be the Softmax activation function, and let be the th Head value vector. Finally, the multi-head output is integrated through a linear transformation. Through the mathematical modeling described above, the system achieves dynamic weight allocation of policy influence at historical moments. This indicates a multi-head output splicing operation. Represented as the first Each attention head output, This is a multi-head output weight matrix.
[0032] The model outputs the predicted annual growth rate of base load over the next 1 to 3 years, the probability distribution of the evolution of time-of-use pricing structure (such as changes in the ratio of peak to off-peak periods and the frequency of peak pricing), and potential policy incentive windows (such as the last 6 months before subsidy reduction). The medium-term trend forecasting model uses a monthly forecast period, and its inputs include historical load data, market electricity price data, and meteorological data from the past 24 months, as well as the annual growth trend and policy window markers output by the long-term trend forecasting model.
[0033] The long-term trend forecasting model employs a hybrid architecture of temporal convolutional networks and gated recurrent units (ROUs). The temporal convolutional networks extract seasonal patterns in load and electricity prices, with kernel sizes covering multiple periods such as 7 days, 30 days, and 90 days. The gated ROUs capture the dynamic lag effect of market response. During training, the model incorporates long-term trend forecasts as soft constraints to ensure that medium-term forecasts do not deviate from the macroeconomic direction. Outputs include typical daily load curves (at 15-minute resolution) for the next 1 to 6 months, average monthly electricity prices, and daily demand forecasts for major ancillary services such as frequency regulation and reserve. The short-term precise forecasting model, with an hourly forecast period, focuses on high-precision load and electricity price forecasts for the next 24 to 48 hours. Its input data includes 15-minute load curves from the past 7 days, real-time rolling market clearing prices, ultra-short-term weather forecasts (updated every 15 minutes for the next 6 hours), and the baseline daily curve output from the medium-term trend forecasting model.
[0034] This mid-term trend forecasting model employs a gradient boosting decision tree algorithm, constructing an ensemble model containing 2000 decision trees, each with a maximum depth of 8 and a learning rate of 0.1. The model makes rolling forecasts in 15-minute time steps, outputting the expected load value and its 95% confidence interval for each future time point, while also predicting the upper and lower limits of electricity price fluctuations. The three forecasting models share the same data lake storage but each has its own independent feature engineering pipeline and model training scheduler, ensuring parallelism and resource isolation in the forecasting tasks.
[0035] The dynamic policy generator is the decision-making center of the system, employing a hierarchical reinforcement learning decision network to achieve collaborative optimization across time scales. This hierarchical reinforcement learning decision network consists of a meta-policy layer and two sub-policy layers (mid-term scheduling sub-policy and real-time control sub-policy), with dynamic coupling between layers achieved through a policy coordinator. The meta-policy layer aims to maximize the system's value throughout its entire lifecycle. Its state space comprises three dimensions: long-term policy guidance strength (ranging from 0 to 1, output by a long-term trend prediction model), market mechanism evolution stage (discrete variables, categorized into initial, growth, maturity, and decline stages), and system asset depreciation status (represented by the percentage of remaining economic life). The action space is defined by three sets of strategic parameters: the weight of the mid-term economic return objective. Long-term policy risk aversion weight Equipment lifespan maintenance weight All three conditions are met. And all of them are non-negative real numbers.
[0036] The meta-policy layer is trained using a deep deterministic policy gradient algorithm. Its policy network consists of three fully connected neural networks: an input layer with a dimension of 128, a hidden layer with a dimension of 256, and an output layer that generates weights after Softmax activation. During training, the meta-policy layer accumulates experience through interaction with sub-policy layers, and its reward signal originates from the long-term cumulative discount value of the composite reward function fed back after the execution of the sub-policy.
[0037] The medium-term scheduling sub-strategy operates under the weights set at the meta-strategy layer. Its state space includes the predicted load curve for the next 7 days (15-minute resolution), the predicted electricity price curve, the current state of charge of the energy storage system, and the battery health index. Its action space comprises the day-ahead market bidding electricity plan for the next 24 hours to 7 days, the daily charging and discharging periods of the energy storage system, and the power ceiling framework. This medium-term scheduling sub-strategy employs a near-end strategy optimization algorithm with a reward function... Defined as: ; This is the normalized value of expected electricity cost savings and ancillary service revenue. Scoring is given based on the alignment between the operational plan and policy guidance. This represents the battery life retention rate estimated based on the depth of charge / discharge.
[0038] For reward function The original calculation logic for its three weight parameters is as follows: The system first obtains the current external environment state variables, including real-time electricity price volatility. Policy change index and battery cycle count. Weights are determined through the following mapping relationship: The Sigmoid function is used to adjust the revenue weights based on the sensitivity to electricity price fluctuations. The slope parameter of the Sigmoid function. This is the threshold (inflection point) parameter for the Sigmoid function; It is positively correlated with the policy propensity value; when the policy uncertainty index is above the threshold... Automatic adjustment ; Calculated in real time based on the percentage of remaining battery cycle life. This refers to the battery's rated total cycle life. This represents the number of battery cycles that have been used.
[0039] The real-time control sub-strategy operates with fine-grained precision within the planning framework provided by the medium-term scheduling sub-strategy. Its state space includes ultra-short-term load and electricity price forecasts for the next two hours (15-minute resolution), real-time state of charge of the energy storage system, current charging / discharging power, and battery temperature. The action space consists of specific charging / discharging power commands (in kilowatts) for the next 15-minute time interval. This real-time control sub-strategy also employs a near-end strategy optimization algorithm, but its reward function introduces an additional penalty term for deviation from the medium-term plan. : ; For immediate economic gain, The Euclidean distance between the current actual situation and the medium-term plan on key indicators. This is the penalty coefficient, dynamically adjusted by the policy coordinator based on the current deviation. The working logic of the policy coordinator is shown in the attached figure. Figure 5 As shown, it continuously monitors the cumulative deviations between the short-term execution trajectory and the medium-term plan at key time points such as peak electricity price periods, peak load periods, and charging / discharging state switching points. When the deviation at any key time point exceeds the tolerance threshold dynamically calculated based on historical volatility for that key time point, the strategy coordinator immediately triggers the online replanning process of the medium-term scheduling sub-strategy, using the latest ultra-short-term forecast data and the actual system state as initial input conditions. After replanning is completed, the new plan is issued to the real-time control sub-strategy, and the event is recorded as a special experience stored in the experience replay buffer of the meta-strategy layer for subsequent meta-strategy robustness evaluation.
[0040] Adaptive execution and feedback loop constitute the closed-loop execution and learning mechanism of the system, and its data flow and interaction logic are shown in the appendix. Figure 4 As shown. The instruction compilation and distribution module receives the charging and discharging power instructions output by the real-time control sub-strategy. It first performs a physical feasibility check, including whether the power exceeds the converter's rated capacity, whether the state of charge is within a safe range (typically 20% to 80%), and whether the temperature rise rate exceeds the limit. After passing the check, the instruction is compiled into a CANopen or Modbus compliant version. The specific control messages of the TCP protocol are sent to the power conversion system via industrial Ethernet with millisecond-level latency. The execution status monitoring module collects data such as actual charging and discharging power, DC side voltage and current, battery cell voltage, and module temperature in real time at a 100-millisecond sampling rate, and compares them point by point with the expected values of the instructions to generate an energy execution deviation report including absolute error, relative error, and cumulative deviation. The multi-dimensional feedback learning module constructs a composite reward function and drives the continuous evolution of the policy network. The economic reward is obtained by connecting to the electricity market settlement system to obtain actual electricity bills and ancillary service revenue vouchers, which are then normalized to obtain a score between 0 and 1.
[0041] The policy alignment reward uses the policy analysis submodule in the long-term trend prediction model to semantically match actual operational behavior with the guiding characteristics of current effective policies, calculating an alignment score. Equipment loss penalty is based on rainflow counting to statistically analyze the depth and number of actual charge-discharge cycles, combined with a battery aging model to estimate the remaining lifespan degradation cost, and then converts this into a negative reward. The composite reward function is a linear combination of these three factors; its coefficients can be set by maintenance personnel in the management interface or dynamically output by the meta-policy layer.
[0042] All reward signals and their corresponding state-action pairs are synchronously written into the experience replay buffer of each policy layer. The system triggers an offline retraining process every 24 hours, using a priority experience replay mechanism to sample high-value samples from the buffer and update the parameters of the policy network. At the same time, after each policy execution, a target network soft update mechanism is used to fine-tune the online policy, achieving incremental optimization of the policy.
[0043] The system operates on a highly reliable integrated hardware and software platform. The hardware platform consists of an industrial-grade server cluster, real-time data acquisition devices, forward and reverse security isolation devices, and a multi-protocol conversion gateway. The server cluster employs a dual-machine hot standby architecture, equipped with redundant power supplies and RAID storage to ensure uninterrupted 24 / 7 operation. The software platform is based on the Kubernetes container orchestration system, encapsulating the data fusion perception layer, multi-timescale prediction engine, dynamic policy generator, and adaptive execution and feedback loop as independent Docker container microservices.
[0044] Each microservice communicates asynchronously via the Apache Kafka messaging middleware. Message topics are categorized by functional domain to ensure the orderliness and traceability of data flow. Service registration and discovery are implemented using Consul, and configuration management is centrally controlled via etcd. The system provides a RESTful API interface to support secure data interaction with the upper-level energy management system and power grid dispatching platform. All sensitive data is encrypted using the national standard SM4 algorithm during transmission and storage. Access control follows the principle of least privilege, and operation logs are audited and retained throughout the entire process.
[0045] In summary, this embodiment achieves a fundamental transformation of energy storage systems from passive response to proactive prediction and from static optimization to dynamic adaptation by constructing a complete technical chain encompassing multi-source data fusion, multi-timescale prediction, hierarchical reinforcement learning decision-making, and closed-loop adaptive execution. The system can not only accurately predict short-term load and electricity prices but also gain insights into medium- and long-term policy and market trends. This forward-looking understanding is transformed into hierarchical collaborative optimization strategies, ultimately evolving continuously through a closed-loop feedback mechanism to ensure that energy storage assets maintain optimal operating conditions and the highest return on investment in a complex and ever-changing energy environment.
[0046] Example 2: Building upon Example 1, this example further refines the specific engineering implementation details of policy text vectorization and feature extraction in the long-term trend prediction model, and enhances the emergency response capability of the meta-policy layer in the dynamic policy generator to sudden events. Please refer to the appendix. Figure 2 This embodiment significantly extends the policy analysis process of the long-term trend forecasting model. When the market and policy data acquisition module captures a new policy document, the system first performs a version comparison. If the policy document is a revision of an existing policy, the differences in the revised content are automatically extracted; if it is a completely new policy, it is marked as "new".
[0047] Before being fed into the language model, all policy texts undergo entity recognition preprocessing: using a BERT-CRF-based named entity recognition model, key entities involved in the policy are extracted, including geographical scope, applicable objects, time points, and numerical indicators. These entities are embedded into the original sentence vectors to form enhanced semantic vectors. Subsequently, the policy topic classifier no longer uses a fixed four-dimensional classification but introduces a dynamic topic evolution mechanism: the system maintains a policy topic knowledge graph, where nodes are policy keywords and edges represent co-occurrence relationships.
[0048] Whenever a new policy is added to the database, the system calculates its similarity to existing topic clusters using a graph neural network. If the similarity is below a threshold of 0.6, a new topic node is automatically created. This mechanism enables the system to identify emerging policy directions, such as "virtual power plant aggregation rules" or "green certificate-carbon market linkage mechanisms." In the time series regression phase, the system not only fits the first and second derivatives of policy intensity but also introduces Fourier transforms to extract the periodic components of policies, such as the annual cycle signal formed by the annual subsidy settlement policies released in the fourth quarter of each year. These frequency domain features are concatenated with the time domain features and then jointly input into a bidirectional long short-term memory network, enhancing the ability to capture policy pulse-like changes.
[0049] Regarding the dynamic strategy generator, this embodiment adds an emergency response submodule to the meta-strategy layer. This emergency response submodule continuously monitors abnormal event signals input from the external environmental data acquisition module, including extreme weather warnings (such as red high-temperature warnings), announcements of major policy changes (such as a sudden 50% drop in the free allocation ratio of carbon quotas), and market price circuit breaker events. Once such an event is detected, the emergency response submodule immediately freezes the weight output of the current meta-strategy and activates the preset emergency strategy template library.
[0050] The emergency strategy template library stores several combinations of emergency strategies that have been validated through historical backtesting. For example, in the scenario of "extreme high temperature and high electricity price," the system automatically and temporarily increases the weight of equipment lifespan maintenance, forcibly limiting the depth of discharge of the energy storage system during the afternoon high-temperature period to protect battery safety. The activation duration of the emergency strategy is 1.5 times the duration of the event's impact, after which the system automatically and smoothly transitions back to the normal strategy. Simultaneously, the entire process data of this emergency event is fully recorded and used for subsequent adversarial training of the meta-policy network, improving its robustness under black swan events.
[0051] The deviation calculation logic of the strategy coordinator has also been enhanced: a time decay factor is introduced when calculating Euclidean distance, making the weight of recent deviations higher than that of long-term deviations; different importance coefficients are assigned to different types of deviations, for example, the power deviation coefficient is set to 1.5 during peak electricity price periods and 0.8 during off-peak periods, reflecting the difference in economic sensitivity. These improvements enable the system to maintain long-term strategic focus while possessing the ability to quickly adapt to sudden disturbances, further improving the overall reliability and economy of operation.
[0052] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0053] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence-based energy storage system load forecasting and optimization management system, characterized in that, include: The data fusion perception layer is used to collect and structure multi-source heterogeneous data in real time. The data fusion perception layer includes a load and operation data collection module, a market and policy data collection module, and an external environment data collection module. A multi-timescale prediction engine is used to run a long-term trend prediction model, a medium-term situation prediction model, and a short-term accurate prediction model in parallel based on the structured data provided by the data fusion perception layer. A dynamic policy generator is used to receive all prediction results output by the multi-timescale prediction engine and generate collaborative optimization policies covering different timescales based on a hierarchical reinforcement learning decision network. An adaptive execution and feedback loop is used to send the final control command output by the dynamic strategy generator to the power conversion system of the energy storage system for execution, and to build a closed-loop learning mechanism.
2. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 1, characterized in that, The load and operation data acquisition module is used to collect historical load curves, real-time load data, state of charge, charging and discharging power, efficiency curves, and equipment health status parameters of the energy storage system. The market and policy data acquisition module is used to collect real-time electricity price data, day-ahead market and real-time market clearing prices, ancillary services market rule texts, energy policy documents, carbon emission quota prices, and renewable energy quota system requirements. The external environment data acquisition module is used to collect meteorological forecast data and macroeconomic prosperity index.
3. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 2, characterized in that, The long-term trend prediction model uses an annual forecast period, with inputs including a compilation of historical policy documents, carbon emission price series, macroeconomic data, and technology cost decline curves. It employs a bidirectional long short-term memory network architecture based on an attention mechanism and outputs a forecast of future base load growth trends, a probability distribution of the evolution of time-of-use electricity price structure, and potential policy incentive windows. The medium-term trend prediction model uses a monthly forecast period. Its inputs are historical load data, market electricity price data, meteorological data, and trend guidance signals output by the long-term trend prediction model. It adopts a hybrid architecture of temporal convolutional network and gated recurrent unit to output the forecast of future monthly typical daily load curve, monthly average electricity price forecast, and major ancillary service demand forecast. The short-term accurate forecasting model uses hourly forecasting periods. Its inputs include refined load data, real-time electricity prices, ultra-short-term weather forecasts, and the baseline curve output by the medium-term situation forecasting model. It employs a gradient boosting decision tree algorithm to continuously forecast future load values and electricity price fluctuation ranges.
4. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 3, characterized in that, The hierarchical reinforcement learning decision network comprises a meta-policy layer and a sub-policy layer; The meta-strategy layer aims to maximize long-term trend prediction results and long-term system returns. Its state space is defined as the intensity of long-term policy guidance, the stage of market mechanism evolution, and the depreciation status of system assets. Its action space consists of weight allocation instructions for medium- and short-term optimization objectives. It is trained through a deep deterministic strategy gradient algorithm and outputs strategic-level decision parameters. The sub-policy layer includes a mid-term scheduling sub-policy and a real-time control sub-policy; The intermediate-term scheduling sub-strategy operates under the target weights set by the meta-strategy layer, with the intermediate-term situation prediction results and the expected returns for the next week as the objectives. The state space consists of the predicted load curve, predicted electricity price curve, and energy storage system status for the next 7 days, while the action space consists of the day-ahead market bidding plan and energy storage system charge and discharge plan framework for the next 24 hours to 7 days. The real-time control sub-strategy operates within the planning framework provided by the medium-term scheduling sub-strategy, with the goal of maximizing short-term accurate prediction results and real-time revenue. Its state space is the load and electricity price prediction for the next 15 minutes to 2 hours and the real-time charge status of the energy storage system. Its action space is the specific charging and discharging power command for the next time interval. The dynamic strategy generator also includes a strategy coordinator, which is used to monitor the deviation between short-term execution results and medium-term plans in real time, and trigger the online replanning of the medium-term scheduling sub-strategy when the deviation exceeds a preset tolerance threshold, and feed back the replanning information to the meta-strategy layer.
5. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 4, characterized in that, The adaptive execution and feedback loop includes an instruction compilation and issuance module, an execution status monitoring module, and a multi-dimensional feedback learning module. The instruction compilation and distribution module is used to compile the charging and discharging power instructions output by the real-time control sub-strategy into specific control commands that conform to the energy storage system communication protocol and then distribute them. The execution status monitoring module is used to collect the actual execution results of the instructions in real time and compare them with the expected values of the instructions to generate an execution deviation report; The multi-dimensional feedback learning module receives the execution deviation report, actual market settlement data, and policy change information. It uses actual revenue, policy fit, and equipment loss as multi-dimensional reward signals, and execution deviation and strategy failure events as penalty signals to form a composite reward function. The composite reward function is then synchronously fed back to the experience replay buffer of each strategy layer in the hierarchical reinforcement learning decision network to drive the network to perform periodic offline retraining and online parameter fine-tuning.
6. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 5, characterized in that, The policy text analysis process in the long-term trend prediction model is as follows: The policy documents are segmented and vectorized using a pre-trained language model, transforming each policy sentence into a high-dimensional semantic vector; the vectors are then categorized into several predefined policy dimensions using a policy topic classifier. For each policy dimension, a time series regression model is used to analyze the trend of policy intensity over time, and the slope and curvature of the trend are extracted as features. The policy characteristics and the economic time-series data of the same period are input into a bidirectional long short-term memory network for joint prediction. The predefined policy dimensions include subsidy intensity, carbon emission constraints, market access, and technical standards.
7. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 6, characterized in that, The strategic decision-making parameters output by the meta-strategy layer specifically include the weights of medium-term economic benefit targets, long-term policy risk aversion, and equipment life maintenance. The reward function of the intermediate scheduling sub-strategy is the weighted sum of each weight and the corresponding sub-reward set by the meta-strategy layer; In addition to immediate economic gains, the reward function of the real-time control sub-strategy also includes a penalty term for deviation from the medium-term plan. The coefficient of this penalty term is dynamically adjusted by the strategy coordinator based on the current deviation.
8. The energy storage system load forecasting and optimization management system based on artificial intelligence according to claim 7, characterized in that, The working logic of the policy coordinator is as follows: The cumulative deviation between the short-term actual execution trajectory and the medium-term scheduling plan trajectory at key time points is continuously calculated. These key time points include peak electricity price periods, peak load periods, and planned charging / discharging state switching points. When the cumulative deviation at any critical time point exceeds the preset dynamic tolerance threshold for that critical time point, a replanning trigger signal is sent to the mid-term scheduling sub-strategy, along with the latest ultra-short-term forecast data and actual state data. After receiving the signal, the intermediate scheduling sub-strategy quickly re-optimizes the remaining planning period starting from the current time, generates a revised scheduling plan, and updates and sends it to the real-time control sub-strategy. The replanning event is recorded and stored as a special experience in the experience replay buffer of the meta-policy layer; The cumulative deviation is calculated using the Euclidean distance formula.
9. The artificial intelligence-based energy storage system load forecasting and optimization management system according to claim 8, characterized in that, The method for constructing the composite reward function in the multi-dimensional feedback learning module is as follows: The economic benefit reward is defined as the normalized value of the sum of actual electricity cost savings and market ancillary service revenue; The policy alignment reward is defined as a function that is directly proportional to the degree to which the system's behavior conforms to the latest policy direction; The device loss penalty is defined as the negative value of the battery life degradation cost estimated based on the actual charge-discharge cycle depth and number of cycles. The composite reward function is a linear combination of economic benefit reward, policy fit reward, and equipment loss penalty. The combination coefficients of the composite reward function are set by the operation and maintenance personnel through the management interface or output by the meta-strategy layer, depending on the different stages of system operation.
10. The energy storage system load forecasting and optimization management system based on artificial intelligence according to claim 9, characterized in that, The system runs on an integrated hardware and software platform; The hardware platform includes an industrial server, a real-time data acquisition device, a security isolation device, and a protocol conversion gateway. The software platform adopts a microservice architecture, encapsulating the data fusion perception layer, the multi-timescale prediction engine, the dynamic strategy generator, and the adaptive execution and feedback loop as independent microservices. The microservices exchange data and are event-driven through message middleware, and are deployed and managed through containerization technology.