A method and system for optimizing energy of optical storage charging transformer based on deep reinforcement learning
By using a hierarchical reinforcement learning control architecture and a high-fidelity physical model, the short-term economic efficiency and long-term security of the power system are optimized in a coordinated manner, solving the problem of nonlinear degradation of equipment in existing technologies and achieving equipment life extension and cost optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 浙江富杰电气有限公司
- Filing Date
- 2025-07-23
- Publication Date
- 2026-04-14
AI Technical Summary
Existing machine learning methods have failed to effectively coordinate the operating states at different time scales in power system optimization, leading to nonlinear degradation of equipment over long time scales and creating safety hazards.
A hierarchical reinforcement learning control architecture is adopted, which includes a high-level policy agent (manager) and a low-level policy agent (executor). The high-level agent is responsible for long-term security policies, while the low-level agent is responsible for immediate economic benefits. The system is optimized collaboratively through a high-fidelity physical model and a reward mechanism.
It significantly improves equipment lifespan, reduces maintenance costs throughout the entire lifecycle, and ensures the safety and economy of the system at different time scales.
Smart Images

Figure CN120855306B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power system energy management technology, and in particular to an energy optimization method and system for photovoltaic-storage-charging transformers based on deep reinforcement learning. Background Technology
[0002] With the increasing prevalence of distributed energy sources and new loads such as photovoltaic power generation, battery energy storage, and electric vehicle charging facilities on the distribution network side, a highly integrated energy system combining photovoltaics, energy storage, charging, and transformers is gradually forming. Efficient energy management and optimized scheduling of such complex systems are crucial for improving system economic efficiency and ensuring the safe and stable operation of the power grid.
[0003] Deep reinforcement learning techniques, due to their ability to handle complex dynamic decision-making problems, have been studied and applied to the optimal scheduling of energy systems. Related technologies primarily focus on maximizing the short-term economic benefits of the system, such as time-of-use pricing arbitrage by controlling the charging and discharging behavior of energy storage systems, or reducing demand charges by cutting peak loads. Another type of application focuses on ensuring the short-term operational stability of the system, for example, by coordinating and managing the charging process of electric vehicles to prevent instantaneous overloads of distribution transformers.
[0004] Existing technologies also employ machine learning and other methods to achieve the aforementioned optimization goals. However, existing technologies generally suffer from a limitation in pursuing short-term optimal solutions. To ensure the computational feasibility of reinforcement learning algorithm models, existing methods typically significantly simplify the physical characteristics and long-term health degradation processes of core power equipment in the system, particularly transformers and energy storage batteries. For example, transformers are often modeled as passive components with a static capacity limit, with the only operational constraint being that the load cannot exceed its rated value, completely ignoring the cumulative and nonlinear effects of different load modes on their insulation lifespan. For energy storage batteries, their complex electrochemical aging mechanisms are usually abstracted into a simplified cost term or cycle count limit, failing to accurately reflect the differentiated damage to their health caused by different charge and discharge strategies (such as charge / discharge rate, depth of discharge, etc.). This insufficient modeling of the equipment's physical characteristics leads to an inherent and significant long-term safety hazard in existing intelligent scheduling strategies. Their optimization goals and decision-making actions occur on a rapid timescale of seconds, minutes, to hours, while the cumulative consequences of these decisions on the health of core assets only become apparent on a slow timescale of months, years, or even decades. This makes the most economical machine learning optimization strategy for a short period of time (such as 24 hours, because this is the minimum cycle of photovoltaic electricity price and charging capacity) actually cause premature aging of devices such as energy storage batteries and transformers that degrade nonlinearly over long time scales, creating huge safety hazards. In other words, traditional short-term (whether it is user charging cycle or real-time machine learning of photovoltaic grid-connected electricity) sensitive models may become traps that cause safety accidents. Summary of the Invention
[0005] (a) Technical issues addressed
[0006] Current machine learning optimization methods for power systems focus on short-term economic efficiency or other indicators, failing to coordinate the system's operational status across various time scales (such as real-time fluctuations in photovoltaic feed-in tariffs, hourly or daily changes in charging load, monthly or yearly variations in energy storage battery performance, and transformer device aging).
[0007] (II) Technical Solution
[0008] To address the aforementioned technical problems, this application provides an integrated energy system lifecycle optimization method based on hierarchical physical information deep reinforcement learning. This method is executed on a computing device and includes the following steps:
[0009] The first step is to construct a hierarchical reinforcement learning control architecture. This architecture includes a high-level policy agent (hereinafter referred to as the "Manager") and a low-level policy agent (hereinafter referred to as the "Executor"). The Manager makes decisions on a slower time scale (e.g., in days, weeks, or months). Its core responsibility is to formulate macro-level, long-term safe operation strategies and boundaries from the perspective of ensuring the safety of the asset throughout its entire lifecycle, preventing irreversible excessive damage to hardware caused by short-term profit-seeking behavior. The Executor makes decisions on a faster time scale (e.g., in 5-minute or 15-minute intervals). Its core responsibility is to maximize immediate economic benefits by executing specific energy scheduling within the safety boundaries set by the Manager. This hierarchical structure, through clear division of responsibilities, fundamentally resolves the conflict between short-term economic pursuits and long-term operational safety assurance.
[0010] The second step is to configure the state space, action space, and reward function driven by the physical characteristics of the power equipment for the managers and executors.
[0011] For the manager, their state space primarily consists of long-term health indicators of the system (which can be viewed as a quantification of macroscopic safety status) and macroscopic environmental predictions. Specifically, this may include: the cumulative equivalent aging of the power transformer calculated based on a high-fidelity physical model, and the state of health (SoH) of the energy storage battery calculated based on a high-fidelity physical model. The manager's actions do not involve outputting specific equipment control commands, but rather issuing a strategic "safety boundary" or "operational permission" to the executor. Preferably, the manager's actions define specific safety constraints for one or more future scheduling cycles. For example, setting the maximum allowable upper limit for the transformer winding hotspot temperature within the next week, or the maximum allowable degradation rate of the energy storage battery's health status within the next month, and the budgeted equivalent full-cycle consumption of the energy storage battery. These indicators collectively constitute the dynamic constraints on the executor's decision space. The manager's reward function aims to incentivize them to make decisions most beneficial to long-term safety; its reward value is directly related to the reduction of asset depreciation costs or the extension of asset lifespan.
[0012] For the executor, its state space mainly consists of the system's real-time operating parameters and the safety boundaries set by the manager. Specifically, this may include: real-time photovoltaic power generation, various load power, state of charge (SoC) of energy storage batteries, real-time time-of-use pricing, and safety constraint indicators received from the manager. The executor's actions are specific energy dispatch instructions, such as setting the charging and discharging power of the energy storage system and adjusting the total charging power of electric vehicle charging stations. The executor's reward function is a composite signal, the main part of which reflects external rewards reflecting immediate economic benefits, such as gains from electricity arbitrage or peak shaving. Simultaneously, when its behavior remains within the safety boundaries set by the manager, it can obtain additional positive rewards; once it exceeds the boundaries, it will be penalized, thereby forcing the executor to strictly adhere to long-term safety principles while pursuing economic efficiency.
[0013] The third step involves executing the closed-loop optimization process of the hierarchical reinforcement learning control architecture. This process iterates continuously. At a slow-timescale decision point, the manager generates a macro-level safety objective for the next scheduling cycle (e.g., the next week or month) based on observed equipment performance trends or long-term health status assessed on an annual basis, and issues this objective as a rigid constraint to the executors. In subsequent scheduling cycles, at each fast-timescale decision point, the executors generate and execute specific energy scheduling actions based on their observed real-time operating status and received safety boundaries. The results of these actions directly generate short-term economic benefits and, through the physical model, influence the long-term safety status of transformers and energy storage batteries. The system feeds back the resulting economic benefits and safety status changes to both the manager and the executors to update their respective policy networks, thereby achieving co-evolution and continuous optimization of the manager's safety strategy and the executors' economic strategy.
[0014] Preferably, to more accurately quantify and set the safety boundary of the transformer, a thermodynamic model based on the IEEE Std C57.91 standard is used when calculating the transformer winding hot spot temperature. This model can comprehensively consider the real-time load factor, ambient temperature, and the transformer's nameplate parameters to dynamically calculate the top oil temperature and winding hot spot temperature. The calculated hot spot temperature is then substituted into the Arrhenius equation to accumulate and calculate the insulation equivalent aging factor, thereby providing a physical basis for managers to set safe temperature upper limits.
[0015] Preferably, to more accurately quantify and define the safety boundary of the energy storage battery, a semi-empirical or electrochemical battery physics model is employed. This model can quantify the combined effects of various operating factors such as high-rate charging and discharging, deep discharging, and operating temperature on the battery. These effects include various degradation mechanisms such as lithium deposition, increased stress in electrode material structures, and abnormal growth of the solid electrolyte interface film. This allows for a more accurate prediction of the battery's SoH decay rate under different scheduling strategies, and the setting of a safe decay rate or equivalent cycle count budget accordingly.
[0016] Preferably, to improve the learning efficiency and constraint compliance of the agent, the method of this application further introduces the concept of physical information reinforcement learning. For hard physical constraints such as the power balance equation, this can be achieved by designing the output of the executor policy network. That is, the executor policy network does not directly output control quantities such as energy storage power, but outputs one or more unconstrained control variables. The final scheduling action is generated by a fixed, differentiable function based on the control variables and other power parameters of the system. This function structurally ensures that the power input and output of the system always satisfy the physical conservation law. For soft but important constraints such as the safety boundary issued by the manager, a constraint reinforcement learning algorithm based on Lagrange relaxation is adopted. A penalty term proportional to the degree of violation of the safety boundary is added to the loss function of the executor, and a learnable Lagrange multiplier is used to dynamically balance economic goals and safety constraints.
[0017] Preferably, in order to further enhance the manager's ability to capture long-term security dynamics, the manager's policy network adopts a Transformer-based neural network architecture, which utilizes its self-attention mechanism to effectively learn and process long-distance dependencies in long-term health data time series, thereby gaining a deeper understanding of the causal relationship between different operating models and asset accumulation and aging, and formulating more forward-looking security strategies.
[0018] Preferably, the implementation of the method described in this application further includes specific data processing and model training procedures. These include an offline pre-training stage and an online fine-tuning stage. In the offline pre-training stage, the neural networks of the manager and executors are fully trained in a simulation environment using historical operational data to obtain an initial strategy with basic scheduling capabilities. After system deployment, the online fine-tuning stage begins, where the model is continuously updated based on actual operational data. The executor model can undergo high-frequency online updates to adapt to short-term changes, while the manager model can use a lower-frequency retraining cycle, such as monthly or quarterly retraining using the latest accumulated health data to optimize its long-term security strategy.
[0019] This application also provides a system for implementing the above method, deployed in a server or embedded controller, comprising: a data interface for receiving short-term and long-term data from various power devices in the energy system and the external environment; one or more processors; and instructions stored on a non-transitory computer-readable medium. When executed by the processor, the instructions can be configured as a manager module and an executor module. The manager module executes the high-level policy agent to generate the strategic security boundary; the executor module executes the low-level policy agent to generate the energy scheduling action under the constraints of the security boundary; wherein the manager module provides the strategic security boundary to the executor module, and the executor module outputs control signals for controlling energy system components.
[0020] (III) Beneficial Effects
[0021] The technical solution provided in this application, by constructing a hierarchical reinforcement learning architecture with upper-layer security policy constraints on lower-layer economic optimization as its core, and deeply integrating a high-fidelity physics model, brings significant beneficial effects:
[0022] 1. High-level managers and low-level implementers are decoupled in terms of the objective function. The high-level managers are only responsible for safety, and they output constraints to the lower levels without participating in the trade-off of real-time economic costs. This sets the framework for the entire optimization and breaks through the existing technology where the state of the power equipment itself only accounts for a part of the weight, thus prioritizing safety.
[0023] 2. The objective function of the high-level optimization strategy is based on important indicators of various power equipment that change nonlinearly over long time scales, which breaks through the potential safety hazards of traditional optimization indicators that are mainly based on electricity price and demand in existing technologies. The reward and penalty functions of the high-level optimization strategy for the lower-level implementers are also specially designed to ensure that the safety strategy based on the nonlinear changes of equipment is rigidly executed rather than just used as a general weighting term in the reward function.
[0024] 3. By employing short-term economic optimization constrained by the long-term expected changes of core fixed assets, the actual service life of these core assets (such as transformers and energy storage batteries) is significantly increased, and the high costs incurred by the system throughout its entire lifecycle due to equipment replacement or major repairs are substantially reduced. Total cost optimization over the entire lifecycle is achieved by using short-term suboptimal economic optimization objectives.
[0025] 4. This application also proposes a loop architecture in which the underlying implementers accumulate real-world data feedback to constrain the manager model over thousands of execution cycles, preventing the tendency of a one-way or top-down constraint system to be too strict and conservative. Attached Figure Description
[0026] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the overall architecture of the photovoltaic-storage-charging-transformer energy optimization system described in the embodiments of this application, illustrating the interconnection between physical devices and the hierarchical control system.
[0028] Figure 2 This is an overall flowchart of the energy optimization method based on hierarchical reinforcement learning described in the embodiments of this application, illustrating the main steps of the method.
[0029] Figure 3 This is a schematic diagram of the hierarchical reinforcement learning control architecture described in the embodiments of this application. It details the interaction logic of state, action and reward information between high-level managers and low-level executors, and is one of the core inventive points of this application.
[0030] Figure 4 This is an information flow diagram of the transformer thermodynamics and aging physics model used for decision-making by senior managers in the embodiments of this application, which reflects the physical information fusion characteristics of this application.
[0031] Figure 5 The information flow diagram of the physical model of the health status of energy storage batteries used for decision-making by senior managers in this application embodiment also reflects the characteristics of physical information fusion.
[0032] Figure 6 In the simulation environment of this application embodiment, a scatter plot comparing the daily equivalent aging amount of transformers and daily economic benefits using the method of this application and the traditional economic optimization method is provided to intuitively demonstrate the beneficial effects of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] Please see Figures 1 to 6 This application provides a technical solution for an integrated energy system lifecycle optimization method based on hierarchical physical information deep reinforcement learning.
[0035] refer to Figure 2 The process shown, the optimized method provided in this application, is executed in a computing device (e.g., a server deployed in the cloud or a local edge computing gateway), specifically including the following steps:
[0036] Step S100: Construct a hierarchical reinforcement learning control architecture.
[0037] The architecture is as follows Figure 3 As shown, the system comprises a high-level strategic agent (hereinafter referred to as the "Manager") and a low-level strategic agent (hereinafter referred to as the "Executor"). The Manager makes decisions on a slower timescale (e.g., in weekly or monthly units). Its core responsibility is to formulate macro-level, long-term security operation strategies and boundaries from the perspective of ensuring the security of the asset throughout its entire lifecycle, preventing irreversible excessive damage to hardware caused by short-term profit-seeking behavior. The Executor makes decisions on a faster timescale (e.g., in 15-minute units). Its core responsibility is to maximize immediate economic benefits by executing specific energy scheduling within the security boundaries set by the Manager. This hierarchical structure, through clear division of responsibilities, fundamentally resolves the conflict between short-term economic pursuits and long-term operational security assurance.
[0038] To better understand this architecture, this embodiment further elaborates on this step:
[0039] Step S101: Define the manager's role and decision-making cycle.
[0040] Managers monitor overall safety across multiple time scales. Their decisions are not based on the hourly electricity price, but rather on whether short-term economic optimization (such as increasing supercharging and fast charging capacity or increasing battery cycle life during peak grid and social electricity demand periods in hot seasons when photovoltaic prices are low) will accelerate sudden hardware aging, or whether to adjust constraints when hardware has been operating at high load for an extended period without deteriorating monitoring indicators. Therefore, their decision-making cycle is set at a slower time scale. For example, at the beginning of each calendar week, managers make a decision based on accumulated equipment health data from the past few weeks or even months, as well as macro-environmental forecasts for the coming week (such as weather forecasts and holiday load patterns). The result of this decision is not a specific power command, but a safety constraint strategy guiding the entire system's operation for the following week. This provides macro-level control over the overall system safety at a slower time scale.
[0041] Step S102: Define the role of the executor and the decision-making cycle.
[0042] The executor makes decisions within the constraints of the policy. It possesses full autonomy within the safety limits defined by the manager, focusing on optimizing the controlled economy of the system. Its decision-making cycle is set to a fast timescale, matching the electricity market trading or typical dispatch cycle, such as every 15 minutes. At each 15-minute decision point, the executor receives the real-time status of the system and quickly calculates a combination of control actions that maximizes the economic benefits during that period (e.g., utilizing peak-valley electricity price differences for charging and discharging). The only hard constraint on its decisions is that it must never exceed the safety boundaries set by the manager. In this way, the executor's agility in pursuing timely economic benefits is ensured, while upper-level constraints prevent it from making actions that compromise long-term safety by solely calculating economics.
[0043] This layered architecture design allows for the decoupling and coordination of system security and economy across different time scales.
[0044] Step S200: Configure the state space, action space, and reward function driven by the physical characteristics of the power equipment for the manager and the executor.
[0045] For the manager, their state space primarily consists of long-term system health indicators (which can be viewed as a quantification of macroscopic safety status) and macroscopic environmental predictions. Specifically, this may include: the cumulative equivalent aging of the power transformer calculated based on a high-fidelity physical model, and the state of health (SoH) of the energy storage battery calculated based on a high-fidelity physical model. The manager's actions involve defining specific safety constraints for one or more future scheduling cycles. For example, setting the maximum allowable upper limit for transformer winding hotspot temperature within the next week, or the maximum allowable degradation rate of the energy storage battery's health status within the next month, and the budgeted equivalent full-cycle consumption of the energy storage battery. The manager's reward function aims to incentivize them to make decisions most beneficial to long-term safety; its reward value is directly related to the reduction of asset depreciation costs or the extension of asset lifespan.
[0046] For the executor, its state space mainly consists of the system's real-time operating parameters and the safety boundaries set by the manager. Specifically, this includes: real-time photovoltaic power generation, various load powers, the state of charge (SoC) of the energy storage battery, real-time time-of-use pricing, and safety constraint indicators received from the manager. The executor's actions are specific energy dispatch instructions, such as setting the charging and discharging power of the energy storage system or adjusting the total charging power of electric vehicle charging stations. The executor's reward function is a composite signal, the main part of which reflects the external reward for immediate economic benefits. Additionally, when its behavior remains within the safety boundaries set by the manager, it receives additional positive rewards; exceeding these boundaries incurs penalties.
[0047] This embodiment provides a detailed description of the core elements involved in this step:
[0048] Step S201, design the manager's State-Action-Reward (SAR) space.
[0049] State space:
[0050] The cumulative equivalent aging amount of the transformer ( This is a dimensionless numerical value representing the cumulative wear and tear of the transformer's insulation material since it was put into operation. For example, the initial value is 0, and when it accumulates to 1, it indicates the end of the insulation's lifespan. This value is determined by... Figure 4 The physical model shown is updated at each decision point on a slow timescale.
[0051] Energy storage battery state of health (SoH): expressed as a percentage, with 100% representing the factory condition. A value below 80% is generally considered the end of the battery's lifespan. This value is determined by... Figure 5 The physical model shown is updated.
[0052] Macroeconomic forecast: This could be a vector that includes the daily total photovoltaic power generation forecast for the coming week, the average temperature forecast, and the electricity price trend (e.g., the expected distribution of high-price periods).
[0053] Action Space: The manager's actions are a vector containing multiple constraint values. ,For example .in This is the upper limit of the transformer winding hot spot temperature for the coming week (unit: °C). This is the allowable degradation budget (in %) for the energy storage battery's SoH. This is the cost budget for the equivalent total number of cycles. Managers can adjust these values to tighten or loosen constraints on implementers.
[0054] Reward function: Manager's reward The formula for calculating at the end of a slow timescale period (such as a week) can be designed as follows: .in and This refers to the actual amount of transformer aging and battery health degradation that occur within the cycle. and It is a very high weighting coefficient representing the value of equipment assets; Under this constraint, the weight of the total economic benefits created by the executor is... The reward is relatively low. This reward function clearly shows that the manager's primary task is to prevent equipment aging, and only secondarily to consider economic efficiency.
[0055] Step S202, design the executor's state-action-reward (SAR) space.
[0056] State space (State): The state vector of the executor At each fast timescale decision point Update The key lies in the actions of the manager. The safety boundary is a static component of the executor's state space, and all decisions of the executor must take it as a prerequisite.
[0057] Action space: The actions of the executor. It is a specific power command vector, for example , representing the charging and discharging power of the energy storage and the total power allocated to the charging pile, respectively.
[0058] Reward function: The reward for the executor. At each fast timescale decision point The calculation can be performed immediately afterward, and the formula can be designed as a compound form: .in, It refers to direct economic benefits (or costs). It is a penalty term, for example, if the current action causes the transformer temperature predicted by the physical model to be higher than expected. Exceeding the manager's setting ,but A large negative value is applied; otherwise, zero or a small positive incentive is applied. This design ensures that implementers will avoid crossing the safety line at all costs, because the penalty far outweighs any potential short-term gains.
[0059] Step S300: Execute the closed-loop optimization process of the hierarchical reinforcement learning control architecture.
[0060] This process iterates continuously. At a slow-timescale decision point, the manager generates a macro-level safety objective for the next scheduling cycle (e.g., the next week) based on observed equipment performance trends or long-term health status assessed on an annual basis, and issues this objective as a rigid constraint to the executors. In subsequent scheduling cycles, at each fast-timescale decision point, the executors generate and execute specific energy scheduling actions based on their observed real-time operating status and received safety boundaries. The results of these actions directly generate short-term economic benefits and, through physical models, influence the long-term safety status of transformers and energy storage batteries. The system feeds back the resulting economic benefits and safety status changes to both the manager and the executors to update their respective policy networks, thereby achieving co-evolution and continuous optimization of the manager's safety strategy and the executors' economic strategy.
[0061] Please refer to Figure 3The specific interaction details of this closed-loop optimization process are as follows:
[0062] Step S301, slow timescale security policy iteration loop (manager loop).
[0063] This cycle operates on a weekly basis. At midnight every Monday, the manager's policy network is activated. It reads the accumulated data from the past few weeks. The SoH data sequence is analyzed to determine its changing trends. If the SoH decay rate exceeded expectations last week, even if it's still within budget, management might conclude that the current constraints are too lenient. Therefore, it outputs a set of stricter safety boundaries through its policy network. (For example, reduce) (The value). This new security boundary. Once issued and locked, it becomes a rigid constraint in the executor's strategy for the entire following week. The manager's strategy network (e.g., a Transformer-based sequence model) is then determined based on the previous week's final reward. Perform a gradient update.
[0064] Step S302, fast timescale economic decision-making iterative loop (executor loop).
[0065] The cycle operates in 15-minute intervals. At any 15-minute decision point during the week, the executor's policy network is activated. It reads the current real-time state. (This includes documents issued and locked by the administrator) Its policy network (e.g., a Soft Actor-Critic, the Actor part of a SAC network) outputs a specific action. The action is sent to the system controller for execution. Fifteen minutes later, the system calculates the economic reward based on the actual grid interaction power and electricity price. The penalty term was calculated through physical model simulation. The total reward is obtained by combining the two. The executor uses this instant reward. and new status This process is repeated 96 times a day to update the policy network and value network. This allows the implementer to quickly adapt to changing operating conditions and electricity prices, optimizing while ensuring safety.
[0066] Step S303: Implementation of co-evolution technology between managers and executors.
[0067] The co-evolution described in this application is achieved through a defined data processing chain and an experience-building mechanism, which specifically includes the following technical steps:
[0068] First, establish a data aggregation and experience building module that synchronizes with the manager's decision-making cycle. This module is activated at the end of a slow-timescale cycle (e.g., one week). Its core function is to process and refine all operational data generated by implementers on a fast timescale within this cycle, preparing input for the manager's strategy updates.
[0069] Second, the module collects and processes the complete interaction sequence of the executor. Within a slow timescale period, the module continuously records the interaction data of all fast timescale decision points generated by the executor (e.g., a total of 96 points / day * 7 days = 672 decision points per week). This data constitutes an interaction sequence that includes at least: the real-time status of each decision point. Actions taken by the executor and the direct economic rewards brought about by this action. .
[0070] Third, the module performs data aggregation and high-level state transition calculations. At the end of the cycle, the module performs the following calculations on the collected interaction sequences:
[0071] • Calculate cumulative economic returns: Sum all economic rewards within the period to obtain the total economic returns. .
[0072] • Calculate cumulative physical losses: Determine the state at each decision point within the period. and actions The values are input into the high-fidelity physical model of the transformer and energy storage battery to calculate the step size of each fast timescale. Equivalent aging increment within and the amount of decline in health status Then, the increments over the entire cycle are summed to obtain the total physical loss index. and .
[0073] Fourth, the module constructs a complete high-level experience tuple and stores it in the manager's experience replay pool. Using the above calculation results, the module constructs an experience tuple specifically for training the manager's policy network, in the format (S... M A M , R M , S' M ),in:
[0074] • High-level status This refers to the device health status at the beginning of this slow timescale period, i.e. .
[0075] • High-level actions It is the strategic security boundary issued by the manager at the beginning of this cycle. .
[0076] • High-level rewards : is the value calculated according to the first reward function described in this application, that is, based on the above calculation. , and The combined value.
[0077] • Next high-level status This refers to the equipment health status at the end of this slow timescale period, i.e. .
[0078] Through the above steps, the macroeconomic and physical consequences of thousands of specific micro-operations performed by the executor over a long period are precisely abstracted and encapsulated into an experiential sample meaningful to the manager for training their long-term strategy. Once this sample is stored in the manager's dedicated experience replay pool, the manager's policy network can perform gradient updates accordingly, thereby learning the mapping relationship between different safety boundary settings (high-level actions) and the final long-term benefits (high-level rewards), achieving iterative optimization of the strategy. This mechanism of converting low-frequency data from the bottom layer to low-frequency effective experience at the top layer constitutes the core technical means for the co-evolution of both.
[0079] In other words, this application does not involve the manager agent unilaterally imposing constraints on the executors from top to bottom. Instead, after N execution cycles, the executors provide feedback on training data to the manager, enabling the manager's model and optimization method to generate more detailed optimizations, which is the synergistic effect described in step 303.
[0080] Simulation tests show that if managers impose constraints unilaterally from top to bottom without feedback from executors based on real-world performance data, managers' optimization strategies may become overly strict.
[0081] Taking the response to prolonged periods of high temperatures as an example, a system based on a traditional single-layer economic optimization method might frequently allow energy storage systems to discharge at high rates to cope with high electricity prices caused by the heat, while simultaneously allowing charging stations to operate at full capacity. This would result in transformers being continuously overloaded in the scorching environment, causing their equivalent aging to increase dramatically within just a few days. However, using the method described in this application, after receiving a forecast of high temperatures for the coming week (as part of the macro-environmental prediction), senior management proactively outputs a stricter safety boundary than under normal operating conditions, such as temporarily lowering the upper limit of transformer winding hotspot temperatures. Upon receiving this tighter constraint, lower-level implementers, even when faced with the temptation of high electricity prices, are forced to adopt more conservative scheduling strategies, such as reducing the energy storage discharge rate and moderately reducing charging load. In this way, this application effectively avoids catastrophic aging of core assets under extreme operating conditions at the cost of a small amount of short-term gains, fully demonstrating the robustness and forward-looking nature of its decision-making.
[0082] In a preferred embodiment of this application, a thermodynamic model based on the IEEE Std C57.91 standard is adopted to more accurately quantify and set the safety boundary of the transformer.
[0083] like Figure 4 As shown, the input to this model includes the real-time load current of the transformer acquired from the SCADA system. and ambient temperature The model includes pre-configured transformer nameplate parameters (such as rated capacity, oil characteristics, cooling method, etc.). Internally, the model dynamically calculates the top-layer oil temperature in real-time by solving a set of differential equations. This allows for the calculation of the hottest temperature of the winding. Subsequently, Substituting into the formula for the insulation aging acceleration factor based on the Arrhenius equation Finally, through the analysis of... Integrate over time to obtain the cumulative equivalent aging amount. This physical model will determine the specific load scheduling behavior of the executor (affecting...). ) and the macro-level safety indicator of transformer long-term lifespan ( A precise, non-linear quantitative relationship was established between them, providing managers with scientific guidance. The constraint provides a physical basis and is far more precise and effective than the "must not exceed rated capacity" constraint in existing technologies.
[0084] In another preferred embodiment of this application, a semi-empirical battery physics model is used to more accurately quantify and set the safety boundary of the energy storage battery.
[0085] A typical semi-empirical battery calendar aging model is as follows:
[0086] Capacity loss due to calendar aging ( It is usually strongly correlated with time and temperature, and can be described using an extended form based on the Arrhenius equation:
[0087]
[0088] in, It refers to the pre-factor. For calendar aging activation energy, It is the ideal gas constant. It is the absolute temperature of the battery. It refers to storage or runtime.
[0089] like Figure 5 As shown, the core of this model is a function capable of predicting the decay of the state of health (SoH). Its inputs include the battery operating current, which is directly determined by the executor's scheduling instructions. The depth of discharge (DoD) obtained by current integration, and the internal battery temperature measured by the battery management system (BMS). The model incorporates empirical formulas for multiple degradation mechanisms. For example, the term describing calendar aging is related to time and temperature, while the term describing cycle aging is related to cycle number, DoD (Domain of Degradation), charge-discharge rate (C-rate), and temperature. By comprehensively calculating the capacity decay caused by these factors, the model can accurately predict that a high-rate deep discharge cycle performed at high temperature will cause significantly more damage to the SoH (Solar Hydraulic Energy) than a shallow charge-discharge cycle performed at a suitable temperature. This allows managers to make informed decisions. or Constraints can truly reflect the operating modes that are beneficial to battery health, guiding operators to make more "gentle" charging and discharging decisions.
[0090] at last, Figure 6 The beneficial effects of the proposed method in simulation experiments are demonstrated. The scatter plot shows the cumulative economic benefits per day on the horizontal axis and the cumulative equivalent aging of the transformer per day on the vertical axis. Two distinct clusters of points are visible in the plot. The point set labeled "traditional methods" (e.g., single-layer reinforcement learning with only economic optimization as the goal) exhibits a wide and high distribution of economic benefits, but also generally high aging amounts, even showing extreme values. In contrast, the point set labeled "the proposed method" shows a slight contraction in economic benefits, but remains very stable overall, and its corresponding aging amount is strictly controlled at an extremely low level. This eloquently demonstrates that the proposed method, by sacrificing negligible short-term economic benefits, achieves a fundamental improvement in the long-term safety of the equipment, ultimately optimizing the total cost of ownership over the entire lifecycle.
[0091] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An energy optimization method based on deep reinforcement learning, applied to an energy system including photovoltaics, energy storage, charging facilities, and transformers, characterized in that, The method includes: Step 1: Construct a hierarchical reinforcement learning control architecture, which includes a high-level policy agent that makes decisions on a slower time scale and a low-level policy agent that makes decisions on a faster time scale. The high-level policy agent is used to ensure the long-term operational security of the core power equipment in the energy system, and the low-level policy agent receives the strategic security boundary as a mandatory constraint in its decision-making environment and pursues the short-term economic efficiency of the energy system under this constraint. Step 2: The high-level strategic agent generates one or more future-oriented strategic security boundaries based on the long-term health status indicators of the core power equipment, and distributes the strategic security boundaries to the low-level strategic agent. Step 3: The underlying policy agent receives the policy security boundary in its state space, and at each decision point on the faster time scale, generates and executes energy scheduling actions to control the energy flow within the energy system based on real-time operating parameters and the policy security boundary. The state space of the high-level strategy agent includes the long-term health status indicators of the core power equipment. The long-term health status indicators include the cumulative equivalent aging of the transformer calculated based on the transformer physical model, and / or the energy storage battery health status SoH calculated based on the battery physical model. The action space of the high-level strategic agent is the set of strategic security boundaries, which include: the upper limit of the allowable temperature of the transformer winding hot spot in a future scheduling cycle, and / or the allowable decay rate budget of the energy storage battery health state SoH, and / or the budget of the equivalent full cycle number that the energy storage battery is allowed to consume. The policy network of the high-level policy agent adopts a Transformer-based neural network architecture to learn and process long-distance dependencies in the time-series data of the long-term health status indicators of the core power equipment using its self-attention mechanism.
2. The method according to claim 1, characterized in that, The state space of the underlying policy agent includes the real-time operating parameters and the policy security boundary received from the higher-level policy agent, wherein the real-time operating parameters include photovoltaic power generation, load power, energy storage battery state of charge (SoC) and real-time electricity price. The action space of the underlying policy agent is a set of energy scheduling actions, which include the charging and discharging power setting value of the energy storage system and / or the total charging power setting value of the charging facility.
3. The method according to claim 1, characterized in that, The high-level policy agent is trained using a first reward function, the value of which is determined by subtracting the asset depreciation cost calculated from the decline in the health status of the core power equipment from the economic benefits generated by the low-level policy agent over a slower time scale period. The underlying policy agent is trained using a second reward function, which is a weighted sum of an economic reward term reflecting short-term economic gains and a boundary penalty term reflecting compliance with the strategic security boundary. The boundary penalty term is activated as a negative value when the behavior of the underlying policy agent causes the system state to violate the strategic security boundary. The boundary penalty term is implemented using a constraint reinforcement learning algorithm based on Lagrange relaxation. Specifically, the boundary penalty term is multiplied by a learnable Lagrange multiplier and then added to the loss function of the underlying policy agent. The Lagrange multiplier is dynamically updated through gradient ascent based on the degree of violation of the policy security boundary to balance economic objectives and security constraints.
4. The method according to claim 1, characterized in that, The strategic security boundary is a set of quantifiable physical state thresholds that are directly related to the physical degradation mechanism of the core power equipment.
5. The method according to claim 3, characterized in that, The transformer physical model is a thermodynamic model based on the IEEE StdC57.91 standard. The model calculates the top oil temperature and winding hot spot temperature of the transformer based on the real-time load factor and ambient temperature, and cumulatively calculates the equivalent aging factor of the insulation material according to the Arrhenius equation, thereby obtaining the cumulative equivalent aging amount of the transformer. The battery physical model is a semi-empirical or electrochemical model. Based on the operating current, depth of discharge, operating temperature, and charge / discharge rate of the energy storage battery, the model quantifies the capacity decay caused by one or more battery degradation mechanisms, such as lithium deposition, increased stress in the electrode material structure, or abnormal growth of the solid electrolyte interface film, thereby obtaining the energy storage battery's state of health (SoH).
6. The method according to claim 5, characterized in that, The calculation of the first reward function is based on the cumulative results over one or more of the slower time scale periods, and its value is dominated by the asset depreciation cost item calculated from the decline in the health status of the core power equipment; The second reward function is calculated based on a single decision point on the faster time scale, and its value is dominated by economic reward terms that reflect short-term economic gains.
7. The method according to claim 1, characterized in that, The policy network of the underlying policy agent is configured not to directly output the energy scheduling action, but to output one or more unconstrained control variables, and then generate the energy scheduling action that satisfies the physical constraints of power balance through a fixed, differentiable function based on the unconstrained control variables and the system power parameters in the real-time operating parameters.
8. The method according to claim 1, characterized in that, The method also includes a collaborative training step, wherein operational data including changes in device health status generated by the lower-level policy agent interacting with the environment within the strategic security boundary is used to evaluate and update the policy of the higher-level policy agent; after the higher-level policy agent is updated, it outputs a new strategic security boundary to guide the subsequent decision-making of the lower-level policy agent.
9. An energy optimization system applied to an energy system comprising photovoltaics, energy storage, charging facilities, and transformers, characterized in that, The system includes: One or more processors; Instructions stored on a non-transitory computer-readable medium; When the instructions are executed by the one or more processors, they enable the system to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Power distribution network dual-time scale voltage control method based on multi-agent deep reinforcement learning
CN118017518A