Microgrid energy management system and energy management method based on multi-scale frame stacking reinforcement learning
By using a multi-scale frame stacking reinforcement learning system, combined with frame stacking technology and multi-scale attention mechanism, the problem of microgrid energy management system in dealing with the intermittency and uncertainty of renewable energy has been solved. This has enabled efficient battery charging and discharging strategies and closed-loop control, improving the operational stability and economy of the microgrid.
Patent Information
- Application Number
- CN202510947109.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-31
AI Technical Summary
Existing microgrid energy management systems lack effective mechanisms for capturing historical time-dependent features and multi-scale energy data processing capabilities, making it difficult to cope with the intermittency and uncertainty of renewable energy.
A microgrid energy management system based on multi-scale frame stacking reinforcement learning is adopted. The microgrid dynamics are modeled by Markov decision process, and the cross-scale time information fusion features are generated by combining frame stacking technology and multi-scale attention mechanism for battery charging and discharging control.
It achieves accurate identification of the time-dependent characteristics of energy systems, improves decision-making accuracy and adaptability, increases battery utilization, reduces constraint violation rates, and has good economic efficiency and reliability, adapting to different renewable energy scenarios.
Smart Images

Figure CN120876157A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microgrid energy management technology, and in particular to a microgrid energy management system and energy management method based on multi-scale frame stacking reinforcement learning. Background Technology
[0002] The large-scale application of renewable energy is showing a significant growth trend. However, its inherent intermittency and uncertainty pose serious challenges to the stability of power grid operation and economic dispatch. Microgrid technology, with its synergistic capabilities in localized energy production, storage, and load management, has become a core solution for the efficient integration of distributed energy resources, while the microgrid energy management system (EMS) plays a crucial role in coordinating diverse resources to meet operational objectives and technical constraints.
[0003] Traditional microgrid energy management methods suffer from multiple technical bottlenecks: metaheuristic algorithms lack theoretical optimality guarantees; mathematical programming methods are prone to the curse of dimensionality in complex scenarios; uncertainty handling techniques such as stochastic programming and robust optimization face the contradiction of high computational costs or overly conservative strategies; and control theory-based methods are highly dependent on prediction accuracy and suffer from dimensionality scalability problems. As the complexity of microgrid systems increases, these methods struggle to effectively address the computational demands of sequential decision-making under uncertain environments. Deep reinforcement learning (DRL), as a model-free optimization paradigm, models energy management as a Markov decision process, achieving optimal policy solutions through interactive learning between agents and the environment. This has led to microgrid applications such as Q-learning battery scheduling, DQN energy storage management, and DDPG continuous control.
[0004] Proximal policy optimization (PPO) algorithms are widely used due to their high sample efficiency and strong training stability. However, existing DRL implementations suffer from temporal inference mechanism defects: they rely on Markov representations that only capture immediate states, lack historical pattern expansion mechanisms, and use uniform temporal granularity to process data, ignoring the multi-scale dynamic characteristics of energy systems. Although some studies have attempted to integrate recurrent neural networks with DRL, they have introduced new problems such as deteriorated training stability and a surge in computational complexity. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem that existing microgrid energy management systems are unable to cope with the intermittency and uncertainty of renewable energy due to the lack of an effective historical time-dependent feature capture mechanism and multi-scale energy data processing capabilities.
[0006] To address the aforementioned technical problems, this invention provides a microgrid energy management system and energy management method based on multi-scale frame stacking reinforcement learning, wherein the microgrid energy management system based on multi-scale frame stacking reinforcement learning includes: a microgrid system modeling module, a multi-scale temporal feature enhancement module, and an energy scheduling decision module;
[0007] The microgrid system modeling module uses a Markov decision process to mathematically model the microgrid dynamics and provides real-time system status information, including power generation and power consumption.
[0008] The multi-scale temporal feature enhancement module receives the system state information and combines frame stacking technology and multi-scale attention mechanism to obtain cross-scale temporal information fusion features.
[0009] The energy dispatch decision module generates a battery charging and discharging control strategy based on the cross-scale time information fusion characteristics, executes the control strategy, and feeds back the execution result and the updated system state to the microgrid system modeling module to form a closed-loop control.
[0010] In one embodiment of the present invention, the multi-scale temporal feature enhancement module is used to receive the system state information and, by combining frame stacking technology and a multi-scale attention mechanism, obtain cross-scale temporal information fusion features, including:
[0011] The system state information acquired at time t is defined as a state vector o including d-dimensional state information. t Set a fixed-length historical buffer of length h to store the state vector o for h consecutive time steps in chronological order. t ;
[0012] The state vector o of h consecutive time steps in the historical buffer. t Stacked into an enhanced matrix
[0013]
[0014] The enhanced state vector is processed by different sampling time intervals. Downsampling is performed to generate features at multiple temporal granularities. Each feature at a temporal granularity is processed independently by a self-attention mechanism. Then, a cross-attention mechanism is used to enable cross-scale information interaction between different feature vectors, resulting in a feature vector that integrates multi-temporal information.
[0015] By adding the feature vector fused with multi-time information to the weighted original state vector of the current time step, the cross-scale time information fusion feature is obtained.
[0016] In one embodiment of the present invention, the update rule of the history buffer is:
[0017] Each time there is a new state vector o t During input, remove the oldest historical state vector o. t-h Then, the new state vector o t Insert at the end of the history buffer.
[0018] In one embodiment of the present invention, the enhanced state vector O is sampled at different time intervals. t Downsampling is performed to generate features at multiple time granularities, including:
[0019] The multi-time granularity features include short-term time granularity features, medium-term time granularity features, and long-term time granularity features;
[0020] When the features of the short-term time granularity are obtained At that time, the state vector o of h consecutive time steps is used directly. t Features constituting the short-term time granularity
[0021]
[0022] When the features of the intermediate time granularity are obtained At that time, set the intermediate sampling interval k. m Through the intermediate sampling interval k m For the enhanced state vector O t Downsampling is performed to obtain the features at the intermediate time granularity.
[0023]
[0024] When the long-term time granularity features are obtained At that time, set the long-term sampling interval k. l Through the long-term sampling interval k l For the enhanced state vector O t Downsampling is performed to obtain the features at the long-term time granularity.
[0025]
[0026] In one embodiment of the present invention, a cross-scale temporal information fusion feature is obtained by adding the feature vector fused with multi-temporal information to the weighted original state vector of the current time step, including:
[0027] The feature vector fused with multi-temporal information is combined with the fully connected weight matrix W r The weighted original state vector o at the current time step t Adding them together yields the cross-scale temporal information fusion feature F. final :
[0028] F final =F fused +W r ·ot .
[0029] In one embodiment of the present invention, the energy dispatch decision module generates a battery charging and discharging control strategy based on the cross-scale time information fusion features, including:
[0030] The cross-scale temporal information fusion feature is used as the state parameter s. t Input to policy network π θ In (a|s), a reward function is established based on economic benefits and operating costs, with the optimization objective being to maximize the reward function value, thereby generating a normalized continuous action a. t ∈[-1,1], mapped to the target charge / discharge quantity Develop a battery charging and discharging control strategy;
[0031] Among them, R max Indicates the battery's maximum charge and discharge power, a t =1 represents the maximum charge, a t =-1 represents the maximum discharge.
[0032] In one embodiment of the present invention, the reward function is:
[0033]
[0034] Where, r t As a reward value, For economic benefits, This indicates revenue from internal power supply services. This represents the cost / revenue of the power grid transaction at time t. Indicates the cost of renewable energy generation. Indicates battery maintenance costs; This is an operational penalty item.
[0035] In one embodiment of the present invention, the operational penalty item The methods for obtaining it include:
[0036] The mathematical model established by the microgrid system modeling module always satisfies the power balance constraint:
[0037]
[0038] The operational penalty is obtained based on the power balance constraint. as follows:
[0039]
[0040] in, Indicates photovoltaic power generation capacity. Indicates wind power generation capacity. D represents the power traded on the power grid. t Indicates load demand, ΔB t Indicates the battery charge / discharge amount, η c Indicates charging efficiency. An indicator function representing the charging status, which takes a value of 1 only when the charging status is active, and 0 in other states; An indicator function representing the discharge state, which takes a value of 1 only during the discharge state and 0 in other states, η d Indicates discharge efficiency; ΔB t >0 indicates that the system is in a charging state, ΔB t <0 indicates that the system is in a discharge state.
[0041] Based on the same inventive concept, the present invention also provides an energy management method, which includes the following steps:
[0042] Step S1: Perform mathematical modeling of the microgrid dynamics using a Markov decision process to provide real-time system status information, including power generation and power consumption;
[0043] Step S2: Receive the system status information, and combine frame stacking technology and multi-scale attention mechanism to obtain cross-scale temporal information fusion features;
[0044] Step S3: Based on the cross-scale time information fusion features, generate a battery charging and discharging control strategy, execute the control strategy, and feed back the execution result and the updated system state to step S1 to form a closed-loop control.
[0045] The present invention also provides a computer storage medium storing a computer software product, the computer software product including a plurality of instructions for causing a computer device to perform the steps of the energy management method described above.
[0046] The technical solution of the present invention has the following advantages compared with the prior art:
[0047] This invention accurately models microgrid dynamics through Markov decision processes, integrates historical observation data using frame stacking technology to capture time-dependent features, and combines a multi-scale attention mechanism to process short-term, medium-term, and long-term energy patterns in parallel, achieving cross-scale information fusion. Finally, reinforcement learning is used to generate the optimal battery charging and discharging strategy, forming a closed-loop control. Its advantages lie in effectively solving the problems of insufficient temporal reasoning and low efficiency in handling uncertainty in traditional methods. It can accurately identify the time-dependent features of energy systems, improve decision-making accuracy and adaptability, significantly improve returns, reduce constraint violation rates, and increase battery utilization in different renewable energy scenarios. Furthermore, its computational efficiency meets the requirements of real-time operation, demonstrating good economic efficiency, reliability, and practicality. Attached Figure Description
[0048] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein...
[0049] Figure 1 This is a schematic diagram of a microgrid energy management system structure based on multi-scale frame stacking reinforcement learning provided in an embodiment of the present invention;
[0050] Figure 2 This is a cross-scale temporal information fusion feature F provided in an embodiment of the present invention. final A flowchart illustrating the acquisition method;
[0051] Figure 3 This is a schematic diagram of a microgrid energy management method provided in an embodiment of the present invention;
[0052] Explanation of reference numerals in the accompanying drawings: 100, Microgrid system modeling module; 200, Multi-scale temporal feature enhancement module; 300, Energy dispatching decision module. Detailed Implementation
[0053] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0054] Example 1:
[0055] like Figure 1 As shown, the present invention provides a microgrid energy management system based on multi-scale frame stacking reinforcement learning, comprising: a microgrid system modeling module 100, a multi-scale temporal feature enhancement module 200, and an energy scheduling decision module 300;
[0056] The microgrid system modeling module 100 uses a Markov decision process to mathematically model the microgrid dynamics and provides real-time system status information, including power generation and power consumption.
[0057] The multi-scale temporal feature enhancement module 200 receives the system state information and combines frame stacking technology and multi-scale attention mechanism to obtain cross-scale temporal information fusion features.
[0058] The energy dispatch decision module 300 generates a battery charging and discharging control strategy based on the cross-scale time information fusion features through a reinforcement learning algorithm, executes the control strategy, and feeds back the execution result and the updated system state to the microgrid system modeling module 100 to form a closed-loop control.
[0059] As can be seen from the above technical solution, this invention employs Markov decision processes to accurately model microgrid dynamics, deeply integrates cross-scale temporal information using frame stacking technology and multi-scale attention mechanisms, and achieves closed-loop control and strategy optimization through reinforcement learning. It can not only accurately perceive the microgrid state and efficiently respond to energy fluctuations, but also adaptively optimize battery charging and discharging strategies, improving renewable energy absorption rates, reducing costs, and extending equipment lifespan. Furthermore, it possesses significant technological innovation, scalability, and practical application value, demonstrating substantial advantages over traditional methods.
[0060] Since a microgrid system consists of four parts: distributed generation units (photovoltaic and wind power), battery energy storage system (BESS), load and grid interaction, the microgrid system modeling module 100 uses a Markov decision process to mathematically model the microgrid dynamics and obtain hourly actual data (renewable energy generation curves, load demand, and electricity prices).
[0061] Its key mathematical models include:
[0062] Distributed generation power G t It is the sum of photovoltaic and wind power, that is Indicates photovoltaic power generation capacity. Indicates wind power generation capacity;
[0063] The state change formula for a battery energy storage system (BESS) is: B t+1 =B t +ΔB t B t Indicates the current state of charge, ΔB t Indicates the charge / discharge amount;
[0064] Power grid trading Must meet Indicates the minimum grid trading power. Indicates the maximum grid trading power;
[0065] Microgrid systems operate while meeting power balance constraints:
[0066] D t Indicates load demand, ΔB t Indicates the battery charge / discharge amount, η c Indicates charging efficiency. An indicator function representing the charging status, which takes a value of 1 only when the charging status is active, and 0 in other states; An indicator function representing the discharge state, which takes a value of 1 only during the discharge state and 0 in other states, η d Indicates discharge efficiency;
[0067] Economic objective: Minimize total cost The optimization goal is It is also the basis for the reward function used by the energy dispatch decision module 300 to generate the battery charging and discharging control strategy; This represents the cost / revenue of the power grid transaction at time t. Indicates the cost of renewable energy generation. This indicates the cost of battery maintenance.
[0068] Furthermore, in this embodiment, as Figure 2 As shown, the multi-scale temporal feature enhancement module 200 is used to receive the system state information and, by combining frame stacking technology and multi-scale attention mechanism, obtain cross-scale temporal information fusion features, including:
[0069] The system state information acquired by the microgrid system modeling module 100 at time t is defined as a state vector o including d=7-dimensional state information. t :
[0070]
[0071] in, Indicates photovoltaic power generation capacity. D represents the wind power generation capacity. t Indicates the internal power demand; Indicates the electricity price. Indicates the electricity purchase price; B t H represents the current state of charge of the battery. t Indicates time information;
[0072] Set a fixed-length history buffer of length h to store the state vectors o for h consecutive time steps in chronological order. t ;
[0073] The state vector o of h consecutive time steps in the historical buffer. t Stacked into an enhanced matrix
[0074]
[0075] The enhanced state vector is processed by different sampling time intervals. Downsampling is performed to generate features at multiple time granularities, including:
[0076] The multi-time granularity features include short-term time granularity features, medium-term time granularity features, and long-term time granularity features;
[0077] When the features of the short-term time granularity are obtained At that time, the original time resolution is used directly, that is, the state vector o for h consecutive time steps.t Features constituting the short-term time granularity
[0078]
[0079] When the features of the intermediate time granularity are obtained At that time, set the intermediate sampling interval k. m Within a one-day timeframe, through the intermediate sampling interval k m For the enhanced state vector O t Downsampling is performed to obtain the features at the intermediate time granularity.
[0080]
[0081] When the long-term time granularity features are obtained At that time, set the long-term sampling interval k. l Through the long-term sampling interval k l For the enhanced state vector O t Downsampling is performed to obtain the features at the long-term time granularity.
[0082]
[0083] Features at each temporal granularity are independently processed through a self-attention mechanism. After processing, the softmax function is used to selectively focus on different parts of the input sequence, enhancing the weights of key temporal features, and obtaining the original branch features F for each time scale. orig Where Q is the query matrix, K is the key matrix, and V is the value matrix, i.e., a linear mapping of the input features; d k Indicates the dimension of the attention matrix;
[0084] Cross-attention mechanisms facilitate information exchange between features at different time scales. For example, attention is calculated by cross-compiling short-term and long-term features, enabling short-term decisions to perceive long-term trends and long-term planning to incorporate short-term dynamics, resulting in a cross-scale fused feature F. cross The original branch feature F orig With cross-scale fusion features F cross After splicing, it is processed through the fully connected layer W f and bias b f Perform a linear transformation, then pass through an activation function. Generate a feature vector F that integrates multi-temporal information. fused :
[0085]
[0086] To ensure that the current system state has a direct impact on decision-making, the feature vector F, which integrates multi-time information, is connected via residual connections. fused With the weight matrix W of the fully connected layer r The weighted original state vector o at the current time step t Linear combination yields the cross-scale temporal information fusion feature F. final :
[0087] F final =F fused +W r ·o t .
[0088] Furthermore, in this embodiment, the update rule for the history buffer is:
[0089] Each time there is a new state vector o t During input, remove the oldest historical state vector o. t-h Then, the new state vector o t Insert at the end of the history buffer.
[0090] Furthermore, in this embodiment, the energy dispatch decision module 300 is based on the cross-scale time information fusion feature F final Generate a battery charge / discharge control strategy, including:
[0091] Based on the Proximal Policy Optimization (PPO) algorithm, the cross-scale temporal information fusion feature F is... final As a state parameter s t Input to policy network π θ In (a|s), a reward function is established based on economic benefits and operating costs, with the optimization objective being to maximize the reward function value, thereby generating a normalized continuous action a. t ∈[-1,1], mapped to the target charge / discharge quantity Develop a battery charging and discharging control strategy;
[0092] Among them, R max Indicates the battery's maximum charge and discharge power, a t =1 represents the maximum charge, a t =-1 represents the maximum discharge.
[0093] Specifically, a reward function is established based on economic benefits and operating costs, as follows:
[0094]
[0095] Where, r t As a reward value, For economic benefits, This indicates revenue from internal power supply services. This represents the cost / revenue of the power grid transaction at time t. Indicates the cost of renewable energy generation. Indicates battery maintenance costs; This is an operational penalty item.
[0096] Furthermore, the aforementioned operational penalty items The methods for obtaining it include:
[0097] The operational penalty term is obtained based on the mathematical model established by the microgrid system modeling module 100, which always satisfies the power balance constraint condition. as follows:
[0098]
[0099] in, Indicates photovoltaic power generation capacity. Indicates wind power generation capacity. D represents the power traded on the power grid. t Indicates load demand, ΔB t Indicates the battery charge / discharge amount, η c Indicates charging efficiency. An indicator function representing the charging status, which takes a value of 1 only when the charging status is active, and 0 in other states; An indicator function representing the discharge state, which takes a value of 1 only during the discharge state and 0 in other states, η d Indicates discharge efficiency; ΔB t >0 indicates that the system is in a charging state, ΔB t <0 indicates that the system is in a discharge state.
[0100] To verify the effectiveness of the proposed method (MF-PPO model), an integrated microgrid test system combining renewable energy and energy storage devices was used for evaluation. This system conforms to a typical medium-sized microgrid architecture, including photovoltaic panels, wind turbines, and a battery energy storage system (BESS), possessing a balanced renewable energy generation capacity and appropriate energy storage capacity. The experimental data used was a 2020 hourly measured dataset from a specific country, covering renewable energy generation curves, load consumption patterns, and electricity market prices. The dataset exhibits significant daily cycle, periodic patterns, and seasonal variations, providing real-world support for multi-scale time-series analysis.
[0101] The MF-PPO model is implemented based on a multi-scale temporal processing framework. It employs three temporal resolutions (1 hour, 6 hours, and 24 hours) to downsample historical data and constructs a multi-scale feature interaction network using a Transformer-based self-attention mechanism. Historical observation data is managed through a circular buffer, with a frame stack length of h and sampling interval parameters k for each time scale. m (Mid-term) and k l (Long-term) , to achieve fusion optimization of features at different resolutions through cross-attention mechanisms.
[0102] Table 1 presents the quantitative comparison results of the MF-PPO model with the standard PPO algorithm and the heuristic strategy under different renewable energy scenarios. The key conclusions are as follows:
[0103] (1) Wind-dominated scenarios:
[0104] MF-PPO Implementation The average daily return compared to the benchmark PPO It improves by 54.4%, and its multi-scale mechanism effectively captures the coupling relationship between short-term wind speed fluctuations and long-term weather cycles, optimizing the energy storage charging and discharging sequence to match the wind power output characteristics.
[0105] (2) Pure solar energy scenario:
[0106] This scenario is extremely challenging due to the concentrated power generation periods and long non-power generation intervals. MF-PPO remains... The positive returns, while the standard PPO and heuristic strategies respectively generated The losses were significant. Furthermore, the constraint violation rate of MF-PPO (23.0±9.0%) was 61.8% lower than that of PPO (60.3±35.9%), validating the effectiveness of multi-scale time awareness in improving operational reliability.
[0107] (3) Seasonal adaptability analysis:
[0108] Summer scenario: The model is synchronized with the daytime solar power generation mode, performing midday charging during the peak photovoltaic period when the electricity price is in the middle range, and accurately discharging during the peak electricity price period in the evening, achieving synergistic optimization of market arbitrage and load shaving.
[0109] Winter scenario: In response to intermittent wind power and fluctuating electricity prices, the model dynamically adjusts the charging strategy, charging with high time accuracy during the off-peak period and discharging during the peak period. Despite the increased volatility of renewable energy, it still maintains a high-efficiency price response capability.
[0110] Table 1
[0111] Scene method Earnings (€ / day) Battery utilization rate (%) Penalty rate (%) pure solar energy MF-PPO 18.0±13.0 54.6±0.8 23.0±9.0 PPO -12.5±5.6 52.8±3.5 60.3±35.9 Heuristic -143.8 39.6 20.5±0.0 pure wind energy MF-PPO 293.7±19.9 45.3±2.7 1.8±0.3 PPO 205.1±49.7 44.5±1.9 1.8±0.4 Heuristic 98.0 27.9 15.3±0.0 Solar-driven MF-PPO 199.0±9.3 55.2±1.0 1.9±0.2 (75% PV) PPO 155.0±13.6 54.0±1.7 2.4±0.2 Heuristic -42.5 28.8 18.2±0.0 Wind power dominant MF-PPO 325.4±27.1 50.5±1.9 1.5±0.1 (75% wind energy) PPO 210.7±78.8 47.2±2.3 1.9±0.4 Heuristic 134.0 29.9 16.8±0.0
[0112] From an engineering implementation perspective, the MF-PPO model achieves a 12-18% increase in revenue (based on a 10MW microgrid system, the additional annual revenue is approximately...). This approach simultaneously reduces the violation rate by 35%. Computational efficiency analysis shows that the model inference latency meets the real-time control requirements of microgrids (≤100ms / time step), supporting deployment in practical control systems. Through multi-scale time-series feature fusion and adaptive strategy optimization, this scheme provides an energy management paradigm that combines economy and reliability for microgrids with a high proportion of renewable energy.
[0113] Example 2:
[0114] Based on the same inventive concept as Embodiment 1, the present invention also provides an energy management method, such as... Figure 3 As shown, the method includes the following steps:
[0115] Step S1: Based on the microgrid system modeling module 100, the microgrid dynamics are mathematically modeled using a Markov decision process to provide real-time system status information, including power generation and power consumption.
[0116] Step S2: The multi-scale temporal feature enhancement module 200 receives the system state information and, by combining frame stacking technology and multi-scale attention mechanism, obtains cross-scale temporal information fusion features;
[0117] Step S3: Based on the cross-scale time information fusion characteristics, the energy scheduling decision module 300 generates a battery charging and discharging control strategy through the policy network of the near-end policy optimization (PPO) algorithm, executes the control strategy, and feeds back the execution result and the updated system state to step S1 to form a closed-loop control.
[0118] Example 3:
[0119] The present invention also provides a computer storage medium storing a computer software product, the computer software product including a plurality of instructions for causing a computer device to execute the steps of the energy management method described in Embodiment 2.
[0120] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0122] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0123] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0124] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A microgrid energy management system based on multi-scale frame stacking reinforcement learning, characterized in that, include: The microgrid system modeling module is used to mathematically model the dynamics of the microgrid through a Markov decision process and provide real-time system status information, including power generation and power consumption. The multi-scale temporal feature enhancement module is used to receive the system state information and combine frame stacking technology and multi-scale attention mechanism to obtain cross-scale temporal information fusion features. The system also includes an energy dispatch decision module, which generates a battery charging and discharging control strategy based on the cross-scale time information fusion features, executes the control strategy, and feeds back the execution results and the updated system state to the microgrid system modeling module to form a closed-loop control.
2. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 1, characterized in that, The multi-scale temporal feature enhancement module is used to receive the system state information and, by combining frame stacking technology and multi-scale attention mechanism, obtain cross-scale temporal information fusion features, including: The system state information acquired at time t is defined as a state vector o including d-dimensional state information. t Set a fixed-length historical buffer of length h to store the state vector o for h consecutive time steps in chronological order. t ; The state vector o of h consecutive time steps in the historical buffer. t Stacked into an enhanced matrix The enhanced state vector is processed by different sampling time intervals. Downsampling is performed to generate features at multiple temporal granularities. Each feature at a temporal granularity is processed independently by a self-attention mechanism. Then, a cross-attention mechanism is used to enable cross-scale information interaction between different feature vectors, resulting in a feature vector that integrates multi-temporal information. By adding the feature vector fused with multi-time information to the weighted original state vector of the current time step, the cross-scale time information fusion feature is obtained.
3. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 2, characterized in that, The update rules for the historical buffer are as follows: Each time there is a new state vector o t During input, remove the oldest historical state vector o. t-h Then, the new state vector o t Insert at the end of the history buffer.
4. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 2, characterized in that, The enhanced state vector O is obtained by sampling at different time intervals. t Downsampling is performed to generate features at multiple time granularities, including: The multi-time granularity features include short-term time granularity features, medium-term time granularity features, and long-term time granularity features; When the features of the short-term time granularity are obtained At that time, the state vector o of h consecutive time steps is used directly. t Features constituting the short-term time granularity When the features of the intermediate time granularity are obtained At that time, set the intermediate sampling interval k. m Through the intermediate sampling interval k m For the enhanced state vector O t Downsampling is performed to obtain the features at the intermediate time granularity. When the long-term time granularity features are obtained At that time, set the long-term sampling interval k. l Through the long-term sampling interval k l For the enhanced state vector O t Downsampling is performed to obtain the features at the long-term time granularity.
5. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 2, characterized in that, By adding the feature vector fused with multi-temporal information to the weighted original state vector of the current time step, a cross-scale temporal information fusion feature is obtained, including: The feature vector fused with multi-temporal information is combined with the fully connected weight matrix W r The weighted original state vector o at the current time step t Adding them together yields the cross-scale temporal information fusion feature F. final : F final =F fused +W r o t .
6. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 1, characterized in that, The energy dispatch decision module generates a battery charging and discharging control strategy based on the cross-scale time information fusion features, including: The cross-scale temporal information fusion feature is used as the state parameter s. t Input to policy network π θ In (a|s), a reward function is established based on economic benefits and operating costs, with the optimization objective being to maximize the reward function value, thereby generating a normalized continuous action a. t ∈[-1,1], mapped to the target charge / discharge quantity Develop a battery charging and discharging control strategy; Among them, R max Indicates the battery's maximum charge and discharge power, a t =1 represents the maximum charge, a t =-1 represents the maximum discharge.
7. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 6, characterized in that, The reward function is: Where, r t As a reward value, For economic benefits, This indicates revenue from internal power supply services. This represents the cost / revenue of the power grid transaction at time t. Indicates the cost of renewable energy generation. Indicates battery maintenance costs; This is an operational penalty item.
8. The microgrid energy management system based on multi-scale frame stacking reinforcement learning according to claim 7, characterized in that, The operational penalties The methods for obtaining it include: The mathematical model established by the microgrid system modeling module always satisfies the power balance constraint: The operational penalty is obtained based on the power balance constraint. as follows: in, Indicates photovoltaic power generation capacity. Indicates wind power generation capacity. F represents the power traded on the power grid. t Indicates load demand, ΔB t Indicates the battery charge / discharge amount, η c Indicates charging efficiency. An indicator function representing the charging status, which takes a value of 1 only when the charging status is active, and 0 in other states; An indicator function representing the discharge state, which takes a value of 1 only during the discharge state and 0 in other states, η d Indicates discharge efficiency; ΔB t >0 indicates that the system is in a charging state, ΔB t <0 indicates that the system is in a discharge state.
9. An energy management method, characterized in that, Includes the following steps: Step S1: Perform mathematical modeling of the microgrid dynamics using a Markov decision process to provide real-time system status information, including power generation and power consumption; Step S2: Receive the system status information, and combine frame stacking technology and multi-scale attention mechanism to obtain cross-scale temporal information fusion features; Step S3: Based on the cross-scale time information fusion features, generate a battery charging and discharging control strategy, execute the control strategy, and feed back the execution result and the updated system state to the microgrid system modeling module to form a closed-loop control.
10. A computer storage medium, characterized in that, The computer storage medium stores a computer software product, which includes several instructions for causing a computer device to execute the steps of the energy management method of claim 9.