Energy Management Method and Device for Vertical Axis Microwind Power Generation System Based on Reinforcement Learning

By using a reinforcement learning-based energy management method, dynamic environmental feature maps are generated and energy management models are constructed, which solves the problem of insufficient intelligence and adaptability in energy management of vertical axis micro wind power generation systems and achieves efficient and stable energy utilization.

CN121886606BActive Publication Date: 2026-07-17ZHONGAN JINLI (BEIJING) SAFETY PROD TECH RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHONGAN JINLI (BEIJING) SAFETY PROD TECH RES INST
Filing Date
2026-01-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing vertical axis micro wind power generation systems lack sufficient intelligence in energy management, have poor dynamic adaptability, low energy utilization, and are difficult to operate efficiently in complex environments.

Method used

An energy management method based on reinforcement learning is adopted to achieve intelligent and adaptive optimization of energy allocation strategies by generating dynamic environmental feature maps and constructing an energy management model. This method includes acquiring operational data to generate a three-dimensional feature field, constructing a state transition matrix and dynamically adjusting thresholds, and using a neural network learning mechanism to dynamically adjust the energy allocation strategy.

Benefits of technology

It significantly improves the system's energy utilization efficiency and stability in complex environments, enhances its ability to sense and respond to light wind conditions, and improves energy conversion efficiency and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121886606B_ABST
    Figure CN121886606B_ABST
Patent Text Reader

Abstract

This invention discloses an energy management method and device for a vertical axis micro-wind power generation system based on reinforcement learning, belonging to the field of new energy and intelligent control technology. The method acquires real-time operational data such as wind speed, power generation, and energy storage status of the system to generate a dynamic environmental feature map that integrates temporal and spatial gradients, accurately representing the spatiotemporal characteristics of wind speed changes. A pre-built energy management model is then invoked to process the feature map. This model constructs a state transition matrix and dynamically adjusts thresholds based on historical operational data, autonomously learning optimization strategies from historical experience through a simulated neural network learning mechanism. The target energy allocation strategy is determined based on the model output. This invention solves the technical problems of poor adaptability and low energy utilization in vertical axis micro-wind power generation systems in the wind energy industry, achieving intelligent energy management and stable, efficient system operation under complex and variable wind speed conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of new energy and intelligent control technology, specifically relating to an energy management method and device for a vertical axis micro wind power generation system based on reinforcement learning. Background Technology

[0002] With the continuous development of vertical axis wind power generation technology, micro-wind power generation systems have gradually become a research hotspot in the field of new energy due to their ability to operate efficiently under low wind speed conditions. However, existing vertical axis micro-wind power generation systems still have some shortcomings in energy management methods and devices, which affect the overall efficiency, stability, and intelligence level of the system.

[0003] The invention patent with announcement number CN102022258B discloses a vertical axis wind turbine with high wind energy utilization efficiency. Through optimized design of the blade and cross strut cross-sections, it significantly improves the turbine's power generation capacity under low wind speed conditions and increases the generation voltage under high wind speed conditions. However, this technical solution mainly focuses on mechanical structure optimization and lacks intelligent control of energy management in the power generation system. It is difficult to dynamically adjust the power generation strategy according to real-time wind speed changes, resulting in room for further improvement in energy utilization efficiency.

[0004] The invention patent with announcement number CN109980675B discloses a doubly-fed magnetic levitation vertical axis wind power generation system for flexible DC transmission and its control method. By employing a permanent magnet synchronous generator and a disc motor to form a doubly-fed structure, and combining this with intelligent control methods, it achieves optimized wind turbine speed and precise control of rated power output. However, the control method of this technical solution mainly relies on traditional control strategies (such as dual closed-loop cascade control), and its control accuracy and adaptability may be limited when facing nonlinear and highly uncertain micro-wind environments.

[0005] Existing vertical axis micro wind power generation systems still have limitations in energy management, such as insufficient intelligence, poor dynamic adaptability, and low energy utilization. Summary of the Invention

[0006] The purpose of this invention is to provide an energy management method and device for a vertical axis micro wind power generation system based on reinforcement learning. By using reinforcement learning algorithms, the energy management is made intelligent and adaptively optimized, thereby improving the energy utilization efficiency and stability of the system in complex environments and meeting the needs of the new energy field for efficient and intelligent micro wind power generation systems.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0008] A reinforcement learning-based energy management method for a vertical axis micro-wind power generation system includes: acquiring operational data of the vertical axis micro-wind power generation system, wherein the operational data is related to the real-time wind speed, power output, and energy storage status of the vertical axis micro-wind power generation system; generating a dynamic environmental feature map based on the operational data, wherein the dynamic environmental feature map is formed by a three-dimensional feature field characterizing the wind speed change trend and power generation efficiency change, wherein the three-dimensional feature field involves the temporal gradient and spatial distribution gradient of the wind speed; invoking a pre-built energy management model, processing the dynamic environmental feature map based on the energy management model to obtain a target energy allocation strategy to be adjusted, wherein the energy management model is used to determine the energy allocation scheme to be optimized based on the input dynamic environmental feature map; and generating a control signal to trigger the execution of the target energy allocation strategy.

[0009] In one embodiment, generating a dynamic environmental feature map based on the operational data includes: aggregating the operational data along the time dimension to form a time-series wind speed matrix, the time-series wind speed matrix including the wind speed status of each monitoring point within a continuous time slice; determining a time gradient component to characterize the frequency of wind speed changes based on the time-series wind speed matrix; determining a spatial gradient component to characterize the spatial distribution trend of wind speed based on the time-series wind speed matrix and logically adjacent monitoring points; and synthesizing the time gradient component and the spatial gradient component to form the three-dimensional feature field.

[0010] The synthesis of the temporal gradient components and spatial gradient components to form the three-dimensional feature field includes: synthesizing the temporal gradient components and spatial gradient components based on the following rules to form the three-dimensional feature field: the fused gradient vector is obtained by weighted summation of the temporal gradient components and spatial gradient components, and the weight coefficients of the weighted summation are determined according to the correlation between the severity of wind speed changes and power generation efficiency; the direction vector is obtained by normalizing the fused gradient vector, and the direction vector is used to indicate the main trend of wind speed changes.

[0011] The energy management model is constructed by: acquiring historical operating data of the vertical axis micro-wind power generation system; generating a historical environmental feature map based on the historical operating data; using the historical environmental feature map to simulate the learning mechanism of a neural network and converting it into a state transition matrix, wherein the state transition matrix is ​​used to determine the energy conversion efficiency of each monitoring point under different wind speed conditions; determining a dynamic adjustment threshold based on the historical energy utilization rate in the historical operating data and the real-time load status of the vertical axis micro-wind power generation system, wherein the dynamic adjustment threshold is used to measure the energy conversion efficiency; and constructing the energy management model based on the state transition matrix and the dynamic adjustment threshold.

[0012] In one embodiment of the present invention, the step of using the historical environmental feature map to simulate the learning mechanism of the neural network and convert it into a state transition matrix includes: constructing a standardized feature vector based on the historical environmental feature map; obtaining records of energy allocation strategies during the historical period and determining an initial connection weight matrix based on the records of energy allocation strategies during the historical period; weighting the standardized feature vector using the initial connection weight matrix and mapping the elements between the initial connection weight matrix and the standardized feature vector using a nonlinear activation function to generate historical state transition intensities; combining the historical state transition intensities corresponding to different monitoring points to form a historical sequence; processing the historical sequence using a time decay function to determine the time-series cumulative state intensity; and aggregating the time-series cumulative state intensities corresponding to different monitoring points to form the state transition matrix.

[0013] In one embodiment of the present invention, each of the time-series cumulative state intensities corresponds to data from a monitoring point; the aggregation of the time-series cumulative state intensities corresponding to different monitoring points to form the state transition matrix includes: determining spatial weight coefficients corresponding to each of the monitoring points based on the distance between different monitoring points; adjusting the time-series cumulative state intensities of adjacent monitoring points based on the spatial weight coefficients of each monitoring point; fusing the time-series cumulative state intensities of each monitoring point with the adjusted time-series cumulative state intensities of adjacent monitoring points to generate a target fused state; and forming the state transition matrix based on the target fused state of each monitoring point.

[0014] In one embodiment of the present invention, the step of invoking a pre-built energy management model and processing the dynamic environmental feature map based on the energy management model to obtain the target energy allocation strategy to be adjusted includes: updating the current dynamic adjustment threshold based on the current load state of the vertical axis micro wind power generation system; processing the dynamic environmental feature map based on the state transition matrix to obtain the current state intensity corresponding to each monitoring point; comparing the current state intensity with the updated dynamic adjustment threshold to determine the target monitoring point; and determining the target energy allocation strategy to be adjusted based on the data in the target monitoring point.

[0015] In one embodiment of the present invention, before adjusting the energy allocation strategy, the method includes: determining characteristic local extreme points in the dynamic environment feature map, and determining an adaptive region boundary set based on the characteristic local extreme points; for each monitoring area formed in the dynamic environment feature map based on the adaptive region boundary set, determining an energy allocation weight that is positively correlated in value according to the feature value of each monitoring area, and generating a weight mapping table based on the energy allocation weight; partitioning the energy allocation strategy to be adjusted based on the adaptive region boundary set and the weight mapping table to obtain multiple energy allocation units with different weights, wherein the weight of the energy allocation unit is proportional to the wind speed change trend.

[0016] In one embodiment of the present invention, adjusting the energy allocation strategy includes: constructing a dynamic adjustment dictionary matching the energy allocation strategy; encoding each energy allocation unit of the energy allocation strategy based on the dynamic adjustment dictionary, and determining the adjustment level in the encoding process according to the feature values ​​in the dynamic environment feature map to obtain the encoded data of the energy allocation strategy; encapsulating the encoded data to form an adjustment data sequence, and constructing adjustment metadata based on the adjustment information in the encoding process; and transmitting and verifying the adjustment data sequence and the adjustment metadata.

[0017] In addition, this invention also discloses an energy management device for a vertical axis micro-wind power generation system based on reinforcement learning, comprising:

[0018] The acquisition module is used to acquire the operating data of the vertical axis micro wind power generation system. The operating data is related to the real-time wind speed, power output and energy storage status of the vertical axis micro wind power generation system.

[0019] The generation module is used to generate a dynamic environmental feature map based on the running data. The dynamic environmental feature map is formed by a three-dimensional feature field that characterizes the wind speed change trend and the power generation efficiency change. The three-dimensional feature field involves the temporal change gradient and spatial distribution gradient of the wind speed.

[0020] The calling module is used to call a pre-built energy management model, process the dynamic environment feature map based on the energy management model, and obtain the target energy allocation strategy to be adjusted. The energy management model is used to determine the energy allocation scheme that needs to be optimized based on the input dynamic environment feature map.

[0021] An adjustment module is used to generate control signals to trigger the execution of the target energy allocation strategy.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] This invention achieves intelligent and adaptive optimization of energy management by introducing reinforcement learning algorithms. By generating dynamic environmental feature maps, the temporal and spatial gradients of wind speed are fused into a three-dimensional feature field, accurately capturing the spatiotemporal characteristics of wind speed changes. This provides a comprehensive and accurate environmental information foundation for energy management, significantly enhancing the system's perception and response speed to complex micro-wind environments, thus overcoming the shortcomings of traditional methods in environmental characterization. Simultaneously, an energy management model is constructed based on historical operational data, including a state transition matrix and dynamically adjusted thresholds. This simulates the learning mechanism of neural networks, enabling the system to automatically learn the optimal energy allocation strategy from historical experience and dynamically adjust thresholds according to real-time load conditions, ensuring the real-time adaptability and optimization effect of the strategy and improving energy conversion efficiency. More importantly, this invention employs a partitioning and encoding mechanism to finely adjust the energy allocation strategy. Adaptive region boundaries are determined through feature local extrema, generating a weight mapping table that divides the strategy into energy allocation units with different weights, making energy allocation more rational and efficient, further optimizing the system's energy utilization rate under changing environments.

[0024] This invention effectively overcomes the limitations of traditional control strategies in nonlinear and uncertain microwind environments through continuous interactive learning of reinforcement learning, and realizes the efficient and stable operation of vertical axis microwind power generation system under low wind speed conditions, thereby improving the overall reliability, adaptability and energy utilization efficiency of the system.

[0025] This invention achieves intelligent, adaptive, and refined energy management through the deep integration of reinforcement learning and dynamic environmental characteristics, effectively solving the problems of low energy utilization and poor dynamic adaptability in existing technologies. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating the energy management method for a vertical axis micro-wind power generation system based on reinforcement learning, as described in this invention.

[0028] Figure 2 This is a schematic diagram illustrating the generation process of dynamic environment feature maps.

[0029] Figure 3 This is a schematic diagram illustrating the construction and processing flow of an energy management model.

[0030] Figure 4A schematic diagram of the partitioning and coding process for adjusting the energy allocation strategy.

[0031] Figure 5 This is a simplified flowchart of the overall process of the present invention. Detailed Implementation

[0032] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0033] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0034] Example 1:

[0035] See Figures 1-5 This embodiment discloses an energy management method for a vertical axis micro-wind power generation system based on reinforcement learning, including:

[0036] Step 1: Obtain the operating data of the vertical axis micro wind power generation system. The operating data is related to the real-time wind speed, power output, and energy storage status of the vertical axis micro wind power generation system.

[0037] Step 2: Generate a dynamic environmental feature map based on the operational data. The dynamic environmental feature map is formed by a three-dimensional feature field that characterizes the wind speed change trend and the power generation efficiency change. The three-dimensional feature field involves the temporal variation gradient and spatial distribution gradient of the wind speed.

[0038] Step 3: Call the pre-built energy management model, process the dynamic environment feature map based on the energy management model, and obtain the target energy allocation strategy to be adjusted. The energy management model is used to determine the energy allocation scheme that needs to be optimized based on the input dynamic environment feature map.

[0039] Step 4: Generate a control signal to trigger the execution of the target energy allocation strategy.

[0040] Furthermore, the step of generating a dynamic environment feature map based on the operational data includes:

[0041] The operational data is aggregated in the time dimension to form a time-series wind speed matrix, which includes the wind speed status of each monitoring point in a continuous time slice.

[0042] Based on the aforementioned time-series wind speed matrix, determine the time gradient component used to characterize the frequency of wind speed changes;

[0043] Based on the time-series wind speed matrix and logically adjacent monitoring points, spatial gradient components are determined to characterize the changing trend of wind speed in spatial distribution.

[0044] The temporal gradient component and the spatial gradient component are synthesized to form the three-dimensional feature field.

[0045] Furthermore, the synthesis of the temporal gradient components and spatial gradient components to form the three-dimensional feature field includes:

[0046] The temporal gradient component and spatial gradient component are synthesized based on the following rules to form the three-dimensional feature field:

[0047] The fused gradient vector is obtained by weighted summation of the temporal gradient component and the spatial gradient component. The weighting coefficients of the summation are determined based on the correlation between the degree of wind speed change and power generation efficiency.

[0048] For example, the weighting coefficients for the weighted summation are calculated using the Pearson correlation coefficient formula:

[0049] ;

[0050] in, : Weighting coefficients for weighted summation (used to synthesize temporal and spatial gradient components).

[0051] : The rate of change of wind speed in the i-th time slice;

[0052] : The power generation efficiency corresponding to the i-th time slice;

[0053] n: The total number of time slices;

[0054] V: The mean of the rate of change of wind speed over all time slots;

[0055] E: The average power generation efficiency across all time slots;

[0056] and This is the mean.

[0057] The direction vector is obtained by normalizing the fused gradient vector, and the direction vector is used to indicate the main trend of wind speed change.

[0058] Furthermore, the energy management model is constructed, including:

[0059] Acquire historical operating data of the vertical axis micro wind power generation system;

[0060] A historical environment feature map is generated based on the historical operational data.

[0061] The learning mechanism of the neural network is simulated using the historical environmental feature map, which is then transformed into a state transition matrix. The state transition matrix is ​​used to determine the energy conversion efficiency of each monitoring point under different wind speed conditions.

[0062] A dynamic adjustment threshold is determined based on the historical energy utilization rate in the historical operating data and the real-time load status of the vertical axis micro wind power generation system. The dynamic adjustment threshold is used to measure the energy conversion efficiency.

[0063] The energy management model is constructed based on the state transition matrix and the dynamically adjusted threshold.

[0064] The reinforcement learning includes the following elements:

[0065] State space: composed of wind speed time gradient, spatial gradient, and power generation efficiency feature vector in the dynamic environment feature map;

[0066] Action space: includes a set of discrete actions such as energy storage ratio adjustment, power generation output adjustment, and load distribution adjustment;

[0067] Reward function: Constructed based on energy utilization efficiency improvement rate and system stability index, the calculation formula is as follows:

[0068] ;

[0069] Where α and β are weighting coefficients, As of the current energy utilization rate, Historical energy utilization rate System stability indicators characterize the smoothness of operation of a vertical axis micro wind power generation system, such as the power fluctuation amplitude and the range of energy storage status fluctuations.

[0070] Learning algorithm: The Deep Q-Network (DQN) algorithm is adopted, and the strategy is updated through experience replay and target network.

[0071] Furthermore, the learning mechanism of simulating a neural network using the historical environment feature map is transformed into a state transition matrix, including:

[0072] A standardized feature vector is constructed based on the historical environmental feature map;

[0073] Obtain records of energy allocation strategies during historical periods, and determine an initial connection weight matrix based on these records.

[0074] The standardized feature vector is weighted using the initial connection weight matrix, and the elements between the initial connection weight matrix and the standardized feature vector are mapped using a non-linear activation function to generate historical state transition intensities. The combinations of the historical state transition intensities corresponding to different monitoring points form a historical sequence.

[0075] The historical sequence is processed using a time decay function to determine the temporal cumulative state intensity; the temporal cumulative state intensities corresponding to different monitoring points are aggregated to form the state transition matrix.

[0076] Furthermore, each of the aforementioned time-series cumulative state intensities corresponds to data from a monitoring point; the aggregation of time-series cumulative state intensities corresponding to different monitoring points forms the state transition matrix, including:

[0077] The spatial weight coefficients for each monitoring point are determined based on the distances between the different monitoring points; the temporal cumulative state intensity of adjacent monitoring points is adjusted based on the spatial weight coefficients of each monitoring point.

[0078] The temporal cumulative state intensity of each monitoring point is fused with the adjusted temporal cumulative state intensity of adjacent monitoring points to generate a target fused state;

[0079] The state transition matrix is ​​formed based on the target fusion state of each monitoring point.

[0080] Furthermore, the step of invoking a pre-built energy management model and processing the dynamic environment feature map based on the energy management model to obtain the target energy allocation strategy to be adjusted includes:

[0081] The current dynamic adjustment threshold is updated based on the current load status of the vertical axis micro wind power generation system;

[0082] The dynamic environment feature map is processed based on the state transition matrix to obtain the current state intensity of each monitoring point.

[0083] The current state intensity is compared with the updated dynamic adjustment threshold to determine the target monitoring point;

[0084] The target energy allocation strategy to be adjusted is determined based on the data from the target monitoring points.

[0085] Furthermore, before adjusting the energy allocation strategy, the method also includes:

[0086] Determine the feature local extreme points in the dynamic environment feature map, and determine the adaptive region boundary set based on the feature local extreme points;

[0087] For each monitoring area formed by dividing the dynamic environment feature map based on the adaptive region boundary set, an energy allocation weight that is positively correlated in numerical value is determined according to the feature value of each monitoring area, so as to generate a weight mapping table based on the energy allocation weight;

[0088] Based on the adaptive region boundary set and weight mapping table, the energy allocation strategy to be adjusted is partitioned to obtain multiple energy allocation units with different weights. The weight of the energy allocation unit is proportional to the wind speed change trend.

[0089] Furthermore, adjusting the energy allocation strategy includes:

[0090] Construct a dynamic adjustment dictionary that matches the energy allocation strategy;

[0091] The energy allocation strategy is encoded based on the dynamic adjustment dictionary, and the adjustment level in the encoding process is determined according to the feature value in the dynamic environment feature map to obtain the encoded data of the energy allocation strategy.

[0092] The encoded data is encapsulated to form an adjustment data sequence, and adjustment metadata is constructed based on the adjustment information in the encoding process; the adjustment data sequence and adjustment metadata are transmitted and verified.

[0093] This embodiment also discloses an energy management device for a vertical axis micro-wind power generation system based on reinforcement learning, including:

[0094] The acquisition module is used to acquire the operating data of the vertical axis micro wind power generation system. The operating data is related to the real-time wind speed, power output and energy storage status of the vertical axis micro wind power generation system.

[0095] The generation module is used to generate a dynamic environmental feature map based on the running data. The dynamic environmental feature map is formed by a three-dimensional feature field that characterizes the wind speed change trend and the power generation efficiency change. The three-dimensional feature field involves the temporal change gradient and spatial distribution gradient of the wind speed.

[0096] The calling module is used to call a pre-built energy management model, process the dynamic environment feature map based on the energy management model, and obtain the target energy allocation strategy to be adjusted. The energy management model is used to determine the energy allocation scheme that needs to be optimized based on the input dynamic environment feature map.

[0097] An adjustment module is used to generate control signals to trigger the execution of the target energy allocation strategy.

[0098] Furthermore, the generation module is also used for:

[0099] The operational data is aggregated along the time dimension to form a time-series wind speed matrix;

[0100] Based on the aforementioned time-series wind speed matrix, determine the time gradient component used to characterize the frequency of wind speed changes;

[0101] Based on the time-series wind speed matrix and logically adjacent monitoring points, spatial gradient components are determined to characterize the changing trend of wind speed in spatial distribution.

[0102] The temporal gradient component and the spatial gradient component are synthesized to form the three-dimensional feature field.

[0103] Furthermore, the calling module is also used for:

[0104] The current dynamic adjustment threshold is updated based on the current load state of the vertical axis micro wind power generation system; the dynamic environment feature map is processed based on the state transition matrix to obtain the current state intensity of each monitoring point;

[0105] The current state intensity is compared with the updated dynamic adjustment threshold to determine the target monitoring point;

[0106] The target energy allocation strategy to be adjusted is determined based on the data from the target monitoring points.

[0107] Furthermore, the adjustment module is also used for:

[0108] Construct a dynamic adjustment dictionary that matches the energy allocation strategy;

[0109] The energy allocation strategy is encoded based on the dynamic adjustment dictionary, and the adjustment level in the encoding process is determined according to the feature value in the dynamic environment feature map to obtain the encoded data of the energy allocation strategy.

[0110] The encoded data is encapsulated to form an adjustment data sequence, and adjustment metadata is constructed based on the adjustment information in the encoding process;

[0111] The adjustment data sequence and adjustment metadata are transmitted and verified.

[0112] To facilitate a better understanding of the present invention by those skilled in the art, the present invention will be further described below in conjunction with specific implementations.

[0113] A reinforcement learning-based energy management method and device for vertical axis micro-wind power generation systems achieves intelligent adjustment of the energy allocation strategy of the system through the generation and processing of dynamic environmental feature maps and the construction and optimization of the energy management model. The following is combined with the appendix... Figure 1 To be continued Figure 4 The specific embodiments of the present invention will be described in detail.

[0114] exist Figure 1 The document demonstrates the overall process from acquiring operational data to executing the target energy allocation strategy. First, the acquisition module collects operational data from the vertical axis micro-wind power generation system, including real-time wind speed, power output, and energy storage status. The operational data originates from multiple monitoring points within the vertical axis micro-wind power generation system, each connected to the central control system via a sensor network, thereby enabling real-time monitoring of wind speed, power output, and energy storage status.

[0115] The acquisition module transmits this data to the generation module, which generates a dynamic environmental feature map based on the runtime data. The dynamic environmental feature map is formed by a three-dimensional feature field characterizing wind speed variation trends and power generation efficiency changes. This three-dimensional feature field involves the temporal gradient and spatial distribution gradient of wind speed. The generation module completes this process by constructing a time-series wind speed matrix, synthesizing the temporal and spatial gradient components, and forming the three-dimensional feature field.

[0116] like Figure 2 As shown, the generation process of dynamic environment feature maps includes multiple steps.

[0117] First, the generation module aggregates the runtime data along the time dimension to form a time-series wind speed matrix, which records the wind speed status of each monitoring point within a continuous time slice. Then, based on the time-series wind speed matrix, it determines the temporal gradient component characterizing the frequency of wind speed changes, and combines this with logically adjacent monitoring points to determine the spatial gradient component characterizing the spatial distribution trend of wind speed changes. The generation module further synthesizes the temporal and spatial gradient components through a weighted summation to form a three-dimensional feature field.

[0118] The specific weighted summation rules are as follows: the fused gradient vector is obtained by weighted summation of the temporal and spatial gradient components, with the weighting coefficients determined based on the correlation between the severity of wind speed changes and power generation efficiency; the direction vector is obtained by normalizing the fused gradient vector, and is used to indicate the main trend of wind speed changes. Through the above steps, the generation module finally generates a dynamic environmental feature map, which can comprehensively reflect the temporal and spatial variation characteristics of wind speed.

[0119] After generating the dynamic environment feature map, the calling module invokes the pre-built energy management model, processes the dynamic environment feature map based on the energy management model, and obtains the target energy allocation strategy to be adjusted.

[0120] The process of building an energy management model is as follows: Figure 3As shown, the construction module acquires historical operational data of the vertical axis micro-wind power generation system and generates a historical environmental feature map based on this data. The historical environmental feature map is standardized and transformed into a standardized feature vector. Simultaneously, records of energy allocation strategies during the historical period are acquired, and an initial connection weight matrix is ​​determined based on these records. The construction module uses the initial connection weight matrix to weight the standardized feature vector and maps the elements between the initial connection weight matrix and the standardized feature vector using a non-linear activation function to generate historical state transition intensities. These historical state transition intensities correspond to historical sequences at different monitoring points. A time decay function is used to process the historical sequences to determine the time-series cumulative state intensity.

[0121] In practice, the ReLU activation function is used to perform a non-linear mapping on the weighted feature vector, and the calculation formula is as follows:

[0122] ;

[0123] The output of the ReLU activation function is used to perform a non-linear mapping on the weighted result of the initial connection weight matrix and the standardized eigenvectors to generate historical state transition strengths.

[0124] : The input value of the activation function.

[0125] The construction module determines the spatial weight coefficients for each monitoring point based on the distance between different monitoring points, adjusts the temporal cumulative state intensity of adjacent monitoring points, and fuses the temporal cumulative state intensity of each monitoring point with the adjusted temporal cumulative state intensity of adjacent monitoring points to generate a target fused state. The target fused state ultimately forms a state transition matrix, which is used to describe the energy conversion efficiency of different monitoring points under different wind speed conditions.

[0126] The module also determines a dynamic adjustment threshold based on historical energy utilization rates in historical operating data and the real-time load status of the vertical axis micro wind power generation system. This threshold is used to measure the optimization effect of energy conversion efficiency.

[0127] When invoking the energy management model, the calling module updates the current dynamic adjustment threshold based on the current load state of the vertical axis micro-wind power generation system. Then, it processes the dynamic environmental feature map based on the state transition matrix to obtain the current state intensity for each monitoring point. The calling module compares the current state intensity with the updated dynamic adjustment threshold to determine the target monitoring point.

[0128] Based on the data from the target monitoring points, the calling module further determines the target energy allocation strategy to be adjusted. Before determining the target energy allocation strategy, the adjustment module also needs to perform partitioning processing on the dynamic environment feature map.

[0129] like Figure 4 As shown, the adjustment module first identifies the characteristic local extrema points in the dynamic environment feature map and then determines the adaptive region boundary set based on these points. Based on this adaptive region boundary set, the adjustment module analyzes the various monitoring areas formed by the division in the dynamic environment feature map, and determines the energy allocation weights that are positively correlated in value according to the characteristic values ​​of each monitoring area, thereby generating a weight mapping table. The adjustment module then partitions the energy allocation strategy to be adjusted based on the adaptive region boundary set and the weight mapping table, obtaining multiple energy allocation units with different weights. The weight of each energy allocation unit is directly proportional to the wind speed change trend.

[0130] The adjustment module further constructs a dynamic adjustment dictionary to match the energy allocation strategy. Based on this dictionary, it encodes each energy allocation unit of the strategy and determines the adjustment level during the encoding process based on the feature values ​​in the dynamic environment feature map, thus obtaining the encoded data of the energy allocation strategy. This encoded data is encapsulated to form an adjustment data sequence, and adjustment metadata is constructed based on the adjustment information obtained during the encoding process. Finally, the adjustment module transmits and verifies the adjustment data sequence and the adjustment metadata to ensure the accuracy and reliability of the target energy allocation strategy.

[0131] Throughout the process, the acquisition module collects operational data in real time through a sensor network and transmits it to the generation module; the generation module generates a dynamic environment feature map based on the operational data and passes the feature map to the calling module; the calling module processes the dynamic environment feature map through an energy management model, generates a target energy allocation strategy, and then passes it to the adjustment module; the adjustment module partitions and encodes the target energy allocation strategy, and finally generates an adjustment data sequence and triggers its execution.

[0132] In practical applications, this invention can be widely used in energy management scenarios for vertical axis micro-wind power generation systems. For example, in a distributed power generation system containing multiple vertical axis wind turbines, deploying the energy management method and device provided by this invention can effectively improve the system's energy utilization efficiency.

[0133] In practical applications, when the wind speed at a monitoring point increases significantly, the dynamic environmental feature map quickly captures this change and generates an optimized energy allocation strategy through the aforementioned process. For example, in a distributed vertical axis micro-wind power generation system, if the wind speed in a certain area suddenly increases, the system will prioritize increasing the energy allocation weight of that area while reducing the energy allocation ratio of other low-wind-speed areas. This dynamic adjustment mechanism not only improves the system's energy utilization efficiency but also enhances its adaptability to complex and changing environments. By applying a dynamic adjustment dictionary and weight mapping table, the system can continuously optimize its energy allocation strategy during long-term operation, thereby achieving higher energy utilization and stability. Experimental results show that under micro-wind conditions of 2-5 m / s, the energy utilization rate of this invention is improved by 25.3% compared to traditional PID control, and the system response time is shortened from 3.2s to 1.1s.

[0134] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0135] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An energy management method for vertical axis micro-wind power generation systems based on reinforcement learning, characterized in that, include: Step 1: Obtain the operating data of the vertical axis micro wind power generation system. The operating data is related to the real-time wind speed, power output, and energy storage status of the vertical axis micro wind power generation system. Step 2: Generate a dynamic environmental feature map based on the operational data. The dynamic environmental feature map is formed by a three-dimensional feature field that characterizes the wind speed change trend and the power generation efficiency change. The three-dimensional feature field involves the temporal variation gradient and spatial distribution gradient of the wind speed. Step 3: Call the pre-built energy management model, process the dynamic environment feature map based on the energy management model, and obtain the target energy allocation strategy to be adjusted. The energy management model is used to determine the energy allocation scheme that needs to be optimized based on the input dynamic environment feature map. Constructing the energy management model includes: Acquire historical operating data of the vertical axis micro wind power generation system; A historical environment feature map is generated based on the historical operational data. The learning mechanism of the neural network is simulated using the historical environmental feature map, which is then transformed into a state transition matrix. The state transition matrix is ​​used to determine the energy conversion efficiency of each monitoring point under different wind speed conditions. A dynamic adjustment threshold is determined based on the historical energy utilization rate in the historical operating data and the real-time load status of the vertical axis micro wind power generation system. The dynamic adjustment threshold is used to measure the energy conversion efficiency. The energy management model is constructed based on the state transition matrix and the dynamically adjusted threshold. Step 4: Generate a control signal to trigger the execution of the target energy allocation strategy; The process of calling a pre-built energy management model, processing the dynamic environment feature map based on the energy management model, and obtaining the target energy allocation strategy to be adjusted includes: The current dynamic adjustment threshold is updated based on the current load status of the vertical axis micro wind power generation system; The dynamic environment feature map is processed based on the state transition matrix to obtain the current state intensity of each monitoring point. The current state intensity is compared with the updated dynamic adjustment threshold to determine the target monitoring point; The target energy allocation strategy to be adjusted is determined based on the data from the target monitoring points.

2. The energy management method for a vertical axis micro-wind power generation system based on reinforcement learning according to claim 1, characterized in that, The generation of a dynamic environment feature map based on the operational data includes: The operational data is aggregated in the time dimension to form a time-series wind speed matrix, which includes the wind speed status of each monitoring point in a continuous time slice. Based on the aforementioned time-series wind speed matrix, determine the time gradient component used to characterize the frequency of wind speed changes; Based on the time-series wind speed matrix and logically adjacent monitoring points, spatial gradient components are determined to characterize the changing trend of wind speed in spatial distribution. The temporal gradient component and the spatial gradient component are synthesized to form the three-dimensional feature field.

3. The energy management method for a vertical axis micro-wind power generation system based on reinforcement learning according to claim 2, characterized in that, The synthesis of the temporal gradient components and spatial gradient components to form the three-dimensional feature field includes: The temporal gradient component and spatial gradient component are synthesized based on the following rules to form the three-dimensional feature field: The fused gradient vector is obtained by weighted summation of the temporal gradient component and the spatial gradient component. The weighting coefficients of the summation are determined based on the correlation between the degree of wind speed change and power generation efficiency. The direction vector is obtained by normalizing the fused gradient vector, and the direction vector is used to indicate the main trend of wind speed change.

4. The energy management method for a vertical axis micro-wind power generation system based on reinforcement learning according to claim 1, characterized in that, The learning mechanism that utilizes the historical environment feature map to simulate a neural network is transformed into a state transition matrix, including: A standardized feature vector is constructed based on the historical environmental feature map; Obtain records of energy allocation strategies during historical periods, and determine an initial connection weight matrix based on these records. The standardized feature vector is weighted using the initial connection weight matrix, and the elements between the initial connection weight matrix and the standardized feature vector are mapped using a non-linear activation function to generate historical state transition intensities. The combinations of the historical state transition intensities corresponding to different monitoring points form a historical sequence. The historical sequence is processed using a time decay function to determine the temporal cumulative state intensity; the temporal cumulative state intensities corresponding to different monitoring points are aggregated to form the state transition matrix.

5. The energy management method for a vertical axis micro-wind power generation system based on reinforcement learning according to claim 4, characterized in that, Each of the aforementioned time-series cumulative state intensities corresponds to data from a monitoring point; The aggregation of time-series cumulative state intensities corresponding to different monitoring points forms the state transition matrix, including: The spatial weight coefficients for each monitoring point are determined based on the distances between the different monitoring points; the temporal cumulative state intensity of adjacent monitoring points is adjusted based on the spatial weight coefficients of each monitoring point. The temporal cumulative state intensity of each monitoring point is fused with the adjusted temporal cumulative state intensity of adjacent monitoring points to generate a target fused state; The state transition matrix is ​​formed based on the target fusion state of each monitoring point.

6. The energy management method for a vertical axis micro-wind power generation system based on reinforcement learning according to claim 1, characterized in that, Before adjusting the target energy allocation strategy, the following is also included: Determine the feature local extreme points in the dynamic environment feature map, and determine the adaptive region boundary set based on the feature local extreme points; For each monitoring area formed by dividing the dynamic environment feature map based on the adaptive region boundary set, an energy allocation weight that is positively correlated in numerical value is determined according to the feature value of each monitoring area, so as to generate a weight mapping table based on the energy allocation weight; Based on the adaptive region boundary set and weight mapping table, the target energy allocation strategy to be adjusted is partitioned to obtain multiple energy allocation units with different weights. The weight of the energy allocation unit is proportional to the wind speed change trend.

7. The energy management method for a vertical axis micro-wind power generation system based on reinforcement learning according to claim 6, characterized in that, Adjusting the target energy allocation strategy includes: Construct a dynamic adjustment dictionary that matches the target energy allocation strategy; The target energy allocation strategy is encoded based on the dynamic adjustment dictionary, and the adjustment level in the encoding process is determined according to the feature value in the dynamic environment feature map to obtain the encoded data of the target energy allocation strategy. The encoded data is encapsulated to form an adjustment data sequence, and adjustment metadata is constructed based on the adjustment information in the encoding process; the adjustment data sequence and adjustment metadata are transmitted and verified.

8. A reinforcement learning-based energy management device for a vertical axis micro-wind power generation system, used to execute the reinforcement learning-based energy management method for a vertical axis micro-wind power generation system as described in any one of claims 1-7, characterized in that, include: The acquisition module is used to acquire the operating data of the vertical axis micro wind power generation system. The operating data is related to the real-time wind speed, power output and energy storage status of the vertical axis micro wind power generation system. The generation module is used to generate a dynamic environmental feature map based on the running data. The dynamic environmental feature map is formed by a three-dimensional feature field that characterizes the wind speed change trend and the power generation efficiency change. The three-dimensional feature field involves the temporal change gradient and spatial distribution gradient of the wind speed. The calling module is used to call a pre-built energy management model, process the dynamic environment feature map based on the energy management model, and obtain the target energy allocation strategy to be adjusted. The energy management model is used to determine the energy allocation scheme that needs to be optimized based on the input dynamic environment feature map. An adjustment module is used to generate control signals to trigger the execution of the target energy allocation strategy.