An Adaptive Distributed New Energy Power Generation Equipment Cooperative Control Method and System
The adaptive control method using deep reinforcement learning and gradient optimization addresses device response inconsistencies in new energy power generation systems, enhancing system adaptability and stability by optimizing collaborative strategies for voltage fluctuations.
Patent Information
- Application Number
- CN202510436363.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In new energy power generation equipment systems, when the grid voltage fluctuates abnormally, the equipment action selection is prone to conflict, resulting in difficulty in collaborative control and difficult to balance the grid stability and equipment operation costs.
Deep reinforcement learning algorithms and gradient optimization strategies are adopted to collect device state characteristics in real time, generate action boundary estimates, build empirical extension data sets, evaluate conflict probability, and iteratively determine collaborative control optimization strategies.
It realizes efficient coordinated control and conflict suppression of new energy power generation equipment in the scenario of abnormal voltage fluctuations, and improves the stability of the power grid and power generation efficiency.
Smart Images

Figure CN119965994B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of renewable energy power generation technology, and in particular to a method and system for collaboratively controlling adaptive distributed renewable energy power generation equipment. Background Art
[0002] In the field of distributed control of new energy power generation equipment, especially in the scenario of abnormal voltage fluctuations in local areas of the power grid, the problems of action conflicts and real-time coordination of new energy power generation equipment systems are particularly prominent, and there are still a series of technical problems that need to be solved urgently.
[0003] Specifically, when the grid voltage fluctuates rapidly due to a sudden load change, multiple renewable energy power generation devices in the area (such as wind turbines or photovoltaic panels, etc.) need to respond quickly. However, these renewable energy power generation devices have differences in physical properties and operating conditions. For example, some renewable energy power generation devices cannot be further adjusted because their output is close to the upper limit, while other renewable energy power generation devices cannot fully exert their power generation capacity due to factors such as excessive temperature. This inconsistency in state causes the action selection of each device to conflict when facing abnormal voltage fluctuations, making it difficult for the system to quickly and effectively coordinate the actions of each device, which brings great difficulties to collaborative control and affects the overall control effect. In addition, when designing control strategies, how to balance the stability of the grid and the operating cost of equipment is a key problem. If the stability of the grid is overemphasized, some equipment may frequently adjust its output, thereby increasing equipment losses and operating costs; on the contrary, if the operating cost of the equipment is given priority, it may not be able to meet the grid's demand for real-time response, resulting in further deterioration of voltage fluctuations. This contradiction has not been effectively resolved in the existing technology.
[0004] In summary, the existing technology has many problems when dealing with complex scenarios such as abnormal voltage fluctuations in local areas of the power grid. Therefore, it is urgent to seek more effective solutions to improve the adaptability and stability of coordinated control of multiple devices in the power grid and achieve optimal regulation under abnormal voltage fluctuations. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a method and system for collaborative control of adaptive distributed renewable energy power generation equipment.
[0006] In a first aspect, the present invention provides a method for cooperatively controlling an adaptive distributed renewable energy power generation device, the method comprising the following steps:
[0007] Extracting the state characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment;
[0008] Based on the state feature representation, use a deep reinforcement learning algorithm to determine the estimated action boundary values of each new energy power generation device in the voltage abnormal fluctuation scenario;
[0009] According to the estimated action boundary values, extract device state-action data pairs from the historical operation database, and based on the device state-action data pairs, generate an empirical expansion data set by simulating different disturbance amplitudes;
[0010] Based on the data sampling frequency and abnormal scenario proportion in the empirical expansion data set, use a weighted method to obtain the device operation reward parameters of the new energy power generation device;
[0011] Based on the device operation reward parameters, extract state-response data pairs with the same time decay coefficient from the pre-constructed distributed shared experience pool to form a shared experience correction set;
[0012] Based on the shared experience correction set and the state feature representation, evaluate the action selection conflict probability of each new energy power generation device, and based on the action selection conflict probability, use a gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.
[0013] In a further implementation, the step of extracting the state feature representation of each new energy power generation device from the real-time collected operation parameters of the new energy power generation device includes:
[0014] Real-time collect grid voltage data and operation parameters of new energy power generation devices, and use a preset data sampling frequency to perform real-time calculation on the grid voltage data to obtain the voltage frequency deviation value;
[0015] Compare the voltage frequency deviation value with a preset frequency deviation threshold. If the voltage frequency deviation value exceeds the preset frequency deviation threshold, extract the state space description of each new energy power generation device from the operation parameters of the new energy power generation device;
[0016] Based on the state space description, use the principal component analysis method to extract the state feature representation of each new energy power generation device.
[0017] In a further implementation, the state space description includes the current disturbance simulation amplitude, device load factor, and output upper limit of the new energy power generation device.
[0018] In a further implementation, the step of using a deep reinforcement learning algorithm to determine the estimated action boundary values of each new energy power generation device in the voltage abnormal fluctuation scenario based on the state feature representation includes:
[0019] Introduce the action space dimension constraint, use the deep Q-network to perform mapping analysis on the state feature representation, and calculate the initial action space range of each new energy power generation device;
[0020] Detect the initial action space range. If it is detected that the initial action space range is in a narrow range, classify the action space limited state of the narrow range to obtain a limited state classification result;
[0021] Taking the perturbation simulation amplitude and the device load factor as clustering features, use the K-means clustering algorithm to analyze the new energy power generation devices with the limited state classification result being the limited state, and obtain the limited level distribution density feature;
[0022] If the limited level distribution density feature exceeds the preset distribution density threshold, input the limited level distribution density feature into the deep Q-network for secondary calculation to generate a dynamic compensation coefficient for the initial action space range;
[0023] Use the dynamic compensation coefficient to perform dynamic boundary secondary optimization on the initial action space range to obtain the action boundary estimation value of each new energy power generation device in the voltage abnormal fluctuation scenario.
[0024] In a further implementation, the step of detecting the initial action space range includes:
[0025] Use the preset time window length to segment and divide the initial action space range, and extract the time series change feature of the power adjustment range and the temperature limit condition fluctuation data within each time window;
[0026] Analyze the dynamic coupling relationship between the time series change feature of the power adjustment range and the temperature limit condition fluctuation data through the linear regression algorithm, and calculate the dynamic change trend coefficient of the two;
[0027] Compare the dynamic change trend coefficient with the preset narrow range threshold. If the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in a narrow range.
[0028] In a further implementation, the step of obtaining the device operation reward parameter of the new energy power generation device by using the weighted method based on the data sampling frequency and the abnormal scenario proportion in the empirical expansion dataset includes:
[0029] Calculate the initial reward function value by using the weighted sum method through the data sampling frequency and the abnormal scenario proportion in the empirical expansion dataset;
[0030] According to the mapping relationship between the distribution density of the initial reward function value and the abnormal scenario proportion, use the density peak clustering algorithm to obtain the reward interval;
[0031] Within each reward interval, calculate the penalty factor reference value according to the load deviation degree between the device load factor and the preset load threshold, and use the moving average method to smooth the penalty factor reference value to obtain the smoothed dynamic cost penalty distribution;
[0032] Compensate and correct the dynamic cost penalty distribution in the high-perturbation reward interval based on the perturbation simulation amplitude to obtain the corrected penalty distribution;
[0033] Perform anti-weight superposition on the corrected penalty distribution and the initial reward function value, and normalize to obtain the device operation reward parameter of the new energy power generation device.
[0034] In a further implementation, the step of extracting state-response data pairs with the same time decay coefficient from the pre-constructed distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set includes:
[0035] Extract the reference time decay coefficient from the device operation reward parameter as the matching reference, and match the reference time decay coefficient with the time decay coefficients of each data pair in the distributed shared experience pool to extract a candidate set of state-response data with the same time decay coefficient;
[0036] Calculate the response delay difference between the action execution response time of each state-response data pair in the candidate set of state-response data and the actual execution timestamp of the device response;
[0037] Based on the response delay difference, use the moving window average method to align the timestamps in the candidate set of state-response data to obtain a time-synchronized set of state-response data;
[0038] Divide the set of state-response data into different time windows, and fuse the state-response data of multiple new energy power generation devices within the same time window to construct a sequentially continuous shared experience correction set.
[0039] In a further implementation, the step of evaluating the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representation includes:
[0040] According to the state-response data in the shared experience correction set, calculate the abnormal cumulative value of each new energy power generation device, and identify high-conflict-risk agents from the new energy power generation devices according to the abnormal cumulative value;
[0041] Taking the main sorting basis as the abnormal voltage fluctuation amplitude and the secondary sorting basis as the time decay coefficient, use the entropy weight method to calculate the action priority of each high-conflict-risk agent to obtain the action priority agent sequence;
[0042] Based on the action priority agent sequence, calculate the overlapping ratio of the power regulation ranges of the high-conflict-risk agents with adjacent action priorities, and determine the potential conflict boundary according to the overlapping ratio of the power regulation ranges;
[0043] Map the real-time action instructions to the power regulation interval of the potential conflict boundary, and count the out-of-bounds frequency and out-of-bounds amplitude of the action instructions of each high-conflict-risk agent;
[0044] Calculate the conflict intensity index according to the product of the out-of-bounds frequency and out-of-bounds amplitude of the action instructions, and normalize the conflict intensity index to obtain the action selection conflict probability of each new energy power generation device.
[0045] In a further embodiment, the step of using the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression according to the action selection conflict probability includes:
[0046] Determine the corresponding initial disturbance frequency value according to the action selection conflict probability of each new energy power generation device and the real-time state characteristics of the power grid;
[0047] Taking minimizing the conflict probability and maximizing the power output efficiency as the optimization objectives, based on the optimization objectives, use the gradient optimization algorithm to calculate the partial derivatives of the initial disturbance frequency value and the power regulation range to obtain the gradient optimization direction;
[0048] Gradually adjust the disturbance frequency value and the power regulation range according to the gradient optimization direction, and through iterative calculation until the optimization objective converges, and output the multi-device collaborative control optimization strategy after conflict suppression.
[0049] In a second aspect, the present invention provides an adaptive distributed new energy power generation device collaborative control system, and the system includes:
[0050] A data acquisition module for extracting the state feature representations of each new energy power generation device from the operation parameters of the new energy power generation devices collected in real time;
[0051] A boundary estimation module for determining the action boundary estimation values of each new energy power generation device in the voltage abnormal fluctuation scenario by using the deep reinforcement learning algorithm based on the state feature representations;
[0052] A data expansion module for extracting device state-action data pairs from the historical operation database according to the action boundary estimation values, and generating an empirical expansion data set by simulating different disturbance amplitudes according to the device state-action data pairs;
[0053] A reward analysis module, configured to obtain the device operation reward parameter of the new energy power generation device by using a weighting method based on the data sampling frequency and the abnormal scenario proportion in the experience expansion dataset;
[0054] An experience correction module, configured to extract state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameter, and form a shared experience correction set;
[0055] A policy generation module, configured to evaluate the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representation, and iteratively determine the multi-device collaborative control optimization policy after conflict suppression by using a gradient optimization algorithm according to the action selection conflict probability.
[0056] The present invention provides an adaptive distributed new energy power generation device collaborative control method and system. The method includes extracting the state feature representation of each new energy power generation device from the real-time collected operation parameters of the new energy power generation device; determining the action boundary estimation value of each new energy power generation device in the voltage abnormal fluctuation scenario by using a deep reinforcement learning algorithm based on the state feature representation; extracting the device state-action data pairs from the historical operation database according to the action boundary estimation value, and generating an experience expansion dataset by simulating different disturbance amplitudes according to the device state-action data pairs; obtaining the device operation reward parameter of the new energy power generation device by using a weighting method based on the data sampling frequency and the abnormal scenario proportion in the experience expansion dataset; extracting state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameter, and forming a shared experience correction set; evaluating the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representation, and iteratively determining the multi-device collaborative control optimization policy after conflict suppression by using a gradient optimization algorithm according to the action selection conflict probability. Compared with the prior art, the method adaptively determines the collaborative control strategy of the new energy power generation device in a complex power grid environment by combining a deep reinforcement learning algorithm and a gradient optimization strategy, realizes the efficient collaborative control and conflict suppression of the new energy power generation device in the voltage abnormal fluctuation scenario, improves the adaptability of the device to voltage abnormal fluctuations, and significantly improves the stability and power generation efficiency of the power grid. Description of the Drawings
[0057] Figure 1 is a schematic flowchart of an adaptive distributed new energy power generation device collaborative control method provided by an embodiment of the present invention;
[0058] Figure 2 is a block diagram of an adaptive distributed new energy power generation device collaborative control system provided by an embodiment of the present invention. Detailed Embodiment
[0059] The embodiments of the present invention will be specifically described below in conjunction with the accompanying drawings. The provided examples are only for illustrative purposes and should not be construed as limiting the present invention. The included drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from its spirit and scope.
[0060] Reference Figure 1 , an embodiment of the present invention provides a method for collaborative control of adaptive distributed new energy power generation equipment. As Figure 1 shown, the method includes the following steps:
[0061] S1. Extract the state feature representations of each new energy power generation equipment from the operation parameters of the new energy power generation equipment collected in real time.
[0062] In some embodiments, the step of extracting the state feature representations of each new energy power generation equipment from the operation parameters of the new energy power generation equipment collected in real time includes:
[0063] Collect the grid voltage data and the operation parameters of the new energy power generation equipment in real time, and perform real-time calculation on the grid voltage data using a preset data sampling frequency to obtain the voltage frequency deviation value;
[0064] Compare the voltage frequency deviation value with a preset frequency deviation threshold. If the voltage frequency deviation value exceeds the preset frequency deviation threshold, extract the state space description of each new energy power generation equipment from the operation parameters of the new energy power generation equipment; the state space description includes the current disturbance simulation amplitude, equipment load factor, and output upper limit of the new energy power generation equipment;
[0065] Based on the state space description, use the principal component analysis method to extract the state feature representations of each new energy power generation equipment.
[0066] Specifically, in this embodiment, the sensor network installed in the new energy power station is used to collect the grid voltage data and the operation parameters of each new energy power generation device in real time, and the fast Fourier transform is used to perform real-time calculation on the grid voltage data according to the preset data sampling frequency to obtain the voltage frequency value of each sampling point. Then, the voltage frequency value of each sampling point is compared with the standard frequency to obtain the voltage frequency deviation value. These voltage frequency deviation values reflect the real-time fluctuation of the grid voltage frequency. In this embodiment, each voltage frequency deviation value is compared with the preset frequency deviation threshold to determine whether the grid voltage frequency is within the normal range. If the voltage frequency deviation value exceeds the preset frequency deviation threshold, it indicates that the grid voltage frequency is abnormal and further analysis of the state of the new energy power generation device is required. At this time, in this embodiment, for each new energy power generation device, key state parameters such as its current disturbance simulation amplitude, device load factor, and output power upper limit are extracted to obtain the state space description of each new energy power generation device. These descriptions include the key operation characteristics of the device under the current grid conditions. Among them, the disturbance simulation amplitude reflects the degree of disturbance suffered by the device under abnormal voltage conditions, and this embodiment can estimate it by comparing parameters such as the change in output power of the device before and after the abnormality; the device load factor represents the ratio of the current load of the device to its maximum designed load, and this embodiment can be calculated by the ratio of the real-time output power to the rated power; the output power upper limit represents the maximum output power that the device can reach under the current voltage conditions. Then, in this embodiment, the state space descriptions of each new energy power generation device extracted are organized into a matrix form, where each row represents the state space description of a device at a certain moment, and the principal component analysis (PCA) method is used to process the data in the state space description to extract the principal components that have the greatest impact on the device operation state. The principal components extracted by PCA are used as the state feature representations of each new energy power generation device, and these state feature representations can effectively reflect the operation state of the device under abnormal voltage conditions.
[0067] S2. Based on the state feature representation, use the deep reinforcement learning algorithm to determine the action boundary estimation value of each new energy power generation device in the voltage abnormal fluctuation scenario.
[0068] In some embodiments, the step of using the deep reinforcement learning algorithm to determine the action boundary estimation value of each new energy power generation device in the voltage abnormal fluctuation scenario based on the state feature representation includes:
[0069] Introduce the action space dimension constraint, use the deep Q-network to perform mapping analysis on the state feature representation, and calculate the initial action space range of each new energy power generation device;
[0070] Detect the initial action space range. If it is detected that the initial action space range is in a narrow range, classify the action space limited state of the narrow range to obtain the limited state classification result;
[0071] Taking the disturbance simulation amplitude and the device load factor as clustering features, the K-means clustering algorithm is used to analyze new energy power generation devices whose classified result of the restricted state is the restricted state, and the distribution density feature of the restricted level is obtained.
[0072] If the distribution density feature of the restricted level exceeds the preset distribution density threshold, the distribution density feature of the restricted level is input into the deep Q-network for secondary calculation to generate a dynamic compensation coefficient for the initial action space range.
[0073] The initial action space range is dynamically optimized by using the dynamic compensation coefficient to obtain the action boundary estimation value of each new energy power generation device in the voltage abnormal fluctuation scenario.
[0074] Specifically, in this embodiment, first, according to the type, performance parameters, and grid connection requirements of the new energy power generation device, the action type of the new energy power generation device in the voltage abnormal fluctuation scenario is determined. For example, power regulation, device protection, etc., and the action space dimension is defined according to the action type. For example, the action space dimension of the photovoltaic inverter may include the output power regulation amplitude, the power factor adjustment range, etc.; the wind turbine generator set may include the pitch angle adjustment range, the generator torque regulation amplitude, etc. In this embodiment, considering the physical characteristics and operation safety of the new energy power generation device, the action space is dimensionally constrained to ensure that all possible actions meet the safety operation standards of the device. For example, the output power regulation amplitude of the photovoltaic inverter is limited within ±20% of its rated power. Then, a deep Q-network (DQN) is constructed. This network is composed of multiple neural networks. The number of neurons in its input layer matches the dimension of the state feature representation. The hidden layer adopts a multi-layer perceptron structure, and the activation function is selected as ReLU. The number of neurons in the output layer corresponds to the dimension of the action space of the new energy power generation device. Based on the dimensional constraint of the action space, in this embodiment, the historical state feature representation data is input into the deep Q-network for training, so that it can accurately map the initial action space range of each new energy power generation device according to the input state feature representation. Through the forward propagation calculation of the network, the initial action space range of each new energy power generation device is obtained. The initial action space range includes the upper and lower limit values of the power regulation range and the threshold interval of the temperature limit condition, and the initial action space range is detected. In some embodiments, the step of detecting the initial action space range includes:
[0075] The initial action space range is segmented by using a preset time window length, and the time series change feature of the power regulation range and the fluctuation data of the temperature limit condition within each time window are extracted.
[0076] Analyze the dynamic coupling relationship between the temporal variation characteristics of the power adjustment range and the fluctuation data of the temperature limit conditions through a linear regression algorithm, and calculate the dynamic change trend coefficient of the two.
[0077] Compare the dynamic change trend coefficient with a preset narrow range threshold. If the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in the narrow range.
[0078] Specifically, in this embodiment, the time window length is set according to the duration characteristics of the abnormal voltage fluctuation of the power grid and the response speed requirements of the new energy power generation equipment. The initial action space range is segmented according to the time window length to obtain the action space sub-ranges within multiple time windows. For each time window, the temporal variation characteristics of the power adjustment range of the new energy power generation equipment are extracted. Taking a wind turbine generator as an example, in this embodiment, the change rate, change trend and other characteristic values of the pitch angle adjustment amplitude and the generator torque adjustment amplitude within each time window can be calculated. At the same time, the fluctuation data of the temperature limit conditions are extracted. The fluctuation data of the temperature limit conditions can specifically include the ambient temperature change range, the temperature change amplitude of the key components inside the equipment (such as the inverter power module, the generator stator winding, etc.). Then, in this embodiment, the temporal variation characteristics of the power adjustment range and the fluctuation data of the temperature limit conditions are used as variables to construct a linear regression model, and the regression coefficient between the two is calculated by methods such as the least squares method to obtain the dynamic change trend coefficient of the two. This dynamic change trend coefficient reflects the synchronism or correlation of the two with time. At the same time, in this embodiment, a narrow range threshold is set according to the operation experience of the new energy power generation equipment and the requirements of the power grid operation stability. Whether the initial action space range is in the narrow range is judged by the narrow range threshold. If the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in the narrow range. At this time, the action selection of the equipment is greatly restricted.
[0079] Then, in this embodiment, the support vector machine is used to classify the state of the action space limited within a narrow range. For example, the limited state is divided into three categories: slightly limited, moderately limited, and severely limited, and the classification result of the limited state is obtained. Then, with the disturbance simulation amplitude and the device load factor as the clustering features, the K-means clustering algorithm is initialized. Through iterative calculation, the cluster centers and the sample belonging categories are obtained, and then the limited level distribution density feature is obtained. The limited level distribution density feature reflects the distribution of devices with different limited levels. If the limited level distribution density feature exceeds the preset distribution density threshold, the limited level distribution density feature is input into the deep Q-network for secondary calculation. In this embodiment, a compensation coefficient calculation module is added to the deep Q-network. The compensation coefficient calculation module takes the limited level distribution density feature as the input and generates a dynamic compensation coefficient for the initial action space range through the forward propagation of the network. The dynamic compensation coefficient is used to adjust the initial action space range to cope with the limitations brought by the narrow range. For example, for a certain new energy power generation device, the generated dynamic compensation coefficient is 1.2, indicating that a moderate expansion compensation needs to be performed on the initial action space range. Finally, in this embodiment, the dynamic compensation coefficient and the initial action space range are subjected to an element-wise multiplication operation to obtain the optimized action boundary estimation value. For example, if the initial output power adjustment amplitude range is [-15%, +18%] and the dynamic compensation coefficient is 1.2, the optimized output power adjustment amplitude range is [-18%, +21.6%]. The optimized action boundary estimation value can more accurately reflect the actual action ability of the new energy power generation device in the voltage abnormal fluctuation scenario, providing strong support for the intelligent control of the device.
[0080] S3. Extract the device state-action data pairs from the historical operation database according to the action boundary estimation value, and generate an empirical expansion dataset by simulating different disturbance amplitudes according to the device state-action data pairs.
[0081] Specifically, in this embodiment, device status data and corresponding action data related to the action boundary estimation value are screened out from the historical operation database to form device status-action data pairs. These device status-action data pairs include the operation parameters of the device in different states and the corresponding operation records. Then, preprocessing operations are performed on the extracted device status-action data pairs. The preprocessing operations include, but are not limited to, data cleaning and data standardization, etc., to ensure data quality. Next, feature dimensions that can represent the diversity of the device operation status are selected from the extracted device status-action data pairs to construct a clustering feature vector, such as the deviation value of the grid voltage frequency, the device load factor, the disturbance simulation amplitude, etc. The clustering feature vector is input into the K-means clustering algorithm for clustering analysis, and the clustering categories of each data pair are obtained through iterative calculation, so as to obtain the classified state distribution characteristics. Furthermore, the state coverage breadth represented by each clustering center is calculated according to the classified state distribution characteristics. Among them, the state coverage breadth can be measured by the number or density of data points around the clustering center.
[0082] In this embodiment, the device status data is divided into different categories according to the state coverage breadth, and the state distribution characteristics of each category are extracted. These characteristics may include the coordinates of the clustering center, the clustering radius, the state density, etc. For each classified state distribution characteristic, it is judged whether its state coverage breadth is lower than the preset coverage breadth threshold. If the state coverage breadth is lower than the coverage breadth threshold, it is considered that the state data of this category is insufficient and needs to be empirically extended. Randomly select seed data pairs from the categories that need to be empirically extended as the basis, and at the same time analyze the type and amplitude distribution of the disturbances suffered by the device in the historical operation data to determine the range and step size of the disturbance simulation amplitude. On the basis of the original device status-action data pairs, according to different ranges and step sizes of the disturbance simulation amplitude, the device status feature representation in the seed data pairs is disturbed and simulated to generate a new device status feature representation, and the corresponding action data of the new device status feature representation is simulated. For example, according to the MPPT (maximum power point tracking) algorithm and the inverter control strategy of the photovoltaic inverter, calculate the change in the output power adjustment amplitude and the power factor adjustment value under the disturbed input power, so as to obtain the action data corresponding to the new state. The generated new state-action data pairs are added to the original data set to form an empirically extended data set.
[0083] S4. Based on the data sampling frequency and the abnormal scenario proportion in the empirically extended data set, use a weighted method to obtain the device operation reward parameter of the new energy power generation device.
[0084] In some embodiments, the step of using a weighted method to obtain the device operation reward parameter of the new energy power generation device based on the data sampling frequency and the abnormal scenario proportion in the empirically extended data set includes:
[0085] Expand the data sampling frequency and the proportion of abnormal scenarios in the dataset through experience, and calculate the initial reward function value using the weighted sum method;
[0086] According to the mapping relationship between the distribution density of the initial reward function value and the proportion of abnormal scenarios, use the density peak clustering algorithm to obtain the reward interval;
[0087] Within each reward interval, calculate the penalty factor reference value according to the load deviation degree between the device load factor and the preset load threshold, and use the moving average method to smooth the penalty factor reference value to obtain the smoothed dynamic cost penalty distribution;
[0088] Compensate and correct the dynamic cost penalty distribution in the high perturbation reward interval based on the perturbation simulation amplitude to obtain the corrected penalty distribution;
[0089] Perform inverse weight superposition on the corrected penalty distribution and the initial reward function value, and normalize to obtain the device operation reward parameter of the new energy power generation device.
[0090] Specifically, in this embodiment, extract the data sampling frequency values at all sampling moments from the experience-expanded dataset, find the maximum frequency value among them, and normalize the data sampling frequency in the experience-expanded dataset according to the maximum frequency value, that is, calculate the ratio of the frequency at each sampling moment to the maximum frequency value to obtain the frequency weight coefficient. At the same time, count the number of data points of various abnormal scenarios (such as voltage sag or frequency fluctuation, etc.) in the experience-expanded dataset, and calculate the proportion of each abnormal scenario according to the ratio between the number of data points of various abnormal scenarios and the total number of data points. According to the proportion of abnormal scenarios in the experience-expanded dataset, use the inverse proportional function to allocate scenario weight coefficients for the data of various abnormal scenarios. The higher the proportion of the abnormal scenario, the lower the corresponding scenario weight coefficient. Then, this embodiment extracts the power output efficiency and voltage deviation value from the operation parameters of the new energy power generation device, and performs weighted summation on the power output efficiency and voltage deviation value based on the frequency weight coefficient and the scenario weight coefficient. Specifically, for each sampling moment, this embodiment multiplies the power output efficiency by the frequency weight coefficient to obtain the power output efficiency weight value, multiplies the voltage deviation value by the scenario weight coefficient to obtain the voltage deviation weight value, and then performs weighted summation on the power output efficiency weight value and the voltage deviation weight value to obtain the initial reward function value at each sampling moment.
[0091] Next, in this embodiment, the initial reward function values are divided according to a preset time window to obtain a sequence of reward function values within multiple time windows, and statistical analysis is performed on the initial reward function values within each time window to obtain the distribution density of the reward function values. For example, the proportion of the number of data points in each reward value interval is statistically analyzed using the histogram method. According to the mapping relationship between the distribution density of the reward function values and the proportion of abnormal scenarios, taking the distribution density of the reward function values as the basis for selecting the clustering center, and combining the gradient change of the proportion of abnormal scenarios, the density peak clustering algorithm is used to divide the initial reward function values into several reward function intervals. Specifically, in this embodiment, taking the distribution density of the reward function values as the clustering center and combining the gradient change of the proportion of abnormal scenarios, the local density of each reward function value point and the distance to the nearest high-density point are calculated; then, points with high local density and large distance are selected as the clustering centers; finally, other points are assigned to the nearest clustering center to form a high-reward interval, a medium-reward interval, and a low-reward interval, and a priority label is assigned to each interval. The high-reward interval corresponds to a low proportion of abnormal scenarios and has the highest priority. Within each reward function interval, according to the load deviation degree of the device load factor from the preset load threshold (the percentage of the absolute difference between the actual load and the threshold to the threshold), the reference value of the penalty factor is calculated. The greater the deviation degree, the higher the penalty value. The sliding average method is used to smooth the reference value of the penalty factor. At each data point, the reference values of the penalty factors of the two adjacent data points are averaged to eliminate the instantaneous load fluctuation noise, and the smoothed dynamic cost penalty distribution is obtained. The voltage fluctuation weight coefficient is multiplied by the disturbance simulation amplitude to generate a boundary adjustment coefficient. In this embodiment, the boundary adjustment coefficient is equal to the voltage fluctuation weight multiplied by the normalized value of the disturbance amplitude, and the normalized value of the disturbance amplitude is the ratio of the current disturbance amplitude to the maximum disturbance amplitude. The high-priority reward interval in the dynamic cost penalty distribution is compensated and corrected according to the boundary adjustment coefficient, and the compensation amount is equal to the penalty value multiplied by the boundary adjustment coefficient. This step preferentially compensates for the deviation of the reward parameters in high-disturbance scenarios by fusing voltage fluctuation and disturbance amplitude information. The corrected penalty distribution and the initial reward function values are inversely weighted and superimposed, specifically, inversely superimposed according to the priority weights. The high-priority weight accounts for 70%, and the medium and low priorities each account for 15%. Through range normalization processing, the superimposed result is mapped to the [0, 1] interval to obtain the device operation reward parameters of the new energy power generation device. These device operation reward parameters can comprehensively reflect the performance of the device under different operating states and provide strong support for the optimal operation and scheduling of the device.
[0092] S5. Based on the device operation reward parameters, state-response data pairs with the same time decay coefficient are extracted from the pre-constructed distributed shared experience pool to form a shared experience correction set.
[0093] In some embodiments, the step of extracting state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameters to form a shared experience correction set includes:
[0094] Extract the benchmark time decay coefficient from the device operation reward parameters as a matching benchmark, and match the benchmark time decay coefficient with the time decay coefficients of each data pair in the distributed shared experience pool to extract a candidate set of state-response data with the same time decay coefficient;
[0095] Calculate the response delay difference between the action execution response time of each state-response data pair in the candidate set of state-response data and the actual execution timestamp of the device response;
[0096] Based on the response delay difference, use the sliding window averaging method to align the timestamps in the candidate set of state-response data to obtain a set of time-synchronized state-response data;
[0097] Divide the set of state-response data into different time windows, and fuse the state-response data of multiple new energy power generation devices within the same time window to construct a chronologically continuous shared experience correction set.
[0098] Specifically, in this embodiment, the benchmark time decay coefficient is extracted from the device operation reward parameters. This benchmark time decay reflects the decay characteristic of the reward parameters over time. For example, in the reward parameters of a photovoltaic inverter, assuming its benchmark time decay coefficient is 0.95, it means that for each unit of time passed, the reward value decays to 95% of the original. Traverse the pre-constructed distributed shared experience pool, which stores a large number of state-response data pairs from different new energy power generation devices. Each data pair contains information such as device state characteristics, action execution records, response timestamps, and corresponding time decay coefficients. For each data pair in the experience pool, calculate the absolute difference percentage between its time decay coefficient and the benchmark time decay coefficient. Then, in this embodiment, a preset absolute difference threshold is set. If the difference percentage is less than or equal to this absolute difference threshold, it is determined that the time decay coefficient of this data pair is consistent with the benchmark time decay coefficient. Filter out all qualified data pairs to form a candidate set of state-response data, and at the same time eliminate invalid data with excessive time decay differences.
[0099] For each data pair in the state-response data candidate set, calculate the difference between the response time after the action execution and the actual execution timestamp of the device response to obtain the response delay difference. Meanwhile, in this embodiment, a dynamic delay threshold is set. If the response delay difference exceeds the dynamic delay threshold, the data pair with the response delay difference exceeding the dynamic delay threshold is determined as misleading data caused by device response delay and is excluded. For the remaining data, this embodiment uses the sliding window averaging method to align timestamps. The length of the sliding window can be set to twice the mean of the delay differences. The timestamps of each data pair are averaged within this window to obtain the new aligned timestamps, thereby eliminating the clock deviation of distributed devices. After alignment processing, a time-synchronized state-response data set is obtained. Then, this embodiment divides the time-synchronized state-response data set according to a preset time window. For the state-response data of multiple devices within the same time window, the Kalman filter algorithm is used for fusion. Specifically, a state space model is constructed with voltage and power as state variables. For example, the state variables of a photovoltaic inverter can include the DC-side voltage, AC-side output power, etc. Then, the state value at the current moment is predicted based on the state estimate value at the previous moment, and the prediction value is corrected using the multi-device data actually observed at the current moment to update the state estimate value and the covariance matrix, thereby fusing multi-device data and suppressing noise interference. A timing continuity label is added to the fused data to ensure that there are no breakpoints on the time axis. Finally, a multi-device fusion data set with continuous timing and noise suppression is obtained, thus forming a shared experience correction set.
[0100] S6. Based on the shared experience correction set and the state feature representation, evaluate the action selection conflict probability of each new energy power generation device, and according to the action selection conflict probability, use the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.
[0101] In some embodiments, the step of evaluating the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representation includes:
[0102] According to the state-response data in the shared experience correction set, calculate the abnormal cumulative value of each new energy power generation device, and based on the abnormal cumulative value, identify high-conflict-risk agents from the new energy power generation devices;
[0103] Taking the main sorting basis as the abnormal voltage fluctuation amplitude and the secondary sorting basis as the time decay coefficient, use the entropy weight method to calculate the action priority of each high-conflict-risk agent to obtain the action priority agent sequence;
[0104] Based on the action priority agent sequence, calculate the overlapping ratio of the power adjustment ranges of adjacent high-conflict-risk agents with action priorities, and based on the overlapping ratio of the power adjustment ranges, determine the potential conflict boundary;
[0105] Map the real-time action instructions to the power regulation intervals of potential conflict boundaries, and count the out-of-bounds frequencies and out-of-bounds amplitudes of the action instructions of each high-conflict-risk agent.
[0106] Calculate the conflict intensity index based on the product of the out-of-bounds frequency and the out-of-bounds amplitude of the action instructions, and normalize the conflict intensity index to obtain the action selection conflict probability of each new energy power generation device.
[0107] Specifically, in this embodiment, state-response data is extracted from the shared experience correction set. These data include the real-time states (such as voltage, current, power, etc.) of new energy power generation devices and corresponding responses (such as adjustment actions, fault information, etc.). According to the state-response data of each new energy power generation device, calculate its abnormal cumulative value within a period of time, which can be specifically calculated by counting the number of occurrences of abnormal states. Identify the new energy power generation devices with abnormal cumulative values exceeding the abnormal threshold as high-conflict-risk agents. Then, in this embodiment, extract the voltage abnormal fluctuation amplitude data of each high-conflict-risk agent from the shared experience correction set as the main sorting basis. The voltage abnormal fluctuation amplitude can be expressed as the absolute value of the difference between the actual voltage and the rated voltage, which reflects the instability of the device voltage state. Extract the time decay coefficient of each high-conflict-risk agent as the secondary sorting basis. The time decay coefficient is used to consider the influence of historical abnormalities on the current state. Perform normalization processing on the voltage abnormal fluctuation amplitude and the time decay coefficient to eliminate the dimension and magnitude differences, and use the entropy weight method to calculate the action priority of each high-conflict-risk agent. The entropy weight method determines its weight by calculating the entropy value of each sorting basis, calculates the comprehensive sorting score of each high-conflict-risk agent according to the weight, and sort the high-conflict-risk agents from high to low according to the comprehensive sorting score to obtain the action priority agent sequence.
[0108] In this embodiment, in the action priority agent sequence, power adjustment range data of high-conflict-risk agents with adjacent action priorities is obtained. The power adjustment range represents the output power interval that the device can adjust in the current operating state. The length of the overlapping part of the power adjustment ranges of high-conflict-risk agents with adjacent action priorities is calculated. Based on the length of the overlapping part of the power adjustment ranges, the overlapping ratio of the power adjustment ranges of high-conflict-risk agents with adjacent action priorities is calculated. The area where the overlapping ratio of the power adjustment ranges exceeds the overlapping threshold is determined as the potential conflict boundary. Then, the real-time action instructions of each high-conflict-risk agent are obtained, and the real-time action instructions are mapped to the power adjustment interval of the potential conflict boundary to evaluate whether the action instructions exceed the boundary. For example, if the power adjustment range of the potential conflict boundary is [150, 250], and the power adjustment value corresponding to the real-time action instruction of a certain agent is 220, then this value falls within the conflict boundary. At the same time, within a certain time range, the out-of-bound frequency and out-of-bound amplitude of the action instructions of each high-conflict-risk agent are counted. The out-of-bound frequency refers to the number of times the action instruction exceeds its own power adjustment range or the power adjustment interval of the potential conflict boundary. The out-of-bound amplitude refers to the absolute value of the difference between the out-of-bound action instruction value and the boundary value of the power adjustment range. Based on the product of the out-of-bound frequency and out-of-bound amplitude of the action instructions, the conflict intensity index of each high-conflict-risk agent is calculated. This conflict intensity index reflects the severity of the out-of-bound action instructions. The conflict intensity indices of all high-conflict-risk agents are normalized to obtain the action selection conflict probabilities of each new energy power generation device. Normalization can ensure that the conflict probability values of all devices are between 0 and 1, which is convenient for comparison and evaluation.
[0109] In some embodiments, the step of using the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression according to the action selection conflict probability includes:
[0110] According to the action selection conflict probabilities of each new energy power generation device and the real-time state characteristics of the power grid, determine the corresponding initial perturbation frequency values;
[0111] Taking minimizing the conflict probability and maximizing the power output efficiency as the optimization objectives, based on the optimization objectives, use the gradient optimization algorithm to calculate the partial derivatives of the initial perturbation frequency values and the power adjustment ranges to obtain the gradient optimization directions;
[0112] Gradually adjust the perturbation frequency values and the power adjustment ranges according to the gradient optimization directions, and through iterative calculations until the optimization objectives converge, and output the multi-device collaborative control optimization strategy after conflict suppression.
[0113] Specifically, in this embodiment, the action selection conflict probability of each new energy power generation device and the real-time state characteristic data of the power grid are collected. Among them, the real-time state characteristics of the power grid may include key parameters such as the power grid voltage deviation rate and the load rate. In this embodiment, according to the action selection conflict probability, the pre-constructed conflict-disturbance mapping relation table is queried to obtain the basic disturbance frequency value, and the basic disturbance frequency value is dynamically compensated according to the real-time state characteristics of the power grid. The basic disturbance frequency value is added to the dynamic compensation value to obtain the initially disturbed frequency value after dynamic adjustment. For example, when the voltage deviation rate increases by 1% each time, the disturbance frequency value is compensated by 0.5 Hz; when the load rate increases by 10% each time, the disturbance frequency value is compensated by 0.3 Hz. Then, this embodiment defines the optimization objective as minimizing the action selection conflict probability and maximizing the power output efficiency to ensure that while reducing conflicts, the power output of the device is increased as much as possible, and a multi-objective optimization function is constructed. Specifically:
[0114]
[0115] In the formula, is the multi-objective optimization function value; is the action selection conflict probability; is the power output efficiency; and are both weight coefficients. The weight coefficients can be dynamically adjusted according to the real-time stability of the power grid voltage and are used to balance the relationship between conflict probability suppression and power output efficiency improvement.
[0116] In this embodiment, the numerical differential method is used to calculate the partial derivatives of the multi-objective optimization function value with respect to the disturbance frequency value and the power adjustment range to obtain the change rate of the objective function value. According to the change rate of the objective function value, the direction in which the objective function value changes fastest is found to determine the gradient descent direction. For example, if the calculation result of the partial derivative of the initially disturbed frequency value is negative, it means increasing the initially disturbed frequency value, and vice versa; similarly, the adjustment direction of the power adjustment range is determined. In this embodiment, the disturbance frequency value and the power adjustment range are gradually adjusted according to the gradient optimization direction. The new initially disturbed frequency value is the difference between the initially disturbed frequency value and the partial derivative of the initially disturbed frequency value, and the new power adjustment range is the difference between the power adjustment range and the partial derivative of the power adjustment range. After each adjustment, the multi-objective optimization function value is recalculated according to the updated initially disturbed frequency value and power adjustment range. When the change rates of the multi-objective optimization function values for three consecutive iterations are all less than the preset change rate threshold, it is determined that the optimization process converges, and the iterative calculation is terminated. When the optimization objective converges, the multi-device collaborative control optimization strategy after conflict suppression is output. This strategy includes the final disturbance frequency value and power adjustment range of each device, as well as the corresponding control instructions. The optimization strategy is deployed to the actual new energy power generation device control system to guide the operation and control of the device, realizing effective conflict suppression and efficient operation of the system.
[0117] An embodiment of the present invention provides an adaptive distributed new energy power generation equipment collaborative control method, and the method includes extracting state feature representations of each new energy power generation equipment from the operation parameters of the new energy power generation equipment collected in real time; based on the state feature representations, using a deep reinforcement learning algorithm to determine action boundary estimated values of each new energy power generation equipment in a voltage abnormal fluctuation scenario; according to the action boundary estimated values, extracting equipment state-action data pairs from a historical operation database, and according to the equipment state-action data pairs, generating an empirical expansion data set by simulating different disturbance amplitudes; based on the data sampling frequency and abnormal scenario proportion in the empirical expansion data set, obtaining an equipment operation reward parameter of the new energy power generation equipment by using a weighting method; based on the equipment operation reward parameter, extracting state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool to form a shared experience correction set; based on the shared experience correction set and the state feature representations, evaluating the action selection conflict probability of each new energy power generation equipment, and according to the action selection conflict probability, iteratively determining a multi-equipment collaborative control optimization strategy after conflict suppression by using a gradient optimization algorithm. Compared with the prior art, through the combination of the deep reinforcement learning algorithm and the gradient optimization strategy, this method adaptively determines the collaborative control strategy of new energy power generation equipment in a complex power grid environment, realizes the efficient collaborative control and conflict suppression of new energy power generation equipment in a voltage abnormal fluctuation scenario, improves the adaptability of the equipment to voltage abnormal fluctuations, and significantly improves the stability and power generation efficiency of the power grid.
[0118] It should be noted that the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or posterior. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0119] In one embodiment, as Figure 2 shown, an embodiment of the present invention provides an adaptive distributed new energy power generation equipment collaborative control system, and the system includes:
[0120] A data acquisition module 101, configured to extract state feature representations of each new energy power generation equipment from the operation parameters of the new energy power generation equipment collected in real time;
[0121] A boundary estimation module 102, configured to, based on the state feature representations, use a deep reinforcement learning algorithm to determine action boundary estimated values of each new energy power generation equipment in a voltage abnormal fluctuation scenario;
[0122] A data expansion module 103, configured to, according to the action boundary estimated values, extract equipment state-action data pairs from a historical operation database, and according to the equipment state-action data pairs, generate an empirical expansion data set by simulating different disturbance amplitudes;
[0123] A reward analysis module 104, configured to obtain device operation reward parameters of new energy power generation equipment by using a weighting method based on the data sampling frequency and the abnormal scenario proportion in the experience expansion dataset;
[0124] An experience correction module 105, configured to extract state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameters, and form a shared experience correction set;
[0125] A policy generation module 106, configured to evaluate the action selection conflict probability of each new energy power generation equipment based on the shared experience correction set and the state feature representation, and use a gradient optimization algorithm to iteratively determine a multi-device collaborative control optimization policy after conflict suppression according to the action selection conflict probability.
[0126] For the specific limitations on an adaptive distributed new energy power generation equipment collaborative control system, reference can be made to the above limitations on an adaptive distributed new energy power generation equipment collaborative control method, which will not be elaborated here. Those of ordinary skill in the art can realize that, in combination with the various modules and steps described in the embodiments disclosed in this application, they can be implemented by hardware, software, or a combination of both. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0127] An embodiment of the present invention provides an adaptive distributed new energy power generation equipment collaborative control system. The system extracts the state feature representations of each new energy power generation equipment from the operation parameters of the new energy power generation equipment collected in real time through a data acquisition module; a boundary estimation module determines the action boundary estimation values of each new energy power generation equipment in the voltage abnormal fluctuation scenario by using a deep reinforcement learning algorithm based on the state feature representations; a data expansion module extracts equipment state-action data pairs from the historical operation database according to the action boundary estimation values, and generates an empirical expansion data set by simulating different disturbance amplitudes according to the equipment state-action data pairs; a reward analysis module obtains the equipment operation reward parameters of the new energy power generation equipment by using a weighting method based on the data sampling frequency and the abnormal scenario proportion in the empirical expansion data set; an empirical correction module extracts state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the equipment operation reward parameters to form a shared experience correction set; a policy generation module evaluates the action selection conflict probability of each new energy power generation equipment based on the shared experience correction set and the state feature representations, and iteratively determines the multi-device collaborative control optimization policy after conflict suppression by using a gradient optimization algorithm according to the action selection conflict probability. Compared with the prior art, the system adaptively determines the collaborative control strategy of the new energy power generation equipment in a complex power grid environment through the combination of the deep reinforcement learning algorithm and the gradient optimization strategy, realizes the efficient collaborative control and conflict suppression of the new energy power generation equipment in the voltage abnormal fluctuation scenario, improves the adaptability of the equipment to the voltage abnormal fluctuation, and significantly improves the stability and power generation efficiency of the power grid.
[0128] The above embodiments only express several preferred embodiments of the present application, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art of this technology, without departing from the technical principle of the present invention, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the protection scope of the claims.
Claims
1. An adaptive distributed new energy power generation equipment collaborative control method, characterized in that It includes the following steps: Extract the state feature representations of each new energy power generation device from the real-time collected operation parameters of the new energy power generation devices; Based on the state feature representations, use the deep reinforcement learning algorithm to determine the estimated action boundaries of each new energy power generation device in the voltage abnormal fluctuation scenario; According to the estimated action boundaries, extract the device state-action data pairs from the historical operation database, and according to the device state-action data pairs, generate an empirical extended dataset by simulating different disturbance amplitudes; Extract the power output efficiency and voltage deviation values from the operation parameters of the new energy power generation devices, and based on the data sampling frequency and abnormal scenario proportion in the empirical extended dataset, perform weighted summation on the power output efficiency and voltage deviation values to obtain the device operation reward parameter of the new energy power generation device; The device operation reward parameter is used to characterize the performance of the device in different operation states; Based on the device operation reward parameter, extract the state-response data pairs with the same time decay coefficient from the pre-constructed distributed shared experience pool to form a shared experience correction set; Based on the shared experience correction set and the state feature representations, evaluate the action selection conflict probability of each new energy power generation device, and according to the action selection conflict probability, use the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression; The step of generating the empirical extended dataset by simulating different disturbance amplitudes includes calculating the action data corresponding to the new state after the disturbance, and adding the generated new state-action data pairs to the original dataset to form the empirical extended dataset; The step of extracting the state-response data pairs with the same time decay coefficient from the pre-constructed distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set includes: Extract the reference time decay coefficient from the device operation reward parameter as the matching reference, and match the reference time decay coefficient with the time decay coefficients of the state-response data pairs in the distributed shared experience pool to extract the candidate set of state-response data pairs with the same time decay coefficient; Fuse the state-response data of multiple new energy power generation devices to construct a temporally continuous shared experience correction set; Among them, the pre-constructed distributed shared experience pool stores the state-response data pairs from different new energy power generation devices; The reference time decay coefficient represents the proportional coefficient of the reward value in the device operation reward parameter decaying to the original value every unit time.
2. The collaborative control method for an adaptive distributed new energy power generation device according to claim 1, wherein The step of extracting the state feature representations of each new energy power generation device from the real-time collected operation parameters of the new energy power generation devices includes: Real-time collect the grid voltage data and the operation parameters of the new energy power generation devices, and use the preset data sampling frequency to perform real-time calculation on the grid voltage data to obtain the voltage frequency deviation value; Compare the voltage frequency deviation value with the preset frequency deviation threshold. If the voltage frequency deviation value exceeds the preset frequency deviation threshold, extract the state space descriptions of each new energy power generation device from the operation parameters of the new energy power generation devices; Based on the state - space description, the state - feature representation of each new - energy power - generation device is extracted using the principal - component analysis method.
3. The collaborative control method for an adaptive distributed new energy power generation device according to claim 2, characterized in that: The state - space description includes the current disturbance - simulation amplitude, the device load factor, and the output upper limit of the new - energy power - generation device.
4. The collaborative control method for an adaptive distributed new energy power generation device according to claim 1, characterized in that, The steps of determining the action - boundary estimated value of each new - energy power - generation device in the voltage - abnormal - fluctuation scenario using the deep - reinforcement - learning algorithm based on the state - feature representation include: Introduce the action - space dimension constraint, and use the deep Q - network to perform mapping analysis on the state - feature representation to calculate the initial action - space range of each new - energy power - generation device. Detect the initial action - space range. If it is detected that the initial action - space range is in a narrow range, classify the action - space restricted state of the narrow range to obtain the restricted - state classification result. Taking the disturbance - simulation amplitude and the device load factor as clustering features, use the K - means clustering algorithm to analyze the new - energy power - generation devices with the restricted - state classification result of the restricted state to obtain the restricted - level distribution - density feature. If the restricted - level distribution - density feature exceeds the preset distribution - density threshold, input the restricted - level distribution - density feature into the deep Q - network for secondary calculation to generate the dynamic compensation coefficient of the initial action - space range. Use the dynamic compensation coefficient to perform dynamic - boundary secondary optimization on the initial action - space range to obtain the action - boundary estimated value of each new - energy power - generation device in the voltage - abnormal - fluctuation scenario.
5. The collaborative control method for an adaptive distributed new energy power generation device according to claim 4, wherein, The steps of detecting the initial action - space range include: Use the preset time - window length to segment and divide the initial action - space range, and extract the time - series change feature of the power - regulation range and the temperature - limit - condition fluctuation data in each time window. Analyze the dynamic coupling relationship between the time - series change feature of the power - regulation range and the temperature - limit - condition fluctuation data through the linear - regression algorithm, and calculate the dynamic - change trend coefficient of the two. Compare the dynamic - change trend coefficient with the preset narrow - range threshold. If the dynamic - change trend coefficient is less than the narrow - range threshold, it is determined that the initial action - space range is in a narrow range.
6. The collaborative control method for an adaptive distributed new energy power generation device according to claim 1, characterized in that, The steps of obtaining the device - operation reward parameter of the new - energy power - generation device using the weighting method based on the data - sampling frequency and the abnormal - scenario proportion in the experience - expansion dataset include: Calculate the initial reward - function value using the weighted - sum method through the data - sampling frequency and the abnormal - scenario proportion in the experience - expansion dataset. According to the mapping relationship between the distribution density of the initial reward - function value and the abnormal - scenario proportion, use the density - peak clustering algorithm to obtain the reward interval. In each reward interval, calculate the penalty - factor reference value according to the load deviation degree between the device load factor and the preset load threshold, and use the moving - average method to smooth the penalty - factor reference value to obtain the smoothed dynamic - cost penalty distribution. Compensate and correct the dynamic - cost penalty distribution in the high - disturbance reward interval based on the disturbance - simulation amplitude to obtain the corrected penalty distribution. Perform anti - weight superposition on the corrected penalty distribution and the initial reward - function value, and normalize to obtain the device - operation reward parameter of the new - energy power - generation device.
7. An adaptive distributed new energy power generation equipment collaborative control method according to claim 1, characterized in that The step of extracting state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set includes: Extract the benchmark time decay coefficient from the device operation reward parameter as the matching benchmark, and match the benchmark time decay coefficient with the time decay coefficients of each data pair in the distributed shared experience pool to extract a candidate set of state-response data with the same time decay coefficient; Calculate the response delay difference between the action execution response time of each state-response data pair in the candidate set of state-response data and the actual execution timestamp of the device response; Based on the response delay difference, use the sliding window averaging method to align the timestamps in the candidate set of state-response data to obtain a set of time-synchronized state-response data; Divide the set of state-response data into different time windows, and fuse the state-response data of multiple new energy power generation devices within the same time window to construct a temporally continuous shared experience correction set.
8. An adaptive distributed new energy power generation equipment collaborative control method according to claim 1, characterized in that, The step of evaluating the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representation includes: Calculate the abnormal cumulative value of each new energy power generation device according to the state-response data in the shared experience correction set, and identify high-conflict-risk agents from the new energy power generation devices according to the abnormal cumulative value; Taking the voltage abnormal fluctuation amplitude as the main sorting basis and the time decay coefficient as the secondary sorting basis, use the entropy weight method to calculate the action priority of each high-conflict-risk agent to obtain an action priority agent sequence; Based on the action priority agent sequence, calculate the overlapping ratio of the power adjustment ranges of adjacent high-conflict-risk agents with action priorities, and determine the potential conflict boundary according to the overlapping ratio of the power adjustment ranges; Map the real-time action instruction to the power adjustment interval of the potential conflict boundary, and count the out-of-bounds frequency and out-of-bounds amplitude of the action instructions of each high-conflict-risk agent; Calculate the conflict intensity index according to the product of the out-of-bounds frequency and out-of-bounds amplitude of the action instruction, and normalize the conflict intensity index to obtain the action selection conflict probability of each new energy power generation device.
9. An adaptive distributed new energy power generation equipment collaborative control method according to claim 1, characterized in that, The step of using the gradient optimization algorithm to iteratively determine the multi-device cooperative control optimization strategy after conflict suppression according to the action selection conflict probability includes: Determine the corresponding initial disturbance frequency value according to the action selection conflict probability of each new energy power generation device and the real-time state characteristics of the power grid; Taking minimizing the conflict probability and maximizing the power output efficiency as the optimization objectives, based on the optimization objectives, use the gradient optimization algorithm to calculate the partial derivatives of the initial disturbance frequency value and the power adjustment range to obtain the gradient optimization direction; Gradually adjust the disturbance frequency value and the power adjustment range according to the gradient optimization direction, and through iterative calculation until the optimization objective converges, and output the multi-device cooperative control optimization strategy after conflict suppression.
10. An adaptive distributed new energy power generation equipment collaborative control system, characterized in that, Applying the adaptive distributed new energy power generation device cooperative control method according to any one of claims 1 to 9, the system includes: A data acquisition module, which is used to extract the state feature representations of each new energy power generation device from the operation parameters of the new energy power generation devices collected in real time; A boundary estimation module, which is used to determine the action boundary estimation values of each new energy power generation device in the voltage abnormal fluctuation scenario based on the state feature representations by using a deep reinforcement learning algorithm; A data expansion module, which is used to extract device state-action data pairs from the historical operation database according to the action boundary estimation values, and generate an empirical expansion data set by simulating different disturbance amplitudes according to the device state-action data pairs; A reward analysis module, which is used to obtain the device operation reward parameters of the new energy power generation device by using a weighting method based on the data sampling frequency and abnormal scenario proportion in the empirical expansion data set; An empirical correction module, which is used to extract state-response data pairs with the same time decay coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameters to form a shared experience correction set; A policy generation module, which is used to evaluate the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representations, and iteratively determine the multi-device collaborative control optimization policy after conflict suppression by using a gradient optimization algorithm according to the action selection conflict probability.
Citation Information
Patent Citations
Continuous reinforcement learning incomplete information game method and device based on afterward review and progressive extension
CN114048834A
Underwater detector cluster adaptive detection method and system based on distributed reinforcement learning
CN119204155A